Method and device for detecting small target in complex environment based on unmanned aerial vehicle image

By using an adaptive feature extraction network and a multi-scale feature fusion network in the image detection of drones, combined with deformable convolution and coordinate attention modules, the accuracy and efficiency of small object detection in complex environments are solved, and a more efficient small object detection effect is achieved.

CN120032212AActive Publication Date: 2025-05-23UNIT 32002 OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411995372.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In complex environments, small target detection based on drone images has problems with accuracy and efficiency, especially misjudgment and misjudgment caused by changes in the scale of the target and complex background.

Method used

Adaptive feature extraction network and multi-scale feature fusion network are used to extract features and add outputs through two extraction branch networks. Combining the deformable convolution and coordinate attention modules, the receptive field is dynamically adjusted and background interference is reduced. At the same time, a multi-scale feature fusion network is used to fuse features at different levels, and multi-scale object detection is performed through multiple detection heads.

Benefits of technology

The accuracy and efficiency of small target detection in complex environments are improved, the occurrence of misjudgment and misjudgment is reduced, and the small targets in the image can be detected more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032212A_ABST
    Figure CN120032212A_ABST
Patent Text Reader

Abstract

The invention discloses a complex environment small target detection method and device based on an unmanned aerial vehicle image, and belongs to the technical field of target detection and artificial intelligence. The method comprises the steps of obtaining an image collected by an unmanned aerial vehicle; performing feature extraction on the image by adopting a self-adaptive feature extraction network; the feature extraction network firstly extracts features through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs the added features; fusing features extracted from different layers of the adaptive feature extraction network by using a multi-scale feature fusion network; carrying out target detection by adopting a target detection network; and the detection head performs smaller target detection based on the fusion feature of the high-resolution feature of the shallow layer of the adaptive feature extraction network and the high-resolution feature of the shallow layer of the multi-scale feature fusion network. According to the invention, more accurate small target detection in a complex environment can be realized based on the aerial image of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target detection and artificial intelligence technology, and specifically relates to a method and device for detecting small targets in complex environments based on drone images. Background Art

[0002] With the increasing application of drones in various fields, especially in military reconnaissance, traffic monitoring, search and rescue operations and other scenarios, target detection technology based on drone platforms has developed rapidly.

[0003] However, due to the constant changes in the height and angle of the drone during aerial photography, the proportion of the target object in the obtained image often changes accordingly. In addition, when detecting small targets, the original small target pixels in the image are relatively small, and the image background will also produce varying degrees of interference during the detection process. These factors make it particularly difficult to accurately and effectively obtain the features of small targets.

[0004] At present, deep learning network methods for detecting drone aerial images mainly include the following two categories:

[0005] The two-stage detection method based on candidate regions first generates corresponding candidate regions using drone aerial images, then extracts target features contained in the candidate regions, and finally classifies and regresses the targets based on the extracted feature information. The advantage of this type of method is that it has superior detection accuracy and precision, but is slow.

[0006] A single-stage detection method based on regression. This method regards the target detection of drone aerial images as a regression problem. It does not need to generate corresponding proposal regions, but directly performs feature extraction and classification regression on the entire image to achieve end-to-end target detection. Representative networks include SSD, YOLO, RetinaNet, etc. Most of these small target detection methods are based on single-scale or fixed feature extraction strategies, which cannot fully cope with the scale changes and complex background problems of the target. For example, although the YOLO series of algorithms can complete small target detection to a certain extent, they are not sensitive enough to the scale of the target, especially small targets in complex scenes are easily misjudged or missed. Summary of the invention

[0007] In view of the above problems, the present invention discloses a small target detection method and device based on complex background UAV images, so as to improve the accuracy of small target detection in a complex environment.

[0008] A first aspect of the present invention provides a method for detecting small targets in complex environments based on drone images, comprising:

[0009] Obtain images collected by drones;

[0010] An adaptive feature extraction network is used to extract features from the image; the feature extraction network first extracts features respectively through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them; wherein the first extraction branch network is composed of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network is composed of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence;

[0011] Using a multi-scale feature fusion network to fuse the features extracted from different layers of the adaptive feature extraction network;

[0012] A target detection network is used for target detection; wherein the target detection network includes multiple detection heads, one of which detects smaller targets based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.

[0013] In other examples of the above method, the adding and outputting the features extracted by the two extraction branch networks includes: scaling the features extracted by the two extraction branch networks to the same size, and linearly adding the feature matrices and outputting them through a CBS module.

[0014] In other examples of the above method, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features of different scales; the feature fusion network includes three fusion branches, each fusion branch is used to fuse features of one scale and their convolution results.

[0015] In other examples of the above method, the multi-scale feature fusion network is used to fuse the features extracted by different layers of the adaptive feature extraction network, and further includes: fusing the features output by the three fusion branches in pairs: F(i, j) = Resize[F(i)]*F(j)

[0016] Among them, F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features of different scales to the same scale, and * represents cross-correlation.

[0017] In other examples of the above method, the plurality of detection heads further include:

[0018] A small target detection head receives high-resolution features from a shallow layer of the multi-scale feature fusion network to perform small target detection;

[0019] A medium target detection head, which receives medium resolution features from the middle layer of the multi-scale feature fusion network to perform medium target detection;

[0020] The large object detection head receives low-resolution features from the deep layer of the multi-scale feature fusion network to perform large object detection.

[0021] A second aspect of the present invention provides a small target detection device in a complex environment based on drone images, comprising:

[0022] An input module, configured to obtain images collected by a drone;

[0023] A feature extraction module is configured to use an adaptive feature extraction network to extract features from the image; the feature extraction network first extracts features respectively through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them; wherein the first extraction branch network is composed of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network is composed of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence;

[0024] A feature fusion module is configured to fuse features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network;

[0025] The target detection module is configured to use a target detection network for target detection; wherein the target detection network includes multiple detection heads, one of which detects smaller targets based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.

[0026] In other examples of the above device, the adding and outputting the features extracted by the two extraction branch networks includes: scaling the features extracted by the two extraction branch networks to the same size, and linearly adding the feature matrices and outputting them through the CBS module.

[0027] In other examples of the above-mentioned device, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features of different scales; the feature fusion network includes three fusion branches, each fusion branch is used to fuse features of one scale and their convolution results.

[0028] In other examples of the above device, the method of fusing the extracted features using a multi-scale feature fusion network to output features of different dimensions further includes: fusing the features output by the three fusion branches in pairs:

[0029] F(i, j) = Resize[F(i)]*F(j)

[0030] Among them, F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features of different scales to the same scale, and * represents cross-correlation.

[0031] In other examples of the above device, the plurality of detection heads further include:

[0032] A small target detection head receives high-resolution features from a shallow layer of the multi-scale feature fusion network to perform small target detection;

[0033] A medium target detection head, which receives medium resolution features from the middle layer of the multi-scale feature fusion network to perform medium target detection;

[0034] The large object detection head receives low-resolution features from the deep layer of the multi-scale feature fusion network to perform large object detection.

[0035] The present invention can realize more accurate small target detection in complex environments based on unmanned aerial vehicle aerial images. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0037] Figure 1 A schematic diagram of an adaptive feature extraction network according to an embodiment of the present invention;

[0038] Figure 2 A schematic diagram of a multi-scale feature fusion network according to an embodiment of the present invention;

[0039] Figure 3 is a schematic diagram of a target detection network according to an embodiment of the present invention;

[0040] Figure 4 Schematic diagram of the process of a method for detecting small targets in complex environments based on drone images according to an embodiment of the present invention;

[0041] Figure 5 Schematic diagram of a small target detection device in a complex environment based on drone images according to an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0043] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0044] When performing target detection tasks based on UAV platforms, the complex environment such as background complexity, dynamic changes in the surrounding environment, lighting and weather conditions, and terrain undulations may seriously affect the effect and accuracy of target detection. In addition, changes in perspective caused by rotation, tilt, and sideways flight during the flight of the drone will change the appearance and posture of the target. The size of the target in the image varies greatly, small targets may be mixed with the background, and large targets may not appear completely in the field of view. In order to cope with these complex environmental challenges, the target detection algorithm needs to have strong robustness and good generalization ability to ensure that it can work reliably in a changing environment.

[0045] According to an embodiment of the present invention, a target detection model is provided, which can be trained to detect small targets in complex environments based on input drone aerial images. The target detection model includes three parts: an adaptive feature extraction network, a multi-scale feature fusion network, and a target detection network, which are described in detail below.

[0046] (I) Adaptive feature extraction network

[0047] Obtain complex environment images collected by drones and perform feature extraction on the images to obtain features (feature maps) of different scales. The adaptive feature extraction network consists of a deformable convolution (DC) module and a coordinate attention (CA) module, which aims to help the network focus on the key features of small targets and eliminate redundant background interference by dynamically adjusting the receptive field.

[0048] Traditional convolution operations sample a fixed local area, while deformable convolution allows dynamic adjustment of the sampling position by introducing a learnable offset, which allows the convolution kernel to adapt to the geometric transformation of the input features and selectively focus on important feature areas. For small targets or irregularly shaped objects, convolutions with fixed receptive fields may not be able to effectively capture their important features, while deformable convolutions adaptively adjust their receptive fields according to the input data, enhancing the network's ability to detect small targets.

[0049] Coordinate attention achieves effective encoding of location information by decomposing spatial information and channel information. It applies coordinated attention weights on the input feature map, enabling the network to better identify spatial structures and long-range dependencies. Through this attention mechanism, the network can highlight the target area in the feature map while suppressing redundant background information, which allows the network to focus more on key areas when detecting small targets and improve the resolution of local features.

[0050] The present invention combines DC and CA. The feature extraction network can not only flexibly adjust the receptive field through DC to focus on local important features, but also further separate important information from unimportant information through CA on a global scale, which can reduce the interference of background noise, so that the network can effectively identify small targets in complex scenarios, thereby improving the overall detection accuracy and efficiency.

[0051] like Figure 1 As shown, according to an embodiment of the present invention, the deformable convolution module and the coordinate attention module are cross-fused. Among them, the input is an image of a complex environment image collected by a drone after preprocessing. In this embodiment, the feature extraction network includes two branches, the first branch is composed of a CBS module, a DC module, and a CA module connected in sequence, and the second branch is composed of a CBS module, a CA module, and a DC module connected in sequence. The CBS module consists of a Conv layer, that is, a convolution layer, a BN layer, that is, a Batch normalization layer, and a Silu layer, which is an activation function, and is a module commonly used in the YOLO series of target detection algorithms.

[0052] It can be seen that the difference between the two branches is that the order of the deformable convolution module and the coordinate attention module is different. This combination avoids feature loss in the branch processing process, helps to extract more accurate feature information, and reduces the interference of complex background.

[0053] Specifically, the different order of modules in the two branches provides diverse feature extraction capabilities. The first branch uses DC first and then CA to focus on local geometric features earlier, and then uses CA to focus on important areas. The second branch uses CA first and then DC to initially filter information through the attention mechanism in the entire feature map, and then uses DC to achieve more accurate feature alignment. In this way, this diversity helps the network capture different types of features in the input data and enhances the generalization ability of the model.

[0054] In addition, the design of applying DC first and then CA in the first branch allows the DC module to flexibly adjust the receptive field first, and then the CA module strengthens the importance of local features. This order helps to build global context from the details. The second branch applies CA first and then DC, which can first identify the globally important areas through the attention mechanism. The subsequent DC can adjust its receptive field more accurately to ensure in-depth analysis of the detailed structure of these areas.

[0055] In addition, the processing paths of the two branches in different orders allow the model to integrate information through multi-path fusion, thereby reducing dependence on individual path anomalies, and can effectively cope with complex and changing input data, especially when processing tasks that contain both local small targets and global context information. It is more robust. In particular, this design can resolve the contradiction between complex environments and limited UAV computing resources, providing deep models with richer representation and adaptability.

[0056] Then, the feature scales extracted by the two branches are scaled to the same size, and the feature matrices are directly linearly added. Finally, they are output through the CBS module.

[0057] The present invention can automatically adjust the convolution operation to adapt to small targets of different sizes and shapes through an adaptive feature extraction network, realizes dynamic adjustment of the receptive field and adaptive allocation of spatial attention, thereby enhancing the ability to focus on small targets and suppress the background, and improving the accuracy and efficiency of detection and tracking.

[0058] (II) Multi-scale feature fusion network

[0059] According to an embodiment of the present invention, the Figure 2 The multi-scale feature fusion network shown.

[0060] like Figure 2 As shown, the left half of the multi-scale feature fusion network uses Feature Pyramid Networks (FPN).

[0061] FPN provides the detector with rich, semantically informative multi-scale features by constructing a feature pyramid from top to bottom. The core idea of ​​this architecture is to use the information of feature maps at different levels in the convolutional neural network (CNN) to form a feature pyramid from high to low levels, where each layer is responsible for capturing objects at a specific scale.

[0062] FPN is usually based on the classic CNN architecture (such as ResNet, VGG, etc.) as the backbone network, which extracts preliminary features from the input image and outputs feature maps at multiple levels.

[0063] In the backbone network, the higher the level, the lower the spatial resolution of the feature map, but the richer its semantic information. FPN takes advantage of this feature and upsamples these high-level feature maps step by step (usually using bilinear interpolation) from top to bottom (Top-down Pathway).

[0064] In addition, each cross-layer connection allows the upsampled result of the high-level feature map to be added element-wise to the low-level feature map. These lateral connections help combine the detailed low-level texture information with the semantic information of the top layer, enhancing the expressiveness of features at different scales.

[0065] Through the combination of top-down pathways and lateral connections, a set of feature maps with the same number of channels but different resolutions are finally generated. These feature maps together form a feature pyramid, and each map is used to detect objects of a specific scale. For example, the top feature map with a smaller resolution is used to detect large objects, while the bottom feature map with a higher resolution is used to detect small objects.

[0066] Therefore, FPN enhances the detection capability of objects of different sizes by combining feature maps of different resolutions, thereby achieving efficient multi-scale detection. Vertical connections allow semantic information to be shared between feature maps, thereby improving the quality of high-level semantic information in the underlying feature maps. In addition, compared with other methods that process different scales through multiple branches and multiple input paths, FPN is lighter and more computationally efficient, greatly improving the detection performance and accuracy of the model in complex scenarios.

[0067] The FPN output represents three features at different layers. A feature fusion network is designed in the right half of the multi-scale feature fusion network. The feature fusion network includes three fusion branches, each of which is used to fuse the features of a scale and its convolution results. In the feature extraction process, the shallow and deep features are adaptively weighted and fused to ensure that the network can effectively capture the details of small targets. The specific method is as follows:

[0068] First, the three fusion branches perform 1x1 convolution on the three features respectively, and then fuse them with the original features (the features output by FPN) to output the fused features. This process fuses the features between each channel of each feature map, which can be expressed as:

[0069] F′ 1 =Add(Conv 1×1 (F 1 ), F 1 )

[0070] Then, the features output by the three fusion branches are fused pairwise, and the fusion method is:

[0071] F(i, j) = Resize[F(i)]*F(j)

[0072] Among them, the Resize method adjusts features of different scales to the same scale, and obtains three sets of cross-correlation feature maps F(1, 2), F(1, 3) and F(2, 3).

[0073] In tasks such as object detection that require multi-scale feature fusion, the Resize operation is used to adjust feature maps of different scales to the same size, so as to achieve feature fusion (such as Concat, Add, etc.) in the future, ensuring that feature maps from different sources are aligned in the spatial dimension. Resizing methods can be implemented, for example, through bilinear interpolation, bicubic interpolation, nearest neighbor interpolation, and pooling operations. By properly selecting Resize, e This method can effectively align the sizes of feature maps from different sources, lay the foundation for feature fusion or analysis in subsequent layers, and ensure that feature maps are fully utilized in deep learning models.

[0074] (III) Object Detection Network

[0075] Features of different scales have been obtained through the previous feature fusion, and the target detection network performs target detection based on the features of different scales.

[0076] In the detection scheme of the YOLO V5 network, three detection heads (Headl-3) are used to detect target information of different scales:

[0077] Small object detection head Head3 receives high-resolution features from earlier layers of the network. Although the semantic information may not be as rich as the deep features, these features retain more detailed information and are used to detect small objects in the image, such as distant or small objects. In this branch, since the feature map has a high spatial resolution, the operation of detecting small objects is directly managed after convolution processing.

[0078] The medium target detection head Head2 receives the feature map from the middle layer, whose features are integrated with the context information in the pyramid and used to locate medium-sized targets in the image. Usually the feature map has medium spatial resolution, which is suitable for processing target objects that are neither too small nor too large.

[0079] The large object detection head Head1 receives feature maps from deeper layers of the network. These feature maps have low resolution but contain rich semantic information for processing large objects in the image. The lower spatial resolution allows this detection head to efficiently process large-scale objects and has lower requirements for underlying details.

[0080] The three detection heads use different feature map resolutions to focus on target objects of different sizes. Their outputs include bounding box regression information, category probability and confidence. Finally, they will go through combination and post-processing steps (such as non-maximum suppression) to obtain the final detection results.

[0081] Although the design of this multi-scale detection head improves the performance of YOLO v5 in processing multi-scale target scenes, so that both small and large targets can be accurately identified and located. However, as the network depth increases, the gradient feature information of small targets in drone aerial images gradually weakens. In order to prevent the loss of the previous layer feature information and enable the network to better detect the feature information of small targets in the image, according to an embodiment of the present invention, Figure 3 As shown in the figure, a smaller target head detection head Head4 is added based on the YOLO v5 network Head detection.

[0082] exist Figure 3 In the target detection model shown, when the feature extraction network on the left extracts features layer by layer, as the sampling gradually increases, the feature map resolution gradually decreases, and the high-dimensional information gradually increases. The feature fusion network in the middle fuses the feature information and divides it into three dimensions: low, medium, and high. The detection head on the far right detects features of different dimensions to achieve multi-scale target detection. The present invention extracts features from the original high-resolution features of the feature extraction network, which contain rich small target information, and fuses them with the low-dimensional features in the feature fusion network, and inputs the fused features into the added smaller target head detection head Head4, that is, the fourth detection head for target detection. In this way, on the one hand, the network can be fully utilized to obtain the feature information of the front layer, and on the other hand, the network's perception ability of multiple small targets can be further improved, so that the network can more accurately detect small targets in the image.

[0083] Based on the above target detection model, according to one embodiment of the present invention, a method for detecting small targets in complex environments based on drone images is provided. Figure 4 As shown, the method comprises the following steps:

[0084] A first aspect of the present invention provides a method for detecting small targets in complex environments based on drone images, comprising:

[0085] S401, acquiring images collected by the drone;

[0086] S402, extracting features from the image using an adaptive feature extraction network; the feature extraction network first extracts features respectively through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them; wherein the first extraction branch network is composed of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network is composed of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence;

[0087] S403, using a multi-scale feature fusion network to fuse features extracted from different layers of the adaptive feature extraction network;

[0088] S404, using a target detection network to perform target detection; wherein the target detection network includes multiple detection heads, one of which detects smaller targets based on fusion features of high-resolution features of a shallow layer of an adaptive feature extraction network and high-resolution features of a shallow layer of a multi-scale feature fusion network.

[0089] In other examples of the above method, the features extracted by the two extraction branch networks are added and then output, including: scaling the features extracted by the two extraction branch networks to the same size, and linearly adding the feature matrices and then outputting them through the CBS module.

[0090] In other examples of the above method, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features of different scales; the feature fusion network includes three fusion branches, each fusion branch is used to fuse features of one scale and their convolution results.

[0091] In other examples, the features output by the three fusion branches are fused pairwise:

[0092] F(i, j) = Resize[F(i)]*F(j)

[0093] Among them, F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features of different scales to the same scale, and * represents cross-correlation.

[0094] In other examples of the above method, the plurality of detection heads further comprises:

[0095] Small object detection head, which receives high-resolution features from the shallow layer of the multi-scale feature fusion network for small object detection;

[0096] Medium object detection head, which receives medium-resolution features from the middle layer of the multi-scale feature fusion network for medium-resolution object detection;

[0097] Large object detection head, which receives low-resolution features from the deep layer of the multi-scale feature fusion network for large object detection.

[0098] According to another embodiment of the present invention, a small target detection device in a complex environment based on drone images is provided. Figure 5 As shown, the device comprises:

[0099] A second aspect of the present invention provides a small target detection device in a complex environment based on drone images, comprising:

[0100] An input module 501 is configured to obtain images collected by a drone;

[0101] The feature extraction module 502 is configured to use an adaptive feature extraction network to extract features from the image; the feature extraction network first extracts features respectively through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them; wherein the first extraction branch network is composed of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network is composed of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence;

[0102] A feature fusion module 503 is configured to fuse features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network;

[0103] The target detection module 504 is configured to use a target detection network for target detection; wherein the target detection network includes multiple detection heads, one of which detects smaller targets based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.

[0104] In other examples of the above device, the features extracted by the two extraction branch networks are added and then output, including: scaling the features extracted by the two extraction branch networks to the same size, and linearly adding the feature matrices and then outputting them through the CBS module.

[0105] In other examples of the above-mentioned device, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features of different scales; the feature fusion network includes three fusion branches, each fusion branch is used to fuse features of one scale and their convolution results.

[0106] In other examples, the features output by the three fusion branches are fused pairwise:

[0107] F(i, j) = Resize[F(i)]*F(j)

[0108] Among them, F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features of different scales to the same scale, and * represents cross-correlation.

[0109] In other examples of the above device, the plurality of detection heads further include:

[0110] Small object detection head, which receives high-resolution features from the shallow layer of the multi-scale feature fusion network for small object detection;

[0111] Medium object detection head, which receives medium-resolution features from the middle layer of the multi-scale feature fusion network for medium-resolution object detection;

[0112] Large object detection head, which receives low-resolution features from the deep layer of the multi-scale feature fusion network for large object detection.

[0113] The present invention improves the YOLO v5 network:

[0114] The adaptive feature extraction network is used to replace the original network's CSP Darknet network for feature extraction, which can dynamically adjust the receptive field, thereby improving the network's performance in small target detection tasks.

[0115] In addition, the feature fusion network uses three fusion branches to replace the PAN network of the original network, which can effectively enhance the network's perception ability of targets of different scales and reduce the occurrence of missed detection and false detection of small targets.

[0116] In addition, a smaller target detection head is added based on the three detection heads of the original network, and small target detection is performed based on the fusion results of the original high-resolution features extracted by the feature extraction network and the high-resolution features of the feature fusion network, so that the network can more accurately detect smaller targets in the image.

[0117] Therefore, the present invention can realize more accurate small target detection in complex environments based on drone aerial images.

[0118] Although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the inventive concept of the present invention, the technical solutions of the embodiments of the present invention may be modified or replaced by equivalents, which shall not depart from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting small targets in complex environments based on drone images, characterized in that: include: Obtain images collected by drones; Using an adaptive feature extraction network to extract features from the image; The feature extraction network first extracts features respectively through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them; wherein the first extraction branch network is composed of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network is composed of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence; Using a multi-scale feature fusion network to fuse the features extracted from different layers of the adaptive feature extraction network; A target detection network is used for target detection; wherein the target detection network includes multiple detection heads, one of which detects smaller targets based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.

2. The method for detecting small targets in complex environments according to claim 1, characterized in that: The adding and outputting the features extracted by the two extraction branch networks comprises: scaling the feature scales extracted by the two extraction branch networks to the same size, performing linear addition on the feature matrices and outputting them through the CBS module.

3. The method for detecting small targets in complex environments according to claim 1, characterized in that: The multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features of different scales; the feature fusion network includes three fusion branches, each fusion branch is used to fuse features of one scale and their convolution results.

4. The method for detecting small targets in complex environments according to claim 3, characterized in that: The method of fusing the features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network also includes: fusing the features output by the three fusion branches in pairs: F(i,j)=Reize[F(i)]*F(j) Among them, F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features of different scales to the same scale, and * represents cross-correlation.

5. The method for detecting small targets in complex environments according to claim 3, characterized in that: The plurality of detection heads also include: A small target detection head receives high-resolution features from a shallow layer of the multi-scale feature fusion network to perform small target detection; A medium target detection head, which receives medium resolution features from the middle layer of the multi-scale feature fusion network to perform medium target detection; The large object detection head receives low-resolution features from the deep layer of the multi-scale feature fusion network to perform large object detection.

6. A small target detection device in a complex environment based on drone images, characterized in that: include: An input module, configured to obtain images collected by a drone; A feature extraction module is configured to extract features from the image using an adaptive feature extraction network; The feature extraction network first extracts features respectively through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them; wherein the first extraction branch network is composed of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network is composed of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence; A feature fusion module is configured to fuse the features extracted by different layers of the adaptive feature extraction network using a multi-scale feature fusion network to output features of different dimensions; The target detection module is configured to use a target detection network for target detection; wherein the target detection network includes multiple detection heads, one of which detects smaller targets based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.

7. The method for detecting small targets in complex environments according to claim 6, characterized in that: The adding and outputting the features extracted by the two extraction branch networks comprises: scaling the feature scales extracted by the two extraction branch networks to the same size, performing linear addition on the feature matrices and outputting them through the CBS module.

8. The method for detecting small targets in complex environments according to claim 6, characterized in that: The multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features of different scales; the feature fusion network includes three fusion branches, each fusion branch is used to fuse features of one scale and their convolution results.

9. The method for detecting small targets in complex environments according to claim 7, characterized in that: The method of fusing the features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network also includes: fusing the features output by the three fusion branches in pairs: F(i,j)=Resize[F(i)]*F(j) Among them, F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features of different scales to the same scale, and * represents cross-correlation.

10. The method for detecting small targets in complex environments according to claim 7, characterized in that: The plurality of detection heads also include: A small target detection head receives high-resolution features from a shallow layer of the multi-scale feature fusion network to perform small target detection; A medium target detection head, which receives medium resolution features from the middle layer of the multi-scale feature fusion network to perform medium target detection; The large object detection head receives low-resolution features from the deep layer of the multi-scale feature fusion network to perform large object detection.

Citation Information

Patent Citations

  • Unmanned aerial vehicle image detection method based on multi-scale feature fusion and context enhancement

    CN117037004A

  • Forest fire detection method based on improved YOLOv8

    CN118411602A

  • Unmanned aerial vehicle target detection method and system based on multi-scale fusion and dynamic perception

    CN118628940A

  • Unmanned aerial vehicle target detection method based on space-frequency feature fusion detection head

    CN118691929A

  • Method for detecting small target in aerial image of unmanned aerial vehicle

    CN118762168A