Weak and small target detection method for unmanned aerial vehicle image

By introducing a context-scale perception module and a multi-scale refinement module into the object detector, the problem of weak rotational object detection in drone images is solved, and higher recognition accuracy and adaptability are achieved.

CN119942370APending Publication Date: 2025-05-06YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311452714.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing object detectors have difficulty effectively detecting weakly rotated targets in drone images, especially in open world scenarios, limited by the high resolution and complex background of the image.

Method used

Using a context-based object detector based on context-scale perception, local geometric information and global semantic features are learned through the context-aware module and the multi-scale refinement module to generate discriminant representations and context information to improve detection performance.

Benefits of technology

The recognition accuracy of weak rotation targets is significantly improved in the drone images, and is suitable for open-world application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942370A_ABST
    Figure CN119942370A_ABST
Patent Text Reader

Abstract

The invention discloses a weak and small target detection method for an unmanned aerial vehicle image. The method has certain universality in the target detection direction. In order to solve the problem that weak and small target detection difficulty is large, the invention provides a context sensing module which utilizes a feature extractor with a deformable convolutional layer and a non-local-based global context (GC) module to jointly learn local geometric information and global semantic features. Meanwhile, a path enhancement network (PANet) is used for learning multi-scale features, then a refined block is applied, and high-level semantic information and low-level position information of different layers are integrated, so that the target positioning capability is improved. And higher recognition precision can be achieved in weak and small target detection based on the unmanned aerial vehicle image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rotating target detection in deep learning, and is aimed at small target detection, especially rotating target detection technology in drone images. Background Art

[0002] Over the past few decades, object detection has attracted extensive attention in remote sensing and computer vision. Many milestone works have contributed to its unprecedented development speed. Many challenges in the field have been addressed by typical methods such as region-based convolutional neural networks (RCN) and feature pyramid networks (FPN). In terms of object detection, both the accuracy and speed have been significantly improved. The goal of remote sensing object detection is to locate geospatial objects and identify them in an open environment. Since the categories and connotations of geospatial objects are very large and have a variety of characteristics, remote sensing object detection is more challenging than other types of detection in natural images.

[0003] With the rapid development of remote sensing imaging technology, many types of imaging sensors are being used, including satellite aerial and unmanned aerial vehicle (UAV) cameras, which significantly improve the spatial resolution of images. Because these high-resolution remote sensing (HRRS) images can show sufficient details of objects and backgrounds. Because these high-resolution remote sensing (HRRS) images can show sufficient details of objects and backgrounds, they are gradually becoming a basic data source for remote sensing applications. At the same time, applications such as geological disaster monitoring and urban monitoring are used to detect small-scale objects with weak feature responses, which have high intra-class variations, non-uniform distributions, and noise. Han et al. gave a clear definition of these weak targets: (1) they account for a small proportion of HRRS images; (2) they have sufficient variations in shape, texture, and color features, and are easily affected by imaging conditions and complex backgrounds. Therefore, effectively detecting these small weak objects is of great significance to meet the requirements of open-world applications.

[0004] Compared with the traditional remote sensing technology of satellite and aerial sensors, UAV is a new fine-grained earth observation method developed in the past few decades, and its imaging is highly variable and can achieve higher spatial resolution than satellite images. The above facts mean that the intra-class variability of objects and complex environments is more obvious in UAV images than in satellite and aerial images, especially in terms of scale variation. Some targets may be very large or very small. More related works should be generated to promote the detection of weak and small targets in UAV. In real-world scenarios, some geospatial objects appear in only a small part of UAV images due to the large distance between the imaging sensor and the object. In addition, the scale, color, texture, and shape characteristics of objects have sufficient variations, and their features may be affected by imaging conditions, surrounding objects, and shadows. These challenges increase the difficulty of object detection in UAV images. Existing detectors for general scenes have been widely used for remote sensing object detection. They were originally developed to detect a few object categories and a few instances from noise-free images; therefore, they are not suitable for detecting weak and small targets in UAV images. Summary of the invention

[0005] In order to detect small and weak objects in the open world in drone images, this paper proposes a context-scale-aware object detector with the following main structure: Figure 1 shown.

[0006] The technical solution adopted by the present invention is:

[0007] Step 1: For a given input image, a context-aware module utilizes a feature extractor with deformable convolutional layers and a non-local based global context (GC) module to jointly learn local geometric information and global semantic features.

[0008] Step 2: The above features are integrated by a multi-scale refinement module, which fuses features from different layers and shares semantic and location information.

[0009] Step 3: The region proposal network (RPN) generates many proposed candidate boxes based on the extracted features, and the universal detection head simultaneously regresses the locations of these proposals and the predicted classifications to obtain the final result.

[0010] The context-aware and multi-scale feature refinement modules jointly explore the discriminative representation and contextual information of each object to improve the detection performance.

[0011] Compared with the prior art, the present invention has the following beneficial effects:

[0012] (1) It can achieve higher recognition accuracy in the detection of small targets based on UAV images; BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 Figure 1: Overall network structure and module structure diagram. DETAILED DESCRIPTION

[0014] The present invention is further described below in conjunction with the accompanying drawings.

[0015] First, the process of extracting features from drone images using the network model is as follows: Figure 1 shown.

[0016] In order to extract global and local context information simultaneously, a GC block is added to cooperate with the deformable convolution. Figure 1 As shown in Figure 2, the global context is modeled by a GC module. It is combined with a simplified non-local module to effectively model long-range dependencies in HRRS images. Then, a bottleneck-like module transforms the attention map with lightweight computation. For a specific feature map, the attention map is the same for all query locations and it can be shared among these locations. Therefore, the non-local representation can be simplified to

[0017]

[0018] Among them, x j The order of index query positions, x j List all possible positions on the feature map one by one, W k and W v is the linear transformation matrix. The global background is modeled as the average of all feature locations and summarized as the original feature. Next, a bottleneck-like module is used to significantly reduce the parameters of the transformation module from K·K to 2·K / r. Specifically, when the reduction rate r is set to 16, its parameters are reduced to 1 / 8 of the normal case.

[0019] In this paper, a path augmentation network (PANet) is used to learn multi-scale features, and then a refinement block is applied to integrate high-level semantic information and low-level location information at different layers. Compared with FPN, PANet supplements an additional bottom-up enhancement path. The augmentation path can effectively propagate low-level information and has a high response to edges. This is very helpful for target positioning.

[0020] like Figure 1 As shown, the output features produced by the pre-trained extractor are denoted as {C2, C3, C4, C5}, and the original branch of FPN is denoted as {P2, P3, P4, P5}. From P2 to P5, the downsampling rate is set to 2 by default. The enhancement path of PANet starts from the lowest layer P2 and gradually reaches P5, denoted as {N2, N3, N4, N5}, which correspond to {P2, P3, P4, P5}. N2 is simply P2 without any operation. In order to get each N i, firstly, the high-resolution map N i-1 Downsampling is performed.

[0021] It is first downsampled and then connected horizontally through a 3×3 convolutional layer. i Fusion. In order to integrate the advantages of multi-level features and retain valuable semantics, {N2, N3, N4, N5} must be further refined. They are first resized to an intermediate feature size (such as N4) and interpolated and pooled. Then, a balanced semantic feature N b can be calculated by a simple averaging operation

[0022]

[0023] L is the number of features extracted by PANet, and l is the index of the feature level, ranging from 2 to 5. Then, Gaussian non-local attention further refines the extracted features to make them more discriminative.

[0024] In this case, features from different levels are fused simultaneously. Finally, the refined features are resampled to the scale of the original features. In this process, each feature obtains equivalent information from other features. In this process, each feature obtains equivalent information from other features. This module helps the detector learn the location and semantic information of multi-scale objects simultaneously.

[0025] Specific methods

[0026] (1) For a given input image, a context-aware module utilizes a feature extractor with deformable convolutional layers and a non-local based global context (GC) module to jointly learn local geometric information and global semantic features.

[0027] (2) The above features are integrated by a multi-scale refinement module, which fuses features from different layers and shares semantic and location information.

[0028] (3) The region proposal network (RPN) generates many proposed candidate boxes based on the extracted features, and the universal detection head simultaneously regresses the locations of these proposals and the predicted categories to obtain the final result.

[0029] The context-aware and multi-scale feature refinement modules jointly explore the discriminative representation and contextual information of each object to improve the detection performance.

Claims

1. A small target detection algorithm for drone images, characterized in that: The following steps are involved: Step 1: For a given input image, a context-aware module utilizes a feature extractor with deformable convolutional layers and a non-local based global context (GC) module to jointly learn local geometric information and global semantic features. Step 2: The above features are integrated by a multi-scale refinement module, which fuses features from different layers and shares semantic and location information. Step 3: The region proposal network (RPN) generates many proposed candidate boxes based on the extracted features, and the universal detection head simultaneously regresses the locations of these proposals and the predicted classifications to obtain the final result. The context-aware and multi-scale feature refinement modules jointly explore the discriminative representation and contextual information of each object to improve the detection performance.

Citation Information

Cited By

  • Unmanned aerial vehicle remote sensing small target detection method and system based on space-frequency domain aggregation and multi-scale feature enhancement

    CN121937920A