Aerial image target detection method based on content awareness, program, equipment and storage medium

By adopting content-aware detection methods in aerial image target detection, and using technical means such as CenterNet and difficult area extraction module, the problem of low detection accuracy of small targets in aerial images is solved, and higher detection performance and small target detection effect are achieved.

CN120126033APending Publication Date: 2025-06-10HARBIN ENG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510197273.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The detection accuracy of small targets in aerial images is low, and the general purpose target detector performs poorly in high-resolution images, resulting in small target information loss and degradation of detection performance.

Method used

Using aerial image object detection method based on content perception, CenterNet is used as the basic detector, combining the difficult area extraction module and the global information injection module, through the focus-cropping-detection framework, the network focus area is defined as the area where the detection errors occur the most, and the Gausserstein distance loss function is introduced to improve detection performance.

Benefits of technology

The accuracy and performance of small object detection in aerial images is improved, especially in detecting small objects, and has achieved better results on the VisDrone2021 dataset compared with the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126033A_ABST
    Figure CN120126033A_ABST
Patent Text Reader

Abstract

The invention discloses an aerial image target detection method based on content awareness, a program, equipment and a storage medium, and belongs to the field of aerial image target detection. According to the method, a focusing-cutting-detection framework is used as a basic framework, firstly, a basic detector is used for executing global rough detection on an input image, and meanwhile, a difficult region extraction module predicts a region in which a difficult target is concentrated; and cutting off an area in which the difficult targets are concentrated, and sending the area into a global information injection module to execute fine detection. The method comprises the following steps: firstly, defining a region which a network should focus as a region where errors frequently occur in the network, and proposing a difficult region prediction branch based on reverse attention so as to focus the difficult region in the defined region; meanwhile, an iterative fusion algorithm is provided, a region with concentrated prediction errors is selected as supervision information of a network focusing region, consistent indication signals are provided for a region prediction branch network, and the target detection capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection in aerial images, and particularly to a method, program, device, and storage medium for object detection in aerial images based on content awareness. Background Art

[0002] Object detection in aerial images is a key technology for intelligent interpretation of aerial images, aiming to locate and identify objects of interest from an aerial perspective. With the development of deep learning, object detection algorithms have made significant progress and demonstrated important advantages in downstream tasks. Due to the unique acquisition platform of aerial images, which provides a special perspective and a wide field of view, it plays an important role in the fields of intelligent monitoring and military. However, this also leads to obvious differences between the objects in aerial images and natural images. The proportion of small objects in aerial images is very large, and these objects often concentrate in specific areas, resulting in large background areas, which poses great challenges to object detection.

[0003] General object detectors have proven their effectiveness in object detection in natural images. However, when directly applying these general detectors to aerial images, the performance often drops significantly. In complex aerial scenes, to avoid excessive memory consumption, general object detectors usually downsample high-resolution images. However, this direct downsampling will lead to significant loss of small object information, reduce the representation ability, and thus affect the performance of object detection. Therefore, to address the high-resolution problem, researchers have designed dedicated networks to improve the performance of object detection. Considering the high resolution and uneven distribution characteristics of aerial images, a natural idea is to select local regions for detection. For example, C. Duan et al. in "Coarse-grained density map guided object detection in aerial images" obtain cluster regions by predicting the distribution of objects and specify these regions as key regions for subsequent detection. Y. Wang et al. in "Object detection using clustering algorithm adaptive searching regions in aerial images" select regions with low-confidence object clusters based on the initial detection results as key regions for subsequent detection.

[0004] These solutions help to improve the performance, proving that focus-crop-detect is an effective method for aerial images. Therefore, a clear definition of the region where the network should focus and accurate object localization and classification become the key to improving the performance of object detection in aerial images. Summary of the Invention

[0005] To solve the problem of low accuracy in aerial image target detection, the purpose of the present invention is to propose a content-aware aerial image target detection method. Using the focus-crop-detection framework as the basic framework and CenterNet as the basic detector, first use the basic detector to perform global rough detection on the input image, while the difficult extraction module predicts the areas in the difficult target set. Crop the areas in the difficult target set and then send them into the detector for fine detection. Under the focus-crop-detection framework, this method first defines the area where the network should focus as the area where the network often makes mistakes, and proposes a difficult area prediction branch network based on reverse attention to focus on the area we defined. At the same time, an iterative fusion algorithm is proposed, which selects the area with concentrated prediction errors as the supervision information for the network focus area, providing a consistent indication signal for the area prediction branch network. Finally, in the fine detection process of the local area, the global information injection module (GLI) is used to fuse global information into the local detection to improve the target detection ability.

[0006] The present invention provides a content-aware aerial image target detection method, including the following steps:

[0007] Step 1: Obtain an aerial image dataset to construct a training set, where each sample in the training set includes an image and the corresponding target detection label;

[0008] Step 2: Train using the training set under the anchor-free detector CenterNet network architecture;

[0009] The anchor-free detector CenterNet establishes a difficult area extraction module for identifying and predicting areas in the difficult target set and a global information injection module for local fine detection, and also introduces a Gaussian Wasserstein distance loss function to unify the difficult area prediction branch of the difficult area extraction module and the original detection head, and measures the prediction deviation of targets at multiple scales from the perspective of distribution;

[0010] Step 3: Input the image to be detected into the trained anchor-free detector CenterNet network architecture to obtain the image target detection result.

[0011] Further, the anchor-free detector CenterNet takes the complete image as the input to obtain global rough detection information; the difficult area extraction module uses the generated global rough detection information to predict the difficult areas, crops the predicted areas, and then sends them into the detector again to perform local fine detection.

[0012] Furthermore, the difficult region extraction module first generates the bounding boxes of error-dense regions through the difficult focus label algorithm; then uses the bounding boxes to supervise the prediction of difficult regions to further identify difficult regions; finally, crops the generated focused regions for fine detection.

[0013] Furthermore, the difficult region extraction module introduces a branch in the anchor-free detector CenterNet detection head, taking the global features generated in the global rough detection process as input; first, the global features are input into two dilated convolutional layers and a residual structure to obtain the global feature information and scene information of objects at different granularity levels, then input into the deconvolution layer to increase the spatial resolution of the global feature information, obtaining the feature map heatmap of the object distribution, and then obtaining the difficult region map of the image through the reverse attention module of the difficult region extraction module;

[0014] X' global = Concat(DCon1(X global ), DConv2(X global ), X global )

[0015] heatmap = TCR(X' global )

[0016] where X' global is the feature processed by the residual structure and dilated convolution; X global is the global feature information; Concat(·) is the integration layer; DConv1(·) and DConv2(·) are convolutional layers; TCR(·) is the deconvolution layer.

[0017] Furthermore, the global information injection module extracts and integrates the information of the sub-image feature F local generated in the local image fine detection process and the original image feature F global ; the original feature image F global is processed by two independent convolutional layers to generate F g_embed and F act , the sub-image feature F local is generated into F l_embed through another convolutional layer, and then F g_embed and F l_embed are fused into F attfuse through the attention mechanism, and finally the information is further extracted and integrated through RepBlock;

[0018] F act = sigmoid(Conv_act(F global ))

[0019] Fg_embed = Conv g embed(F global )

[0020] F l_embed = Conv l embed(F local )

[0021] F att_fuse = Att_fuse(F l_embed * F act + F g_embed )

[0022] F out = RepBlock(F att_fuse )

[0023] Among them, sigmoid(·) is the activation function; Conv g embed(·) is the convolutional layer for encoding the original image features; Conv l embed(·) is the convolutional layer for encoding the sub-images; Att_fuse(·) is the attention mechanism.

[0024] Furthermore, the loss function L of the CenterNet det :

[0025] L det = L k + L gwd

[0026] L gwd = 1 - 1 / τ + f(W 2 ), τ ≥ 1

[0027]

[0028] Among them, L k is the focal loss; L gwd is the Gaussian Wasserstein distance loss function; τ is the temperature coefficient; f(x) = ln(x + 1) is a non-linear function; W is the Gaussian distribution of the bounding box; u is the mean of the Gaussian distribution; Σ is the covariance matrix; F is the Frobenius norm; (x, y) is the center coordinate of the bounding box; h and w are the length and width of the bounding box respectively.

[0029] The present invention also provides a computer device / system, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the steps of the content-aware aerial image target detection method described in any one of the above are implemented.

[0030] The present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the content-aware aerial image target detection method described in any one of the above are implemented.

[0031] The present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the content-aware aerial image target detection method described in any one of the above are implemented.

[0032] The beneficial effects of the present invention are as follows:

[0033] First, the present invention proposes a content-aware aerial image target detection method, which uses the anchor-free detector CenterNet as the basic detector for small target detection in aerial images, and introduces the Gaussian Wasserstein distance loss function to better adapt to small target detection. Compared with existing methods, higher performance can be achieved on the aerial image dataset.

[0034] Second, the present invention proposes a difficult region prediction branch network. First, the definition of the focus region is improved, and the focus region of the network is defined as the region where the most detection errors occur. A difficult focus label algorithm is designed to generate the labels of the region where the network should focus, providing consistent supervision information for the difficult prediction branch, so that the network can accurately predict the difficult region. The difficult region prediction branch network first uses dilated convolution to increase the receptive field of the target, and then uses inverse attention to coordinate the inconsistency of the training objectives, and can more accurately focus on the region where the difficult-to-detect target is located.

[0035] Third, the present invention proposes a global information injection module, which incorporates global information into the local detection process, provides useful context clues, and improves the detection ability of targets in difficult regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a framework diagram of a content-aware aerial image target detection method of the present invention;

[0037] Figure 2 is a structural diagram of the difficult region prediction branch network of the present invention;

[0038] Figure 3 is a structural diagram of the global information injection module of the present invention;

[0039] Figure 4 is a comparison result diagram of the present invention and other aerial image target detection methods. DETAILED DESCRIPTION OF THE INVENTION

[0040] The present invention will be further described in detail below in conjunction with the drawings and the specific embodiments.

[0041] As Figure 1 shown, a content-aware aerial image target detection method, which includes the following steps:

[0042] S1. Obtain an aerial image dataset and divide it into a training set, a validation set, and a test set;

[0043] S2. Use the anchor-free detector CenterNet as the basic detector. When the complete image is used as the input, it is called the global rough detection process, and when the local image is used as the input, it is called the local fine detection process;

[0044] S3. Establish a difficult region extraction module (CRE) for predicting the concentrated regions of difficult-to-detect targets. Use the complete image as the input to obtain the global rough detection result. At the same time, use the global information generated in this process as the input of the difficult region prediction branch of the difficult region extraction module to perform difficult region prediction. Crop the predicted region and send it into the detector again to perform local fine detection;

[0045] S4. During the local detection process, establish a global information injection module (GLI) for extracting the global information beneficial to detecting local region targets from the global information and injecting it into the local fine detection process. Use the global information and the local information generated in the local sub-image fine detection process as the input of the global information injection module to provide favorable global information for local detection;

[0046] S5. Improve the loss function in CenterNet, introduce the Gaussian Wasserstein distance loss function to unify the difficult region prediction branch and the original detection head, and measure the prediction deviation of targets at multiple scales from the perspective of distribution;

[0047] S6. Use the joint action of the two modules and the Gaussian Wasserstein distance loss function to constrain the network training process to obtain the detection result.

[0048] Furthermore, the difficult region extraction module in S3 is supervised by the region bounding box labels of the error-dense regions generated by the difficult focusing label algorithm, so that the difficult region extraction module can accurately predict the regions where the wrong targets are dense. The difficult focusing label algorithm uses an improved TIDE tool to evaluate the global detection results of the basic network and returns the region bounding boxes for the wrongly detected targets. Then, these bounding boxes are expanded by ω (Ω = 20) pixels. Identify the overlapping bounding boxes by calculating their IoU, and calculate the largest enclosed bounding box, and iterate this process K times. Use the bounding boxes of these difficult regions to supervise the prediction of the difficult prediction branch network to ensure that the network can consistently identify difficult regions. Finally, crop the generated focused regions for subsequent fine detection. The specific algorithm flow is shown in Table 1 below.

[0049] Table 1 Difficult Focus Label Algorithm Process

[0050]

[0051] As Figure 2 shown, the difficult region extraction module in S3 introduces a branch in the anchor-free detector CenterNet's anchor-free detection head, which uses the global feature X generated during the global rough detection process global as input. First, two dilated convolutions and a residual structure are used to increase the receptive field of the target, obtaining the global position information and scene information of the target at different granularity levels, and stitching them together to enhance the representation ability of the global information. Subsequently, deconvolution is used to increase the spatial resolution of the global feature information to obtain the heatmap, and the difficult map is obtained through the reverse attention module. At this time, the difficult map not only contains the distribution of difficult regions but also contains background information. Finally, the spatial resolution of the difficult map is increased to be the same as that of the original image to display more detailed information about the difficult regions. Under the action of the supervision information, the irrelevant background information is suppressed, and the characteristics of the difficult regions are understood more deeply to capture the location of the difficult target regions. The calculation formulas for the global information and the heatmap are as follows:

[0052] X′ global = Concat(DConv1(X global ), DConv2(X global ), X global )

[0053] heatmap = TCR(X′ global )

[0054] where X' global is the feature processed by the residual structure and dilated convolution; X global is the global feature information; Concat(·) is the integration layer, DConv1(·) and DConv2(·) are convolutional layers, and TCR(·) is the deconvolution layer.

[0055] As Figure 3 shown, the input of the global information injection module in S4 is the sub-image feature (local feature: F local ) generated during the local image fine detection process and the global feature of the original image (global feature: F global ). In the global information injection module, F global is processed through two independent convolutional layers with different parameters but the same structure for F global , generating F g_embed and F act , Flocal Generate F through another convolutional layer l_embed These features are then fused through an attention mechanism to produce F attfuse Finally, through the RepBlock, information is further extracted and integrated. The entire process can be summarized by the following formula:

[0056] F act = sigmoid(Conv_act(F global ))

[0057] F g_embed = Conv g embed(F global )

[0058] F l_embed = Conv 1 embed(F local )

[0059] F att_fuse = Att_fuse(F l_embed * F act + F g_embed )

[0060] F out = RepBlock(F att_fuse )

[0061] where sigmoid(·) is the activation function; Conv g embed(·) is the convolutional layer for encoding the original image features; Conv l embed(·) is the convolutional layer for encoding the sub-images; Att_fuse(·) is the attention mechanism.

[0062] Furthermore, the definition of the Gaussian Wasserstein distance loss function in S5 is as follows:

[0063] L det = L k + L gwd

[0064] L gwd = 1 - 1 / τ + f(W 2 ), τ ≥ 1

[0065]

[0066] where L det is the CenterNet loss, L k is the focal loss, τ is the temperature coefficient, L gwdis the Gaussian Wasserstein distance loss function, u is the mean of the Gaussian distribution, ∑ is the covariance matrix, f(x) = ln(x + 1) is a non-linear function that can smooth the Wasserstein distance and improve the optimization performance, W is the Gaussian distribution of the bounding box, F is the Frobenius norm, (x, y) are the center coordinates of the bounding box, and h and w are the length and width of the bounding box respectively.

[0067] Example 1

[0068] The present invention uses the VisDrone2021 dataset for experiments. This dataset is the largest and most representative drone dataset, consisting of 8,599 images spanning 10 object categories. It includes 6,471 images for training, 548 for validation, and 1,580 for testing. These 10 object categories are awning, bicycle, bus, car, motorcycle, pedestrian, person, truck, tricycle, and minivan. This dataset covers a variety of complex environments and scenarios, such as urban areas, rural areas, highways, etc., across different weather conditions, flight altitudes, times of day, and camera angles.

[0069] A comparative experiment was conducted by taking the method of the present invention and other detection methods, such as Figure 4 As shown, the test results of the comparative experiment are shown in Table 2.

[0070] In the performance comparison on VisDrone, "o" represents the original image, "u" represents uniform cropping (6 sub-images), "fd" represents focus cropping detection, and "aug" represents data augmentation. "*" represents multi-scale inference. Image refers to the number of images input to the detector during the test.

[0071] Table 2 Comparative experiment results

[0072]

[0073] It can be seen from the data that the focus-crop-detect method proposed by the present invention has better results than the uniform cropping method, and it has obvious improvements in all object scales (AP s 、AP m and AP l ), especially in detecting small objects. This proves the effectiveness of this method in capturing difficult regions and detecting small objects.

[0074] The present invention also provides a computer device / equipment / system, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, it implements the steps of the content-aware aerial image target detection method described in any one of the above.

[0075] The present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the content-aware aerial image target detection method described in any one of the above are implemented.

[0076] The present invention also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the content-aware aerial image target detection method described in any one of the above are implemented.

[0077] In summary, the present invention proposes a content-aware aerial image target detection method, which simulates the human visual hierarchical difficulty processing mechanism by updating the definition of the focus area, uses the area where the network prediction error is concentrated as the supervision information, provides clear guidance for the model, and enables it to fully understand the features of the focused area. In the present invention, a difficult area extraction module and a global information injection module are proposed to locate the difficult areas and inject global information during the further fine detection process. The present invention fully considers the uneven distribution of aerial image targets and the large proportion of small targets. Compared with some excellent methods, the best performance is achieved on the VisDrone2021 dataset, proving that the present invention can well handle the target detection task of aerial images.

[0078] The specific embodiments in this specification are only examples and should not be regarded as a limitation to the present invention. Those skilled in the art can modify and optimize the embodiments on the basis of understanding the spirit and scope of the present invention, so as to form various equivalent embodiments, all of which are within the protection scope of the present invention. The protection boundary is subject to the claims, and any technical solution that does not exceed its scope is regarded as the protected object of the present invention.

Claims

1. A content-aware aerial image target detection method, characterized in that: The following steps are involved: Step 1: Obtain the aerial image dataset to build a training set. Each sample in the training set includes an image and a corresponding target detection label. Step 2: Use the training set for training under the anchor-free detector CenterNet network architecture; The anchor-free detector CenterNet establishes a difficult region extraction module for identifying and predicting the concentrated areas of difficult-to-detect targets and a global information injection module for local fine detection. It also introduces a Gauss-Wasserstein distance loss function to unify the difficult region prediction branch of the difficult region extraction module and the original detection head, and measures the prediction deviation of targets of multiple scales from the perspective of distribution. Step 3: Input the image to be detected into the trained anchor-free detector CenterNet network architecture to obtain the image target detection result.

2. The method for aerial image target detection based on content perception according to claim 1, characterized in that: The anchor-free detector CenterNet takes the complete image as input to obtain global rough detection information; the difficult area extraction module uses the generated global rough detection information to predict the difficult area, crops the predicted area, and sends it to the detector again to perform local fine detection.

3. The method for aerial image target detection based on content perception according to claim 2, characterized in that: The difficult region extraction module first generates an error-intensive region bounding box through a difficult focus label algorithm; then uses the region bounding box to supervise the difficult region for prediction and further identifies the difficult region; and finally crops the generated focus region for fine detection.

4. The method for aerial image target detection based on content perception according to claim 2, characterized in that: The difficult region extraction module introduces a branch in the detection head of the anchor-free detector CenterNet, and takes the global features generated in the global rough detection process as input; firstly, the global features are input into two dilated convolutional layers and a residual structure to obtain the global feature information and scene information of targets at different granularity layers, and then input into the deconvolution layer to increase the spatial resolution of the global feature information, and obtain the feature map heatmap of the target distribution, and then the reverse attention module of the difficult region extraction module obtains the difficult region map of the image; X' global =Concat(DCon1(X global ),DConv2(X global ),X global ) heatmap=TCR(X' global ) Among them, X' global is the feature processed by residual structure and dilated convolution; X global is the global feature information; Concat(·) is the integration layer; DConv1(·) and DConv2(·) are convolutional layers; TCR(·) is the deconvolution layer.

5. The method for aerial image target detection based on content perception according to claim 1, characterized in that: The global information injection module generates sub-image features F generated by the local image fine detection process. local and the original image features F global Extract and integrate information; original feature image F global After two independent convolutional layers, F g_embed and F act , sub-image feature F local After another convolutional layer, F l_embed , then F g_embed and F l_embed By integrating the attention mechanism into the generated F attfuse ,Finally, RepBlock is used to further extract and integrate information; F act =sigmoid(Conv_act(F global )) F g_embed =Conv g embed(F global ) F l_embed =Convlembed(F local ) F att_fuse =Att_fuse(F l_embed *F act +F g_embed ) F out =RepBlock(F att_fuse ) Among them, sigmoid(·) is the activation function; Conv g embed(·) is the convolutional layer that encodes the features of the original image; Convlembed(·) is the convolutional layer that encodes the sub-image; Att_fuse(·) is the attention mechanism.

6. The method for aerial image target detection based on content perception according to claim 1, characterized in that: The loss function L of CenterNet det : L det =L k +L gwd L gwd =1-1 / τ+f(W 2 ),τ≥1 Among them, L k is the focal loss; L gwd is the Gaussian Wasserstein distance loss function; τ is the temperature coefficient; f(x) = ln(x+1) is a nonlinear function; W is the Gaussian distribution of the bounding box; u is the mean of the Gaussian distribution; Σ is the covariance matrix; F is the Frobenius norm; (x, y) is the center coordinate of the bounding box; h and w are the length and width of the bounding box, respectively.

7. A computer device / equipment / system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Unmanned aerial vehicle air-to-air detection method and system based on space and time sequence information, and medium

    CN121505491A

  • Unmanned aerial vehicle air-to-air detection method and system based on spatial and temporal information, and medium

    CN121505491B