Lightweight remote sensing image aircraft target detection method based on HS-YOLO

By constructing the HS-YOLO model, the problems of slow feature extraction speed and low accuracy of small target detection in remote sensing image aircraft target detection are solved. The model is lightweight and efficient, improving detection accuracy and real-time performance, and is suitable for resource-constrained environments.

CN120997479APending Publication Date: 2025-11-21DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511095604.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing remote sensing image aircraft target detection methods suffer from slow feature extraction speed, low accuracy in detecting small targets, and poor real-time performance due to large model parameter count, making it difficult to meet the high-precision detection requirements in resource-constrained scenarios.

Method used

The lightweight remote sensing image aircraft target detection method based on HS-YOLO enhances feature representation capabilities, optimizes feature extraction and downsampling, and reduces computational load by constructing an HS-YOLO model, including a C2f_EE module, an ADown convolution module, and a shared convolutional detection head (ESCD). This reduces computational load and achieves lightweight model and efficient detection.

Benefits of technology

It improves detection accuracy, especially for small targets and complex backgrounds, reduces the number of model parameters, is suitable for deployment in resource-constrained environments, and meets real-time detection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997479A_ABST
    Figure CN120997479A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight remote sensing image aircraft target detection method based on HS-YOLO, and belongs to the field of computer vision and remote sensing image processing. According to the method, firstly, an HS-YOLO model is constructed, a backbone network (BackBone) of the HS-YOLO model replaces an original C2f module with a C2fEE module, and the feature representation capability is enhanced; an ADown convolution module is introduced into the down-sampling link, and the feature extraction and down-sampling efficiency is optimized; a detection head (Head) adopts a shared convolution detection head (ESCD), so that light-weight and efficient detection is realized; the neck network (Neck) integrates multi-scale features through feature fusion and up-sampling operation, and supports a subsequent detection task. And then lightweight remote sensing image aircraft target detection is carried out by using the constructed HS-YOLO model. The method can improve the detection precision and efficiency, achieves the light weight of the model, and meets the deployment requirements of a resource-limited scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and remote sensing image processing, specifically relating to a lightweight aircraft target detection method (HS-YOLO) based on improved YOLOv8, applicable to scenarios such as aviation safety monitoring. This method enables high-precision and high-efficiency detection of aircraft targets in remote sensing images, providing key technical support for civil aviation safety management. Background Technology

[0002] Aircraft target detection in remote sensing images is of great value in the aviation field. Traditional remote sensing target detection methods rely on manual feature selection and design, which requires a large amount of prior knowledge and expert experience. Furthermore, the features lack robustness and adaptability, making it difficult to cope with complex dynamic scenes. They also perform poorly in multi-target detection, and the complex detection process leads to high computational resource consumption, making it difficult to balance accuracy and speed.

[0003] With the development of deep learning, target detection algorithms based on convolutional neural networks (CNN) are widely used in the aviation field. They are mainly divided into two categories: (1) Two-stage detection algorithms (such as Faster R-CNN and Mask R-CNN): have high detection accuracy but slow speed; (2) Single-stage detection algorithms (such as YOLO series and SSD): have moderate accuracy and excellent speed, suitable for real-time tasks.

[0004] The YOLO series, as a groundbreaking method, directly regresses the bounding box position and category through a single neural network, simultaneously achieving target localization and classification. Existing related improvements include: Luo et al. proposed the YOLOv5-Aircraft enhanced detection model, which improved detection accuracy and inference efficiency through batch normalization module calibration, smoothed KL divergence loss function optimization, and CSandGlass module to reduce information loss; Zhang et al. improved the YOLOv5 algorithm by introducing a coordinate attention module, improving multi-scale feature fusion, and adopting the EIoU loss function, thereby improving target localization and regression accuracy; Dong et al. introduced the SPD-Conv module in YOLOv8 to retain fine-grained information, improving detection accuracy and robustness in small targets and low-resolution scenes.

[0005] However, existing methods for aircraft detection in remote sensing images still suffer from problems such as slow feature extraction speed, low accuracy in locating small targets, insufficient real-time performance, and large number of model parameters, making it difficult to meet the high-precision detection requirements in resource-constrained scenarios. Summary of the Invention

[0006] To address the problems of low feature extraction efficiency, insufficient accuracy in small target detection, and poor real-time performance due to large model parameters in existing technologies for aircraft target detection in remote sensing images, this invention proposes a lightweight remote sensing image aircraft target detection method based on HS-YOLO. The aim is to improve detection accuracy and efficiency while achieving model lightweighting to meet the deployment requirements of resource-constrained scenarios.

[0007] The technical solution of the present invention:

[0008] A lightweight remote sensing image aircraft target detection method based on HS-YOLO is proposed. This method is based on the YOLOv8n framework and reconstructs the HS-YOLO lightweight detection model. Performance is improved through three innovative modules, as follows:

[0009] Step 1: Construct the HS-YOLO model

[0010] The HS-YOLO model includes a C2f_EE module, an ADown convolution module, a shared convolutional detector head (ESCD), and a neck network.

[0011] The HS-YOLO model's backbone network replaces the original C2f module with a C2f_EE module to enhance feature representation capabilities. An ADown convolutional module is introduced in the downsampling stage to optimize feature extraction and downsampling efficiency. The head uses a shared convolutional detection head (ESCD) to achieve lightweight and efficient detection. The neck network integrates multi-scale features through feature fusion and upsampling operations to support subsequent detection tasks.

[0012] Step 2: Use the constructed HS-YOLO model to perform lightweight remote sensing image aircraft target detection.

[0013] Furthermore, in step one, the C2f_EE module replaces the original Bottleneck with EE (EdgeEnhance), introducing edge information into feature extraction. The input image is first processed through the Sobel edge branch to extract edge information, highlighting the aircraft's outline edges, which is crucial for target localization in object detection tasks. Simultaneously, the input image undergoes a 3x3 convolution operation through the Conv branch to extract a conventional feature map. Then, the Sobel edge features and convolutional features are concatenated. The concatenated feature map is then processed through a 1x1 convolution operation, reducing the number of channels while fusing features from the edges and convolution branches. The 1x1 convolution facilitates feature integration and compression, and enhances feature representation by linearly combining information from different channels. The features processed by the 1x1 convolution operation are residually connected to the original input features. This connection method preserves the original input information and avoids feature loss during multiple convolutions. Finally, a final convolution is performed, outputting an enhanced feature map. The C2f_EE module...

[0014] The block combines the efficient multi-branch feature extraction of C2f with the edge enhancement capabilities of EE, making it particularly suitable for small target detection in complex backgrounds.

[0015] Furthermore, in step one, layers 3, 5, 7, 17, and 19 in the model are replaced with ADown convolutional modules. These modules use max pooling to focus on key features and average pooling to retain overall information. Combined with 1×1 convolution to compress the number of channels, the computational load is reduced.

[0016] Furthermore, in step one, the shared convolutional detection head (ESCD) workflow is as follows:

[0017] After the three feature layers output from the neck region enter the detection head, they are first adjusted for channel count using 1×1 convolutional layers to ensure that the channel count of each feature layer is uniformly the size of the intermediate layer. Then, all feature layers are merged into a shared convolutional module for feature extraction, which uses 3×3 convolutional kernels. The use of shared convolutions reduces the number of model parameters and computational cost. Subsequently, the model processes regression and classification branches separately. In the regression branch, 1×1 convolutional layers predict the coordinate offset of the bounding box. To address the differences in target scale processed by different detection heads, the regression branch output is scaled by a scaling layer to adjust the feature scale, ensuring accurate localization of aircraft targets of different sizes. The classification branch predicts the probability of each category using 1×1 convolutional layers. The convolutional layers of the two branches have independent weights, allowing the model to learn localization and classification tasks separately.

[0018] The beneficial effects of this invention are:

[0019] (1) Improved detection accuracy: HS-YOLO achieved an mAP50 of 0.979 and an mAP50-95 of 0.788 on the MAR20 dataset, which are 0.75% and 0.91% higher than YOLOv8n, respectively. The improvement in detection accuracy is particularly significant for small targets (such as A4 and A8 models) and complex backgrounds.

[0020] (2) Lightweight model: The number of parameters is only 1.7M, which is 37% less than YOLOv8n (2.7M) and 15% less than YOLOv9t (2.0M), making it more suitable for deployment on edge devices or in resource-constrained environments.

[0021] (3) Efficiency optimization: By using the efficient downsampling of the ADown module and the shared convolution design of ESCD, the processing speed is improved while reducing the computational overhead, thus meeting the real-time detection requirements.

[0022] (4) Robustness enhancement: The edge feature enhancement of the C2f_EE module and the noise resistance of the ADown module enable the model to maintain stable performance in complex weather, shadow occlusion and other scenarios. Attached Figure Description

[0023] Figure 1 HS-YOLO overall structure diagram;

[0024] Figure 2 C2f_EE module structure diagram (showing Sobel edge branches, regular convolution branches, feature fusion, and residual protection design);

[0025] Figure 3 ADown module structure diagram (showing combinations of average pooling, max pooling, and 1×1 convolution);

[0026] Figure 4 : ESCD detection head structure diagram (showing the process of feature standardization, shared convolution extraction, and task-specific processing). Detailed Implementation

[0027] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0028] 1. C2f_EE module

[0029] Design concept: To address the issue of insufficient edge feature extraction of aircraft in complex backgrounds by the original C2f module, the EE (EdgeEnhance) module is used to replace the original Bottleneck structure to enhance edge feature extraction.

[0030] The structure consists of: Sobel edge branch (extracting image edge information to highlight the aircraft outline), regular convolution branch (3×3 convolution to extract standard feature maps), feature fusion mechanism (edge ​​features and convolution features are concatenated and then fused by 1×1 convolution for dimensionality reduction), and residual protection design (preserving the original input features and avoiding information loss).

[0031] Function: To provide more refined feature representation for small target detection and improve the accuracy of small target detection in complex backgrounds.

[0032] 2. ADown Convolution Module

[0033] Design concept: To address the issues of spatial information loss, fixed receptive field, and susceptibility to noise in standard convolution, we integrate average pooling and max pooling operations to optimize the downsampling process.

[0034] Structural features: It focuses on key features through max pooling, preserves overall information through average pooling, and reduces the computational load by combining 1×1 convolution to compress the number of channels.

[0035] Function: To achieve more efficient feature extraction and downsampling, improve the accuracy and robustness of the model in multi-scale aircraft target (especially long-distance small targets) detection, and improve computational efficiency.

[0036] 3. High-efficiency shared convolutional detection head (ESCD)

[0037] Design concept: To address the issue of large parameter count caused by the decoupled structure of the YOLOv8 standard detection head, the number of parameters and computational complexity are reduced by sharing convolutional layers and optimizing the architecture.

[0038] Workflow:

[0039] (1) Feature standardization: The channel dimension of the multi-scale feature layer is unified by 1×1 convolution;

[0040] (2) Shared feature extraction: All feature layers jointly extract features through a shared 3×3 convolution module, reducing the number of parameters by 37%;

[0041] (3) Task-specific processing: regression branch (1×1 convolution predicts bounding box offset, and the feature scale is adjusted by a scaling layer), classification branch (independent 1×1 convolution predicts class probability).

[0042] Function: While reducing the number of parameters and computational overhead, it enhances multi-scale detection capabilities, improves model processing speed, and maintains high detection accuracy.

[0043] 4. Overall Structure of HS-YOLO

[0044] Backbone: The original C2f module is replaced with the C2f_EE module to enhance feature representation capabilities;

[0045] Downsampling stage: Introducing the ADown convolution module to optimize feature extraction and downsampling efficiency;

[0046] Detection Head: Employs an ESCD detection head for lightweight and efficient detection;

[0047] Neck network: Through feature fusion and upsampling operations, it integrates multi-scale features to support subsequent detection tasks.

[0048] (I) Experimental Environment

[0049] 1. Hardware platform: GPU is NVIDIA GeForce RTX 4090, GPU acceleration platform is CUDA 11.8.0, CPU is Intel(R) Core(TM) i7-8750H CPU@2.20GHz.

[0050] 2. Software platform: The operating system is Ubuntu 22.04, the development language is Python 3.8, the deep learning framework is PyTorch 2.0.1, and the integrated development environment is VSCode.

[0051] (II) Dataset and Parameter Settings

[0052] 1. Dataset: The MAR20 dataset is used, which contains 3,842 images, 20 aircraft models, and 22,341 instances. It supports horizontal bounding box and directed bounding box annotation.

[0053] 2. Training parameters: 150 training epochs, using SGD optimizer (initial learning rate 0.01), batch size 16, input image size 640×640, and using early stopping for 10 epochs of mosaic data augmentation; save the training set and validation set losses, and finally evaluate them on the test set using the optimal weights.

[0054] 3. Evaluation metrics: Accuracy (P), recall (R), mAP50, mAP50-95, and number of parameters (Params) are used to evaluate model performance.

[0055] (III) Experimental Results and Verification

[0056] (1) Experiments before and after model improvement

[0057] The table shows the detection results of the improved model on various aircraft datasets before and after the upgrade. The results indicate that both the unimproved and improved models have good mAP50 detection accuracy for most aircraft, but poor accuracy for A8 and A19 aircraft. This may be due to the high internal similarity between these aircraft types and other aircraft types, leading to false positives. After the upgrade, the mAP50 for A4 improved from 0.951 to 0.975, and the mAP50-95 improved from 0.736 to 0.775, showing a significant improvement in accuracy for these aircraft. The improved model also shows significant performance improvements for certain aircraft types, especially in the mAP50-95 range, performing better under high IoU conditions such as A8, A14, and A17. Overall, the improved model maintains high accuracy for most aircraft types, particularly in high IoU scenarios where the improvement in mAP50-95 is more pronounced, indicating that the new model performs better in refined detection tasks.

[0058] Detection results of multiple aircraft types before and after model improvement

[0059]

[0060] (2) Comparison and verification of different algorithms:

[0061]

[0062] To further verify the superiority of the proposed HS-YOLO algorithm, this study conducted a comparative analysis with YOLOv5, YOLOv8, YOLOv9, YOLOv10, and YOLO11 models. The comparative experimental results show that HS-YOLO exhibits superior detection performance among all compared models, especially achieving the highest values ​​in both mAP50 and mAP50-95 metrics. Specifically, HS-YOLO's mAP50 reaches 0.979, an improvement of approximately 1.35% compared to YOLOv5s, 0.75% compared to YOLOv8n, and 0.5% compared to YOLO11. In the more stringent mAP50-95 metric, HS-YOLO leads YOLOv5s by approximately 2.44% with a score of 0.788, leads YOLOv8n by approximately 0.91%, and leads YOLO11 by approximately 0.4%. Furthermore, HS-YOLO has only 1.7M parameters, which is 15% less than YOLOv9t, the model with the fewest parameters in the comparison, while maintaining higher detection accuracy.

[0063] In summary, HS-YOLO has significant advantages in both detection accuracy and model lightweighting, making it very suitable for high-precision detection applications in resource-constrained environments.

[0064] (3) Ablation test

[0065]

[0066]

[0067] Ablation experiments were conducted to investigate the impact of three modules—C2f-EE, ESCD, and ADown—on model performance, either individually or in combination. The results showed that Experiment C, which introduced the ADown module alone, performed better in terms of mAP50 (0.977) and mAP50-95 (0.793), with Recall increasing to 0.946. Experiment B, which introduced the ESCD module alone, increased Recall to 0.931 and slightly improved Precision. Experiment A, which introduced the C2f-EE module alone, slightly improved mAP50 and Recall, but slightly decreased Precision. In the combined module experiments, Experiment D (C2f-EE+ESCD) achieved a precision of 0.967, while Experiment F (ESCD+ADown) had an mAP50-95 of 0.790. The HS-YOLO model that introduced all three modules simultaneously showed the best overall performance, with an mAP50 of 0.979, an mAP50-95 of 0.794, a precision of 0.967, and a recall of 0.950. This indicates that the synergistic effect of the three modules can effectively improve the detection precision and recall of the model.

Claims

1. A lightweight remote sensing image aircraft target detection method based on HS-YOLO, characterized in that, Specifically as follows: Step 1: Construct the HS-YOLO model The HS-YOLO model includes a C2f_EE module, an ADown convolution module, a shared convolutional detector head ESCD, and a neck network; The HS-YOLO model's backbone network, BackBone, replaces the original C2f module with a C2f_EE module to enhance feature representation capabilities. An ADown convolutional module is introduced in the downsampling stage to optimize feature extraction and downsampling efficiency. The detection head uses a shared convolutional detection head, ESCD, to achieve lightweight and efficient detection. The neck network integrates multi-scale features through feature fusion and upsampling operations to support subsequent detection tasks. Step 2: Use the constructed HS-YOLO model to perform lightweight remote sensing image aircraft target detection.

2. The lightweight remote sensing image aircraft target detection method based on HS-YOLO according to claim 1, characterized in that, In step one, the C2f_EE module replaces the original Bottleneck in C2f with EE. The input image is first processed through the Sobel edge branch to extract edge information. Simultaneously, the input image undergoes a 3x3 convolution operation through the Conv branch to extract a regular feature map. Then, the Sobel edge features and convolution features are concatenated. The concatenated feature map undergoes a 1x1 convolution operation to reduce the number of channels while fusing features from the edges and convolution branches. The features processed by the 1x1 convolution operation are residually connected with the original input features. Finally, a convolution is performed to output an enhanced feature map.

3. The lightweight remote sensing image aircraft target detection method based on HS-YOLO according to claim 1, characterized in that, In step one, layers 3, 5, 7, 17, and 19 in the model are replaced with ADown convolutional modules. These modules use max pooling to focus on key features and average pooling to retain overall information. Combined with 1×1 convolution to compress the number of channels, the computational load is reduced.

4. The lightweight remote sensing image aircraft target detection method based on HS-YOLO according to claim 1, characterized in that, In step one, the shared convolutional detection head (ESCD) workflow is as follows: After the three feature layers output from the neck enter the detection head, the number of channels is first adjusted by a 1×1 convolutional layer so that the number of channels in each feature layer is the same as the size of the middle layer; then all feature layers are gathered into a shared convolutional module for feature extraction, which uses a 3×3 convolutional kernel. Subsequently, the model processes the regression branch and the classification branch separately; In the regression branch, the coordinate offset of the bounding box is predicted through a 1×1 convolutional layer. To cope with the differences in target scale processed by different detector heads, the regression branch output is scaled through a scale layer to adjust the feature scale and ensure accurate localization of aircraft targets of different sizes. The classification branch uses a 1×1 convolutional layer to predict the probability of each category; The two branches of the convolutional layer have independent weights so that the model can learn localization and classification tasks separately.