Satellite-borne lightweight remote sensing target detection algorithm and system
By optimizing the YOLOv5s object detection model, replacing the key parts are FasterNet, GSconv and BoTNet, and using the EIoU loss function, the problems of large amount of remote sensing object detection and huge model on satellite-borne devices are solved, and efficient and lightweight remote sensing object detection is achieved.
Patent Information
- Application Number
- CN202311359841.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-19
- Publication Date
- 2025-07-22
AI Technical Summary
The existing remote sensing object detection methods are computationally expensive and inefficient on satellite devices, and the deep learning method model is huge, making it difficult to deploy in real-time on resource-constrained satellite platforms.
By optimizing the YOLOv5s object detection model, replacing the Backbone part with the FasterNet module, the Head part with the GSconv module, and the Neck part with the BoTNet module, and using global self-attention replacement spatial convolution in the last three Bottleneck parts of ResNet, combined with the EIoU loss function, a lightweight remote sensing object detection algorithm is constructed.
While improving detection accuracy, the model size is significantly reduced, adapting to the resource limitations of satellite-based equipment, and achieving efficient deployment on satellite-based equipment.
Smart Images

Figure CN120356108A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a lightweight remote sensing target detection algorithm and system for spaceborne applications. Background Art
[0002] As one of the most popular research directions in the field of computer vision, target detection has been widely applied in multiple fields, including face recognition, logistics patrol, autonomous driving, medical diagnosis, and remote sensing images. Due to the rapid development of remote sensing technology, the quantity and quality of remote sensing images that can be obtained are continuously improving, providing powerful data support for remote sensing target detection. Remote sensing image target detection utilizes the image data obtained by remote sensing technology and, with the aid of computer vision and machine learning methods, automatically identifies and extracts target objects in the images. How to quickly and accurately extract the required information from complex remote sensing images has become a research hotspot in recent years.
[0003] The real-time target detection technology for spaceborne remote sensing images has the ability to extract high-value remote sensing targets in real time on satellites. Compared with the processing solution of transmitting the original data back to the ground for detection and analysis, this technology can effectively reduce latency and data transmission bandwidth, providing valuable data support for rapid response and decision-making. Although with the improvement of the computing power of computer graphics processors in recent years, deep learning methods, especially those based on convolutional neural networks (CNNs), are no longer restricted by computing resources, and their more powerful feature extraction capabilities are demonstrated compared with traditional methods. However, the hardware computing power of edge computing still cannot meet the huge computational overhead, and deep learning algorithms are difficult to be deployed on resource-constrained platforms such as satellites, and the detection speed also cannot meet the real-time requirements.
[0004] In the prior art, traditional target detection methods have problems such as huge computational complexity, low efficiency, and high window redundancy. Although the accuracy of deep learning methods has been greatly improved, due to the existence of numerous convolutional operations and multi-channel arranged pooling and other structures in their models, they occupy a large amount of computing resources. Even single-stage target detection methods for fast inference still need to be implemented on the GPU side and are difficult to be deployed and run on terminal device chips. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a lightweight remote sensing target detection algorithm and system for spaceborne applications, which solves the problem of deploying deep learning methods on mobile computing terminals by optimizing parameters, reconstructing models, etc., combines improving accuracy with achieving model lightweight, and significantly reduces the model size while ensuring performance.
[0006] To solve the above technical problems, the first aspect of the embodiments of the present invention provides a lightweight remote sensing target detection algorithm for spaceborne applications, including the following steps:
[0007] Based on the historical remote sensing image data, construct a remote sensing target detection dataset and divide it into a training set and a test set according to a preset ratio;
[0008] Construct an improved YOLOv5s target detection model. Replace the Conv module in the original Backbone part of the YOLOv5s target detection model with the FasterNet module, replace the Conv module in the Head part with the GSconv module, and replace the C3 module in the Neck part with the BoTNet module. Replace the spatial convolution with global self-attention in the last three Bottleneck parts of ResNet;
[0009] Train the improved YOLOv5s target detection model based on the training set;
[0010] Obtain the target image to be detected and perform target recognition based on the improved YOLOv5s target detection model.
[0011] Furthermore, the historical remote sensing image data includes: NWPU VHR-10 dataset, RSOD dataset, and DOTA dataset.
[0012] Furthermore, before constructing the remote sensing target detection dataset, it also includes:
[0013] Perform format conversion processing on the DOTA dataset and convert the format of the DOTA dataset into the VOC2007 dataset format based on Python;
[0014] Perform cutting processing on the format-converted DOTA dataset to obtain pictures of a preset size;
[0015] Generate an annotation information xml file corresponding to the picture.
[0016] Furthermore, after generating the annotation information xml file corresponding to the picture, it also includes:
[0017] Delete the annotation information xml files that do not meet the preset requirements;
[0018] The preset requirements include: the annotation target is empty, the difficult of all annotation targets is 1, or the annotation target is out of bounds.
[0019] Furthermore, before performing cutting processing on the format-converted DOTA dataset, it also includes:
[0020] Perform visualization processing on the ground truth of the pictures in the DOTA dataset.
[0021] Furthermore, the improved YOLOv5s object detection model further includes: replacing the original CIoU loss function with the EIoU loss function in the calculation of the intersection over union (IoU) of the overlap degree between the detection box and the ground truth box.
[0022] Furthermore, the preset ratio for dividing the training set and the test set is 8:2.
[0023] Furthermore, the preset parameter values of the improved YOLOv5s object detection model include: the initial learning rate is 0.01, the momentum parameter is 0.937, and the weight decay is 0.0005.
[0024] Correspondingly, a second aspect of the embodiments of the present invention provides a lightweight spaceborne remote sensing object detection system, including:
[0025] A dataset construction module, which is used to construct a remote sensing object detection dataset based on historical remote sensing image data and divide it into a training set and a test set according to a preset ratio;
[0026] An improved model construction module, which is used to construct an improved YOLOv5s object detection model, replace the Conv module in the original Backbone part of the YOLOv5s object detection model with the FasterNet module, replace the Conv module in the Head part with the GSconv module, replace the C3 module in the Neck part with the BoTNet module, and replace the spatial convolution with global self-attention in the last three Bottleneck parts of ResNet;
[0027] A model training module, which is used to train the improved YOLOv5s object detection model based on the training set;
[0028] An object recognition module, which is used to obtain an image of the object to be detected and perform object recognition based on the improved YOLOv5s object detection model.
[0029] Furthermore, the evaluation metrics for the lightweight spaceborne remote sensing object detection system deployed on the GPU side are: mean average precision, accuracy, recall rate, and / or the number of model parameters;
[0030] The evaluation metrics for the lightweight spaceborne remote sensing object detection system deployed on the chip side are: detection rate, mean average precision, and / or operating power consumption.
[0031] Correspondingly, a third aspect of the embodiments of the present invention provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned spaceborne-oriented lightweight remote sensing target detection method.
[0032] Correspondingly, a fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the above-mentioned spaceborne-oriented lightweight remote sensing target detection method is implemented.
[0033] The above technical solutions of the embodiments of the present invention have the following beneficial technical effects:
[0034] By optimizing parameters, reconstructing the model, etc., the problem of deploying deep learning methods on spaceborne mobile computing terminals is solved. While improving the accuracy, the improvement of accuracy is combined with the realization of model lightweight, and on the premise of ensuring performance, a significant reduction in model size is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a diagram of the spaceborne-oriented lightweight remote sensing target detection method provided by the embodiments of the present invention;
[0036] Figure 2 is a schematic diagram of the network structure of the improved YOLOv5s target detection model provided by the embodiments of the present invention;
[0037] Figure 3 is a schematic diagram of the PConv structure provided by the embodiments of the present invention;
[0038] Figure 4 is a schematic diagram of the GSConv structure provided by the embodiments of the present invention;
[0039] Figure 5 is a structural diagram of the BotNet provided by the embodiments of the present invention;
[0040] Figure 6 is a block diagram of the spaceborne-oriented lightweight remote sensing target detection system module provided by the embodiments of the present invention.
[0041] Figure 7 is a diagram of the training process of the improved algorithm of the present invention on the RSOD dataset;
[0042] Figure 8a is the visualization detection result of the improved algorithm on the RSOD dataset Figure 1 ;
[0043] Figure 8b is the visualization detection result of the improved algorithm on the RSOD datasetFigure 2 ;
[0044] Figure 8c It is the visualization detection result of the improved algorithm on the RSOD dataset Figure 3 ;
[0045] Figure 8d It is the visualization detection result of the improved algorithm on the RSOD dataset Figure 4 ;
[0046] Figure 8e It is the visualization detection result of the yolov5s algorithm on the RSOD dataset Figure 1 ;
[0047] Figure 8f It is the visualization detection result of the yolov5s algorithm on the RSOD dataset Figure 2 ;
[0048] Figure 8g It is the visualization detection result of the yolov5s algorithm on the RSOD dataset Figure 3 ;
[0049] Figure 8h It is the visualization detection result of the yolov5s algorithm on the RSOD dataset Figure 4 ;
[0050] Reference signs:
[0051] 1. Dataset construction module, 2. Improved model construction module, 3. Model training module, 4. Target recognition module. Detailed implementation manners
[0052] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the specific implementation manners and with reference to the accompanying drawings. It should be understood that these descriptions are exemplary and are not intended to limit the scope of the present invention. In addition, in the following descriptions, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0053] Existing traditional methods for remote sensing image target detection: Traditional methods in the field of remote sensing target detection have problems such as large computational complexity, low efficiency, and high window redundancy. For example, the template matching method requires manual classification and extraction of target features, and the template matching method does not have rotational invariance. In the task background of multi-tiny object detection in remote sensing image target detection, both efficiency and accuracy are relatively low. Although the HOG (Histogram of oriented gradients) feature descriptor has a certain detection ability for targets of different scales and poses, its speed is slow, the real-time performance is average, and it is sensitive to occlusion and illumination changes. Another type of traditional method is to combine machine learning algorithms represented by SVM and AdaBoost for target detection. After feature extraction using feature extractors such as the histogram of oriented gradients within a certain range, classifiers such as SVM and AdaBoost are used to classify the targets to be detected. However, the HOG+SVM method needs to describe a large number of features using complex mathematical statistical methods, and its detection performance is significantly insufficient, and the detection time is relatively long.
[0054] Deep learning methods, by using multi-level convolutional and fully connected neural network structures, can automatically learn feature representations from remote sensing images and perform target detection. At present, there are two mainstream methods for target detection in the field of deep learning, namely single-stage target detection and two-stage target detection. Two-stage target detection realizes target detection through two stages of processing. First, in the first stage, a large number of candidate boxes are generated through an algorithm, and then in the second stage, all candidate regions are sent into a classifier to determine whether each candidate region contains a target, and the position of the target is accurately located after screening and correction. The two-stage target detection R-CNN (Region-based Convolutional Neural Networks) algorithm was the first to apply a deep learning network to target detection and achieved success, and improved R-CNN, Fast R-CNN, Faster R-CNN, etc. were successively proposed. Single-stage target detection algorithms are usually end-to-end and complete target detection in a fast inference process. Some well-known single-stage target detection algorithms include: the YOLO (You Only Look Once) series: YOLOv1, YOLOv2, YOLOv3, YOLOv4, YOLOv5, etc., SSD (Single Shot MultiBox Detector), and RetinaNet. The two target detection methods each have their own advantages. However, in terms of hardware deployment and real-time performance, single-stage target detection is superior with its accuracy no less than that of two-stage detection, smaller model volume, and faster detection speed.
[0055] Please refer to Figure 1 and Figure 2, in the first aspect of the embodiments of the present invention, a lightweight remote sensing target detection algorithm for spaceborne applications is provided, including the following steps:
[0056] Step S200: Based on the historical data of remote sensing images, construct a remote sensing target detection data set and divide it into a training set and a test set according to a preset ratio.
[0057] Step S300: Construct an improved YOLOv5s target detection model. Replace the Conv module in the original Backbone part of the YOLOv5s target detection model with the FasterNet module, replace the Conv module in the Head part with the GSconv module, and replace the C3 module in the Neck part with the BoTNet module. Replace the spatial convolution with global self-attention in the last three Bottleneck parts of ResNet.
[0058] Considering the task requirements of satellite terminals, the present invention is based on the more concise YOLOv5s-6.0 version with fewer parameters. The improved network structure is as Figure 2 shown. Introduce the FasterNet module with fewer parameters based on the new convolution in the backbone network to further compress the model volume without affecting the accuracy of various vision tasks. Replace it with GSConv, whose computational cost is about 60% - 70% of the standard convolution, in the feature fusion layer. Maximize the retention of the hidden connections between each channel while reducing redundant repeated information. The C3 module in the original network Neck is replaced by BoTNet. By replacing the spatial convolution with global self-attention only in the last three bottleneck blocks of ResNet, the BoTNet improves the baseline structure and reduces the latency overhead.
[0059] Optimize the Backbone part of the YOLOv5 target detection model and replace the traditional Conv with the FasterNet module. The core design of the PConv method is to systematically apply the traditional convolution operation (Conv) only on some input channels while keeping the rest unchanged by using the redundant information existing in the feature map. This strategy makes PConv essentially have lower FLOPs than the traditional convolution operation and higher floating-point operation rate (FLOPS) than the depthwise separable convolution (DWConv) method. Compared with the past methods that tried to reduce FLOPS, the outstanding feature of PConv is that it can effectively reduce the frequent memory access times, thus reducing FLOPS without introducing the additional data operations such as increased memory access that often accompany it. Through this method, lower additional data operations can be achieved. Its design idea is as Figure 3As shown, the FasterNet method was further constructed on this basis, successfully overcoming the FLOPs in previous attempts while maintaining a higher floating-point operation rate, providing a new and effective way for the fast operation of neural networks.
[0060] FasterNet introduced a novel partial convolution (PConv) method, whose main advantage is to more efficiently perform spatial feature extraction by reducing computational redundancy and memory access simultaneously. The core design of the partial convolution method is to utilize the redundant information existing in the feature map and systematically apply traditional convolution operations only on some input channels while keeping the remaining channels unchanged. This strategy makes the partial convolution essentially have a lower number of floating-point operations than traditional convolution operations and also has a higher floating-point operation rate than the depthwise separable convolution method. Compared with the methods that tried to reduce the floating-point operation rate in the past, the outstanding feature of the partial convolution is that it can effectively reduce the frequent memory access times, thus reducing the floating-point operation rate without introducing the problem of additional data operations such as the accompanying increase in memory access. Through this method, lower additional data operations can be achieved. On this basis, the FasterNet method was further constructed, maintaining a higher floating-point operation rate, providing a new and effective way for the fast operation of neural networks.
[0061] Optimize the Head part of the YOLOv5 object detection model and replace the Conv module in the original model with the GSconv module. According to Figure 4 As shown, GSConv first undergoes a normal convolution for downsampling, then uses DWConv for depth convolution operation, and concatenates the results of the standard convolution and depthwise separable convolution operations. Finally, a channel shuffle operation is performed, that is, the number of channels is rearranged so that the channels corresponding to the previous two convolution operations are adjacent. The purpose of this step is to make the output of DSC as close as possible to that of SC, which is also the main goal of GSConv. By using the shuffle operation, the information from SC (generated through dense convolution operation) is penetrated into each part of the information generated by DSC. This method allows the information from SC to be fully mixed into the output of DSC. This method allows the information from SC to be fully mixed into the output of DSC by evenly exchanging local feature information on different channels, minimizing the negative impact of DSC defects on the model and effectively utilizing the advantage of DSC with a smaller volume.
[0062] GSConv first undergoes a normal convolution for downsampling, then uses depthwise separable convolution for depth convolution operation, and concatenates the results of the standard convolution and depthwise separable convolution operations.
[0063] The GSconv partial convolution structure can better utilize the computing power on the device by reducing the redundant information in the feature map. It is also effective in extracting spatial features. The partial convolution only applies the conventional convolution to a part of the input channels for spatial feature extraction, and the remaining channels remain unchanged. For continuous or conventional memory access, only the first or the last continuous channel is considered as the representative of the entire feature map for calculation. Without loss of generality, it is assumed that the input and output feature maps have the same number of channels.
[0064] Based on this, FasterNet has no other complex redundant structures. It runs fast and is very effective for many vision tasks. And it keeps the architecture as simple as possible, making it more adaptable to the deployment on the hardware side in general, meeting the task requirements of on-orbit high-performance computing in the context of resource constraints. Finally, a channel rearrangement operation is performed to make the channels corresponding to the previous two convolution operations adjacent. The purpose of this step is to make the output of the channel sparse convolution as close as possible to that of the channel dense convolution, which is also the main goal of GSConv. By using the channel rearrangement operation, the information generated from the dense convolution operation is penetrated into each part of the information generated from the channel sparse convolution. By uniformly exchanging local feature information on different channels, it is fully mixed into the output of the channel sparse convolution, minimizing the negative impact of the channel sparse convolution defect on the model and effectively utilizing its advantage of smaller volume.
[0065] The computational cost of GSConv is about 50% of that of the standard convolution, but its contribution to the model's learning ability is comparable to that of the latter. In the feature fusion layer of YOLOv5s-6.0, the number of parameters of GSConv in the three-scale detection heads is 392,448, while the number of parameters of Conv before replacement is 771,072; the former is only 52.3% of the latter, indicating that this lightweight hybrid convolution meets the volume and accuracy requirements of on-board edge computing.
[0066] Optimize the Head part of the YOLOv5 object detection model and replace the C3 module in the original model with the BoTNet module. The BoT (Bottleneck Transformers) structure is a conceptually simple but powerful backbone architecture that combines self-attention for multiple computer vision tasks, including image classification, object detection, and instance segmentation. By replacing the spatial convolution with global self-attention only in the last three Bottleneck parts of ResNet and making no other changes, the baseline was significantly improved in instance segmentation and object detection, while also reducing the parameters with minimal latency overhead. The design of BoTNet allows for various ways to modify the intermediate ResNet Bottleneck parts. Consider the ResNet Bottleneck with self-attention as a Transformer block. As Figure 5 shown, in the design of the Bot network structure, only the last three bottleneck blocks of ResNet are replaced with BoT blocks, that is, only the last three 3×3 convolutions are replaced with the MHSA (Multi-Head Self-Attention) layer, reducing the training and inference overhead while linking to the Transformer module. At the same time, the self-attention mechanism also enhances the network's feature extraction ability. Finally, the network structure composed of blocks like BoT is called BotNet.
[0067] The BoT (Bottleneck Transformers) structure is a conceptually simple but powerful backbone architecture that combines self-attention for multiple computer vision tasks. By replacing the spatial convolution with global self-attention only in the last three bottleneck blocks of the residual network, the baseline was significantly improved in object detection, while also reducing the parameters and having minimal latency overhead. The design of BoT allows for various ways to modify the intermediate residual network bottleneck parts. In the design of the BoT network structure, only the last three 3*3 convolutions are replaced with MHSA (Multi-Head Self-Attention), reducing the training and inference overhead while linking to the Transformer module. At the same time, the self-attention mechanism also enhances the network's feature extraction ability. Finally, the network structure composed of combinations like BoT is called BoTNet.
[0068] Dividing the model's input into multiple heads to form multiple subspaces allows the model to focus on different aspects of information. Using multi-head attention expands the model's ability to focus on different locations, thus giving the attention mechanism multiple sub-expressions. Adding attention mechanisms such as CBAM and SE to the YOLOv5s backbone network requires an additional layer of network architecture and increases network parameters. Compared with the former, this method replaces the convolution of a specific part with a multi-head self-attention mechanism module through the BoT structure design, avoiding the computational overhead of adding additional network parameters.
[0069] Step S400, training the improved YOLOv5s target detection model based on the training set.
[0070] Step S500, obtaining a target image to be detected, and performing target recognition based on the improved YOLOv5s target detection model.
[0071] In addition, the improved YOLOv5s target detection model also includes: replacing the original CIoU loss function with the EIoU loss function in the calculation of the intersection over union (IoU) of the overlap between the detection box and the true box.
[0072] EIoU breaks down the aspect ratio based on CIoU and explicitly measures the differences in three geometric factors, namely the overlapping area, the center point, and the side length. At the same time, Fcoal loss is introduced to solve the problem of imbalance between difficult and easy samples.
[0073] In view of the problems in current target detection such as large amount of computation, low efficiency, high window redundancy, large model of deep learning method and high computing resource consumption, the present invention will solve the problem of deploying deep learning methods in mobile computing terminals by optimizing parameters, reconstructing models, etc. It not only focuses on improving accuracy, but more importantly, combines this improvement with the lightweight model, thereby achieving a significant reduction in model size without sacrificing performance.
[0074] The technical problem to be solved by the present invention is that most existing YOLO improvement methods often rely on powerful GPU performance, resulting in huge model parameters, which is not conducive to real-time detection on mobile terminals with limited resources, and is also difficult to adapt to on-orbit high-performance computing in the aviation field with high resource and power constraints.
[0075] Specifically, the historical remote sensing image data include: NWPU VHR-10 dataset, RSOD dataset and DOTA dataset.
[0076] Furthermore, before step S200, constructing the remote sensing target detection data set, the method further includes:
[0077] Step S110: Perform format conversion processing on the DOTA dataset, and convert the format of the DOTA dataset into the VOC2007 dataset format based on Python.
[0078] Step S120: Perform cutting processing on the DOTA dataset after format conversion to obtain images of a preset size.
[0079] Step S130: Generate an XML file of annotation information corresponding to the images.
[0080] Furthermore, after generating the XML file of annotation information corresponding to the images in Step S130, it further includes:
[0081] Step S140: Delete the XML files of annotation information that do not meet the preset requirements.
[0082] Among them, the preset requirements include: the annotation target is empty, the difficult of all annotation targets is 1, or the annotation target is out of bounds.
[0083] Furthermore, before performing cutting processing on the DOTA dataset after format conversion in Step S120, it further includes:
[0084] Step S111: Perform visualization processing on the ground truth of the images in the DOTA dataset.
[0085] To make the detection robustness of the model stronger, the open-source remote sensing image datasets used in the present invention are the NWPU VHR-10 dataset, the RSOD dataset, and the DOTA dataset respectively. The NWPU VHR-10 dataset consists of 715 RGB images and 85 sharpened color infrared images. Among them, 715 RGB images are collected from Google Earth, and the spatial resolution ranges from 0.5m to 2m. 85 pan-sharpened infrared images with a spatial resolution of 0.08m come from the Vaihingen data. This dataset contains a total of 10 geospatial object classes and 3775 object instances in total. The RSOD dataset used in this experiment has a total of 936 images, including four categories, namely 4993 airplanes in 446 images, 191 playgrounds in 189 images, 191 playgrounds in 189 images, and 1586 fuel tanks in 165 images. The DOTA dataset includes 2806 remote sensing images with a size of approximately 4000*4000, 188282 instances, and a total of 15 categories.
[0086] Among them, for the DOTA dataset, it is necessary to use Python to convert the format of the DOTA dataset into the format of the VOC2007 dataset, and visualize the ground truth of the images in the DOTA dataset. Since some of the images in the DOTA dataset have too large aspect ratios and cannot be directly used for subsequent training, it is necessary to cut the DOTA dataset. Cut the images in the dataset into images of a fixed size of 600*600, and generate corresponding annotation information xml files for the cut images. After processing, delete the xml files that do not meet the requirements. There are three cases for xml files that do not meet the requirements: 1. The annotation target is empty; 2. The difficult of all annotation targets is 1; 3. There is a problem of out-of-bounds for the annotation target (note: there are six cases of annotation out-of-bounds: xmin<0, ymin<0, xmax>width, ymax>height, xmax<xmin, ymax<ymin).
[0087] Specifically, the preset ratio for dividing the training set and the test set is 8:2.
[0088] Specifically, the preset parameter values of the improved YOLOv5s object detection model include: the initial learning rate is 0.01, the momentum parameter is 0.937, and the weight decay is 0.0005.
[0089] The same above experimental parameters are used in all experiments. To verify the effectiveness of the improved module of the present invention, ablation experiments are designed on the RSOD remote sensing dataset. The same above experimental parameters are used in all experiments. The results of the ablation experiments are shown in Table 1. Among them, improved model 1 means changing the Conv module of the Backbone to the FasterNet strategy based on partial convolution, improved model 2 means replacing the Conv module in the feature fusion layer with an improved channel recombination new convolution, improved model 3 means using a multi-head attention mechanism embedded model to replace the C3 module of the original feature fusion layer, and improved model 4 means using the EIOU loss function.
[0090] Table 1 Ablation experiments of the model on RSOD
[0091]
[0092] As can be seen from the data in Table 1, compared with the original YOLOv5s algorithm, the average detection accuracy of the improved algorithm has increased by 8.4%. Using partial convolution PConv, the average detection accuracy of the model reaches 97.8%, proving that using PConv can improve the detection accuracy when training on this dataset. After adopting the GSconv module, the detection accuracy has also been greatly improved, reaching 97.4%. When only using the improvement of replacing the standard convolution with the attention mechanism built into the residual module, the improvement in accuracy is small. However, due to uniformly replacing the C3 module of the original three-scale detection heads in the feature fusion layer, this improvement alone reduces the model parameters by 14.6% and the floating-point computation by 13.1%. Given the task background of on-board edge computing, this lightweight improvement is very necessary. After using EIOU as the loss function, the average detection accuracy has increased by 7.5% and the model volume remains unchanged. Finally, by combining the improvements in extracting the standard convolution channel part, the new depthwise separable convolution model, the attention mechanism built-in residual module, and optimizing the loss function, the detection accuracy can be improved on the premise of lightweighting the original model volume, reaching 98.1%.
[0093] Table 2 Comparison of detection results of different algorithms on the RSOD dataset
[0094]
[0095] In this paper, the test results on the RSOD dataset are visually displayed. Figure 8 is the comparison diagram before and after the improved algorithm. Among them, (a), (b), (c), and (d) are the detection result diagrams of the YOLOv5s algorithm, and (e), (f), (g), and (h) are the detection result diagrams of the improved algorithm in this paper. By comparing the two groups of aircraft target detection result diagrams of (a), (e) and (b), (f), it can be seen that there is one missed detection in the detection results of the YOLOv5s algorithm in the two images respectively. However, in the improved algorithm FGBE-YOLOv5s of the present invention, the aircraft targets with smaller volumes in the images are detected. Comparing (c), (g) of the YOLOv5s algorithm, there are still missed detections, and the correct results detected by the improved algorithm in this paper are more complete. By comparing (d) and (h), it can be seen that the confidence of the improved algorithm is higher, and the detection box fits more closely to the remote sensing target to be detected.
[0096] Table 3 Comparison of detection results of different algorithms on DOTA
[0097] Table 3 results of different algorithms on DOTA
[0098]
[0099]
[0100] The improved algorithm of the present invention is experimented on DOTA-v1.0 with more image data and compared with other mainstream object detection algorithms. The comparison results are shown in Table 3.
[0101] In terms of the average precision of all categories, the improved algorithm of the present invention is 3.9% higher than YOLOv5s and slightly higher than other improved YOLO-based algorithms. This algorithm has higher detection accuracy in categories such as small vehicles (SV) and ships (SH) among 15 object detection categories, effectively solving the problem of difficult small object detection in remote sensing object detection.
[0102] Table 4 Comparison of model volume and inference speed
[0103]
[0104] Table 4 provides important data for the performance comparison of the object detection model before and after improvement. The improved algorithm model shows a smaller model volume in different formats, which gives it an advantage in resource-constrained environments. In addition, the improved algorithm model shows a faster inference speed on GPU, CPU, and NPU, highlighting its excellent performance in lightweight object detection tasks. These data provide strong support for the actual deployment of the present invention, indicating that the improved algorithm model performs better than YOLOv5s in terms of model volume and inference speed.
[0105] The present invention makes four improvements and optimizations to the YOLOv5s object detection algorithm. It introduces the FasterNet module with fewer parameters based on a new type of convolution in the backbone network to further compress the model volume without affecting the accuracy of various visual tasks. The GSConv with a computational cost of about 60% - 70% of the standard convolution is used to replace the feature fusion layer, maximizing the retention of hidden connections between each channel while reducing redundant and repetitive information. The C3 module in the original network Neck is replaced by BoTNet, which improves the baseline structure by replacing spatial convolutions with global self-attention only in the last three bottleneck blocks of ResNet, reducing the latency overhead. Finally, the EIoU loss function is used to replace the original CIoU loss function in the calculation of the intersection over union (IoU) between the detection box and the ground truth box. EIoU disassembles the aspect ratio on the basis of CIoU, clearly measuring the differences in three geometric factors, namely the overlapping area, the center point, and the side length, and at the same time introducing the Focal loss to solve the problem of imbalance between easy and difficult samples.
[0106] The present invention adopts FasterNet with fewer new convolution parameters, lightweight convolutional Gsconv, BoTNet with an embedded multi-head self-attention mechanism, and the EIOU function with better intersection over union calculation, etc. Based on these implementation methods, the present invention has not only made a breakthrough in achieving high-precision object detection, but more importantly, has achieved a lightweight effect in terms of model volume at the same time. This means that the present invention has excellent application potential in on-orbit high-performance computing in the aerospace field.
[0107] The present invention adopts an innovative method to enhance the object detection model, especially in the YOLOv5s backbone network. Compared with the conventional approach, adding attention mechanisms such as CBAM and SE usually requires additional independent network layers, resulting in additional network parameters and computational overhead. However, the method of the present invention adopts a more precise way to avoid these unnecessary complexities and computational burdens.
[0108] Specifically, the present invention replaces a specific convolutional layer in the residual network structure with a multi-head self-attention mechanism module. This unique method not only avoids adding extra parameters to the network, but also provides multiple advantages. First of all, it improves the computational efficiency of the network because the attention mechanism module can capture the context information of specific regions without adding too many parameters. This helps to improve the overall speed and lightness of the model, especially suitable for deployment on embedded and mobile devices.
[0109] Secondly, the method of the present invention introduces an attention mechanism to enhance the object detection ability of the model for small objects. Small objects usually have limited context information, so the introduction of the multi-head self-attention mechanism helps the model to better understand the context background of the object, improves the accuracy of small object detection, and improves the poor performance of existing solutions in this regard.
[0110] In summary, the uniqueness of the present invention lies in not only avoiding the computational overhead of adding extra network parameters, but also introducing an effective way to enhance the detection of small objects and improve the model performance. This method maintains the efficiency of the model while improving the accuracy, which is an important innovation.
[0111] In addition, the improved object detection algorithm has taken the lead in successfully realizing efficient deployment on the on-board rk3568 chip. This achievement is to meet the high requirements for the object detection system in the embedded and edge computing environments, which requires taking into account multiple key indicators such as model volume, detection speed, and operating power consumption while ensuring the detection accuracy.
[0112] First, the present invention breaks through the limitations of traditional object detection algorithms on devices with limited computing resources, achieving precise object detection. At the same time, the model size is optimized to ensure the efficient deployment of the model on the rk3568 chip, which is particularly crucial for spaceborne systems as spaceborne devices are usually limited by storage and computing resources. Second, the present invention fully considers the detection speed to ensure fast object detection on spaceborne devices. This not only improves the response time but also meets the requirements for fast dynamic scenarios such as automatic navigation and monitoring. Most importantly, the present invention optimizes the operating power consumption as a key metric to ensure the long-term stable operation of spaceborne devices in an environment with limited energy. This is crucial for the sustainability and reliability of spaceborne systems.
[0113] The present invention not only achieves high precision of the object detection algorithm in an embedded environment but also comprehensively considers various factors such as model size, detection speed, and operating power consumption, providing unprecedented performance for the object detection application of spaceborne devices. This comprehensive innovation is expected to promote the multi-field applications of spaceborne devices, including fields such as earth observation, communication, and navigation, providing strong support for future technological progress and exploration.
[0114] Correspondingly, please refer to Figure 6 , the second aspect of the embodiment of the present invention provides a lightweight remote sensing object detection system for spaceborne applications, including:
[0115] A dataset construction module 1, which is used to construct a remote sensing object detection dataset based on historical remote sensing image data and divide it into a training set and a test set according to a preset ratio;
[0116] An improved model construction module 2, which is used to construct an improved YOLOv5s object detection model, replace the Conv module in the original Backbone part of the YOLOv5s object detection model with a FasterNet module, replace the Conv module in the Head part with a GSconv module, and replace the C3 module in the Neck part with a BoTNet module, by replacing the spatial convolution with global self-attention in the last three Bottleneck parts of ResNet;
[0117] A model training module 3, which is used to train the improved YOLOv5s object detection model based on the training set;
[0118] An object recognition module 4, which is used to obtain an image of the object to be detected and perform object recognition based on the improved YOLOv5s object detection model.
[0119] Specifically, the evaluation metrics for the lightweight remote sensing object detection system deployed on the GPU side are: mean Average Precision (mAP), precision, recall, and / or the number of model parameters; the evaluation metrics for the lightweight remote sensing object detection system deployed on the chip side are: detection rate, mean Average Precision (mAP), and / or operating power consumption.
[0120] In this invention, mean Average Precision (mAP), precision, recall, and the number of model parameters are used as evaluation metrics on the GPU side, and detection rate, mean Average Precision (mAP), and operating power consumption are used as evaluation metrics for chip-side deployment. The operating system used in this experiment is Ubuntu 18.04 LTS, the GPU is NVIDIA GeForce RTX 3090, the deep learning framework version is pytorch1.9, and CUDA is 12.0. The ablation experiment and the comparative experiment both use the pre-trained weights of yolov5s version 6.0. The initial learning rate is 0.01, and the momentum parameter and weight decay are 0.937 and 0.0005 respectively. Training is carried out for 200 epochs with a batch size of 32. The experiment is divided into a training set and a test set at a ratio of 8:2. The hardware deployment device uses RK3588S launched by Rockchip at the end of 2021. The GPU graphics processor is Mali-G610, and its NPU has three cores, with a maximum support of 6Tops (INT8).
[0121] Each module of the above satellite-borne lightweight remote sensing object detection system can be further divided into several functional units to execute each step in the above method, which will not be elaborated here.
[0122] Correspondingly, a third aspect of the embodiments of the present invention provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above satellite-borne lightweight remote sensing object detection method.
[0123] Correspondingly, a fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the above satellite-borne lightweight remote sensing object detection method is implemented.
[0124] An embodiment of the present invention aims to protect a lightweight remote sensing target detection algorithm and system for spaceborne applications. The method includes: constructing a remote sensing target detection dataset based on historical remote sensing image data and dividing it into a training set and a test set according to a preset ratio; constructing an improved YOLOv5s target detection model by replacing the Conv module in the original Backbone part of the YOLOv5s target detection model with a FasterNet module, replacing the Conv module in the Head part with a GSconv module, and replacing the C3 module in the Neck part with a BoTNet module, and replacing the spatial convolution with global self-attention in the last three Bottleneck parts of ResNet; training the improved YOLOv5s target detection model based on the training set; obtaining an image of the target to be detected and performing target recognition based on the improved YOLOv5s target detection model. The above technical solution has the following effects:
[0125] By optimizing parameters, reconstructing the model, etc., the problem of deploying deep learning methods on spaceborne mobile computing terminals is solved. While improving the accuracy, the improvement of accuracy is combined with the realization of model lightweight, and on the premise of ensuring performance, a significant reduction in model size is achieved.
[0126] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0128] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A lightweight remote sensing target detection method for spaceborne applications, characterized in that, It includes the following steps: Based on the historical remote sensing image data, construct a remote sensing target detection dataset and divide it into a training set and a test set according to a preset ratio; Construct an improved YOLOv5s target detection model. Replace the Conv module in the original Backbone part of the YOLOv5s target detection model with the FasterNet module, replace the Conv module in the Head part with the GSconv module, and replace the C3 module in the Neck part with the BoTNet module. Replace the spatial convolution with global self-attention in the last three Bottleneck parts of ResNet; Train the improved YOLOv5s target detection model based on the training set; Obtain the target image to be detected and perform target recognition based on the improved YOLOv5s target detection model.
2. The spaceborne lightweight remote sensing target detection method according to claim 1, wherein The historical remote sensing image data includes: NWPU VHR-10 dataset, RSOD dataset, and DOTA dataset.
3. The spaceborne lightweight remote sensing target detection method according to claim 2, characterized in that Before constructing the remote sensing target detection dataset, it further includes: Perform format conversion processing on the DOTA dataset, and convert the format of the DOTA dataset into the VOC2007 dataset format based on Python; Perform cutting processing on the format-converted DOTA dataset to obtain pictures of a preset size; Generate an annotation information xml file corresponding to the pictures.
4. The spaceborne-oriented lightweight remote sensing target detection method according to claim 3, characterized in that After generating the annotation information xml file corresponding to the pictures, it further includes: Delete the annotation information xml files that do not meet the preset requirements; The preset requirements include: the annotation target is empty, the difficult of all annotation targets is 1, or the annotation target has an out-of-bounds situation.
5. The spaceborne-oriented lightweight remote sensing target detection method according to claim 3, wherein Before performing the cutting processing on the format-converted DOTA dataset, it further includes: Perform visualization processing on the ground truth of the pictures in the DOTA dataset.
6. The spaceborne lightweight remote sensing target detection method according to any one of claims 1-5, wherein The improved YOLOv5s target detection model further includes: replacing the original CIoU loss function with the EIoU loss function in the calculation of the intersection over union IoU between the detection box and the ground truth box.
7. The spaceborne lightweight remote sensing target detection method according to any one of claims 1-5, wherein The preset ratio for dividing the training set and the test set is 8:
2.
8. The spaceborne lightweight remote sensing target detection method according to any one of claims 1-5, wherein The preset parameter values of the improved YOLOv5s target detection model include: the initial learning rate is 0.01, the momentum parameter is 0.937, and the weight decay is 0.0005.
9. A lightweight spaceborne remote sensing target detection system, characterized in that, It includes: A dataset construction module, which is used to construct a remote sensing target detection dataset based on the historical remote sensing image data and divide it into a training set and a test set according to a preset ratio; An improved model construction module, which is used to construct an improved YOLOv5s object detection model. The Conv module in the original Backbone part of the YOLOv5s object detection model is replaced with a FasterNet module, the Conv module in the Head part is replaced with a GSconv module, and the C3 module in the Neck part is replaced with a BoTNet module. Global self-attention is used to replace spatial convolution in the last three Bottleneck parts of ResNet; A model training module, which is used to train the improved YOLOv5s object detection model based on a training set; An object recognition module, which is used to obtain an image of an object to be detected and perform object recognition based on the improved YOLOv5s object detection model.
10. The spaceborne lightweight remote sensing object detection system according to claim 9, wherein The evaluation metrics for the lightweight remote sensing object detection system deployed on the GPU side are: mean average precision, accuracy, recall rate, and / or the number of model parameters; The evaluation metrics for the lightweight remote sensing object detection system deployed on the chip side are: detection rate, mean average precision, and / or operating power consumption.