A lightweight unmanned aerial vehicle small target detection method, system, device and storage medium

CN122821401APending Publication Date: 2026-09-25GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610890073.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

因此,本发明提供了一种轻量化无人机小目标检测方法,用于改善现有目标检测模型在无人机小目标场景中存在的模型复杂度高、小目标特征易丢失以及复杂背景干扰较强的问题

Benefits of technology

1、本发明针对无人机航拍场景目标微小且密集、背景复杂、机载算力受限的问题,对YOLO系列模型进行了重构。通过协同优化骨干网络和颈部网络,本发明在削减模型参数量和浮点运算量的同时,提升了对微小目标的召回率和平均精度(mAP),适用于无人机等边缘计算设备的实时部署需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821401A_ABST
    Figure CN122821401A_ABST
Patent Text Reader

Abstract

The application discloses a kind of lightweight unmanned aerial vehicle small target detection method, system, equipment and storage medium, comprising, in backbone network, introduce feature enhancement unit based on feature complementary mapping, through channel segmentation and bidirectional cross attention weighting, realize the complementary fusion of high-dimensional semantics and shallow spatial position information, enhance small target feature expression;In neck network construction micro small target special detection architecture, truncation 32 times deep layer down-sampling path, retain 4 times to 16 times high-resolution shallow feature pyramid, and move forward spatial pyramid pooling module to compensate receptive field;Lightweight hybrid convolution is fixed-point deployed in the bottom-up path of neck network, to process high-resolution features at a lower computational cost.The application reduces the model parameter quantity and computational complexity while improving the detection accuracy of small targets under the perspective of unmanned aerial vehicles, suitable for edge real-time deployment requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision inspection, and in particular to a lightweight method, system, device, and storage medium for small target detection on unmanned aerial vehicles (UAVs). Background Technology

[0002] In recent years, with breakthroughs in deep learning theory, object detection algorithms centered on convolutional neural networks (CNNs) have made significant progress. These algorithms can be broadly categorized into two paradigms: two-stage detectors, represented by the R-CNN series, and single-stage detectors, represented by YOLO and SSD. Single-stage detectors, with their excellent balance between detection speed and accuracy, dominate many real-time applications. Especially in the field of unmanned aerial vehicles (UAVs), real-time object detection using their onboard visual sensors has become a key technology for tasks such as aerial surveillance, power line inspection, emergency search and rescue, and precision agriculture. Existing high-performance detection models, such as the YOLO series, extract and fuse image features by constructing deep network structures and complex multi-scale feature pyramid networks (FPNs), demonstrating powerful detection performance on standard datasets and laying a solid foundation for the development of UAV visual perception technology.

[0003] However, directly applying existing general-purpose object detection algorithms to UAV platforms still faces two major challenges. First, UAV onboard computing platforms typically have strict limitations on computing power, power consumption, and memory. Current mainstream high-precision detection models often have a huge number of parameters and high computational complexity, making it difficult to achieve real-time and efficient deployment on resource-constrained edge devices. This severely restricts the autonomous perception capabilities of UAVs. Second, UAVs usually fly at high altitudes, resulting in ground targets appearing extremely small in scale and with low pixel counts in aerial images. They are also susceptible to interference from complex backgrounds, lighting changes, and occlusion. Traditional detection models expand the receptive field and extract high-level semantic information through continuous downsampling operations in the deep layers of the network. However, this inevitably leads to a significant loss of spatial detail information for small targets, resulting in weak feature representations that are easily ignored in deep networks, leading to missed detections and false detections of small targets. Furthermore, when standard feature pyramid networks fuse deep semantic information with shallow detail information, for small targets, the coarse information from the extremely deep feature maps may introduce noise, affecting the final localization accuracy. Existing network structures still have room for improvement in the efficiency of feature extraction and fusion, and have not fully taken into account the dual requirements of lightweight design and feature fidelity for small targets.

[0004] Therefore, designing an algorithm specifically suited for drone scenarios that can simultaneously achieve lightweight models and maintain high detection accuracy for small targets is a pressing technical challenge in this field. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0006] In view of the problems existing in the prior art, the present invention is proposed. Therefore, the present invention provides a lightweight UAV small target detection method to improve the problems of high model complexity, easy loss of small target features, and strong interference from complex backgrounds in existing target detection models in UAV small target scenarios.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a lightweight unmanned aerial vehicle (UAV) small target detection method, comprising: Acquire drone aerial images; A lightweight object detection network model is constructed, comprising a backbone network, a neck network, and a detection head. The network model is constructed as follows: In the backbone network, the preset feature fusion module used to output the second and third level features is replaced with a feature enhancement unit based on feature complementary mapping. In the neck network, the deep feature layer corresponding to the maximum downsampling factor is removed, and the spatial pyramid pooling module is moved to the end of the third-level feature layer, while reconstructing a shallow feature pyramid containing the first, second, and third-level features. In the bottom-up feature propagation path of the neck network, the downsampling convolutions used to connect the first-level features to the second-level features and the second-level features to the third-level features are replaced with lightweight hybrid convolution modules. The lightweight target detection network model is trained to obtain the trained target detection model; During the inference phase, the trained backbone network is used to extract feature maps of the first, second, and third levels from the aerial images. The feature maps of the three levels are input into the neck network, and cross-layer feature fusion is performed through the shallow feature pyramid to generate a fused multi-scale feature map. The fused multi-scale feature map is input into the detection head, and target detection is performed on the feature maps at the three levels respectively to obtain the detection result.

[0008] Secondly, the present invention also provides a lightweight UAV small target detection system, comprising: The data acquisition module is configured to acquire aerial images taken by the drone; The feature extraction module is configured to extract feature maps of three levels—first, second, and third—from the aerial image through a backbone network. The feature fusion module is configured to perform cross-layer feature fusion on the feature maps of the three layers through the neck network; The target detection module is configured to perform target detection on the fused feature map using a detection head to obtain detection results; The backbone network used in the feature extraction module is configured to replace the preset feature fusion module that outputs second-level and third-level features with a feature enhancement unit based on feature complementary mapping. The neck network used in the feature fusion module is configured to remove the deep feature layer corresponding to the maximum downsampling factor, and move the spatial pyramid pooling module to the end of the third-level feature, while reconstructing a shallow feature pyramid containing the three levels of features. The bottom-up feature delivery path of the neck network is configured to replace the downsampled convolutions connecting the first-level to the second-level features and the second-level to the third-level features with lightweight hybrid convolution modules.

[0009] As a preferred embodiment of the lightweight UAV small target detection method of the present invention, the working process of the feature enhancement unit based on feature complementary mapping includes: The input features are segmented along the channel dimension into a first sub-feature and a second sub-feature; The first sub-feature is subjected to convolutional transformation to extract semantic features, and the second sub-feature is subjected to convolutional transformation to preserve spatial location features; Channel attention weights are generated based on the semantic features, and spatial attention weights are generated based on the spatial location features; The channel attention weights are weighted with the spatial location features, and the spatial attention weights are weighted with the semantic features to achieve bidirectional cross-weighting. The weighted features are aggregated to generate enhanced output features.

[0010] As a preferred embodiment of the lightweight UAV small target detection method of the present invention, the working process of the lightweight hybrid convolution module includes: The input features are processed by a standard convolution to obtain the first intermediate features. The first intermediate feature is input into a depthwise separable convolution for processing to obtain the second output feature. The first intermediate feature and the second output feature are concatenated along the channel dimension; Channel shuffling is performed on the concatenated features to fuse information from standard convolutions and depthwise separable convolutions.

[0011] As a preferred embodiment of the lightweight UAV small target detection method of the present invention, the step of removing the deep feature layer corresponding to the maximum downsampling ratio includes: Remove the fifth-level feature layer with a downsampling rate of 32x; The shallow feature pyramid formed by the reconstruction includes a first level, a second level, and a third level of features, with the first level, the second level, and the third level corresponding to 4x, 8x, and 16x downsampling rates, respectively.

[0012] As a preferred embodiment of the lightweight UAV small target detection method of the present invention, the method further includes a step of training the lightweight target detection network model, the training step including: The training process employs a comprehensive loss function that is a weighted combination of bounding box regression loss, distribution focus loss, and classification loss.

[0013] Thirdly, the present invention also provides an electronic device, characterized in that it includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores a computer program, which, when executed by the at least one processor, implements any step of the above-described lightweight UAV small target detection method.

[0014] As a preferred embodiment of the electronic device described in this invention, the electronic device is an airborne computing platform for unmanned aerial vehicles, an edge computing device, or a ground control station.

[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements any step of the above-described lightweight UAV small target detection method.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention addresses the challenges of small and dense targets, complex backgrounds, and limited onboard computing power in drone aerial photography scenarios by reconstructing the YOLO series models. Through collaborative optimization of the backbone and neck networks, this invention improves recall and mean accuracy (mAP) for small targets while reducing the number of model parameters and floating-point operations, making it suitable for the real-time deployment needs of edge computing devices such as drones.

[0017] 2. The C2f-FCM unit designed in the backbone network of this invention adopts a channel segmentation strategy to decouple features into semantic branches and spatial location branches. By calculating channel attention and spatial attention and performing bidirectional cross-weighting (i.e., using spatial weights to guide semantic features and channel weights to guide spatial features), the network can adaptively focus on the key regions of small targets, effectively suppressing the interference of background noise and enhancing the model's ability to extract and represent features of small targets.

[0018] 3. Traditional target detection networks typically include a 32x downsampling layer (P5). For small targets like UAVs with extremely low pixel counts, continuous downsampling significantly weakens their spatial detail features. Therefore, this invention removes the P5 layer and its related structures, reconstructing a shallow high-resolution feature pyramid containing P2 (4x), P3 (8x), and P4 (16x) downsampling layers. This improvement reduces the parameter overhead caused by deep redundant structures and provides fine-grained geometric and positional information for small targets by preserving the high-resolution features of the P2 layer. Simultaneously, the SPPF module is moved to the end of the P4 layer to compensate for the global receptive field loss caused by truncating deep networks.

[0019] 4. Since this invention retains high-resolution feature maps such as the P2 layer, continuing to use standard convolutions for feature fusion in the neck network would result in high computational overhead. Therefore, this invention replaces the GSConv module at specific points in the bottom-up feature propagation path of the neck network. The GSConv module combines standard convolutions and depthwise separable convolutions, supplemented by a channel shuffling mechanism, to reduce the computational complexity of processing high-resolution feature maps while ensuring multi-scale feature fusion capabilities, thereby improving the model's inference efficiency. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating the overall process of a lightweight UAV small target detection method according to an embodiment of the present invention. Figure 2 This is an overall structural diagram of a lightweight UAV small target detection system according to an embodiment of the present invention; Figure 3 This is a network topology diagram of a lightweight UAV small target detection method according to an embodiment of the present invention; Figure 4This is a structural diagram of the C2f-FCM feature enhancement unit in a lightweight UAV small target detection method according to an embodiment of the present invention; Figure 5 This is a structural diagram of the feature complementary mapping unit of the lightweight UAV small target detection method according to an embodiment of the present invention; Figure 6 This is a structural diagram of the GSConv module of a lightweight UAV small target detection method according to an embodiment of the present invention; Figure 7 This is a comparison of the training convergence curves of the lightweight UAV small target detection method according to one embodiment of the present invention. Detailed Implementation

[0021] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0022] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0023] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0024] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0025] Example 1 Reference Figures 1 to 6 This is the first embodiment of the present invention, which provides a lightweight unmanned aerial vehicle (UAV) small target detection method, including: S1. Acquire drone aerial image data.

[0026] It should be noted that the aerial image data of this drone comes from the image acquisition equipment (such as visible light camera, infrared camera, etc.) carried by the drone.

[0027] Specifically, the acquired drone aerial image data is used to form an image dataset, which contains multiple aerial images with target categories and bounding box annotations, including a large number of small targets (such as pedestrians, vehicles, etc.).

[0028] Furthermore, after obtaining the image dataset, the dataset is divided into a training set, a validation set, and a test set.

[0029] In one possible implementation of this embodiment, the image size in the acquired image dataset is uniformly scaled to 640×640 pixels.

[0030] In an alternative implementation of this embodiment, the image size can be flexibly set to different resolutions such as 320×320, 416×416, 512×512, 800×800, or 1280×1280, depending on the actual application scenario, in order to achieve a balance between different accuracy and speed requirements.

[0031] S2. Preprocess the acquired drone aerial image data.

[0032] Furthermore, the image data in the training dataset is preprocessed.

[0033] Specifically, the preprocessing method includes uniform scaling of image size (in this embodiment, uniform scaling to 640×640 pixels), pixel value normalization (normalization from [0,255] to [0,1]), data augmentation (including Mosaic data augmentation, random horizontal flipping, random scaling, random perturbation of HSV color space, etc.), and label format conversion.

[0034] Furthermore, the preprocessed image data can be represented as: in, This represents the original input drone aerial image data; This represents the combined transformation function used in the above preprocessing operations; This represents the tensor data output after preprocessing, which will be used as the input to the model.

[0035] S3. Construct a lightweight target detection network model.

[0036] Furthermore, a lightweight object detection network model is constructed, referring to... Figure 3The network model consists of three parts: a backbone network, a neck network, and a detection head. It should be explained that the backbone network is mainly used to extract multi-level features from the input image, and it is composed of multiple cascaded convolutional layers and downsampling layers. Based on this, this embodiment improves upon YOLOv8s with a combined structure, as follows: (i) Integrating C2f-FCM feature enhancement units in the backbone network.

[0037] It should be noted that, in this embodiment, the C2f-FCM feature enhancement unit mainly includes a Conv standard convolutional layer, a Split channel segmentation operation, N cascaded feature complementarity mapping units, a feature concatenation operation, and an output convolutional layer.

[0038] Specifically, the standard convolutions and the first C2f module in the shallow layers of the backbone network are retained. The pre-defined feature fusion modules (i.e., the second and third C2f modules in the original network) used to output features of the second (P3) and third (P4) layers are replaced with feature enhancement units based on complementary feature mapping (C2f-FCM). (Refer to...) Figure 4 and Figure 5 The working process of this unit is as follows: The input feature map is first processed by a standard Conv convolutional layer for channel transformation, and then a Split operation is performed along the channel dimension to divide the feature map into two branches. The first branch is directly fed into the concatenation stage to form a short-circuit connection, and the second branch is fed into N cascaded feature complement mapping units for progressive feature enhancement. Finally, the feature maps of the two branches are concatenated along the channel dimension and then processed by the output convolutional layer for channel compression and fusion.

[0039] It should be noted that the working method of this unit inherits the cross-stage gradient propagation advantage of the original C2f module, while enhancing the complementary expression of spatial details and semantic information at the P3 and P4 feature output positions adopted in this invention.

[0040] It is particularly important to emphasize that a more detailed description of the above unit's working process is that the input features are first processed along the channel dimension ( ,in It is the number of channels. It is the height of the feature map (or image). The feature map (or image) is divided into a first sub-feature and a second sub-feature according to the proportions: , in, This represents the feature map input to the FCM unit; This indicates a segmentation operation along the channel dimension; For the channel segmentation ratio, this embodiment preferably sets it to... This is used to achieve a balanced distribution of information between the semantic enhancement branch and the position preservation branch; This is the first sub-feature, used for subsequent extraction of semantic information; This is the second sub-feature, used to preserve spatial location information later.

[0041] Next, to separately obtain semantic and spatial information and ensure tensor dimension alignment during subsequent feature fusion, convolutional transformations are applied to the two sub-features. Specifically, a convolutional transformation is performed on the first sub-feature to extract semantic features, and a convolutional transformation is performed on the second sub-feature to preserve spatial features. in, Represents semantic features, Represents spatial location characteristics, and , . express Standard convolution operation; express Pointwise convolution operation.

[0042] In addition, in order to overcome the feature misalignment problem between semantic and spatial information caused by independent extraction of two branches, the FCM module performs information interaction between the channel dimension and the spatial dimension around the aforementioned semantic features and spatial features respectively, generates complementary channel attention weights and spatial attention weights, and performs cross-weighted fusion.

[0043] Specifically, in the channel interaction branch, channel-wise convolution (DWConv) is first applied to semantic features to independently model spatial dependencies channel by channel. Then, channel attention weights are generated sequentially through an adaptive global average pooling layer and a sigmoid activation function, denoted as... This weight reflects the importance of each feature channel to the current detection task.

[0044] Specifically, in the spatial interaction branch, spatial features are sequentially processed through 1×1 pointwise convolutions for channel compression (from C to 1), batch normalization, and a sigmoid activation function to generate spatial attention weights, denoted as . This weight reflects the importance of each location in the feature map space to the current detection task.

[0045] Finally, by weighting the channel attention weights with spatial location features, and then weighting the spatial attention weights with semantic features, bidirectional cross-weighting is achieved. The weighted features are then aggregated to generate the enhanced output features. in, This represents element-wise multiplication with a broadcast mechanism (Hadamard product). This indicates element-wise addition.

[0046] It should be noted that, through the above processing, semantic features can obtain spatial location guidance through cross-multiplication with spatial attention weights, making it easier to locate the target region using high-dimensional semantic information. Simultaneously, spatial features, through cross-multiplication with channel attention weights, obtain the importance level of channel semantics, making shallow spatial information more explicitly serve small target detection. It is important to note that this C2f-FCM feature enhancement unit in this invention needs to be used in conjunction with P5 truncation, SPPF transfer, and GSConv neck deployment to achieve synergistic improvement in small target detection against complex backgrounds.

[0047] (ii) Construct a TSD-Arch micro-target-specific detection architecture in the neck network.

[0048] It should be noted that this neck network mainly adopts the TSD-Arch (Tiny Object-Specific Detection Architecture) proposed in this invention to achieve efficient feature fusion for small objects. This detection architecture includes two core designs: a deep truncation strategy and shallow high-resolution pyramid reconstruction.

[0049] Furthermore, the deep truncation strategy structurally removes the P5 stage (i.e., the deep feature layer with a maximum downsampling ratio of 32x) from the original YOLOv8s at the macro-topology level, limiting the maximum downsampling rate of the network model to 16x (i.e., the P4 layer). In this embodiment, with a 640×640 input, the P4 feature map size is 40×40 and the number of channels is 256. This deletion action is jointly constituted by SPPF migration, P2 / P3 / P4 detector head reconstruction, and GSConv downsampling path, as described later.

[0050] It should be clarified that for a large number of tiny targets smaller than 16×16 pixels in aerial images, after five consecutive downsampling passes (each with a stride of 2), their theoretical mapping area on the P5 feature map (size 20×20) has shrunk to the sub-pixel level (mathematically less than 0.5×0.5 pixels). This means that the P5 layer not only fails to provide effective distinguishable features for small targets because the spatial information of the small targets has been completely compressed, but also introduces a large amount of background noise and redundant computation due to its high channel dimension (512 channels). Therefore, in order to compensate for the loss of multi-scale global receptive field caused by deep network truncation, this detection architecture places the Spatial Pyramid Pooling-Fast Version (SPPF) module at the end of the P4 stage, that is, after the C2f-FCM module that outputs P4 features and before the top-down fusion path of the neck network. When the SPPF module receives a 40×40×256 P4 feature map, it constructs equivalent receptive fields of different scales through continuous 5×5 max pooling operations, and splices and fuses the original features and multi-level pooling features to output semantically enhanced features that match the number of P4 channels. These features are used to compensate for the contextual expressiveness and receptive field range provided by the deep branches after removing the P5 stage.

[0051] In summary, based on the aforementioned deep truncation strategy, the deep feature layer corresponding to the maximum downsampling factor is removed from the neck network. Specifically, the fifth-level (P5) feature layer with a 32x downsampling rate and its corresponding downsampled Conv layer, C2f module, and related concatenation paths are removed. Simultaneously, the Spatial Pyramid Pooling (SPPF) module is moved from the end of the original P5 layer to the end of the third-level feature layer (P4) to compensate for the multi-scale global receptive field loss caused by the deep network truncation.

[0052] It should be noted that, unlike existing methods that use an additive design to stack complex modules on top of deep networks in exchange for accuracy, this invention adopts a subtractive structural reorganization combined with shallow reconstruction: by removing the P5 stage to reduce redundant parameters and background noise in deep layers, computational resources are concentrated on the fusion of high-resolution features in shallow layers P2, P3, and P4.

[0053] Furthermore, for shallow high-resolution pyramid reconstruction, in order to compensate for the degradation of small target features during forward propagation, the detection architecture introduces the high-resolution P2 stage into the path aggregation network (PANet), reconstructing the neck network into a shallow {P2,P3,P4} feature pyramid structure.

[0054] Specifically, for an input image size of H×W, the spatial dimensions of feature maps P2, P3, and P4 correspond to approximately H / 4×W / 4, H / 8×W / 8, and H / 16×W / 16, respectively; and 160×160, 80×80, and 40×40 for a 640×640 input. During feature interaction, the top-down path uses upsampling to progressively inject deep semantic information from P4 (40×40) to P3 (80×80) and then to P2 (160×160), providing rich category semantic guidance to the shallow, high-resolution feature maps. In the bottom-up path, a downsampling operation with a stride of 2 progressively transfers fine-grained spatial details from P2 to P3 and then to P4, enabling the deep feature maps to obtain accurate spatial localization information.

[0055] Furthermore, to strictly control the computational overhead introduced by processing large-size high-resolution feature maps in the bottom-up path, the detection architecture of this invention only replaces the 3×3 downsampling standard convolutions with a stride of 2 at P2→P3 and P3→P4 in the reconstructed neck network with lightweight GSConv modules with a stride of 2, while keeping the standard convolutions of other backbone networks unchanged. Finally, the three decoupled detection heads operate and predict directly on the three shallow high-resolution feature maps P2, P3, and P4. In this process, taking a 640×640 input as an example, P2 (160×160) mainly serves the localization of small targets, P3 (80×80) mainly serves small to medium-scale targets, and P4 (40×40) mainly serves larger-scale targets; the above scale division can be adjusted according to the target bounding box size statistics in the training dataset and the detection head sample allocation strategy, and does not constitute an absolute limitation on the target size.

[0056] In short, the goal of shallow high-resolution pyramid reconstruction is to reconstruct a shallow feature pyramid containing three levels of features: the first level (P2, 4x downsampling), the second level (P3, 8x downsampling), and the third level (P4, 16x downsampling). By adding an upsampling path from P3 to P2, the shallow high-resolution feature map gains rich categorical semantic guidance.

[0057] (iii) Deploy lightweight hybrid convolutional modules (GSConv) at fixed points in the neck network.

[0058] It should be noted that the GSConv module is mainly integrated into the reconstructed neck network to constrain the parameters and computational overhead caused by high-resolution feature maps while maintaining feature extraction capabilities.

[0059] Furthermore, refer to Figure 6To strictly limit the computational overhead of processing high-resolution feature maps, in the bottom-up feature propagation path of this neck network, only the downsampled standard convolution with a stride of 2 used to connect the first-level features to the second-level features (P2→P3) and the second-level features to the third-level features (P3→P4) are replaced with a lightweight hybrid convolution module (GSConv).

[0060] Specifically, the workflow of this lightweight hybrid convolution module is as follows: Let the number of channels of the input feature map X be... The expected number of output channels is .

[0061] (1) The input feature map X is first processed by a standard convolution (SC, 3×3 convolution kernel): The standard convolution output is used as the first intermediate feature to obtain... The first characteristic of the / 2 channel is denoted as This preserves the dense channel semantic information extracted by standard convolution. Subsequently, the intermediate features from the first path are further input into a depthwise separable convolution (DSC, composed of 3×3 channel-wise convolutions and 1×1 pointwise convolutions) to extract lightweight spatial structure information, resulting in... The second characteristic of the / 2 channel is denoted as .

[0062] (2) The first intermediate feature output by the standard convolution and the second feature output by the depthwise separable convolution are concatenated along the channel dimension to obtain the following result. Channel splicing feature diagram.

[0063] (3) Perform channel shuffling on the spliced ​​feature map to evenly interweave the channels from SC and DSC, so that subsequent network layers can achieve interaction and fusion of the two types of feature information.

[0064] Specifically, the formula for this lightweight hybrid convolution module is as follows: in, This indicates a channel shuffling operation, used to uniformly interweave and arrange the channels of the spliced ​​features, fusing information from standard convolution and depthwise separable convolution. This indicates the operation of splicing the output features of the first and second paths along the channel dimension. This represents the standard convolution operation. This indicates a depthwise separable convolution operation.

[0065] It should be noted that by using a hybrid structure design that extracts dense channel information through standard convolution, supplements lightweight spatial information through depthwise separable convolution, and promotes information interaction through channel shuffling, the lightweight hybrid convolution module can maintain good feature extraction capabilities while reducing computational costs.

[0066] S4. Model training and loss function optimization.

[0067] Furthermore, during the model training phase, the training process employs a comprehensive loss function that is a weighted combination of bounding box regression loss, distribution focus loss, and classification loss for parameter optimization. This comprehensive loss function is expressed as: in, This represents the total loss of the network model; The bounding box regression loss (CIoULoss in this embodiment) is used to optimize the positional deviation between the predicted box and the ground truth box. The distributed focus loss is used to optimize the continuous distribution probability of the bounding box position and improve the boundary localization accuracy of small targets. The classification loss (BCE Loss in this embodiment) is used to optimize multi-class classification prediction; The weighting coefficients for the three losses mentioned above are respectively (e.g., ).

[0068] Furthermore, during training, a stochastic gradient descent (SGD) optimizer is used, combined with a cosine annealing learning rate scheduling strategy and an early stopping mechanism to complete the iterative convergence of the model.

[0069] In one feasible implementation of this embodiment, the initial learning rate for model training is 0.01; the momentum coefficient is 0.937; the weight decay is 0.0005; and the learning rate scheduling strategy employs a cosine annealing strategy, gradually decaying from the initial value to 1% of the initial value. Simultaneously, Mosaic data augmentation (mosaic=1.0) and Automatic Mixed Precision (AMP) training strategies are used to improve training efficiency. Furthermore, an early stopping mechanism is enabled, with a patience value set to 100 epochs. That is, if the performance on the validation set no longer improves for 100 consecutive epochs, training is automatically terminated to prevent overfitting and avoid unnecessary computational overhead.

[0070] In an alternative implementation of this embodiment, the optimizer may also employ any one or a combination of optimization algorithms such as Adam, AdamW, and RMSprop. Furthermore, the learning rate scheduling strategy may employ step decay, exponential decay, OneCycleLR, or a combination thereof. The initial learning rate can be adjusted within the range of 0.1 to 0.0001 depending on the specific optimizer and network architecture.

[0071] S5. Extract multi-level features using the trained backbone network.

[0072] Specifically, during the inference phase, the preprocessed aerial image (e.g., 640×640×3) is input into the trained backbone network. Through standard convolutional layers, the first C2f module, and the C2f-FCM module located at the P3 and P4 output positions, feature maps of the first, second, and third layers are extracted layer by layer. The specific outputs are: P2 layer feature map (160×160×64), P3 layer feature map (80×80×128), and P4 layer feature map (40×40×256).

[0073] S6. Use the neck network for cross-layer feature fusion.

[0074] Specifically, during the inference phase, feature maps from three levels are input into the neck network, and cross-layer feature fusion is performed through a shallow feature pyramid. The semantic information of P4 is gradually injected into P3 and P2 through a top-down path; the spatial details of P2 are gradually passed to P3 and P4 through GSConv through a bottom-up path, ultimately generating fused multi-scale feature maps (P2', P3'', P4'').

[0075] S7. Output the detection results using the detection head.

[0076] Furthermore, the fused multi-scale feature map is input into the detection head, and target detection is performed on the feature maps at three different levels. Each decoupled detection head predicts the target's class confidence and bounding box coordinates. After confidence thresholding and non-maximum suppression (NMS) post-processing, the final detection result is obtained. The detection result output can be formally represented as: in, This represents the final set of detection results output. This represents the total number of detected targets in the final output. Indicates the first The category index of each detected target; Indicates the first The confidence score of each detected target (within the range [0,1]) is only true when... The target is retained if the confidence level is greater than a preset confidence threshold (e.g., 0.5). Indicates the first The bounding box coordinates of each detected target, where The normalized coordinates of the bounding box center point in the image (values ​​[0,1]). The normalized bounding box width and height (with values ​​[0,1]).

[0077] Further reference Figure 2This embodiment also provides a lightweight UAV small target detection system, including: The data acquisition module is configured to acquire aerial images taken by the drone; The feature extraction module is configured to extract feature maps of three levels—first, second, and third—from aerial images through a backbone network. The feature fusion module is configured to perform cross-layer feature fusion on feature maps from three layers through the neck network; The target detection module is configured to perform target detection on the fused feature map using a detection head to obtain detection results; The backbone network used in the feature extraction module is configured to replace the preset feature fusion module that outputs second-level and third-level features with a feature enhancement unit based on feature complementarity mapping. The neck network used in the feature fusion module is configured to remove the deep feature layer corresponding to the maximum downsampling factor and move the spatial pyramid pooling module to the end of the third-level feature layer, while reconstructing a shallow feature pyramid containing three levels of features. The bottom-up feature delivery path of the neck network is configured to replace the downsampled convolutions connecting the first-level to the second-level features and the second-level to the third-level features with lightweight hybrid convolution modules.

[0078] Furthermore, this feature enhancement unit based on feature complementarity mapping is configured as follows: The input features are segmented into a first sub-feature for extracting semantic features and a second sub-feature for preserving spatial location features; Channel attention weights are generated based on semantic features, and spatial attention weights are generated based on spatial location features. The channel attention weights are cross-weighted with spatial location features, and the spatial attention weights are cross-weighted with semantic features; The weighted features are aggregated to generate enhanced output features.

[0079] Furthermore, this embodiment also provides an electronic device suitable for lightweight unmanned aerial vehicle (UAV) small target detection methods, including: At least one processor; and a memory communicatively connected to the at least one processor; the memory stores a computer program that, when executed by the at least one processor, implements the lightweight UAV small target detection method proposed in the above embodiments. The electronic device includes, but is not limited to, an UAV onboard computing platform, an edge computing device, or a ground control station.

[0080] This embodiment also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the lightweight UAV small target detection method proposed in the above embodiment.

[0081] The storage medium proposed in this embodiment belongs to the same inventive concept as the method proposed in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0082] Example 2 Reference Figure 7 This is the second embodiment of the present invention, which mainly compares the method of the present invention with the following two representative existing methods. Comparison Method 1 (General YOLO series object detection algorithms): YOLO-NAS, YOLOv10s, YOLOv11s, YOLOv12s, YOLOv13s. These models have excellent detection performance and robustness on general object detection benchmark datasets (e.g., MSCOCO) and are often used as general baseline models for various visual detection tasks. Comparison Method 2 (Dedicated improved models for small object detection): PC-YOLO11s, EdgeYOLO-S, Drone-YOLO-N, ST-YOLO, WCDB-YOLO, Modified-YOLOv8, PVswin-YOLOv8s, TPH-YOLO. These models are based on the existing state-of-the-art (SOTA) baseline architecture and specifically integrate small object optimization strategies, exhibiting stronger perception capabilities for small-scale objects and representing the mainstream cutting-edge methods in this field.

[0083] Based on the comparison methods 1 and 2 above, this embodiment adopts the evaluation protocol of the target detection field standard (following the MSCOCO evaluation protocol) to perform quantitative evaluation using the following indicators: Precision: The proportion of correctly predicted positive samples out of all predicted positive samples; it measures the accuracy of the model's predictions. The formula is: Where TP represents the number of correct detections and FP represents the number of incorrect detections.

[0084] Recall: The proportion of correctly predicted positive samples out of all actual positive samples; it measures the model's ability to detect the target. The formula is: , where FN is the number of missed detections.

[0085] Mean Accuracy mAP@0.5: The mean accuracy (AP) of all N target categories when the IoU (Intersection over Union) threshold is 0.5. AP is the area under the Precision-Recall curve.

[0086] Mean precision mAP@0.5:0.95: The average mAP value across 10 different IoU thresholds from 0.5 to 0.95 (step size 0.05).

[0087] Number of parameters: The total number of all trainable parameters in the model, measured in millions (M), which directly determines the model's storage space and memory usage.

[0088] Computational complexity: The number of floating-point operations required for one forward inference of the model, measured in billions, which directly determines the computational complexity and inference time of the model.

[0089] Based on the above evaluation, this embodiment further demonstrates the beneficial effects of the present invention through ablation experiments. The ablation experiments were conducted on the VisDrone2019 validation set (548 images). The VisDrone2019 dataset was used in this embodiment according to the official division: 6471 images for training, 548 images for validation, and 3190 images for testing. Unless otherwise stated, both the ablation and comparison experiments were evaluated on the validation set. Each model was trained and tested using the same input size, number of training epochs, optimizer configuration, pre-training weight initialization method, and evaluation script. To reduce the impact of data augmentation, sample order, and randomness in the training process on the experimental results, five random seeds (7, 42, 521, 1314, 2026) were used for independent training under the same training configuration. Table 1 shows the representative optimal results obtained from multiple independent training sessions, illustrating the influence of each module on detection performance and model complexity. The experiment started with the YOLOv8s baseline model and gradually introduced the TSD-Arch, C2f-FCM and GSConv modules.

[0090] Table 1 Based on Table 1 above, the ablation experiment results are analyzed as follows: In the case of adding only TSD-Arch: the structural pruning of TSD-Arch contributed the most significant single impact, namely, the number of parameters dropped from 11.1M to 3.3M (a reduction of 70.3%), mAP@0.5 jumped from 41.7% to 45.2% (+3.5 percentage points), and accuracy improved from 51.4% to 55.9% (+4.5 percentage points). These improvements validate the effectiveness of the deep truncation strategy (removing the P5 layer, which has no benefit for small targets) and the shallow high-resolution reconstruction (introducing the P2 layer). However, GFLOPs increased from 28.5 to 29.8 (+4.6%), which is due to the inevitable additional computation introduced by processing the high-resolution P2 feature map (160×160).

[0091] In the TSD-Arch+C2f-FCM case: Based on TSD-Arch, the standard C2f in the backbone network is replaced with C2f-FCM. GFLOPs decrease from 29.8 to 28.4 (-4.7%), the number of parameters further decreases from 3.3M to 3.1M, while mAP@0.5 remains unchanged at 45.2%. This verifies that the simplified decoupled structure of C2f-FCM (DWConv replacing part of the standard convolution) can effectively suppress backbone computational overhead without compromising detection accuracy.

[0092] The TSD-Arch+GSConv case: Based on TSD-Arch, the downsampling convolution of the neck network was replaced with GSConv. GFLOPs remained stable at 29.6 (a decrease compared to 29.8 in pure TSD-Arch), and mAP@0.5 was 45.0% (still 3.3 percentage points higher than the baseline), verifying GSConv's ability to maintain competitive accuracy while constraining computational cost.

[0093] LFC-YOLO (a complete three-module collaboration, namely TSD-Arch, C2f-FCM, and GSConv): When the three modules are fully integrated, mAP@0.5 and mAP@0.5:0.95 reach 45.7% and 27.9% respectively (reference). Figure 7 The ablation parameters are 2.9M, and the computational cost is 28.0 GFLOPs. Within the scope of this ablation study, the three modules exhibit a positive synergistic and complementary relationship: TSD-Arch provides the structural foundation, C2f-FCM enhances the feature quality of P3 / P4, and GSConv constrains the computational cost of the bottom-up path, thus achieving a better trade-off between accuracy and efficiency.

[0094] Based on the above ablation experiments, the proposed method was compared with the general YOLO detection algorithm (comparison method 1). The results are shown in Table 2.

[0095] Table 2 As shown in Table 2, within the scope of this comparison, the LFC-YOLO of this invention achieves high detection accuracy (45.7% mAP@0.5) with a relatively small number of parameters (2.9M). It should be noted that the GFLOPs (28.0) of this invention are higher than some general YOLO models, but its detection accuracy is significantly improved and the number of parameters is lower, indicating that the performance gain mainly comes from structural optimization for small target scenarios of UAVs, rather than simply stacking parameters.

[0096] Similarly, the proposed solution is compared with a dedicated improved model for small target detection (Comparison Method 2). The results are shown in Table 3.

[0097] Table 3 As shown in Table 3, Drone-YOLO-N has only 3.1M parameters, closest to the lightweight level of this invention (2.9M), but its mAP@0.5 is only 38.1%, 7.6 percentage points lower than this invention, indicating significantly insufficient detection accuracy. This shows that a purely lightweight design cannot simultaneously guarantee detection accuracy. EdgeYOLO-S achieved a relatively high accuracy of 44.8%, but its parameter count is as high as 40.5M, and its computational cost is as high as 109.1 GFLOPs, requiring high hardware resources and making it unsuitable for deployment on resource-constrained UAV edge devices. TPH-YOLO, as a classic improved model for small target detection on UAVs, has 60.4M parameters and 145.7 GFLOPs of computation, but its mAP@0.5 is only 36.2%, indicating that this type of superimposed structure has high overhead in terms of both parameter count and computational cost. WCDB-YOLO achieves 45.5% mAP@0.5, very close to the present invention, but its parameter count (19.8M) is 6.8 times that of the present invention (2.9M), and its computational cost (60.1 GFLOPs) is 2.1 times that of the present invention. At similar accuracy, the difference in parameter count and computational cost indicates that the subtractive optimization route of the present invention has a better efficiency advantage within the comparison range.

[0098] In summary, within the experimental results and comparison range described above, the LFC-YOLO of this invention achieves 45.7% mAP@0.5 with 2.9M parameters and 28.0 GFLOPs, demonstrating superior overall performance with lower parameter count and higher mAP. This result illustrates the synergistic improvement effect of TSD-Arch, C2f-FCM, and GSConv in their combined configuration for small target detection tasks on UAVs.

[0099] In addition, this embodiment also analyzes YOLOv8s and LFC-YOLO through generalization verification, and the results are shown in Table 4.

[0100] Table 4 To verify this generalization ability, this embodiment further evaluates the adaptability of the invention in different drone scenarios. This is achieved through cross-dataset validation experiments on the publicly available UAVDT dataset. The UAVDT dataset is designed for vehicle detection and tracking scenarios from a complex drone perspective. It contains approximately 80,000 representative frames extracted from original drone videos and approximately 840,000 labeled vehicle instances. It is characterized by high vehicle density, complex road traffic environments, and large variations in flight altitude, differing from the data distribution of VisDrone2019 (general aerial photography scenarios, 10 categories). To ensure fair comparison, both the baseline YOLOv8s and the LFC-YOLO of this invention are initialized using pre-trained YOLOv8s weights and maintain consistent training configurations, and are trained and evaluated on the UAVDT dataset. Table 4 shows that LFC-YOLO achieves improvements over the baseline model on this dataset, specifically, mAP@0.5 increases from 58.3% to 61.5% (+3.2 percentage points), and mAP@0.5:0.95 increases from 41.4% to 44.0% (+2.6 percentage points), with precision and recall improving by 1.3% and 2.9%, respectively. Therefore, under this validation analysis, the feature complementarity ensemble and shallow high-resolution localization architecture proposed in this invention have certain scene adaptability and can improve the detection performance of small targets in scenarios with dense vehicle occlusion and scale variations.

[0101] Furthermore, to verify the independent role and combined effect of each structural module on the UAVDT dataset, this embodiment adopts the same module combination method as the VisDrone2019 ablation experiment and conducts supplementary ablation experiments on the UAVDT dataset. The results are shown in Table 5.

[0102] Table 5 As shown in Table 5, the overall trend on the UAVDT dataset is basically consistent with that of VisDrone2019. When TSD-Arch is introduced alone, mAP@0.5 increases from 58.3% to 60.3%, and mAP@0.5:0.95 increases from 41.4% to 42.8%, while the number of parameters decreases from 11.1M to 3.3M, indicating that shallow high-resolution detection structures are also effective for UAV vehicle scenarios. C2f-FCM, when introduced alone, can improve the mAP metric while reducing the number of parameters and computational cost. GSConv, when introduced alone, has a relatively small gain, but when combined with TSD-Arch, it can constrain the computational cost of the neck network and achieve a high mAP@0.5. The complete LFC-YOLO achieves 70.3% precision, 56.1% recall, 61.5% mAP@0.5, and 44.0% mAP@0.5:0.95 on UAVDT, indicating that the three modules still have a certain synergistic effect in cross-dataset vehicle detection scenarios.

[0103] Based on all the above experiments, we can conclude that compared with the baseline YOLOv8s, this invention improves detection accuracy (P+4.3%, R+2.4%, mAP@0.5+4.0%, mAP@0.5:0.95+2.8%), while reducing the number of parameters by 73.9% and the computational cost by 1.8%. This demonstrates that this invention can improve small target detection performance while reducing model redundancy. Furthermore, compared with general YOLO series models (YOLOv10s~v13s), this invention achieves a higher mAP@0.5 with a smaller number of parameters within the comparison range, indicating that its performance improvement comes from structural improvements for UAV small target detection. Moreover, compared with dedicated improved models for small target detection, this invention achieves higher detection accuracy while maintaining a lower number of parameters, indicating that TSD-Arch, C2f-FCM at P3 / P4 positions, and GSConv at P2→P3 / P3→P4 positions have a synergistic effect. Finally, both the UAVDT generalization verification and the UAVDT ablation experiment show that the present invention has a certain scene adaptability under the above experimental conditions. The three structural modules can still form an effective cooperation in cross-dataset vehicle detection scenarios and can be used in diverse real-world UAV deployment scenarios.

[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A lightweight unmanned aerial vehicle (UAV) small target detection method, characterized in that, include: Acquire drone aerial images; A lightweight object detection network model is constructed, comprising a backbone network, a neck network, and a detection head. The network model is constructed as follows: In the backbone network, the preset feature fusion module used to output the second and third level features is replaced with a feature enhancement unit based on feature complementary mapping. In the neck network, the deep feature layer corresponding to the maximum downsampling factor is removed, and the spatial pyramid pooling module is moved to the end of the third-level feature layer, while reconstructing a shallow feature pyramid containing the first, second, and third-level features. In the bottom-up feature propagation path of the neck network, the downsampling convolutions used to connect the first-level features to the second-level features and the second-level features to the third-level features are replaced with lightweight hybrid convolution modules. The lightweight target detection network model is trained to obtain the trained target detection model; During the inference phase, the trained backbone network is used to extract feature maps of the first, second, and third levels from the aerial images. The feature maps of the three levels are input into the neck network, and cross-layer feature fusion is performed through the shallow feature pyramid to generate a fused multi-scale feature map. The fused multi-scale feature map is input into the detection head, and target detection is performed on the feature maps at the three levels respectively to obtain the detection result.

2. The lightweight UAV small target detection method as described in claim 1, characterized in that, The working process of the feature enhancement unit based on complementary feature mapping includes: The input features are segmented along the channel dimension into a first sub-feature and a second sub-feature; The first sub-feature is subjected to convolutional transformation to extract semantic features, and the second sub-feature is subjected to convolutional transformation to preserve spatial location features; Channel attention weights are generated based on the semantic features, and spatial attention weights are generated based on the spatial location features; The channel attention weights are weighted with the spatial location features, and the spatial attention weights are weighted with the semantic features to achieve bidirectional cross-weighting. The weighted features are aggregated to generate enhanced output features.

3. The lightweight UAV small target detection method as described in claim 1, characterized in that, The operation of the lightweight hybrid convolution module includes: The input features are processed by a standard convolution to obtain the first intermediate features. The first intermediate feature is input into a depthwise separable convolution for processing to obtain the second output feature. The first intermediate feature and the second output feature are concatenated along the channel dimension; Channel shuffling is performed on the concatenated features to fuse information from standard convolutions and depthwise separable convolutions.

4. The lightweight UAV small target detection method as described in claim 1 or 2, characterized in that, The removal of the deep feature layer corresponding to the maximum downsampling factor includes: Remove the fifth-level feature layer with a downsampling rate of 32x; The shallow feature pyramid formed by the reconstruction includes a first level, a second level, and a third level of features, with the first level, the second level, and the third level corresponding to 4x, 8x, and 16x downsampling rates, respectively.

5. The lightweight UAV small target detection method as described in claim 1, characterized in that, The method further includes a step of training the lightweight object detection network model, the training step comprising: The training process employs a comprehensive loss function that is a weighted combination of bounding box regression loss, distribution focus loss, and classification loss.

6. A lightweight unmanned aerial vehicle (UAV) small target detection system, characterized in that, include: The data acquisition module is configured to acquire aerial images taken by the drone; The feature extraction module is configured to extract feature maps of three levels—first, second, and third—from the aerial image through a backbone network. The feature fusion module is configured to perform cross-layer feature fusion on the feature maps of the three layers through the neck network; The target detection module is configured to perform target detection on the fused feature map using a detection head to obtain detection results; The backbone network used in the feature extraction module is configured to replace the preset feature fusion module that outputs second-level and third-level features with a feature enhancement unit based on feature complementary mapping. The neck network used in the feature fusion module is configured to remove the deep feature layer corresponding to the maximum downsampling factor, and move the spatial pyramid pooling module to the end of the third-level feature, while reconstructing a shallow feature pyramid containing the three levels of features. The bottom-up feature delivery path of the neck network is configured to replace the downsampled convolutions with a stride of 2 connecting the first-level features to the second-level features and the second-level features to the third-level features with a lightweight hybrid convolution module; the lightweight hybrid convolution module includes standard convolution, depthwise separable convolution and channel shuffling operation.

7. The lightweight UAV small target detection system as described in claim 6, characterized in that, The feature enhancement unit based on complementary feature mapping is configured as follows: The input features are segmented into a first sub-feature for extracting semantic features and a second sub-feature for preserving spatial location features; Channel attention weights are generated based on the semantic features, and spatial attention weights are generated based on the spatial location features; The channel attention weights are cross-weighted with the spatial location features, and the spatial attention weights are cross-weighted with the semantic features; The weighted features are aggregated to generate enhanced output features.

8. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that, when executed by the at least one processor, implements the method as described in any one of claims 1 to 5.

9. The electronic device according to claim 8, characterized in that, The electronic device is an unmanned aerial vehicle (UAV) onboard computing platform, edge computing device, or ground control station.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method described in any one of claims 1 to 5.