Adaptive sparse convolution guided cross-space unmanned aerial vehicle target detection method and system
Through the cross-space drone target detection method guided by adaptive sparse convolution, the adaptive pixel-level mask unit and context-enhanced sparse convolution module are used to solve the problem of computing resource consumption and slow detection speed caused by high target sparseness in the complex background of drone images, and achieve efficient and low-cost target detection.
Patent Information
- Application Number
- CN202510145508.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Under the complex background of drone images, the high target sparseness leads to the problem of target detection algorithms consuming a lot of computing resources and slow detection speed.
A cross-space drone target detection method with adaptive sparse convolution guidance is designed. Through adaptive pixel-level masking units and context-enhanced sparse convolution modules, efficient detection of target areas is achieved and computational costs are reduced.
It significantly reduces the calculation cost of object detection, improves detection speed and efficiency, and maintains detection accuracy, and is suitable for various object detection algorithms.
Smart Images

Figure CN119992062A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle image processing, and relates to a cross-space unmanned aerial vehicle target detection method and system guided by adaptive sparse convolution. Background Art
[0002] As one of the key areas of remote sensing image application, drone target detection has long been the focus of intelligent remote sensing information processing research. Its core task is to identify the location and category of targets in images through feature extraction and analysis. The rapid development and widespread application of drone image-based target detection technology has brought tremendous changes and development opportunities to all walks of life, playing an indispensable and important role in the country's social economy and national defense security, and has been highly valued by countries around the world.
[0003] There are significant differences between drone images and natural images. The background of natural images is simple, the target area accounts for a large proportion, and the detection task is relatively simple. However, drone remote sensing images have a long acquisition distance, large image area, complex and high background, and small and sparse target area, which significantly increases the difficulty of target detection. Although deep neural networks (such as CNN and Transformer) perform well on natural image datasets (such as COCO), target detection in drone images is still challenging. Existing studies mostly design complex models to improve the detection accuracy of small targets in high-resolution images, but the rich background and large background area lead to high computing resource requirements. At the same time, drone hardware resources are limited, and lightweight models are urgently needed to achieve fast reasoning and low latency.
[0004] Scholars at home and abroad have conducted in-depth research on this issue. Lu Yongxi et al. proposed an adaptive target detection method, which intelligently allocates computing resources through neighbor prediction and scaling prediction, focuses on small object areas, and achieves efficient detection. Yang Fan et al. designed a clustered region extraction network, used ScaleNet to accurately estimate the target scale, and integrated the cropped image with the global image detection results to improve performance. Deng Sutao et al. combined global coarse detection with local fine detection, adaptively cropped dense areas, used local super-resolution networks to enhance details, and achieved accurate target detection through a global-local detection network.
[0005] Although these methods have shown promising results in UAV target detection, the complexity in designing network details, such as cluster region extraction and super-resolution reconstruction, increases the detection time and limits their application potential in multiple scenarios. Summary of the invention
[0006] The purpose of the present invention is to provide a cross-space UAV target detection method and system guided by adaptive sparse convolution, so as to solve the problem in the prior art that the target sparsity is high under the complex background of UAV images, resulting in the target detection algorithm consuming a lot of computing resources and slow detection speed. In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present application discloses an adaptive sparse convolution-guided cross-space UAV target detection method, comprising: Obtain a UAV target detection data set, perform multi-scale feature extraction and multi-scale feature fusion on the UAV target detection data set in turn, and obtain multi-scale features; The multi-scale features are input into the preset adaptive sparse detection head for processing to obtain classification features and regression features; The target detection results are obtained based on classification features and regression features.
[0007] Preferably, the step of acquiring a drone target detection dataset and sequentially performing multi-scale feature extraction and multi-scale feature fusion on the drone target detection dataset to obtain multi-scale features specifically includes: S101: Acquire UAV remote sensing dataset; S102: randomly crop the UAV remote sensing dataset to a uniform size to obtain a detection dataset; S103: Divide the detection data set into a training set, a validation set and a test set and convert them into coco data format; S104: Perform multi-scale feature extraction on the detection data set to obtain basic features; S105: Perform multi-scale feature fusion on the basic features to obtain multi-scale features.
[0008] Preferably, the step of inputting the multi-scale features into a preset adaptive sparse detection head for processing to obtain classification features and regression features specifically includes: S201: The multi-scale features are processed by global feature generation to obtain global features, and the multi-scale features are processed by adaptive pixel-level masking to obtain mask features; S202: multi-scale features, global features and mask features are simultaneously subjected to a number of context-enhanced sparse convolutions to obtain a sparse feature map with enhanced expressiveness; S203: The sparse feature map with enhanced expressiveness is processed through a plurality of sparse convolutions to obtain classification features and regression features.
[0009] Preferably, the S201 specifically includes: S2011: Multi-scale features generate global features at each scale through point-by-point convolution; S2012: Multi-scale feature input mask unit processing to obtain a soft feature mask; a binary Gumbel-Softmax method is used on each element of the soft feature mask to convert the soft feature mask into a pixel-level hard feature mask as the mask feature output.
[0010] Preferably, the S202 specifically includes: S2021: Multi-scale features and mask features are processed by sparse convolution to obtain sparse features; S2022: Sparse features and global features are processed by context-enhanced group normalization to obtain a sparse feature map; S2023: Add the sparse feature map to the global feature to obtain a sparse feature map with enhanced expressiveness.
[0011] Preferably, the adaptive sparse detection head comprises: The global feature generation module is used to process multi-scale features through point-by-point convolution to obtain global features; An adaptive pixel-level mask unit, used to perform mask processing on multi-scale features to obtain mask features; Context-enhanced sparse convolution module, which is used to perform feature enhancement based on multi-scale features, global features, and mask features to obtain sparse feature maps with enhanced expressiveness; Sparse convolution module, used to optimize the sparse feature map with enhanced expressiveness to obtain classification features and regression features; The loss function module is used to calculate the loss of classification features and regression features. After iterative training, the adaptive sparse detection head with the best weight is obtained.
[0012] Preferably, the loss function module performs loss calculation by the following formula:
[0013]
[0014] Classification Loss for:
[0015]
[0016] Regression Loss for:
[0017] Normalized loss for:
[0018] Mask loss for:
[0019]
[0020] in, and For balance and Hyperparameters of For the The predicted value of the category; For the The true value of the category; Represents the intersection and union ratio of the predicted value and the true value; is the feature map convolved on the complete input feature map; Sparse feature maps for enhanced expressiveness; Represents mask features; Represents the number of positive samples; represents the true value of the mask; Represents the predicted value of the pixel-level mask unit.
[0021] In the second aspect, the present application discloses an adaptive sparse convolution-guided cross-space UAV target detection system, comprising: An acquisition unit is used to acquire a UAV target detection data set, and sequentially perform multi-scale feature extraction and multi-scale feature fusion on the UAV target detection data set to obtain multi-scale features; A data processing unit, used for inputting multi-scale features into a preset adaptive sparse detection head for processing to obtain classification features and regression features; The output unit is used to obtain the target detection result based on the classification features and regression features.
[0022] In a third aspect, the present application discloses an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the adaptive sparse convolution-guided cross-space UAV target detection method described in any one of the above items are implemented.
[0023] In a fourth aspect, the present application discloses a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the adaptive sparse convolution-guided cross-space UAV target detection method described in any one of the above items.
[0024] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention designs an adaptive pixel-level mask unit that can adaptively extract the target area in the image and realizes the detection of the target area by combining the sparse convolution principle, which significantly reduces the computational cost of target detection and improves the detection speed.
[0025] (2) The present invention designs a context-enhanced group normalization strategy based on the residual structure, which effectively enhances the context information of the target area and reduces the loss of effective feature information, so as to reduce the computational cost without reducing the accuracy of target detection.
[0026] (3) The present invention designs a plug-and-play target detection head that can be applied to current mainstream target detectors, reducing the computational cost of mainstream algorithms and improving detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0028] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a network structure diagram of the cross-space UAV target detection method guided by adaptive sparse convolution in the present invention; Figure 3 This is a diagram of the structure of a common sparse convolutional network in the present invention; Figure 4 This is a diagram of the network structure of the adaptive pixel-level mask unit in the present invention; Figure 5 A structural diagram of a context-enhanced sparse convolutional network in the present invention; Figure 6 This is a network structure diagram of the adaptive sparse convolution detection head in the present invention; Figure 7 This is the detection result after the present invention is applied to the common algorithm and the corresponding pixel mask visualization result. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0030] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0032] In the description of the embodiments of the present invention, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. indicate an orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0033] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0034] In the description of the embodiments of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0035] The present invention is further described in detail below in conjunction with the accompanying drawings: See also Figure 1 , an adaptive sparse convolution-guided cross-space UAV target detection method, comprising: S1: Obtain a drone target detection dataset, perform multi-scale feature extraction and multi-scale feature fusion on the drone target detection dataset in sequence, and obtain multi-scale features; S2: The multi-scale features are input into the preset adaptive sparse detection head for processing to obtain classification features and regression features; the preset plug-and-play adaptive sparse detection head is suitable for various target detection algorithms and significantly reduces the computational cost.
[0036] S3: Obtain target detection results based on classification features and regression features.
[0037] The present invention first obtains a drone target detection data set and performs preprocessing. Then, an adaptive sparse detection head is designed. Then, the adaptive sparse detection head is transplanted into the mainstream one-stage and two-stage target detection models. Then, the improved model is trained and verified, and the improved method is evaluated. Finally, the drone image is input into the trained model to obtain the detection result. The present invention solves the problem that the target detection algorithm consumes a lot of computing resources and has a slow detection speed due to the high sparsity of targets in the complex background of drone images. The adaptive sparse detection head designed by the present invention can be plug-and-play, is applicable to various target detection algorithms, and significantly reduces the computing cost.
[0038] In some embodiments, the step of acquiring a drone target detection dataset and sequentially performing multi-scale feature extraction and multi-scale feature fusion on the drone target detection dataset to obtain multi-scale features specifically includes: S101: Acquire UAV remote sensing dataset; S102: randomly crop the UAV remote sensing dataset to a uniform size to obtain a detection dataset; S103: Divide the detection data set into a training set, a validation set and a test set and convert them into coco data format; S104: Perform multi-scale feature extraction on the detection data set to obtain basic features; S105: Perform multi-scale feature fusion on the basic features to obtain multi-scale features.
[0039] In some embodiments, the inputting of the multi-scale features into a preset adaptive sparse detection head for processing to obtain classification features and regression features specifically includes: S201: The multi-scale features are processed by global feature generation to obtain global features, and the multi-scale features are processed by adaptive pixel-level masking to obtain mask features; S202: multi-scale features, global features and mask features are simultaneously subjected to a number of context-enhanced sparse convolutions to obtain a sparse feature map with enhanced expressiveness; S203: The sparse feature map with enhanced expressiveness is processed through a plurality of sparse convolutions to obtain classification features and regression features.
[0040] In some embodiments, the multi-scale features are processed by global feature generation and adaptive pixel-level mask processing to obtain global features and mask features, respectively, specifically: S2011: Multi-scale features generate global features at each scale through point-by-point convolution; S2012: Multi-scale features are processed through adaptive pixel-level masking to generate mask features at each scale; that is, multi-scale features are input into mask units for processing to obtain soft feature masks; a binary Gumbel-Softmax method is used on each element of the soft feature mask to convert the soft feature mask into a pixel-level hard feature mask as the mask feature output.
[0041] In some embodiments, the multi-scale features, global features and mask features are subjected to context-enhanced sparse convolution and loss function calculation to obtain a sparse feature map with enhanced expressiveness, specifically including: S2021: Multi-scale features and mask features are processed by sparse convolution to obtain sparse features; S2022: Sparse features and global features are processed by context-enhanced group normalization to obtain a sparse feature map; S2023: Add the sparse feature map to the global feature to obtain a sparse feature map with enhanced expressiveness.
[0042] In some embodiments, the adaptive sparse detection head comprises: The global feature generation module is used to process multi-scale features through point-by-point convolution to obtain global features; An adaptive pixel-level mask unit, used to perform mask processing on multi-scale features to obtain mask features; Context-enhanced sparse convolution module, which is used to perform feature enhancement based on multi-scale features, global features, and mask features to obtain sparse feature maps with enhanced expressiveness; Sparse convolution module, used to optimize the sparse feature map with enhanced expressiveness to obtain classification features and regression features; The loss function module is used to calculate the loss of classification features and regression features. After iterative training, the adaptive sparse detection head with the best weight is obtained.
[0043] The present invention designs a plug-and-play target detection head, which is applicable to various target detection algorithms and significantly reduces the computational cost. An adaptive pixel-level mask unit is designed through the task alignment strategy to generate a target pixel-level mask, and a sparse convolution based on the global context information enhancement strategy is used to achieve cross-space detection of the target, focusing only on the target area detection while eliminating the interference of complex backgrounds, while ensuring the detection accuracy, reducing the amount of calculation and improving the detection efficiency.
[0044] In some embodiments, an adaptive sparse convolution-guided cross-space drone target detection method comprises the following steps: Step 1: Obtain the drone target detection dataset and perform preprocessing.
[0045] The common UAV remote sensing dataset VisDrone dataset is used; preprocessing includes random cropping to a uniform size and dividing the dataset into training set, verification set and test set; the dataset is converted into coco data format.
[0046] Step 2: Design of adaptive sparse detection head, which includes: ordinary sparse convolution, global feature generation, adaptive pixel-level mask unit, context-enhanced sparse convolution module, and loss function design.
[0047] The adaptive sparse detection head design, the generation process of which includes the following sub-steps: Step 2.1: For the input drone image, multi-scale feature extraction is performed through the backbone network in the basic algorithm, ResNet18 is used here, and it is recorded as {C3, C4, C5}. Then multi-scale feature fusion is performed in the feature pyramid network to obtain multi-scale features {P3, P4, P5, P6, P7}. It is further input into the adaptive sparse detection head designed by the present invention for sparse detection.
[0048] Step 2.2: The ordinary sparse convolution used is submanifold sparse convolution. The convolution output will only be calculated when the position covered by the center of the sparse convolution kernel is in an activated state. The position where the value of the mask corresponding to the feature map is 1 is in an activated state.
[0049] Step 2.3: Global feature generation uses point-by-point convolution to generate global features .
[0050] Step 2.4: Adaptive pixel-level masking unit, defines the spatial location processed by the convolution, using a small masking unit Acting on multi-scale features , output the corresponding soft feature mask, . Mask Cell The 1×1 convolution is used. The Gumbel-Softmax module is used. , by Apply the binary Gumbel-Softmax technique to each element of Convert to pixel-level hard feature masks :
[0051] Hard feature masks are mask features Will serve as an execution mask, indicating the spatial locations at which operations are to be performed.
[0052] Step 2.5: Context-enhanced sparse convolutional module to input features , mask features , global features As input, context-enhanced sparse convolution operations are performed on the input features.
[0053] Step 2.6: Loss function design, which includes classification loss , regression loss , normalized loss , and mask loss , the total loss function is as follows:
[0054] in and To balance and Hyperparameters of .
[0055] Step 2.7: Design the adaptive sparse convolution detection head, for the input features , which is processed by the mask unit to generate mask features , and generate global features after point-by-point convolution ,After being processed by four context-enhanced sparse convolutions, two ordinary sparse convolutions are used to obtain the classification features and regression features.
[0056] In some embodiments, step 1: transplant the preset adaptive sparse detection head into the mainstream one-stage and two-stage target detection models. The mainstream one-stage and two-stage target detection models obtain a drone target detection dataset, perform multi-scale feature extraction and multi-scale feature fusion on the drone target detection dataset, and obtain multi-scale features; the one-stage and two-stage target detection models used include Faster R-CNN, RetinaNet, FSAF, GFL_V1, and TOOD. These five models include not only one-stage and two-stage target detection models, but also target detection models based on anchor boxes and anchor-free boxes, which are representative. The sparse detection head is transplanted into these five models, named Faster R-CNN-SC, RetinaNet-SC, FSAF-SC, GFL_V1-SC, and TOOD-SC.
[0057] Step 2: Train and verify the improved model and evaluate the improved method.
[0058] The basic network models Faster R-CNN, RetinaNet, FSAF, GFL V1, TOOD and the improved network models based on sparse detection head, Faster R-CNN-SC, RetinaNet-SC, FSAF-SC, GFL_V1-SC, TOOD-SC were trained on the VisDrone dataset. The trained models were evaluated. The average precision AP [50:95] was used as the accuracy evaluation index, and the floating point operation number GFLOPs and frame rate FPS were used to evaluate the computational efficiency of the model.
[0059] Step 3: Input the drone image into the trained model to obtain the detection results.
[0060] [Example] The specific steps of the embodiment of the present invention are as follows: Step 1: Obtain the drone target detection dataset and perform preprocessing. The common drone remote sensing dataset VisDrone dataset is used. The dataset used is reasonably divided into training set, validation set and test set, which contain 6,471 images for model training, 548 images for model parameter adjustment and verification, and 3,190 images for final performance evaluation and test analysis. Preprocessing includes random cropping to a uniform size, input size 1333 × 800, and converting the dataset format to coco data format.
[0061] Step 2: Design of adaptive sparse detection head, which includes: ordinary sparse convolution, global feature generation, adaptive pixel-level mask unit, context-enhanced sparse convolution module, and loss function design. Specifically, it includes the following steps: Step 2.1: Figure 2 The network structure diagram of the cross-space UAV target detection method guided by adaptive sparse convolution. For the input UAV image, the backbone network in the basic algorithm, ResNet18, is used here to extract multi-scale features, which are recorded as {C3, C4, C5}. Then, multi-scale feature fusion is performed in the feature pyramid network to obtain multi-scale features {P3, P4, P5, P6, P7}. It is further input into the adaptive sparse detection head designed by the present invention for sparse detection.
[0062] Step 2.2: The common sparse convolution used is submanifold sparse convolution, such as Figure 3 As shown in the figure, the convolution output will only be calculated when the position covered by the center of the sparse convolution kernel is in an activated state, and the position where the value of the mask corresponding to the feature map is 1 is in an activated state.
[0063] Step 2.3: Global feature generation, based on multi-scale features {P3, P4, P5, P6, P7}, global features are generated at each scale through point-by-point convolution Specifically, the convolution kernel of the point-by-point convolution , the generated global features .
[0064] Step 2.4: Adaptive pixel-level mask unit, input the generated multi-scale features {P3, P4, P5, P6, P7} into the adaptive pixel-level mask unit designed by the present invention, and generate mask features at each scale , please refer to Figure 4 The specific principles are as follows: Given the first i Feature map of the layer , feature processing is performed by introducing a mask network. The mask network consists of a shared convolution kernel Composition, of which B , C , H , W represent batch size, number of channels, height and width respectively, and k The present invention adopts point-by-point convolution, that is, k =1 to reduce the computational complexity. exist Convolution is performed on it to produce a soft feature , which is further converted into a hard feature mask matrix by using the Gumbel-Softmax technique , which is expressed as follows:
[0065] in, and represents two random Gumbel noises, is the Sigmoid function, is the corresponding temperature parameter in Gumbel-Softmax.
[0066] Step 2.5: The context-enhanced sparse convolution module performs feature enhancement by combining focal information and global context information, thereby improving detection accuracy and enhancing the stability of subsequent calculations. The network structure diagram is referenced in Figure 5 , the principle is as follows: As a key component in the context-enhanced sparse convolution module, the global context information It is embedded into the Sparse Convolution-GN-ReLU layer. Since the sparse convolution is only To operate on some elements in After multiple rounds of sparse convolution processing, it gradually stabilizes, and the process adds almost no additional computational cost, ensuring both network efficiency and accuracy. The input of the context-enhanced sparse convolution module includes feature maps , sparse mask And global features ,in j Indicates j A context-enhanced sparse convolution module. Different from the traditional sparse convolution, the context-enhanced sparse convolution module does not use activation elements to calculate the group normalization statistics, but adopts The mean and standard deviation of are normalized to compensate for the missing context. Yes The output feature map after applying sparse convolution is then normalized by context-enhanced grouping to obtain the feature , as follows:
[0067] in, and denote the mean and standard deviation respectively, and and is a learnable parameter.
[0068] Finally, before the activation layer, in order to better preserve the context information, a residual structure is used to convert the global context features Directly add it to the sparse feature map, that is, operation.
[0069] Step 2.6: Loss function design. It includes classification loss , regression loss , normalized loss , and mask loss , specifically: It is used to measure the error between the category of the prediction box and the true value category. The predicted value of the category is The corresponding truth value is , The calculation of is as follows:
[0070] It is used to measure the error between the predicted box positioning and the true value box position. Using IoU loss calculation:
[0071] In order to effectively reduce the information loss that may occur during sparse convolution and ensure the stability of the training process, the present invention introduces conventional dense convolution during training. Specifically, in addition to sparse convolution, a feature map convolved on the complete input feature map is generated. This extra step not only helps capture more detailed information, but also provides a reliable basis for subsequent feature optimization. Subsequently, the feature map is optimized by optimizing the mean square error loss function. With sparse feature maps Combined to enhance the expressiveness of sparse feature maps, as follows:
[0072] in, L Represents the number of outputs of the feature pyramid.
[0073] Mask loss of the present invention The binary cross entropy loss function BCEWithLogitsLoss is used to calculate the soft feature mask generated by the adaptive pixel mask module. For supervision, the corresponding mask truth value is expressed as ,but It can be expressed as:
[0074] The total loss function is as follows:
[0075] in and To balance and The hyperparameters of the present invention are: and Set to 1 and 10.
[0076] Step 2.7: The designed adaptive sparse convolution detection head has a network structure as shown below: Figure 6 As shown, for multi-scale features as input features , which is processed by the mask unit to generate mask features , and generate global features after point-by-point convolution , further, , , As input, after being processed by four context-enhanced sparse convolutions, it is processed by two ordinary sparse convolutions to obtain classification features and regression features.
[0077] Step 3: Transplant the adaptive sparse detection head into the mainstream one-stage and two-stage target detection models. Specifically, the one-stage and two-stage target detection models used include Faster R-CNN, RetinaNet, FSAF, GFL_V1, and TOOD. These five models include not only one-stage and two-stage target detection models, but also target detection models based on anchor boxes and anchor-free boxes, which are representative. The sparse detection head is transplanted into these five models, named FasterR-CNN-SC, RetinaNet-SC, FSAF-SC, GFL_V1-SC, and TOOD-SC.
[0078] Step 4: Train and verify the improved model, and evaluate the improved method. Specifically, the basic network models Faster R-CNN, RetinaNet, FSAF, GFL V1, TOOD and the improved network models based on sparse detection head, Faster R-CNN-SC, RetinaNet-SC, FSAF-SC, GFL_V1-SC, TOOD-SC are trained on the VisDrone dataset respectively. The trained model is evaluated. The average precision AP[50:95] is used as the accuracy evaluation index, and the floating point operation number GFLOPs and frame rate FPS are used to evaluate the computational efficiency of the model. Its quantitative evaluation index is shown in Table 1: The results show that after the method of the present invention is applied to the classic target detection model, the detection results on key indicators such as AP[50:95], FPS, GFLOPs, etc. show obvious improvements, which fully verifies the effectiveness and superiority of the method of the present invention in the target detection task of UAV remote sensing image, and proves its significant advantages in sparse detection, computational efficiency and detection accuracy.
[0079] Table 1. Quantitative experimental results of the basic model and the improved model based on the method of the present invention on the VisDrone dataset.
[0080]
[0081] Step 5: Input the drone image into the trained model to obtain the detection result. Specifically, the trained model is used to perform adaptive sparse convolution-guided cross-space detection on the drone image to perform efficient detection. The detection results are as follows: Figure 7As shown, through the analysis of the detection results and pixel mask visualization results in different complex scenes, it can be seen that the four target detection methods after applying the method of the present invention can accurately detect the targets under the UAV remote sensing image, can adaptively generate the target area mask, and only identify the target area, thereby greatly reducing the model calculation amount. The detection results and pixel mask visualization results fully confirm the advantages of the method of the present invention in reducing the model calculation amount. Through the cross-space sparse detection strategy, the method of the present invention only focuses on the detection of the target area, while eliminating the interference of the complex background, while ensuring the detection accuracy, reducing the calculation amount and improving the detection efficiency.
[0082] The present application also discloses an adaptive sparse convolution-guided cross-space UAV target detection system, comprising: An acquisition unit is used to acquire a UAV target detection data set, and sequentially perform multi-scale feature extraction and multi-scale feature fusion on the UAV target detection data set to obtain multi-scale features; A data processing unit, used for inputting multi-scale features into a preset adaptive sparse detection head for processing to obtain classification features and regression features; The output unit is used to obtain the target detection result based on the classification features and regression features.
[0083] The present application discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the cross-space UAV target detection method guided by adaptive sparse convolution are implemented as described in any one of the above items.
[0084] The present application discloses a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the steps of the cross-space UAV target detection method guided by adaptive sparse convolution are implemented.
[0085] In summary, the present invention first obtains the drone target detection data set and preprocesses it. Then, an adaptive sparse detection head is designed. Then, the adaptive sparse detection head is transplanted into the mainstream one-stage and two-stage target detection models. Then, the improved model is trained and verified, and the improved method is evaluated. Finally, the drone image is input into the trained model to obtain the detection result. The present invention solves the problem that the target detection algorithm consumes a lot of computing resources and has a slow detection speed due to the high sparsity of targets in the complex background of drone images. The adaptive sparse detection head designed by the present invention can be plug-and-play, is suitable for various target detection algorithms, and significantly reduces the computational cost. An adaptive pixel-level mask unit is designed through the task alignment strategy, which can generate a target pixel-level mask, and realizes cross-space detection of the target through sparse convolution based on the global context information enhancement strategy, focusing only on the target area detection, while eliminating the interference of the complex background, while ensuring the detection accuracy, reducing the amount of calculation and improving the detection efficiency. It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0086] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0087] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An adaptive sparse convolution-guided cross-space UAV target detection method, characterized in that: include: Obtain a UAV target detection data set, perform multi-scale feature extraction and multi-scale feature fusion on the UAV target detection data set in turn, and obtain multi-scale features; The multi-scale features are input into the preset adaptive sparse detection head for processing to obtain classification features and regression features; The target detection results are obtained based on classification features and regression features.
2. According to claim 1, the method for cross-space UAV target detection guided by adaptive sparse convolution is characterized in that: The method of obtaining a drone target detection data set and sequentially performing multi-scale feature extraction and multi-scale feature fusion on the drone target detection data set to obtain multi-scale features specifically includes: S101: Acquire UAV remote sensing dataset; S102: randomly crop the UAV remote sensing dataset to a uniform size to obtain a detection dataset; S103: Divide the detection data set into a training set, a validation set and a test set and convert them into coco data format; S104: Perform multi-scale feature extraction on the detection data set to obtain basic features; S105: Perform multi-scale feature fusion on the basic features to obtain multi-scale features.
3. According to claim 1, the method for cross-space UAV target detection guided by adaptive sparse convolution is characterized in that: The multi-scale features are input into a preset adaptive sparse detection head for processing to obtain classification features and regression features, specifically including: S201: The multi-scale features are processed by global feature generation to obtain global features, and the multi-scale features are processed by adaptive pixel-level masking to obtain mask features; S202: multi-scale features, global features and mask features are simultaneously subjected to a number of context-enhanced sparse convolutions to obtain a sparse feature map with enhanced expressiveness; S203: The sparse feature map with enhanced expressiveness is processed through a plurality of sparse convolutions to obtain classification features and regression features.
4. According to claim 3, the method for cross-space UAV target detection guided by adaptive sparse convolution is characterized in that: The S201 specifically includes: S2011: Multi-scale features generate global features at each scale through point-by-point convolution; S2012: Multi-scale feature input mask unit processing to obtain a soft feature mask; a binary Gumbel-Softmax method is used on each element of the soft feature mask to convert the soft feature mask into a pixel-level hard feature mask as the mask feature output.
5. According to claim 3, the method for cross-space UAV target detection guided by adaptive sparse convolution is characterized in that: The S202 specifically includes: S2021: Multi-scale features and mask features are processed by sparse convolution to obtain sparse features; S2022: Sparse features and global features are processed by context-enhanced group normalization to obtain a sparse feature map; S2023: Add the sparse feature map to the global feature to obtain a sparse feature map with enhanced expressiveness.
6. The method for cross-space UAV target detection guided by adaptive sparse convolution according to claim 1 is characterized in that: The adaptive sparse detection head comprises: The global feature generation module is used to process multi-scale features through point-by-point convolution to obtain global features; An adaptive pixel-level mask unit, used to perform mask processing on multi-scale features to obtain mask features; Context-enhanced sparse convolution module, which is used to perform feature enhancement based on multi-scale features, global features, and mask features to obtain sparse feature maps with enhanced expressiveness; Sparse convolution module, used to optimize the sparse feature map with enhanced expressiveness to obtain classification features and regression features; The loss function module is used to calculate the loss of classification features and regression features. After iterative training, the adaptive sparse detection head with the best weight is obtained.
7. The method for cross-space UAV target detection guided by adaptive sparse convolution according to claim 6 is characterized in that: The loss function module uses the following formula to calculate the loss: Classification Loss for: Regression Loss for: Normalized loss for: Mask loss for: in, and For balance and Hyperparameters of For the The predicted value of the category; For the The true value of the category; Represents the intersection and union ratio of the predicted value and the true value; is the feature map convolved on the complete input feature map; Sparse feature maps for enhanced expressiveness; Represents mask features; Represents the number of positive samples; represents the true value of the mask; Represents the predicted value of the pixel-level mask unit.
8. An adaptive sparse convolution-guided cross-space UAV target detection system, characterized in that: include: An acquisition unit is used to acquire a UAV target detection data set, and sequentially perform multi-scale feature extraction and multi-scale feature fusion on the UAV target detection data set to obtain multi-scale features; A data processing unit, used for inputting multi-scale features into a preset adaptive sparse detection head for processing to obtain classification features and regression features; The output unit is used to obtain the target detection result based on the classification features and regression features.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the adaptive sparse convolution-guided cross-space UAV target detection method described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the adaptive sparse convolution-guided cross-space UAV target detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Unmanned aerial vehicle tracking method based on adaptive target frame optimization
CN114820712A
Sparse SAR ship target detection method and device based on attention feature fusion
CN117372676A
Unmanned aerial vehicle aerial image target detection method based on selective attention and adaptive feature fusion
CN118570676A
Unmanned aerial vehicle target detection method based on space-frequency feature fusion detection head
CN118691929A