Unmanned aerial vehicle target detection method and system based on dense shielded network
By introducing dense occluded networks in UAV target detection, combined with OKA, PSI, SDEA and PDF algorithms, the problem of low target detection accuracy in complex environments is solved, and higher detection accuracy and real-time performance are achieved.
Patent Information
- Application Number
- CN202411726224.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-27
AI Technical Summary
The existing drone target detection algorithms have reduced detection accuracy in complex environments, especially under dense distribution and occlusion conditions, making it difficult to effectively identify and locate small targets.
A drone object detection method based on dense occluded network is proposed. Basic features are extracted through OKA module, PSI module performs feature integration, SDEA module performs feature aggregation, and the two features are fused by PDF algorithm, and finally target detection is used using RT-DETR decoder.
It significantly improves the target detection accuracy of the drone in complex scenarios, especially under dense occlusion conditions, which can more accurately identify and locate small targets while maintaining high real-time performance.
Smart Images

Figure CN120047849A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of UAV target detection, and particularly relates to a UAV target detection method and system based on a dense occluded network. Background Art
[0002] With the continuous development of UAV technology, UAVs are increasingly widely used in various application scenarios, such as agricultural monitoring, urban planning, environmental protection, etc. Through its flexible flight ability and high-definition camera, a UAV can obtain high-resolution images from multiple angles, providing rich data for target detection. However, due to the dense target distribution and occlusion problems in the images captured by UAVs, the performance of traditional target detection algorithms in these scenarios is often unsatisfactory. Especially in complex environments, the occlusion and dense distribution of targets lead to a decrease in detection accuracy, which becomes a major challenge for UAV target detection.
[0003] In recent years, the development of deep learning technology has brought significant progress to target detection. Classic target detection algorithms can be divided into single-stage algorithms such as the SSD series, CenterNet, YOLO series, etc. and two-stage algorithms such as the R-CNN series. The R-CNN series has improved the detection accuracy and speed from R-CNN to Faster R-CNN by gradually improving the region extraction and feature extraction methods, especially achieving a significant improvement in accuracy. The YOLO series transforms the target detection task into a regression problem, providing an efficient real-time detection solution. And further optimizing the network structure and training strategy, improving the detection ability for complex scenarios and small targets. Single-stage detectors have the performance advantage of end-to-end, but their accuracy in locating and recognizing small targets is relatively low. Relatively speaking, two-stage target detectors adopt the strategy of first locating and then recognizing, and usually have better accuracy than single-stage detectors, but they are less real-time.
[0004] DETR was introduced by Facebook in 2022, which first introduced the Transformer architecture into the field of object detection. It can perform object detection and recognition in an end-to-end manner, omitting traditional prior boxes and non-maximum suppression techniques. The emergence of the DETR algorithm has attracted extensive attention in the field of object detection and achieved remarkable results. With the in-depth research, the DETR algorithm has been continuously evolving and has surpassed traditional convolutional neural network (CNN) methods in many scenarios. In particular, RT-DETR significantly improves the real-time speed by designing an efficient hybrid encoder module and outperforms similar YOLO detectors in terms of speed and accuracy, demonstrating the great potential of the DETR algorithm in object detection. However, although the DETR algorithm effectively handles long-range dependencies through the self-attention mechanism and simplifies the detection process, the problem of missed detection of small targets in drone images under dense distribution and occlusion conditions makes the existing methods insufficient in object extraction and recognition. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a drone object detection method and system based on a dense occluded network, which solves the problems in the prior art.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A drone object detection method based on a dense occluded network, comprising the following steps:
[0008] Collect aerial image data by a drone, resize the image, and then preprocess the image;
[0009] Based on the preprocessed image, use the OKA module to extract basic features;
[0010] Based on the basic features, use the PSI module for feature integration;
[0011] Based on the integrated features, use the SDEA module for feature aggregation;
[0012] Use the PDF algorithm to fuse the features integrated by the PSI module and the features aggregated by the SDEA module, fuse the two-way feature outputs by element-wise addition, and finally use the RT-DETR decoder for object detection to output the detection results;
[0013] Post-process and optimize the obtained detection results.
[0014] Furthermore, the OKA module combines KAGN adaptive convolution and GRAM polynomial to expand the feature space. First, it normalizes the input features to the interval [-1, 1], then expands the feature representation through the Legendre polynomial basis, optimizes the feature extraction through the KAN network, and provides the flexibility of feature extraction by parameterizing the learnable activation function and spline function.
[0015] Furthermore, the process of expanding the feature representation through the Legendre polynomial basis is as follows:
[0016] 1) Calculate the Legendre polynomial basis
[0017]
[0018] where N is the highest order of the polynomial, x′ is the independent variable of the polynomial, and n is the current order of the polynomial;
[0019] 2) Combine the Legendre polynomial basis with the learned polynomial weights to obtain the expanded feature y:
[0020]
[0021] Furthermore, the PSI module first performs spatial alignment on the multi-scale feature maps, unifies different-scale features through adaptive average pooling, identity mapping, and bilinear interpolation; then applies partial convolution for feature selection, controls the convolution calculation using a mask matrix, and reorganizes the feature channels through channel shuffle technology.
[0022] Furthermore, the SDEA module uses convolution kernels with different dilation rates to extract multi-scale features, introduces the SimAM attention mechanism to dynamically adjust the feature weights, and uses the structural reparameterization technology to optimize the inference efficiency.
[0023] Furthermore, the post-processing optimization process includes:
[0024] Perform confidence filtering on the detection results to screen high-confidence targets;
[0025] Output the final target category and location coordinate information, and generate a visualization image of the detection results.
[0026] An unmanned aerial vehicle target detection system based on a dense occluded network, comprising:
[0027] Image data acquisition and processing module: Collect aerial image data through an unmanned aerial vehicle, resize the image, and then preprocess the image;
[0028] Base feature extraction module: Based on the preprocessed image, use the OKA module to extract base features;
[0029] Feature integration module: Based on the base features, use the PSI module to perform feature integration;
[0030] Feature aggregation module: Based on the integrated features, use the SDEA module to perform feature aggregation;
[0031] Feature fusion module: Use the PDF algorithm to fuse the features integrated by the PSI module and the features aggregated by the SDEA module. Fuse the two-way feature outputs by element-wise addition. Finally, use the RT-DETR decoder for object detection and output the detection results;
[0032] And, result optimization module: Post-process and optimize the obtained detection results.
[0033] A computer storage medium stores a readable program that, when the program runs, can execute the above-mentioned UAV target detection method based on a dense occluded network.
[0034] An electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete mutual communication through the communication bus;
[0035] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned UAV target detection method based on a dense occluded network.
[0036] A DRONet model for UAV target detection includes: a backbone network, an encoder, and an RT-DETR decoder;
[0037] The backbone network includes an OKA module for feature extraction. The OKA module combines KAGN adaptive convolution and GRAM polynomial to expand the feature space. First, normalize the input features to the [-1, 1] interval, then expand the feature representation through the Legendre polynomial basis, and optimize the feature extraction through the KAN network. And use learnable activation functions and spline function parameterization to provide flexibility in feature extraction;
[0038] The encoder includes a PSI module and an SDEA module for feature integration and aggregation respectively; the PSI module first performs spatial alignment on the multi-scale feature maps, unifying features of different scales through adaptive average pooling, identity mapping, and bilinear interpolation; subsequently, it applies partial convolution for feature selection, controls the convolution calculation using a mask matrix, and reorganizes the feature channels through the channel shuffle technique; the SDEA module extracts multi-scale features using convolution kernels with different dilation rates, introduces the SimAM attention mechanism to dynamically adjust the feature weights, and uses the structural reparameterization technique to optimize the inference efficiency;
[0039] And a PDF algorithm is proposed by combining the PSI module and the SDEA module. After fusing the integrated features of the PSI module and the aggregated features of the SDEA module through the PDF algorithm, the RT-DETR decoder is used for object detection to output the detection results.
[0040] Advantages of the present invention:
[0041] 1. The present invention proposes a novel real-time multi-object detection algorithm for unmanned aerial vehicles, effectively integrating the deep feature extraction advantage of the convolutional neural network and the global context modeling ability of the Transformer architecture, and improving the multi-object detection accuracy of unmanned aerial vehicles.
[0042] 2. The present invention designs an Occlusion-Aware KAGN Block (OKA) module, which integrates a novel and efficient KAGN convolution into ResNet, and uses adaptive adjustment of the convolution kernel and GRAM multiple selection, significantly improving the efficiency and feature extraction ability of the model.
[0043] 3. The present invention proposes a Perceptual Dilated Fusion (PDF) feature fusion algorithm. This algorithm combines the advantages of the Perceptual Spatial Integration (PSI) and Scalable Dilated Efficient Aggregation (SDEA) modules, and through techniques such as integrated feature map adjustment, partial convolution, channel shuffle, and dilated convolution, realizes the precise extraction and efficient fusion of multi-level deep features. This feature fusion mechanism significantly improves the feature aggregation efficiency of the model in complex scenarios and the object recognition accuracy under dense occlusion conditions. Description of the Drawings
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0045] Figure 1 It is the architecture diagram of the DRONet model of the present invention;
[0046] Figure 2 It is the architecture diagram of the KAN of the present invention;
[0047] Figure 3 It is the structure diagram of the OKA module of the present invention;
[0048] Figure 4 It is the flow chart of the PDF algorithm of the present invention;
[0049] Figure 5 It is the action diagram of the PConv of the present invention;
[0050] Figure 6 It is the combined action diagram of PConv and channelshuffle of the present invention;
[0051] Figure 7 It is the structure diagram of the PSI module of the present invention;
[0052] Figure 8 It is the structure diagram of the DRB of the present invention;
[0053] Figure 9 It is the structure diagram of the DEB of the present invention;
[0054] Figure 10 It is the structure diagram of the SDEA module of the present invention;
[0055] Figure 11 It is the change curve of some important evaluation indexes of the DRONet model and RTDETR-r18 during the training process.
[0056] Figure 12 It is the confusion matrix and normalized confusion matrix of RTDETR-r18;
[0057] Figure 13 It is the confusion matrix and normalized confusion matrix of DRONet;
[0058] Figure 14a It is the original image for inference of two models, DRONet and RTDETR-r18, on a specific image;
[0059] Figure 14b It is the inference result diagram of the RTDETR-r18 model on a specific image; Figure 14c It is the inference result diagram of the DRONet model on a specific image; Figure 15 It is the inference heat map result of two models, DRONet and RTDETR-r18, on a specific image. Detailed Implementation Manner
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0061] Embodiment 1
[0062] In this embodiment, a Dense Receptive Occlusion Network (DRONet) model is proposed. As Figure 1 shown, the architecture consists of three parts: a backbone network, an encoder, and an RT-DETR (Real-Time Detection Transformer) decoder. In the backbone part, an Occlusion-Aware KAGN Block (OKA) module for feature extraction is proposed; this module can efficiently extract multi-scale context information in a dense target scene and enhance the model's feature extraction ability. In the encoder part, a novel Perceptual Dilated Fusion (PDF) algorithm is designed, which includes a Perceptual Spatial Integration (PSI) module and a Scalable Dilated Efficient Aggregation (SDEA) module. The PSI module helps the model obtain multi-layer feature information of the target by screening feature groups and shuffling the permutation order through channel shuffling, improving the detection performance for occluded targets. The SDEA module increases the model's spatial receptive field and efficiently aggregates the depth features of the image, enhancing the model's analysis ability in multi-scale complex scenes. After fusing the features integrated by the PSI module and the features aggregated by the SDEA module through the PDF algorithm, the RT-DETR decoder is used for target detection, and the detection results are output.
[0063] Based on the above DRONet model, this embodiment also proposes a UAV target detection method based on a dense occluded network, including the following steps:
[0064] S1, collect aerial image data through a UAV, adjust the image to a size of 640×640 as the input image, and then preprocess the input image;
[0065] The process of preprocessing the input image includes:
[0066] 1) First, standardize the input image and normalize the pixel values to the range [0, 1].
[0067] 2) Then, perform data augmentation on the image, including geometric transformations such as random rotation, flipping, and scaling, as well as optical enhancements such as brightness, contrast, and hue. Additionally, include occlusion enhancement methods such as random occlusion and mixed cropping.
[0068] S2. Based on the preprocessed image in S1, use the OKA module to extract basic features.
[0069] As Figure 3 shown, the OKA module combines KAGN adaptive convolution and GRAM polynomial to expand the feature space. First, normalize the input features to the range [-1, 1], then expand the feature representation through the Legendre polynomial basis, optimize the feature extraction through the KAN network, and use learnable activation functions and spline function parameterization to provide flexibility in feature extraction.
[0070] 1) The GRAM polynomial expands the spatial representation of the input features through the Legendre polynomial to enhance the model's expressive power. Specifically, first normalize the input feature x to the range [-1, 1], and then calculate the Legendre polynomial basis where N is the highest order of the polynomial.
[0071] The definition of the Legendre polynomial basis is as follows:
[0072]
[0073] In the formula, N is the highest order of the polynomial, x' is the independent variable of the polynomial, and n is the order of the current polynomial.
[0074] After that, combine the Legendre polynomial basis with the learned polynomial weights to obtain the expanded feature y:
[0075]
[0076] These expanded features are then passed to the subsequent layers of the network for further processing.
[0077] 2) Different from traditional neural networks that use fixed activation functions, KAN adopts learnable activation functions at the edges of the network and replaces the weight parameters in KAN with univariate functions, which are usually parameterized in the form of spline functions, thus providing extremely high flexibility and being able to simulate complex functions with fewer parameters, enhancing the interpretability of the model. Since the output of the aggregation function of each node in KAN does not use non-linear transformation, the problems of inherently low parameter efficiency and limited interpretability of multi-layer perceptrons (MLPs) are overcome. As Figure 2 shown, similar to MLP, the N-layer KAN can be expressed as the nesting of multiple KAN layers:
[0078]
[0079] where, Φ N represents the function vector of the activation function of the Nth layer of the entire KAN network, where each neuron has a different activation function. Each KAN layer has an n in -dimensional input and an n out -dimensional output, and Φ includes n in × n out learnable activation functions:
[0080] Φ = {φ q,p}, p = 1, 2, …, n in , q = 1, 2…, n out , (4)
[0081] KAN can have any number of layers, and each layer can have any number of nodes. Assume that the input of the KANs in the lth layer is and the output is The result calculation formula is M n+1 = Φ n M n , where,
[0082]
[0083] 3) A new type of KAGN convolutional layer is designed by combining KAN with GRAM polynomials, as Figure 3As shown in the figure. KAN ensures the effective extraction of basic features and has good robustness even in the case of occlusion, while the GRAM polynomial enhances the learning ability of complex patterns by expanding the feature space. The KAGN convolutional layer extracts basic features through basic convolutional operations and expands these features through the GRAM polynomial, significantly improving the model's detection performance for occlusion and small targets. By integrating the KAGN convolutional layer into ResNet to form the OKA model, the OKA model first uses the basic convolutional layer to extract basic features (referred to as basis). Then, the Legendre polynomial basis is calculated and combined with the corresponding weights to generate a new feature representation. Finally, these new features are combined with the basic feature basis and processed through normalization and activation functions to produce the final feature representation. The OKA module helps the model to better identify and locate targets in a dense target scenario, and can maintain a high detection accuracy even when the target distribution is dense and there is occlusion. OKA not only improves the robustness of the model, but also enhances the model's adaptability and generalization ability to complex scenarios.
[0084] S3, based on the basic features extracted in S2, uses the PSI module for feature integration;
[0085] The PSI module first performs spatial alignment on the multi-scale feature maps, unifying different-scale features through adaptive average pooling, identity mapping, and bilinear interpolation; subsequently, partial convolution (PConv) is applied for feature selection, using a mask matrix to control the convolution calculation, and the feature channels are reorganized through channel shuffle technology to enhance the expressive ability of the features.
[0086] In the prior art, the most common concat operation is used for the fusion operation of feature maps of different sizes and levels. However, the simple concat operation simply concatenates multiple feature maps along the channel dimension. Since UAV aerial images contain complex backgrounds and small targets, such processing will introduce a large number of non-target features, interfering with the model's judgment. And due to the insufficient cross-level interaction of the concat operation, the model cannot fully learn the spatial and semantic information between the feature maps, making it more difficult to detect small targets and distinguish dense targets. To solve these problems, the PSI module is proposed. While increasing the extraction of feature space and semantic information, it uses partial convolution and channel confusion technology for feature selection and fusion, enabling the model to have stronger adaptability to dense occlusion and complex scenarios.
[0087] As Figures 5 to 6As shown, in the PSI module, PConv is used for spatial extraction of features, which can reduce redundant calculations and minimize the memory access frequency at the same time. PConv selectively applies Conv to a subset of the input channels for feature extraction while keeping the remaining channels unchanged. For continuous or regular memory access patterns, the first or last adjacent set of channels of c p is regarded as a representative sample of the entire feature mAP for calculation. PConv treats the input and output feature mAP channels as equivalent, thus improving the computational efficiency and reducing the need for frequent memory access. PConv can dynamically activate specific parts of the feature map according to the feature importance of a specific area, which helps to filter out irrelevant information in complex backgrounds, thereby enhancing the focusing ability on small targets. The calculation formula of the floating-point operation FLOP of PConv is as follows:
[0088]
[0089] In the formula, h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c p is the number of channels.
[0090] PConv focuses on enhancing the object detection ability by selectively activating valuable feature regions, especially filtering out irrelevant information in complex backgrounds. However, the scope of action of PConv is limited to the spatial information of the current feature map. By introducing channel shuffling, the PSI module can rearrange the features of different channels, enabling more sufficient interaction of feature information from different levels and scales between channels, as Figure 6 shown. Such hybrid interaction not only enhances the information fusion in the spatial dimension but also enables PConv to better utilize the cross-channel semantic information, thereby enhancing the overall feature expression ability and optimizing the adaptability of the model to small targets and dense scenes.
[0091] The algorithm pseudo-code of the PSI module is shown in Table 1 below.
[0092] Table 1 Algorithm Pseudo-code of the PSI Module
[0093]
[0094] As Figure 7 shown, the steps of using the PSI module for feature integration include:
[0095] 1) Using the hierarchical feature maps generated by the encoding of DRONet where represents the initial feature map generated by the i-th layer encoder, H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map; three multi-scale feature maps are input for fusion, and 1×1 convolution is used to align to the channels of c, and the resulting feature map is denoted as
[0096] 2) In the i-th level of the decoder of the PSI module, represents the target reference. Then, the size of the feature map at each j-th level is adjusted to match the size of, as shown in Equation (7):
[0097]
[0098] where, represents the output feature after the i-th layer in the second stage of the SDI module processes the j-th layer feature. i is the index of the current layer, and j is the index of the feature map layer to be processed. F d , F I and F U represent adaptive average pooling, identity mapping, and bilinear interpolation. (H i , W i ) represent the height and width dimensions of the target feature map.
[0099] 3) After the feature map sizes are aligned, PConv is applied for fusion:
[0100]
[0101] where, W is the convolution kernel, M is the mask matrix that controls which pixels participate in the convolution calculation, and ∑M represents the sum of the masks within the convolution window, which is used for normalization.
[0102] 4) The channels of the feature map are rearranged through channel shuffle. The input feature map is divided into G groups, and the number of channels in each group is as shown in Equation (9):
[0103]
[0104] Channel rearrangement is performed within each group g:
[0105]
[0106] The groups after channel rearrangement are recombined into a complete feature map:
[0107]
[0108] Finally, is the feature map after channel shuffle.
[0109] 5) After converting all the i-th level feature maps to the same resolution, the element-wise Hadamard product is applied to all the resized feature maps (as shown in the following equation) to obtain the feature map to enhance feature fusion, while forwarding to the decoder at the i-th level for further resolution reconstruction and segmentation:
[0110]
[0111] S4. Based on the integrated features of S4, the SDEA module is used for feature aggregation,
[0112] The SDEA module uses convolutional kernels with different dilation rates to extract multi-scale features, introduces the SimAM attention mechanism to dynamically adjust the feature weights, and uses the structural reparameterization technique to optimize the inference efficiency, thereby enhancing the model's detection ability for small targets and occluded targets.
[0113] In a convolutional neural network (CNN), combining large-kernel convolutions with parallel small-kernel convolutions helps to extract feature information at different scales simultaneously. The outputs of these convolutions are added after passing through their respective batch normalization (BN) layers. By using the structural reparameterization technique, the BN layer can be integrated with the convolutional layer, and after training is completed, the small-kernel convolution can be effectively merged into the large-kernel convolution to optimize the computational efficiency in the inference stage.
[0114] As Figure 8 shown, the core design of the DRB (Dilated Re-param Block) is to enhance the performance of the large-kernel convolutional layer through parallel dilated small-kernel convolutional layers. By simultaneously using a large-kernel convolutional layer and multiple convolutional layers with different dilation rates in parallel, the dilated reparameterization block can capture local details and sparse features that are far apart, enabling the model to better understand the multi-level structure in the input data. During inference, these convolutional layers are merged into a non-dilated convolutional layer to reduce the computational overhead. Specifically, this process includes converting the dilated convolution to an equivalent non-dilated convolution and integrating the outputs of each layer. The performance of this module can be optimized by flexibly adjusting the large-kernel size K, dilation rate r, and small-kernel size k. For example, Figure 8 the settings shown in
[0115] According to the advantages of the DRB, the Dilated Efficient Block (DEB) module is designed, as Figure 9As shown in the figure. The DEB module continues the design concept of the DRB. By combining large kernel convolutions with small kernel convolutions of different dilation rates, it enhances the diversity of feature extraction. The DEB not only maintains its core advantage, that is, enhancing the performance of large kernel convolutions through parallel dilated small kernel convolutions. At the same time, it also introduces a lightweight SimAM (Simple Attention Module) attention mechanism, enabling the network to dynamically adjust the importance of features at different scales. Specifically, after each parallel small kernel dilated convolution, the SimAM mechanism is added. This mechanism enhances key features through simple mathematical operations while reducing redundant information processing. The SimAM mechanism can improve the model's attention to important features without additional parameters, thereby capturing local details and long-range dependencies.
[0116] In addition, the DEB module achieves efficient inference through structural reparameterization techniques, such as Figure 9 shown in the figure. During the training process, each dilated convolution path remains independent to capture diverse features. While in the inference stage, these paths will be reparameterized and merged into a single large kernel convolution, which can not only provide powerful feature extraction capabilities during training but also ensure excellent computational efficiency during network deployment. The selection of parameters K (large kernel convolution size), r (set of dilation rates), and k (small kernel convolution size) can be adjusted according to specific application scenarios and computing resources. For example, for high-resolution image processing, larger K values and more combinations of dilation rates r can be selected, while real-time applications tend to choose smaller K values and fewer dilation rate options to meet the requirements of real-time performance.
[0117] The original model used traditional convolution operations in obtaining multi-scale context information. This method usually only focuses on local feature extraction and fails to fully integrate features under different receptive fields, resulting in information loss or confusion, especially when dealing with drone aerial images with complex backgrounds and multi-scale targets. To address these issues, the present invention proposes a Scalable Dilated Efficient Aggregation (SDEA) module, as Figure 10 shown in the figure. The SDEA module divides and recombines features by introducing dilated convolutions and an efficient feature aggregation strategy, and processes them using hierarchical convolutions, comprehensively improving the extraction and fusion effects of multi-scale features. It uses convolution kernels with multiple dilation rates to capture features under different receptive fields and reduces information loss through an efficient feature aggregation mechanism at each level. Thereby significantly enhancing the model's detection ability for small and dense targets and improving the model's adaptability in complex scenarios.
[0118] The SDEA module adopts multiple parallel dilated small kernel convolutional layers. Each convolutional layer has a different dilation rate, which enables SDEA to capture features at different scales. In this way, SDEA can not only handle local details but also capture context information over a larger range. This design makes the model more robust when dealing with complex backgrounds and multi-scale objects. Especially in drone aerial images, it can more accurately identify and locate target objects.
[0119] The feature extraction ability is enhanced by using small kernel convolutions with different dilation rates in parallel. After each parallel small kernel dilated convolution, a SimAM lightweight attention mechanism is added to dynamically adjust the importance of features at different scales and reduce redundant information processing. This design enables SDEA to not only capture local details and long-range dependencies but also highlight key features through the attention mechanism.
[0120] In addition, the SDEA module incorporates the GELAN structure, which is a network architecture specifically designed for efficient feature extraction and aggregation. The GELAN structure enhances feature representation through the segmentation and recombination of features and the use of hierarchical convolution processing. When combined with the SDEA module, the GELAN structure can better handle multi-scale features and maintain high detection performance in complex backgrounds.
[0121] In terms of feature aggregation, the SDEA module achieves efficient inference through the structural reparameterization technique. During the training process, each dilated convolution path remains independent to facilitate the capture of diverse features; while in the inference stage, these paths will be reparameterized and merged into a single large kernel convolution, thus ensuring excellent computational efficiency during network deployment while maintaining the model's strong feature extraction ability.
[0122] By introducing the SDEA module and combining it with the GELAN structure, not only are the limitations of traditional convolutions in multi-scale feature extraction addressed, but the adaptability and robustness of the model in complex scenarios are also improved. The selection of parameters K (large kernel convolution size), r (set of dilation rates), and k (small kernel convolution size) can be adjusted according to specific application scenarios and computing resources to achieve an optimal performance balance. For example, in high-resolution image processing, larger K values and more combinations of dilation rates r can be selected, while in real-time applications, smaller K values and fewer dilation rate options may be preferred to meet real-time requirements. Through these improvements, the SDEA module provides a more efficient and reliable solution for dealing with complex backgrounds and multi-scale objects.
[0123] S5. The PDF algorithm is used to fuse the features integrated by the PSI module and the features aggregated by the SDEA module. The two-way feature outputs are fused by element-wise addition. Finally, the RT-DETR decoder is used for object detection, and the detection results are output, including: object category prediction and bounding box position information.
[0124] A new feature fusion algorithm, called Perceptual Dilated Fusion (PDF), is proposed by combining the PSI and SDEA modules. This algorithm further improves the detection performance of multi-scale objects by combining the fine feature integration ability of the PSI module and the efficient feature aggregation ability of the SDEA module.
[0125] S6. Post-processing optimization is performed on the detection results obtained in S5.
[0126] The post-processing optimization includes: first, confidence filtering is performed on the detection results to screen out high-confidence objects, and then the final object category and position coordinate information are output, and a visualization image of the detection results is generated.
[0127] Based on a similar inventive concept, an embodiment of the present invention also provides a computer storage medium storing a readable program, which can execute the above-mentioned unmanned aerial vehicle object detection method based on a dense occluded network when the program runs.
[0128] Based on a similar inventive concept, an embodiment of the present invention provides an electronic device, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete mutual communication through the communication bus.
[0129] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned unmanned aerial vehicle object detection method based on a dense occluded network.
[0130] Based on a similar inventive concept, an embodiment of the present invention also provides a computer program product, including computer instructions, and the computer instructions instruct a computing device to execute the operations corresponding to the above-mentioned unmanned aerial vehicle object detection method based on a dense occluded network.
[0131] Embodiment 2
[0132] The superiority of the DRONet model proposed by the present invention is verified through the following experimental tests. The data source, implementation details, evaluation metrics, comparative experiments, and experimental analysis of the experiments are as follows:
[0133] 1) In this embodiment, experiments were conducted on the VisDrone dataset and the CARPK dataset. The VisDrone dataset consists of aerial images taken by various drone models under different scenarios, weather conditions, and lighting conditions. The VisDrone dataset was released by the AISKYEYE team of the Machine Learning and Data Mining Laboratory at Tianjin University and consists of aerial images taken by drones in various environments in 4 different cities, including urban and rural areas, under various weather and lighting conditions. The dataset contains 10,209 static images (the training set contains 6,471 images, the validation set contains 548 images, and the test set contains 3,190 images). The dataset includes 10 categories: pedestrians, people, cars, vans, buses, trucks, motorcycles, bicycles, rickshaws, and tricycles, with approximately 2.6 million target instance samples.
[0134] CARPK is the first and largest parking lot dataset in the drone dataset. It includes data collected from four different parking lots, with nearly 90,000 cars observed by drones (UAVs). These images were taken from the drone at a height of approximately 40 meters, and a bounding box was labeled for each car. The maximum number of cars in a single scene is 188. CARPK contains a total of 1,448 static images. In this embodiment, 989 images were used as the training set and 459 images were used as the test set. The label category of this dataset is only "car", and the resolution of all images is 1280×720 pixels.
[0135] 2) The experiments were conducted on a computer with an NVIDIA 3090 GPU with a 2.6GHz CPU, 64GB RAM, and 24GB GPU memory. The corresponding software configuration includes the Ubuntu 20.04 system, PyTorch 1.11.0, CUDA 11.3, and Python 3.8.
[0136] Due to limited computing resources, the input size of 640 and a batch size of 4 were used for algorithm training. During the training process, the initial learning rate was set to 0.0001 and the AdamW optimizer was used, with a fixed momentum of 0.9 and a weight decay coefficient of 0.0001. All training did not use pre-trained weights, and other strategies and hyperparameters followed the baseline.
[0137] 3) In this paper, multiple metrics were used to evaluate the performance of the object detection algorithm, including mean average precision (mAP), GFLOPs, parameter size, and frames per second (FPS). GFLOPs can quantify the computational complexity of the model, the parameters can measure the size of the model, and FPS represents the actual inference speed of the model. The calculation formula of GFLOP is shown as the equation:
[0138] GFLOPs=[(C i×k w ×k h )+(C i ×k w ×k h -1)+1]×C o ×W×H×10 9 (11)
[0139] Among them, C i and C o represent the input and output channels, k represents the size of the convolutional kernel, and H and W represent the size of the output feature map.
[0140] AP (Average Precision) is divided into 10 different intervals according to the IoU threshold, and AP is calculated every 0.05 in [0.5, 0.95]. mAP (mean Average Precision) is the average value of all APs. Generally, a higher mAP indicates better detection performance. mAP 50 is measured using an IoU threshold of 0.5. The calculation steps of mAP are as follows:
[0141]
[0142]
[0143] Among them, TP (True Positives) is the true positive class; TN (True Negatives) is the true negative class; FP (False Positives) is the false positive class; FN (False Negatives) is the false negative class. Precision Pre represents the proportion of objects that are actually positive among all objects predicted as positive by the model. Recall Rec represents the proportion of objects that are correctly predicted as positive by the model among all objects that are actually positive. N represents the total number of categories.
[0144] AP (Average Precision) is the area under the curve calculated by the precision and recall rate curve (PR curve) under the selected decision threshold above, and the calculation formula is:
[0145]
[0146] mAP (mean Average Precision) is the average value of the average precision (AP) values of each category, and its calculation formula is as follows:
[0147]
[0148] 4) Comparative experiment
[0149] To comprehensively evaluate the detection performance of the DRONet model proposed in this embodiment, experiments were conducted on the VisDrone and CARPK datasets respectively. And it was compared with various state-of-the-art object detectors in recent years, such as YOLOv 5L, TPH-YOLOv5, YOLOv 6L, YOLOv7, and YOLOv8L, these YOLO series detection models; DETR and its various variant models such as DN-DETR and DINO; and the two-stage Faster R-CNN detection model. Compared with other networks, the proposed model DRONet in this paper achieved the best accuracy, number of parameters, and computational complexity.
[0150] In the experiment, self-distillation technology was also used to improve the performance of DRONet in multi-object detection for drones. Self-distillation is a method of self-supervised learning for models. By using the prediction results of the model during the training process to generate soft labels, the model can not only rely on the true labels during learning but also learn its own knowledge. In this way, the model can better capture the potential features in the data, thereby enhancing its generalization ability.
[0151] In this paper, the method of feature distillation was adopted for self-distillation. Its core idea is to further optimize the parameters of the model by comparing the differences between the soft labels and the true labels output at different levels of the model. In the experiment, the feature distillation process was integrated into the training stage of DRONet, enabling the model to continuously use the information generated by itself to improve its performance during the training process. This method is especially suitable for dense targets and occlusion situations, and helps to improve the detection ability of small targets. In this paper, five feature distillation methods, namely SPKD
[32] , ChSim
[33] , CWD
[34] , MGD
[35] , and Mimic, were used to conduct self-distillation experiments on the VisDrone dataset respectively, and the results are shown in Table 2. Self-distillation improves the accuracy of the model while keeping the number of model parameters and the amount of calculation unchanged. Among them, the accuracy obtained by using CWD distillation is the best, which is used as a reference for subsequent experimental comparison.
[0152] Table 2 Object detection results of different distillation methods on VisDrone
[0153]
[0154] Figure 11 Shows the change curves of some important evaluation indicators of the DRONet model and RTDETR-r18 proposed in this embodiment during the training process. Among them, Figure 11 (a) in it is the training curve of DRONet and RTDETR-r18 on mAP, Figure 11 (b) in it is the training curve of DRONet and RTDETR-r18 on precision, Figure 11Among them, (c) shows the training curves of DRONet and RTDETR-r18 in terms of recall. It can be seen from the figure that the DRONet model outperforms RTDETR-r18 in three detection metrics: precision, recall, and mAP50. After approximately 20 epochs of training from the beginning, the DRONet model starts to stabilize after about 50 epochs of training. Compared with RTDETR-r18, this method has a faster training speed and better detection effect.
[0155] This embodiment lists the network structures used by the backbones of all models. The backbone of Faster-RCNN uses ResNet50 and adds the FPN (Feature Pyramid Network) structure. The backbone of DynaMask adds a dynamic feature extraction module on the basis of Faster-RCNN to dynamically respond to different targets. The backbone of YOLOv5 uses CSPDarknet53 with a large number of C3 modules. YOLOv6 is improved by introducing EfficientNet on the basis of YOLOv5. The backbone of YOLOv7 introduces the E-ELAN (Extended-ELAN) structure and optimized convolutional calculation modules such as the Conv-Bn-SiLU combination and depthwise separable convolution. The backbone of YOLOv8, Enhanced ELAN, inherits the ELAN structure of YOLOv7 and simultaneously optimizes CSPNet and the convolutional calculation unit. The backbones of DETR and its variant models uniformly use ResNet50. The backbone of the benchmark model RT-DETR selected in this paper consists of ResNet18.
[0156] Table 3 shows the results of the comparative experiments on the VisDrone dataset. The DRONet model of this embodiment achieves a higher accuracy compared to other models. Compared with the baseline model RT-DETR, the mAP50 of the DRONet model of this embodiment has increased by 3.1 percentage points. The model achieves a performance of 60 FPS in real-time detection. And compared with other models, the DRONet model of this embodiment has fewer GFLOPs while having a similar parameter count to other models, but the accuracy has been greatly improved. Due to the addition of multiple components, the GPU memory required during the training of the model is relatively high, exceeding the requirements of most models of the same type. Through the comparison, it shows that the DRONet model designed in this paper has the characteristics of high accuracy, lightweight, and real-time, which is beneficial for deployment on small-computing-power drones with edge computers.
[0157] Table 3 Comparative Experiments of Different Algorithms on the VisDrone Validation Set
[0158]
[0159] At the same time, a migration experiment was conducted on the CARPK dataset to verify the robustness and generalization of the model.
[0160] Table 4 shows that the DRONet model also achieved higher accuracy on the CARPK dataset. Compared with the basic model RT-DETR, the mAP50 increased by 0.7%. Compared with other models in the table, it has higher accuracy, fewer parameters and lower complexity. And the model maintains an FPS of 60, still meeting the real-time detection standard.
[0161] Table 4 Comparative experiments of different algorithms on the CARPK dataset
[0162]
[0163] 5) Ablation experiment
[0164] To verify the effectiveness of each improvement strategy proposed in this embodiment, an ablation experiment was conducted on the baseline model using the VisDrone 2019 dataset. The experimental results are shown in Table 5. "√" indicates that this improved strategy was used, and "-" indicates that this improved strategy was not used.
[0165] Table 5 Detection results after introducing different improvement strategies
[0166]
[0167] The experimental results in Table 5 show that when applied to the baseline model, each improvement strategy has improved the detection performance to varying degrees. The Occlusion-Aware KAGN Block (OKA) module, the Perceptual Spatial Integration (PSI) module, and the Scalable Dilated Efficient Aggregation (SDEA) module are added in sequence. By introducing the multi-adaptive kernel normalization convolution (KAGN) into ResNet to integrate the Occlusion-Aware KAGN Block (OKA) module, the feature extraction ability of the model is improved, and the mAP50 is increased by 1.7%. Due to the significant increase in the number of parameters, the FPS is reduced to 45 frames per second, and the computational volume does not increase but decreases by 5.4G. When only using Perceptual Spatial Integration, the mAP50 is increased by 1.2%, and the computational volume is decreased by 2.7G. When only using Scalable Dilated Efficient Aggregation, the mAP50 is increased by 1.1%, and the computational volume is decreased by 7.4G. The results show that the Perceptual Spatial Integration and Scalable Dilated Efficient Aggregation modules maintain a similar number of parameters to the baseline model while increasing the mAP50 and reducing the computational volume. The reduction in FPS can be ignored in the face of other positive benefits.
[0168] 6) Comprehensive analysis
[0169] A significant feature of deep learning models is their poor interpretability, which poses challenges in understanding the model's decision-making process. This lack of transparency not only reduces the credibility of the model but also limits its application in critical fields. To intuitively and conveniently draw the detection effect diagram of the model, comparative experimental analysis of the model's detection performance is carried out from aspects such as the confusion matrix, the model inference results, and the heat Figure 3 map. Finally, to verify the practicality of this method, an inference experiment is carried out using the image data collected by the self-developed drone in this embodiment.
[0170] To intuitively show the ability of the method proposed in this paper to predict the target category. The confusion matrix and the normalized confusion matrix of DRONet and RTDETR-r18 are drawn, as Figure 12 shown in Figure 13 and Figure 12 . Among them, (a) in Figure 13Among them, (a) is the confusion matrix diagram of DRONet, and (b) is the normalized confusion matrix diagram of DRONet. The rows of the confusion matrix represent the true classes, and the columns represent the predicted classes. The values in the diagonal region indicate the number of correct predictions, while the values in other regions indicate the number of incorrect predictions. The normalized confusion matrix converts these values into proportions, so that each value represents the proportion of that class in the total predictions, facilitating the comparison of the performance of different classes.
[0171] From Figure 12 and Figure 13 it can be seen that the color of the diagonal region of the confusion matrix in DRONet is darker than that in RTDETR-r18, indicating that the ability of the model in this paper to correctly predict object classes has been enhanced. The number of bicycle-like vehicles is small and they usually exist in a dense and occluded form. The increase in the percentage of correct predictions of the model in this paper for bicycles indicates that the model has a good recognition effect on occluded targets. For small targets such as "bicycles", "tricycles", and "awning tricycles", the proportion of objects judged to be higher in RTDETR-r18 means that most of these classes are missed during the detection process. The model in this paper reduces the missed detection rate of these classes and effectively recognizes small targets.
[0172] To intuitively verify the detection effect of this method, inference experiments were carried out using RTDETR-r18 and DRONet. Pictures containing a large number of multi-scale and occluded targets in the VisDrone dataset and the CARPK dataset were selected as experimental data, which are suitable for inference experiments. The detection results are as shown in Figure 14a , Figure 14b , Figure 14c . Among them, Figure 14a is the original input image, Figure 14b is the inference result of RTDETR-r18, Figure 14c is the inference result of DRONet;
[0173] The results show that compared with RTDETR-r18 and DRONet, Figure 14b shows that RTDETR-r18 fails to detect multiple targets and has misdetections, while Figure 14c indicates that the DRONet model in this embodiment successfully detects all targets. The RTDETR-r18 model will cause missed detection and misdetection problems when facing small targets and occluded targets. The DRONet algorithm in this embodiment enhances the detection of multi-scale targets and occluded targets, greatly reducing the problems of the basic model. The DRONet proposed in this embodiment improves the missed detection rate of occluded and dense objects and reduces the misdetection rate, effectively improving the detection performance.
[0174] Generate heatmaps of RTDETR-r18 and DRONet using Grad-CAM++ to visually reflect the regions of the feature maps that the model focuses on. Pixels with lower gradients are represented by darker blue shadows. Conversely, pixels with higher gradients in the feature map are represented by darker red shadows in the heatmap. The experimental results are as Figure 15 shown. Among them, Figure 15 in (a) is the original input image, (b) is the heatmap result of RTDETR-r18, and (c) is the heatmap result of DRONet. It can be seen from Figure 15 that RTDETR-r18 ignores some small targets and occluded targets. In contrast, the DRONet model proposed in this embodiment can more accurately focus on the boundaries and contours of the targets, has a good suppression effect on background noise, and generates heatmaps that overlap and are more relevant to the target area. Through fine feature matching, the predicted bounding boxes of the model are more accurate, which not only enhances the scale invariance of the model but also improves the localization accuracy of the features, making the focus points closer to the true areas of the targets. These results intuitively demonstrate that the DRONet model has greater adaptability for small target tasks and occluded target tasks in complex drone aerial images, showing the advancement of the model.
[0175] In summary, it can be seen that a small target detection algorithm for DRONet proposed in the present invention is specifically aimed at the problems of dense target distribution and occlusion in drone aerial photography scenes. By introducing an occlusion-aware KAGN module (OKA), a perceptual space integration module (PSI), and a scalable dilated efficient aggregation module (SDEA), the detection performance of the model in complex scenarios is effectively improved. The experimental results verify that the DRONet algorithm proposed in the present invention has high detection accuracy for small targets under occlusion conditions. While maintaining high real-time performance, it can run on a drone platform at a speed of 60 frames per second, meeting the requirements of practical applications. Future work will also expand the application of DRONet in other drone tasks, such as target tracking and semantic segmentation, and explore its target detection capabilities in extreme environments such as night, rain, and fog.
[0176] Embodiment 3
[0177] In this embodiment, a drone target detection system based on a dense occluded network is proposed, including:
[0178] Image data acquisition and processing module: Collect aerial image data through a drone, resize the image, and then preprocess the image;
[0179] Basic feature extraction module: Based on the preprocessed image, use the OKA module to extract basic features;
[0180] Feature integration module: Based on the basic features, use the PSI module to perform feature integration;
[0181] Feature aggregation module: Based on the integrated features, use the SDEA module to perform feature aggregation;
[0182] Feature fusion module: Use the PDF algorithm to fuse the features integrated by the PSI module and the features aggregated by the SDEA module, fuse the outputs of the two-way features by element-wise addition, and finally use the RT-DETR decoder for object detection to output the detection results;
[0183] And, result optimization module: Post-process and optimize the obtained detection results.
[0184] The method of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and to be stored in a local recording medium and downloaded through a network, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0185] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A method for detecting drone targets based on dense occluded networks, characterized in that: The following steps are involved: Collect aerial image data through drones, adjust the image size, and then pre-process the image; Based on the preprocessed image, the OKA module is used to extract basic features; Based on the basic features, the PSI module is used to perform feature integration; Based on the integrated features, the SDEA module is used to perform feature aggregation; The PDF algorithm is used to fuse the features integrated by the PSI module with the features aggregated by the SDEA module. The two feature outputs are fused by element-by-element addition. Finally, the RT-DETR decoder is used to perform target detection and output the detection results. The obtained detection results are post-processed and optimized.
2. The method for detecting drone targets based on dense occluded networks according to claim 1 is characterized in that: The OKA module combines KAGN adaptive convolution and GRAM polynomial to expand the feature space. It first normalizes the input features to the interval [-1,1], then expands the feature representation through the Legendre polynomial basis, optimizes feature extraction through the KAN network, and uses learnable activation functions and spline function parameterization to provide flexibility in feature extraction.
3. The method for detecting drone targets based on dense occluded networks according to claim 2, characterized in that: The process of expanding the feature representation through the Legendre polynomial basis is: 1) Calculate the Legendre polynomial basis Where N is the highest order of the polynomial, x ′ is the independent variable of the polynomial, n is the order of the current polynomial; 2) Combine the Legendre polynomial basis with the learned polynomial weights Combined, we get the expanded feature y:
4. The method for detecting drone targets based on dense occluded networks according to claim 1, characterized in that: The PSI module first spatially aligns multi-scale feature maps, unifies features of different scales through adaptive average pooling, identity mapping and bilinear interpolation; then applies partial convolution for feature selection, uses a mask matrix to control convolution calculation, and reorganizes feature channels through channel shuffling technology.
5. The method for detecting drone targets based on dense occluded networks according to claim 1, characterized in that: The SDEA module uses convolution kernels with different expansion rates to extract multi-scale features, introduces the SimAM attention mechanism to dynamically adjust feature weights, and uses structural reparameterization technology to optimize reasoning efficiency.
6. The method for detecting drone targets based on dense occluded networks according to claim 1, characterized in that: The post-processing optimization process includes: Confidence filtering of detection results to select high-confidence targets; Output the final target category and location coordinate information, and generate a visualization image of the detection results.
7. A UAV target detection system based on dense occluded networks, characterized in that: include: Image data acquisition and processing module: collect aerial image data through drones, adjust the image size, and then pre-process the image; Basic feature extraction module: Based on the preprocessed image, the OKA module is used to extract basic features; Feature integration module: Based on the basic features, the PSI module is used to perform feature integration; Feature aggregation module: Based on the integrated features, the SDEA module is used to perform feature aggregation; Feature fusion module: The PDF algorithm is used to fuse the features integrated by the PSI module with the features aggregated by the SDEA module, and the two feature outputs are fused by element-by-element addition. Finally, the RT-DETR decoder is used to perform target detection and output the detection results. And, the result optimization module: performs post-processing optimization on the obtained detection results.
8. A computer storage medium storing a readable program, characterized in that: When the program is running, it can execute the drone target detection method based on a dense occluded network as described in any one of claims 1-6.
9. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the method for detecting unmanned aerial vehicle targets based on a dense occluded network as described in any one of claims 1-6.
10. A DRONet model for drone target detection, characterized in that: include: Backbone network, encoder and RT-DETR decoder; The backbone network includes the OKA module for feature extraction. The OKA module combines KAGN adaptive convolution and GRAM polynomial to expand the feature space. It first normalizes the input features to the interval [-1, 1], then expands the feature representation through the Legendre polynomial basis, and optimizes feature extraction through the KAN network. It also uses learnable activation functions and spline function parameterization to provide flexibility in feature extraction. The encoder includes a PSI module and a SDEA module for feature integration and aggregation respectively; the PSI module first spatially aligns multi-scale feature maps, unifies features of different scales through adaptive average pooling, identity mapping and bilinear interpolation; then applies partial convolution for feature selection, uses a mask matrix to control convolution calculation, and reorganizes feature channels through channel shuffling technology; the SDEA module uses convolution kernels with different expansion rates to extract multi-scale features, introduces the SimAM attention mechanism to dynamically adjust feature weights, and uses structural reparameterization technology to optimize inference efficiency; A PDF algorithm is proposed by combining the PSI module and the SDEA module. The integrated features of the PSI module and the aggregated features of the SDEA module are fused through the PDF algorithm, and the RT-DETR decoder is used for target detection to output the detection results.
Citation Information
Cited By
Roadside radar and camera fused three-dimensional target detection method and device based on nonlinear feature extraction, and medium
CN121074864A