Unmanned aerial vehicle micro-target feature decoupling detection method, system, device and medium

CN122821418APending Publication Date: 2026-09-25EAST CHINA JIAOTONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611281921.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明的首要目的在于解决现有无人机工业巡检目标检测网络中,因深度网络极深层下采样导致的极小目标特征严重丢失、静态固定感受野无法适应尺度剧变与背景遮挡、以及现有交并比与回归损失函数在密集微小目标场景下导致的梯度冲突与模型灾难性遗忘的技术问题

Benefits of technology

(1)降低边缘端部署的算力成本。本发明提出的截断式特征金字塔将网络中计算密集的极深层物理剥离。完美抵消了引入高分辨率浅层(P2检测头)所带来的额外算力开销,使得模型在保持对微小目标极高敏感度的同时,参数量实现大幅度压缩。这有效缓解了边缘设备(如无人机机载算力盒)显存受限与内存带宽不足的问题,大幅提升了模型的单帧推理速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821418A_ABST
    Figure CN122821418A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of computer vision and unmanned aerial vehicle intelligent inspection, and discloses a feature decoupling detection method, system, equipment and medium for a micro target of an unmanned aerial vehicle, the method acquires a multi-scale aerial photograph image and performs preprocessing; is input to a backbone network of YOLO26n to perform forward reasoning, intercepts high-resolution shallow physical features and discards extremely deep features, and constructs a truncated feature pyramid; micro-macro decoupling reconstruction is performed on the feature pyramid, the shallow features are introduced into a dynamic receptive field module to perform channel competition to extract micro features, and the deep features are introduced into a large kernel attention module to perform convolution decoupling reconstruction of macro global topology; a prediction bounding box tensor is extracted, is calculated and substituted into an adaptive normalized Wasserstein distance loss function to generate a regression error and perform backward propagation smooth optimization, and is deployed to an edge device. The application breaks through the detection bottleneck of a micro target, reduces the algorithm power consumption, and completely solves the gradient conflict and feature forgetting problem in a dense micro target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and intelligent inspection technology for unmanned aerial vehicles (UAVs), and particularly to a method, system, device, and medium for feature decoupling detection of small targets on UAVs. Background Technology

[0002] With the deep integration of UAV (Unmanned Aerial Vehicle) technology and computer vision, using UAVs equipped with high-definition gimbals for industrial-grade inspections (such as railways, power lines, and bridges) has become an important means of achieving intelligent operation and maintenance. In complex outdoor inspection scenarios, the system needs to detect and locate various defects and targets in real time from aerial images. Taking railway inspection as an example, these targets include not only elastic clips and bolts in the track system, and insulators, positioning pipes, and cotter pins in the overhead contact system, but also pedestrians, falling rocks, and engineering vehicles that have intruded into the perimeter. These application scenarios require target detection algorithms to maintain extremely high recall rates and pixel-level positioning accuracy even under complex background interference and extremely high vertical viewing angles.

[0003] Compared to conventional natural scenes (such as general open-source datasets), industrial drone inspection images face significant challenges in terms of physical distribution and visual representation. Existing mainstream general-purpose object detection networks (such as the YOLO26n object detection network) exhibit the following theoretical limitations and engineering bottlenecks when directly deployed in this scenario: First, deep spatial downsampling leads to severe loss of features for tiny targets. In industrial inspection scenarios, the physical scale of the targets to be detected varies exponentially. Large targets may occupy most of the image area, while tiny parts (such as vehicles viewed from above or missing cotter pins) often occupy only a very small number of pixels in the input image (e.g., less than 8×8 pixels). Existing YOLO26n target detection networks rely on deep feature pyramid architectures (such as layers P3-P5, where layer P5 is extremely deep). After continuous high-rate spatial compression (downsampling by 32 times), the high-frequency geometric contours of tiny targets suffer severe physical spatial collapse and feature aliasing, and are then completely covered by complex background noise (such as gravel and vegetation), resulting in serious missed detections. However, directly introducing high-resolution shallow prediction heads (such as layer P2) into the existing architecture to salvage tiny targets would cause the computational overhead (FLOPs) of the network on edge devices to explode exponentially, failing to meet the real-time requirements of industrial inspection.

[0004] Second, static, single receptive fields struggle to adapt to multi-scale variations and complex occlusion environments. Existing feature extraction modules often employ conventional convolutional kernels of fixed sizes (e.g., 3×3). In actual flight inspections, the relative distance between the UAV and the target changes dynamically in real time, and the background environment is extremely complex. When facing extremely small targets, static receptive field mechanisms struggle to adaptively shrink to filter redundant noise from the surrounding environment (lacking microscopic focusing capability); while when facing large-area occlusion of lesions or large targets, static small convolutional kernels cannot effectively expand to capture the topological connectivity of the global context (lacking macroscopic perception capability). This static mechanism leads to mutual interference between microscopic and macroscopic features within the same network layer, resulting in severe feature contamination.

[0005] Third, the global static penalty mechanism of the distance metric function is prone to causing gradient calculation imbalance. Bounding box regression in object detection relies on the loss function to provide optimized gradients. When facing extremely crowded scenes from the perspective of drones (such as densely packed vehicles or components), traditional models have fatal flaws: on the one hand, the traditional Intersection over Union (CIoU) loss imposes excessively strict penalties on the boundaries of small targets (overly strict edge detection), which can lead to difficulty in model convergence when the target is blurred; on the other hand, although the Normalized Wasserstein Distance (NWD) introduced in recent years has improved the fuzziness tolerance through Gaussian distribution, its core formula uses a globally fixed static scale factor. When dealing with dense small targets at high resolution levels, an excessively large fixed Gaussian kernel can cause the receptive field boundaries of adjacent targets to become overly smooth, resulting in severe "predicted box sticking". More seriously, when static NWD and CIoU are coupled at a fixed ratio, they will produce severe gradient conflicts at the bottom layer when faced with extremely small and dense prediction boxes. This chaotic error backpropagation will destroy the pre-trained weights of the backbone network, causing "catastrophic forgetting", which will cause the generalization accuracy of the model to decrease instead of increase after long-term training. Summary of the Invention

[0006] The primary objective of this invention is to address the technical problems in existing UAV industrial inspection target detection networks, such as severe loss of features of extremely small targets due to downsampling at extremely deep layers of deep networks, inability of static fixed receptive fields to adapt to drastic scale changes and background occlusion, and gradient conflicts and catastrophic model forgetting caused by existing crossover ratio and regression loss functions in dense small target scenarios.

[0007] To achieve the above objectives, in a first aspect, the present invention discloses a feature decoupling detection method for small targets on unmanned aerial vehicles (UAVs), comprising the following steps: Acquire multi-scale aerial images and perform specialization enhancement and preprocessing to generate multi-dimensional input image tensors; The multidimensional input image tensor is fed into the backbone network of the YOLO26n target detection network for forward inference, and the high-resolution shallow features of the target are truncated and the extremely deep feature extraction branch is discarded to construct a truncated high-resolution feature pyramid set. The decoupling and reconstruction operation of micro- and macro-physical properties is performed on the truncated high-resolution feature pyramid set. Specifically, the decoupling and reconstruction operation includes importing the shallow feature layers in the truncated high-resolution feature pyramid set into a dynamic receptive field module, performing parallel regular convolution and dilated convolution, and dynamically suppressing background noise through channel-level probability allocation to output micro-local adaptive features; and importing the deep feature layers in the truncated high-resolution feature pyramid set into a large kernel attention module, generating a spatial attention mask through a cascade mechanism of depthwise convolution and depthwise dilated separable convolution, and using residual multiplication to enhance the global topological macro-contextual information of the image. The predicted bounding boxes output by the prediction head in the YOLO26n object detection network are extracted, and combined with the physical width and physical height of the real label data, they are substituted into the penalty denominator in the dynamic adaptation for constraint, the regression loss is calculated, and the network parameters of the YOLO26n object detection network are optimized by backpropagation iterative smoothing.

[0008] As an optional implementation of the first aspect of this application, a truncated high-resolution feature pyramid set is constructed by truncating shallow features of the target high resolution and discarding extremely deep feature extraction branches. Specifically, this includes: discarding extremely deep feature extraction branches with a downsampling rate of 32x and moving the deep boundary of the detection head forward; extracting feature layers in the backbone network that have downsampling rate parameters of 4x, 8x, and 16x downsampling, and using the output feature layer with a downsampling rate parameter of 4x downsampling as the shallowest feature layer; and integrating the shallowest feature layer and the output feature layers corresponding to the remaining downsampling rates to generate the truncated high-resolution feature pyramid set.

[0009] As an optional implementation of the first aspect of this application, the shallow feature layers in the truncated high-resolution feature pyramid set are imported into a dynamic receptive field module, and parallel regular convolution and dilated convolution are performed. Background noise is dynamically suppressed through channel-level probability allocation to output microscopic local adaptive features. Specifically, this includes: using regular standard convolution kernels and dilated convolution kernels with a pre-set dilation rate in parallel in the dynamic receptive field module to extract features from the shallow feature layers, respectively obtaining a first feature tensor with local details and a second feature tensor with a global receptive field; fusing the first feature tensor and the second feature tensor by adding their principal components to obtain a fused feature map; applying global average pooling and a multilayer perceptron to the fused feature map to generate a channel-level compact descriptor vector; calculating the dynamic selection weight scalar of the two features through a fully connected layer, and introducing an exponential function mechanism to ensure competition and mutual exclusion; and performing soft selection weighted fusion on the first feature tensor and the second feature tensor according to the calculated dynamic selection weight scalar to generate the microscopic local adaptive features.

[0010] As an optional implementation of the first aspect of this application, the deep feature layers in the truncated high-resolution feature pyramid set are imported into a large kernel attention module. A spatial attention mask is generated through a cascade mechanism of depthwise convolution and dilated separable convolution. The global topological macro-contextual information of the image is enhanced by residual multiplication. Specifically, this includes: extracting local spatial connectivity information from the input deep feature layers using a depthwise separable convolution operator to generate intermediate layer feature tensors; processing the intermediate layer feature tensors using a depthwise dilated separable convolution to expand the physical receptive field and capture long-distance contextual topological information; performing cross-channel information interaction on the captured long-distance contextual topological information using pointwise convolution and mapping it to the spatial attention mask matrix with values ​​in the range [0,1] through an activation function; performing element-wise multiplication of the original deep feature layers with the spatial attention mask matrix, and adding the multiplication result with the original deep feature layers by residual addition to complete the reconstruction of the global topological macro-contextual information of the image.

[0011] As an optional implementation of the first aspect of this application, the predicted bounding boxes output by the prediction head in the YOLO26n object detection network are extracted, and combined with the physical width and physical height of the real label data, they are substituted into the penalty denominator in dynamic adaptation for constraint to generate a regression loss. Specifically, this includes: modeling the predicted center point coordinates, predicted width, and predicted height of the predicted bounding box tensor output by the network as a first two-dimensional Gaussian distribution; modeling the real center point coordinates, real width, and real height of the corresponding real label bounding box as a second two-dimensional Gaussian distribution; and calculating the squared second-order Wasserstein distance between the first two-dimensional Gaussian distribution and the second two-dimensional Gaussian distribution. ;in, This represents the squared second-order Wasserstein distance between two-dimensional Gaussian distributions. The coordinates of the predicted center point of the bounding box. The coordinates of the true center point of the box. The predicted width and predicted height of the bounding box. The true width and true height of the bounding box are given; a dynamic denominator is constructed using the true width and true height, and the second-order Wasserstein squared distance value is combined with the dynamic denominator. The distance is mapped to a normalized metric parameter through the natural exponential function, and a regression loss value is generated based on the normalized metric parameter.

[0012] As an optional implementation of the first aspect of this application, a dynamic denominator is constructed using the actual width and the actual height, and the specific calculation logic is as follows: ;in, The denominator is dynamic; It is an adjustable scale hyperparameter used to control the sensitivity of the penalty; The actual width and actual height of the frame; Here, we take the underflow protection constant. .

[0013] As an optional implementation of the first aspect of this application, during the backpropagation iterative smoothing optimization of the network parameters of the YOLO26n target detection network, a domain-adaptive Gaussian kernel shrinkage and master-slave decoupling constraint mechanism are also executed: when the network features are backpropagated to the shallow high-resolution prediction branch processing dense small targets, the adjustable scale hyperparameter is adaptively shrunk to half or less than half of the initially set threshold to sharpen the Gaussian kernel shape; an asymmetric hybrid decoupling ratio is configured between the intersection-union ratio loss function and the scale-adaptive normalized distance loss function, wherein the weight assigned to the intersection-union ratio loss function is... The weights assigned to the scale-adaptive normalized distance loss function are: .

[0014] Secondly, this application discloses a feature decoupling detection system for small targets on unmanned aerial vehicles, comprising: The data acquisition and preprocessing module is used to acquire multi-scale aerial images and perform specialization enhancement and preprocessing to convert and generate multi-dimensional input image tensors. The truncated feature extraction module is used to feed the multidimensional input image tensor into the backbone network of the YOLO26n target detection network for forward inference, truncate the target's high-resolution shallow features and discard the extremely deep feature extraction branches, and construct a truncated high-resolution feature pyramid set. The micro-macro feature decoupling and reconstruction module is used to perform micro-macro physical property decoupling and reconstruction operations on the truncated high-resolution feature pyramid set. The decoupling and reconstruction operations specifically include importing the shallow feature layers in the truncated high-resolution feature pyramid set into the dynamic receptive field module, performing parallel regular convolution and dilated convolution, and dynamically suppressing background noise through channel-level probability allocation to output micro-local adaptive features; and importing the deep feature layers in the truncated high-resolution feature pyramid set into the large kernel attention module, generating a spatial attention mask through a cascade mechanism of depthwise convolution and depthwise dilated separable convolution, and using residual multiplication to enhance the global topological macro-contextual information of the image. The adaptive bounding box regression module is used to extract the predicted bounding boxes output by the prediction head in the YOLO26n object detection network, combine them with the physical width and physical height of the real label data, substitute them into the penalty denominator in the dynamic adaptation for constraint, calculate the regression loss, and perform backpropagation iterative smoothing optimization on the network parameters of the YOLO26n object detection network.

[0015] Thirdly, this application discloses an electronic device, including: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method as described in the first aspect.

[0016] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Reduce the computational cost of edge deployment. The truncated feature pyramid proposed in this invention physically strips away the computationally intensive deep layers in the network. This perfectly offsets the additional computational overhead brought about by introducing high-resolution shallow layers (P2 detector head), enabling the model to maintain extremely high sensitivity to small targets while significantly compressing the number of parameters. This effectively alleviates the problem of limited video memory and insufficient memory bandwidth of edge devices (such as UAV onboard computing boxes), and greatly improves the single-frame inference speed of the model.

[0018] (2) Overcoming the bottleneck in detecting extremely small and complex targets. Thanks to the extremely high shallow spatial mapping resolution retained by the newly added P2 high-resolution detection head, and the dynamic suppression capability of background redundancy noise by the soft competition mechanism in the Dynamic Receptive Field (DRF) module, this method exhibits extremely excellent feature representation capabilities when processing targets with extremely small shapes (such as those below 8×8 pixels). It not only significantly improves the recall rate of small defects and dense targets, but also effectively reduces the false negative and false positive rates.

[0019] (3) Macroscopic semantics and global topology reconstruction with low computational overhead. To address the potential loss of global perspective and contextual discontinuity caused by truncating extremely deep layers, this invention utilizes a large kernel attention (LKA) module for perfect mathematical compensation. Through the spatial decoupling theorem of large kernel convolution, a huge equivalent physical receptive field is obtained without increasing the number of additional parameters. This achieves "perfect micro-macro-decoupling" from the microscopic features of P2.

[0020] (4) Enhanced spatial generalization and robustness in multi-scale scenarios. This invention achieves an adaptive evolution upgrade of the loss function with scale and density awareness. On the one hand, through the adaptive Gaussian kernel shrinkage strategy, the scale adjustment factor is dynamically shrunk when detecting extremely small and dense targets, which accurately avoids receptive field overflow and target boundary adhesion; on the other hand, through asymmetric hybrid constraint decoupling, the problem of internal gradient conflict that is easily caused by high-resolution detection heads in dense scenarios is completely overcome, and the problem of "catastrophic feature forgetting" and feature pollution that is prone to occur in conventional models when extreme viewpoint migration is eliminated. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0022] Figure 1 This is a flowchart of a feature decoupling detection method for small targets on a UAV, as proposed in an embodiment of the present invention. Figure 2This is a diagram of the overall truncated architecture of the YOLO26n target detection network proposed in this embodiment of the invention; Figure 3 This is a flow chart of the multi-scale convolution and adaptive soft competitive feature selection mechanism of the DRF module in this embodiment of the invention; Figure 4 This is a schematic diagram of the large kernel convolution decomposition, spatial interaction, and residual attention reconstruction of the LKA module in this embodiment of the invention; Figure 5 This is a comparison diagram of the calculation mechanism of the SA-NWD loss function and the adaptive evolution of multi-scale target gradient in the embodiments of the present invention; Figure 6 This is a schematic diagram of the structure of a feature decoupling detection system for small targets on a UAV according to an embodiment of the present invention; Figure 7 This is a structural diagram of an electronic device disclosed in this invention. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0025] Example 1 like Figure 1 As shown, the present invention provides a feature decoupling detection method for small targets on unmanned aerial vehicles, comprising: S11: Acquire multi-scale aerial images and perform specialization enhancement and preprocessing to generate multi-dimensional input image tensors.

[0026] Specifically, a multi-scale aerial image dataset from the inspection task is acquired, and data augmentation techniques (such as copy-paste random stitching and occlusion erasure) are used to enhance extremely small defect samples in the images to balance the physical scale distribution of the samples and increase the prior density of features. Subsequently, the images are preprocessed to a fixed resolution and converted into multi-dimensional tensors for input into the YOLO26n target detection network.

[0027] S12: The multidimensional input image tensor is fed into the backbone network of the YOLO26n target detection network for forward inference, and the high-resolution shallow features of the target are truncated and the extremely deep feature extraction branches are discarded to construct a truncated high-resolution feature pyramid set.

[0028] like Figure 2 As shown, the image tensor is fed into the backbone network for forward inference. When the network reaches the P2 throat layer with a downsampling rate of 4x, the high-resolution shallow feature tensor of this layer is truncated to preserve the fine-grained geometric edge information of extremely small targets (such as those smaller than 8×8 pixels). At the same time, the invalid and redundant convolution calculations of extremely deep layers (more than 16x downsampling, such as the P5 layer) are terminated, and the extracted P2, P3, and P4 feature results are used to construct a truncated high-resolution feature pyramid set.

[0029] Specifically, this invention performs structural-level truncation and reconstruction on the backbone of the YOLO26n object detection network. By discarding the P5 ultra-deep feature extraction branch with a downsampling rate of 32x, the deep boundary of the detection head is moved forward, constructing a hierarchical set focused on preserving high-resolution features, thus improving the downsampling rate parameter. .

[0030] Define the multidimensional input image tensor as ,in Represents multidimensional input image data; 3 represents the set of real numbers; 3 represents the number of channels in the input image (usually RGB three channels); H and W represent the physical height and width of the input image, respectively.

[0031] The first in the backbone network The layer feature output is defined as: in Indicates the first The output feature map of the layer; This represents the output feature map of the previous layer; Identify the composite convolution operator (which typically includes a two-dimensional convolutional layer, a batch normalization layer, and a non-linear activation function). This indicates that the spatial stride of the convolution operation is 2, which means that the spatial downsampling is achieved by 2 times.

[0032] The target high-resolution shallow feature parameters of the shallowest feature layer (P2 layer) retained in this invention are defined as follows: in, and These represent the height and width of the feature map in layer P2, respectively. This layer's features are downsampled only twice, which effectively preserves the high-frequency edge texture and spatial geometric features of small targets.

[0033] Finally, the set of multi-scale output layers of the Feature Pyramid Network (PANet) is defined as follows: in, It is a truncated high-resolution feature pyramid set, used to feed into the subsequent detection head; Represents the channel dimension of the feature map; Downsampling rates corresponding to different feature levels.

[0034] This truncation operation effectively removes redundant channel calculations deep within the model, significantly reducing the number of model parameters and computational complexity, and ensuring from the bottom layer of the network architecture that the physical mapping of small targets is not over-compressed.

[0035] S13: Perform a decoupling and reconstruction operation on the micro-macro physical properties of the truncated high-resolution feature pyramid set. The decoupling and reconstruction operation specifically includes importing the shallow feature layers in the truncated high-resolution feature pyramid set into the dynamic receptive field module, performing parallel regular convolution and dilated convolution, and dynamically suppressing background noise through channel-level probability allocation to output micro-local adaptive features.

[0036] The feature pyramid undergoes decoupling and reconstruction of micro- and macro-physical properties: high-resolution shallow features from layer P2 are imported into the DRF (Dynamic Receptive Field) module, where parallel regular and dilated convolutions are performed. Channel-level probability allocation is achieved through global average pooling and the Softmax formula, dynamically suppressing complex background noise and outputting adaptive features focused on microscopic local targets. Simultaneously, deep features from layers P3 and P4 are imported into the LKA (Large Kernel Attention) module. A spatial attention mask is generated through a cascaded mechanism of depthwise convolutions and deep dilated separable convolutions, and residual multiplication is used to enhance the global topological and macro-contextual information of the image.

[0037] Specifically, in the shallow P2 branch responsible for small target detection, this invention introduces a Dynamic Receptive Field (DRF) module, such as... Figure 3As shown, this module comprises three stages: "multi-path parallel feature extraction," "cross-scale compressed excitation," and "Softmax channel competition selection."

[0038] Assume the feature tensor of this module is ( and (The spatial dimensions of the current feature map).

[0039] Step 1: Multi-path parallel feature extraction The modules use regular convolution and dilated convolution in parallel to extract features under different receptive fields: in Indicates through regular The first feature tensor with high-frequency local details is extracted by convolution; Indicates the void ratio of Second feature tensors with large receptive fields (resistant to occlusion) are extracted by dilated convolution; and These represent standard convolution kernel and dilated convolution kernel operations, respectively. Represents a nonlinear activation function (such as ReLU or SiLU); This represents batch normalization, used to stabilize feature distribution.

[0040] Step 2: Transscale compression excitation The two features are merged by adding the primary element together. Subsequently, global average pooling (GAP) and a multilayer perceptron are used to generate channel-level compact descriptors. : in, Represents the fused feature map In spatial coordinates Pixel value at; This represents a global average pooling operation performed in the spatial dimension, compressing the feature map into channel statistics. Represents a dimension-reduced fully connected layer (or The weight matrix of the convolution is used to learn the correlation between channels; This represents a channel-level compact descriptor vector.

[0041] Step 3: Softmax Channel Contention Calculation Channel-level compact descriptor The dynamic selection weights of the two features are calculated through a fully connected layer, and a Softmax function is introduced to ensure contention and mutual exclusion. in: and They respectively represent the features and The learnable weight matrix; and The calculated scalar weights, and satisfying This represents the network's attentional tendency towards local and large receptive fields at the channel level. This represents the natural exponential function.

[0042] Step 4: Feature Soft Selection Fusion in, This is the final output microscopic local adaptive feature. This mechanism makes the network inclined to assign microscopic local adaptive features when detecting unobstructed, independent, small targets. Higher weight is assigned when detecting occluded targets. Higher weight.

[0043] S14: The deep feature layers in the truncated high-resolution feature pyramid set are imported into the large kernel attention module. A spatial attention mask is generated through the cascade mechanism of depthwise convolution and depthwise dilatational separable convolution. The residual multiplication is used to enhance the global topological macroscopic contextual information of the image.

[0044] To compensate for the loss of macroscopic contextual information caused by truncating the extremely deep P5 layer, this invention configures a large kernel attention (LKA) module in the deeper layers (P3 and P4 layers), such as... Figure 4 As shown, this module utilizes the principle of convolution decoupling to reconstruct large-size convolution kernels with high computational overhead into a lightweight sequence of depthwise separable operators.

[0045] For input features : Step 1: Extraction of local spatial connectivity Extracting local spatial information using depthwise convolution: in, Indicates the kernel size as Depth convolution operator, This is the feature tensor of the intermediate layer.

[0046] Step 2: Long-distance context topology capture Deep holes can be used to separate convolutions and expand the receptive field: in, Indicates the core size is void ratio The depthwise dilated convolution. The equivalent physical receptive field range acquired by this cascade mechanism can be calculated by the formula: , This provides long-distance contextual topology information.

[0047] Step 3: Channel Aggregation and Spatial Mask Generation in, This represents point-wise convolution, used for cross-channel information interaction; This represents the Sigmoid activation function, mapped to a spatial attention mask matrix with values ​​in the range [0,1]. .

[0048] Step 4: Residual Reconstruction in, This indicates element-wise multiplication. This represents the final output feature obtained after four steps: "local spatial connectivity extraction → long-range context capture → channel aggregation and mask generation → residual reconstruction". Through residual connections, the module preserves the input features. At the same time, using attention masks It enhances the macroscopic topological region relevant to the target.

[0049] S15: Extract the predicted bounding boxes output by the prediction head in the YOLO26n object detection network, combine them with the physical width and physical height of the real label data, substitute them into the penalty denominator in the dynamic adaptation for constraint, calculate the regression loss, and perform backpropagation iterative smoothing optimization on the network parameters of the YOLO26n object detection network.

[0050] The predicted bounding box tensors output by each prediction head of the network (especially the P2 high-resolution branch) are extracted and used in the backpropagation optimization stage. Combining the physical width and height of the real label data, the improved "scale- and crowding-aware dynamic hybrid loss function (Dynamic SA-NWD + CIoU decoupling)" is applied. When dealing with small and dense P2-level targets, the scale adjustment factor adaptively shrinks the Gaussian distribution. To prevent receptive field overflow and adhesion to the target boundary at high resolution, asymmetric hybrid constraints are implemented to increase the rigid geometric cutting weights of CIoU and reduce the auxiliary weights of NWD distribution constraints. The comprehensive regression error is calculated and an inverse gradient matrix is ​​generated; this mechanism completely avoids gradient conflicts and catastrophic feature forgetting caused by dense predictions at the underlying physical level. Finally, an optimizer (such as AdamW or SGD) is used to perform smooth iterative updates on the network parameters until the model converges stably.

[0051] Specifically, to address the problem of unbalanced regression loss gradients for multi-scale objectives, this invention proposes an adaptive penalty mechanism for the SA-NWD (Scale-Adaptive) loss function, such as... Figure 5 As shown.

[0052] Step 1: Modeling the Gaussian distribution of the bounding box The predicted bounding box output by the prediction head in the YOLO26n object detection network is defined as... ,in, The coordinates of the predicted center point of the bounding box. The predicted width and predicted height of the bounding box are defined as follows; the true label bounding box is defined as... ,in, The coordinates of the true center point of the box. This represents the actual width and actual height of the frame.

[0053] This method models the two bounding boxes as first-dimensional Gaussian distributions. With the second two-dimensional Gaussian distribution The mean vector of a Gaussian distribution With covariance matrix The parameter is defined as follows: in Here are the coordinates of the center point of the box, and T represents the transpose. Represents a diagonal matrix function; Step 2: Derivation of the second-order Wasserstein distance Derive and simplify the squared second-order Wasserstein distance between two two-dimensional Gaussian distributions: in This represents the squared second-order Wasserstein distance between two two-dimensional Gaussian distributions. This represents the square operation of the second norm (Euclidean distance). The first two terms measure the spatial offset distance of the center point, and the last two terms measure the difference in width and height dimensions.

[0054] Step 3: Construct a dynamic adaptive penalty denominator Unlike existing technologies that use static constants, this invention utilizes the geometric properties of the actual label box to construct a dynamic denominator: in, The denominator is dynamic; and The width and height of the bounding box allow this constant to be dynamically adjusted as the actual physical size of the target changes. It is an adjustable scale hyperparameter used to control the sensitivity of the penalty; For a minimal constant (e.g.) The function of this function is to provide underflow protection during denominator calculation, preventing mathematical division by zero errors when the actual bounding box width and height are extremely small or distortion occurs, thus ensuring the stability of gradient calculation.

[0055] Step 4: Nonlinear distance metric mapping Mapping distance to normalized metric parameters using an exponential function: in, It is a natural exponential function. When the predicted bounding box completely overlaps with the ground truth bounding box... It is 0 at this time. Take the maximum value of 1.

[0056] Step 5: Regression Loss Generation Generate the final loss function value used for network backpropagation optimization: The optimization algorithm minimizes the regression loss value. This drives the network to adjust the weight parameters of the feature extraction layer.

[0057] Building upon this, the present invention further provides a domain-adaptive Gaussian kernel shrinkage and master-slave decoupling constraint mechanism: Adaptive Gaussian kernel shrinkage strategy: in the constant formula of SA-NWD In the middle, make The factor possesses the ability to perceive target resolution and layer density. Specifically, when network features are backpropagated to high-resolution shallow layers (such as the P2 head) to process dense, tiny point-like targets, the system adaptively adjusts the scale factor. Shrink to half or less of the original threshold (e.g., from a general scenario). Adaptive decay to minimize the target scene This strategy makes the Gaussian kernel shape sharper, like a scalpel, precisely avoiding the overflow of receptive fields between dense targets.

[0058] Crowding-aware asymmetric hybrid decoupling: In loss function reconstruction, this invention decouples the symmetric status of CIoU and SA-NWD. When dealing with extremely high-altitude congestion from a vertical perspective, the system dynamically assigns higher decision weights to rigid geometric constraints (e.g., ), and reduce distributed fault tolerance constraints to auxiliary (such as This mechanism addresses the catastrophic feature forgetting problem in high-resolution prediction heads in dense, small-object scenarios by addressing the underlying gradient space.

[0059] Furthermore, structural optimizations (such as operator fusion and folding of convolutional and batch normalization layers) can be performed on the trained and finalized network model, and the static computation graph can be compiled into an accelerated format and burned into an onboard AI computing device.

[0060] After the massive YOLO26n object detection network completes all adaptive optimization and reaches its convergence extreme value on a high-performance cloud training cluster, this invention provides a set of deep compiler-level structural optimization and deployment processes to meet the extremely stringent real-time frame rate and low power consumption requirements of industrial environments. First, to squeeze out the last bit of runtime performance, the system initiates a reparameterization folding process. This process perfectly merges and recombines the scattered mean, variance, scaling weights, and bias parameters from all the independently operating two-dimensional convolutional operation units during training, along with those from the subsequent batch normalization (BN) layers, into a single, independently computed and completely equivalent simple fused convolutional weight matrix through linear algebra matrix operations. This operator folding action directly eliminates the extremely time-consuming feature statistical mean calibration step during the inference phase, significantly reducing the bandwidth pressure on the storage bus. Next, targeting the floating-point computing characteristics of the UAV's onboard neural acceleration unit (such as an NPU or lightweight embedded graphics card), the system smoothly compiles and transforms the original 32-bit full-precision (FP32) static computing topology, which originally occupied a large amount of video memory, into a 16-bit half-precision (FP16) or even 8-bit integer (INT8) hardware-accelerated execution sequence through an extremely rigorous quantization and calibration strategy. Finally, the lightweight yet extremely powerful model, after extreme slimming, is burned into the edge computing box carried on the belly of the multi-rotor UAV, enabling it to perform millisecond-level precise identification and location monitoring of even the smallest safety hazards in real time in the blue sky.

[0061] In summary, in this embodiment: 1. The proposed truncated feature pyramid physically strips away the computationally intensive, extremely deep layers of the network. Experimental analysis shows that this scheme perfectly offsets the additional computational overhead introduced by the high-resolution shallow layer (P2 detection head), enabling the model to maintain extremely high sensitivity to small targets while significantly compressing the number of parameters. This effectively alleviates the problems of limited GPU memory and insufficient memory bandwidth on edge devices (such as UAV-borne computing boxes), significantly improving the model's single-frame inference speed and meeting the stringent real-time requirements for defect detection in industrial settings.

[0062] 2. Benefiting from the extremely high shallow spatial mapping resolution retained by the newly added P2 high-resolution detection head, and the dynamic suppression capability of the Softmax soft competition mechanism in the DRF module for background redundancy noise (such as complex backgrounds like gravel and vegetation), this method exhibits extremely excellent feature representation capabilities when processing targets with extremely small morphologies (such as those smaller than 8×8 pixels). Compared with the benchmark model, this invention can accurately separate the target from the background, significantly improving the recall rate for small lesions and dense targets, and more effectively reducing the false negative and false positive rates, resulting in a robust and breakthrough improvement in detection accuracy evaluation metrics (such as mean accuracy, mAP).

[0063] 3. To address the potential loss of global receptive field and contextual discontinuity caused by truncating extremely deep layers, this invention utilizes the LKA module for perfect mathematical compensation. Through the spatial decoupling theorem of large kernel convolution (decomposing it into depthwise separable and dilated convolution sequences), this method achieves a huge equivalent physical receptive field with almost no additional parameters. This achieves "perfect micro-macro-decoupling" from P2 micro-features, enabling the model to perform accurate region repair and localization even when facing defective targets with drastic scale changes, large areas, or severe occlusion, relying on its powerful macro-contextual topology awareness.

[0064] 4. This invention achieves a scale- and density-aware adaptive evolution upgrade of the SA-NWD loss function. On one hand, through an adaptive Gaussian kernel shrinkage strategy, the scale adjustment factor (such as attenuation) is dynamically shrunk when detecting extremely small and dense targets. The parameters, like a "surgical scalpel," precisely avoid receptive field overflow and target boundary adhesion. On the other hand, through asymmetric hybrid constraint decoupling, the decision weights of rigid geometric constraints (CIoU) and distributed fault-tolerant constraints (NWD) are dynamically adjusted. This mechanism completely overcomes the problem of internal gradient conflicts that are easily caused by high-resolution detectors in dense scenes, enabling the model to have excellent geometric alignment robustness under long distances and drastic scale changes.

[0065] 5. The micro-macro decoupling architecture and dynamic adaptive penalty mechanism constructed in this invention not only significantly resolve the feature extraction conflict caused by extreme imbalance of multi-scale samples, but also fundamentally solve the problems of "catastrophic forgetting" and feature contamination that are prone to occur in conventional models when transferring extreme viewpoints from the underlying gradient space. This allows the model to reach convergence smoothly and extremely quickly during the optimization phase. Because its underlying mathematical reconstruction does not rely on prior knowledge of a specific viewpoint, the system has extremely strong cross-domain transfer capabilities. Without large-scale modifications to the network topology, it can be seamlessly extended from conventional UAV tilted aerial photography to many industrial-grade computer vision tasks with the pain points of "multi-scale and small dense targets," such as extremely high-altitude vertical top-down views, railway inspection, power grid inspection, and autonomous driving perception.

[0066] Example 2 like Figure 6 As shown, the present invention also discloses a feature decoupling detection system for small targets on unmanned aerial vehicles (UAVs), which includes the following key modules: The data acquisition and preprocessing module 11 is used to acquire multi-scale aerial images and perform specialization enhancement and preprocessing to convert and generate multi-dimensional input image tensors; The truncated feature extraction module 12 is used to send the multidimensional input image tensor into the backbone network of the YOLO26n target detection network for forward inference, truncate the shallow high-resolution features of the target and discard the extremely deep feature extraction branches, and construct a truncated high-resolution feature pyramid set. The micro-macro feature decoupling and reconstruction module 13 is used to perform micro-macro physical property decoupling and reconstruction operations on the truncated high-resolution feature pyramid set. The decoupling and reconstruction operations specifically include importing the shallow feature layers in the truncated high-resolution feature pyramid set into the dynamic receptive field module, performing parallel regular convolution and dilated convolution, and dynamically suppressing background noise through channel-level probability allocation to output micro-local adaptive features; and importing the deep feature layers in the truncated high-resolution feature pyramid set into the large kernel attention module, generating a spatial attention mask through a cascade mechanism of depthwise convolution and depthwise dilated separable convolution, and using residual multiplication to enhance the global topological macro-contextual information of the image. The adaptive bounding box regression module 14 is used to extract the predicted bounding boxes output by the prediction head in the YOLO26n object detection network, combine them with the physical width and physical height of the real label data, substitute them into the penalty denominator in the dynamic adaptation for constraint, calculate the regression loss, and perform backpropagation iterative smoothing optimization on the network parameters of the YOLO26n object detection network.

[0067] Furthermore, embodiments of this application also disclose an electronic device, Figure 7This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0068] Figure 7 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the feature decoupling detection method for small targets of a UAV disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0069] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0070] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0071] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0072] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. It can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing a feature decoupling detection method for small targets of a UAV executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.

[0073] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned method for feature decoupling detection of small targets in a UAV. The specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0074] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0075] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs (Compact Disc-Read Only Memory), or any other form of storage medium known in the art.

[0076] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0077] The solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A feature decoupling detection method for small targets on unmanned aerial vehicles (UAVs), characterized in that, Includes the following steps: Acquire multi-scale aerial images and perform specialization enhancement and preprocessing to generate multi-dimensional input image tensors; The multidimensional input image tensor is fed into the backbone network of the YOLO26n target detection network for forward inference, and the high-resolution shallow features of the target are truncated and the extremely deep feature extraction branch is discarded to construct a truncated high-resolution feature pyramid set. The decoupling and reconstruction operation of micro- and macro-physical properties is performed on the truncated high-resolution feature pyramid set. The decoupling and reconstruction operation specifically includes importing the shallow feature layers in the truncated high-resolution feature pyramid set into the dynamic receptive field module, performing parallel regular convolution and dilated convolution, and dynamically suppressing background noise through channel-level probability allocation to output micro-local adaptive features. The deep feature layers in the truncated high-resolution feature pyramid set are imported into the large kernel attention module. A spatial attention mask is generated through a cascade mechanism of depthwise convolution and depthwise dilatational separable convolution. Residual multiplication is used to enhance the global topological macroscopic contextual information of the image. The predicted bounding boxes output by the prediction head in the YOLO26n object detection network are extracted, and combined with the physical width and physical height of the real label data, they are substituted into the penalty denominator in the dynamic adaptation for constraint, the regression loss is calculated, and the network parameters of the YOLO26n object detection network are optimized by backpropagation iterative smoothing.

2. The method according to claim 1, characterized in that, By extracting high-resolution shallow features from the target and discarding extremely deep feature extraction branches, a truncated high-resolution feature pyramid set is constructed, specifically including: The extremely deep feature extraction branch with a downsampling rate of 32 times was discarded, and the deep boundary of the detection head was moved forward; The feature layers corresponding to 4x downsampling, 8x downsampling and 16x downsampling in the backbone network are extracted, and the output feature layer corresponding to the 4x downsampling parameter is taken as the shallowest feature layer. The shallowest feature layer and the output feature layers corresponding to the remaining downsampling rates are integrated to generate the truncated high-resolution feature pyramid set.

3. The method according to claim 1, characterized in that, The shallow feature layers in the truncated high-resolution feature pyramid set are imported into the dynamic receptive field module, where parallel regular convolution and dilated convolution are performed. Background noise is dynamically suppressed through channel-level probability allocation to output microscopic local adaptive features. Specifically, this includes: In the dynamic receptive field module, conventional standard convolution kernels and dilated convolution kernels with a pre-set dilation rate are used in parallel to perform feature extraction on shallow feature layers, respectively obtaining a first feature tensor with local details and a second feature tensor with a global receptive field. The first feature tensor and the second feature tensor are fused by adding the principal elements to obtain a fused feature map. Global average pooling and multilayer perceptron are then applied to the fused feature map to generate channel-level compact descriptor vectors. Based on the channel-level compact descriptor vector, the dynamic selection weight scalar of the two features is calculated through a fully connected layer, and an exponential function mechanism is introduced to ensure contention and mutual exclusion. Based on the calculated dynamic selection weight scalar, soft selection weighted fusion is performed on the first feature tensor and the second feature tensor to generate the microscopic local adaptive features.

4. The method according to claim 1 or 3, characterized in that, The deep feature layers in the truncated high-resolution feature pyramid set are imported into the large kernel attention module. A spatial attention mask is generated through a cascade mechanism of depthwise convolution and depthwise dilated separable convolution. Residual multiplication is used to enhance the global topological macroscopic contextual information of the image, specifically including: The depthwise separable convolution operator is used to extract local spatial connectivity information from the input deep feature layer to generate intermediate layer feature tensors; The intermediate layer feature tensors are processed using deep-dilated separable convolutions to expand the physical receptive field and capture long-range contextual topology information; The captured long-distance context topology information is processed by pointwise convolution to perform cross-channel information interaction, and is mapped to the spatial attention mask matrix with values ​​in the range [0,1] by an activation function; The original deep feature layer is multiplied element-wise with the spatial attention mask matrix, and the result of the multiplication is added to the original deep feature layer as a residual to complete the reconstruction of the global topological macroscopic context information of the image.

5. The method according to claim 1, characterized in that, Extract the predicted bounding boxes output by the prediction head in the YOLO26n object detection network, combine them with the physical width and height of the real label data, and substitute them into the penalty denominator in the dynamic adaptation for constraint to generate the regression loss, specifically including: The predicted center point coordinates, predicted width, and predicted height of the predicted bounding box tensor output by the network are modeled as a first two-dimensional Gaussian distribution. The true center point coordinates, true width, and true height of the corresponding true label bounding box are modeled as a second two-dimensional Gaussian distribution; Calculate the squared second-order Wasserstein distance between the first two-dimensional Gaussian distribution and the second two-dimensional Gaussian distribution: ; in, This represents the squared second-order Wasserstein distance between the first two-dimensional Gaussian distribution and the second two-dimensional Gaussian distribution. The coordinates of the predicted center point of the bounding box. The coordinates of the true center point of the box. The predicted width and predicted height of the bounding box. The actual width and actual height of the frame; A dynamic denominator is constructed using the true width and the true height. The second-order Wasserstein squared distance value is combined with the dynamic denominator, and the distance is mapped to a normalized metric parameter through the natural exponential function. A regression loss value is generated based on the normalized metric parameter.

6. The method according to claim 5, characterized in that, A dynamic denominator is constructed using the actual width and the actual height. The specific calculation logic is as follows: ; in, The denominator is dynamic; It is an adjustable scale hyperparameter used to control the sensitivity of the penalty; The actual width and actual height of the frame; Here, we take the underflow protection constant. .

7. The method according to claim 6, characterized in that, During the backpropagation iterative smoothing optimization of the network parameters of the YOLO26n target detection network, domain adaptive Gaussian kernel shrinkage and master-slave decoupling constraint mechanisms are also performed: When the network features are backpropagated to the shallow high-resolution prediction branch that processes dense small targets, the adjustable scale hyperparameter is adaptively shrunk to half or less than half of the initially set threshold to sharpen the Gaussian kernel shape. Configure the asymmetric hybrid decoupling ratio of the intersection-union ratio (CIRR) loss function and the scale-adaptive normalized distance loss function, wherein the weight assigned to the CIRR loss function is... The weights assigned to the scale-adaptive normalized distance loss function are: .

8. A feature decoupling detection system for small targets on unmanned aerial vehicles (UAVs), characterized in that, include: The data acquisition and preprocessing module is used to acquire multi-scale aerial images and perform specialization enhancement and preprocessing to convert and generate multi-dimensional input image tensors. The truncated feature extraction module is used to feed the multidimensional input image tensor into the backbone network of the YOLO26n target detection network for forward inference, truncate the target's high-resolution shallow features and discard the extremely deep feature extraction branches, and construct a truncated high-resolution feature pyramid set. The micro-macro feature decoupling and reconstruction module is used to perform micro-macro physical property decoupling and reconstruction operations on the truncated high-resolution feature pyramid set. The decoupling and reconstruction operations specifically include importing the shallow feature layers in the truncated high-resolution feature pyramid set into the dynamic receptive field module, performing parallel regular convolution and dilated convolution, and dynamically suppressing background noise through channel-level probability allocation to output micro-local adaptive features; and importing the deep feature layers in the truncated high-resolution feature pyramid set into the large kernel attention module, generating a spatial attention mask through a cascade mechanism of depthwise convolution and depthwise dilated separable convolution, and using residual multiplication to enhance the global topological macro-contextual information of the image. The adaptive bounding box regression module is used to extract the predicted bounding boxes output by the prediction head in the YOLO26n object detection network, combine them with the physical width and physical height of the real label data, substitute them into the penalty denominator in the dynamic adaptation for constraint, calculate the regression loss, and perform backpropagation iterative smoothing optimization on the network parameters of the YOLO26n object detection network.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of a feature decoupling detection method for small targets in a UAV as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of a feature decoupling detection method for micro-targets of a UAV as described in any one of claims 1 to 7.