An aerial image extreme aspect ratio target-oriented task decoupling detection method and system
Patent Information
- Application Number
- CN202610769125.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-31
- Publication Date
- 2026-09-29
AI Technical Summary
然而,这些方法要么依赖不加区分的上下文聚合,引入大量背景干扰,要么缺乏对几何先验的显式建模,无法充分解决上下文扩展与特征纯度之间的内在矛盾
[0035](1)提出任务解耦的空间-通道协同框架(TD-SCS),通过明确地将分类和定位的特征提取解耦,使每个任务能在其自身的最佳条件下运行,定位分支保留原始提案几何形状以进行准确地边界框回归,避免特征错位问题;
Smart Images

Figure CN122841815A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of image or video recognition or understanding, and in particular to a task decoupling detection method and system for targets with extreme aspect ratios in aerial images, applicable to the fields of computer vision and image analysis. Background Technology
[0002] In recent years, aerial imagery, such as images and videos from drone aerial photography and other sources, has played a crucial role in directional target detection across a wide range of applications, including urban planning, maritime monitoring, and disaster assessment. Compared to natural scene images, aerial images possess unique characteristics such as a bird's-eye view, complex backgrounds, dense distribution, and arbitrary orientation. To address these challenges, researchers have proposed numerous deep learning-based orientation detectors.
[0003] Despite progress, detecting targets with extreme aspect ratios (EAR), such as bridges, large ships, and ports, remains a persistent challenge. The core difficulty lies in the insufficient contextual awareness of standard Regions of Interest (RoI) features. For highly elongated targets, traditional RoIAlign primarily captures internal appearance patterns, lacking rich context from the surrounding region. Consequently, EAR targets are frequently confused with structurally similar background elements, such as roads or coastlines, leading to degraded classification performance.
[0004] Expanding RoIs to incorporate additional contextual information is an intuitive solution. However, this introduces a fundamental conflict between classification and localization—classification benefits from the expanded receptive field and context invariance, while bounding box regression requires features that are closely aligned with the target boundary. Applying RoI expansion to a shared detector head inevitably compromises the boundary-sensitive features required for localization, leading to feature misalignment and decreased regression accuracy. Existing methods typically attempt to reconcile these conflicting requirements in a unified feature representation, but fail to explicitly model geometric priors, limiting their effectiveness for EAR targets.
[0005] Several works have explored the classification-localization conflict mentioned above. For example, ARS-DETR introduces an aspect ratio-sensitive mechanism into a Transformer-based detector to adaptively adjust feature perception; Double-Head separates classification and regression branches to reduce feature interference; and in the field of orientation detection, methods such as Fourier Angle Alignment (FAA) also emphasize the importance of task-specific feature alignment. However, these methods either rely on indiscriminate context aggregation, introducing a large amount of background interference, or lack explicit modeling of geometric priors, failing to fully resolve the inherent contradiction between context expansion and feature purity. Therefore, a detection framework that can simultaneously balance rich contextual information and fidelity of boundary features is urgently needed. Summary of the Invention
[0006] This invention solves the problems existing in the prior art and provides a task decoupling detection method and system for targets with extreme aspect ratios in aerial images.
[0007] The technical solution adopted in this invention is a task decoupling detection method for targets with extreme aspect ratios in aerial images, comprising the following steps:
[0008] Construct a spatial-channel collaborative target detection model with task decoupling;
[0009] Obtain a dataset of samples from distant locations and train the model;
[0010] The aerial image to be detected is input into the trained model to generate rotating candidate boxes. The target detection head is decoupled into independent localization and classification branches. The rotating candidate boxes are localized and the classification features are extracted based on context enhancement, respectively. The outputs of the localization and classification branches are fused to obtain the directional target detection results.
[0011] Preferably, the task-decoupled spatial-channel collaborative target detection model includes a backbone network, a feature pyramid network, a region generation network, and a target detection head, wherein the target detection head includes decoupled localization and classification branches;
[0012] The localization branch uses the rotation region of interest alignment module to extract features of the rotation candidate boxes, and performs bounding box regression based on the rotation candidate boxes;
[0013] The classification branch is sequentially equipped with a task-decoupled geometric perception region expansion module and a geometrically guided channel calibration module, which are used to perform geometric perception spatial expansion and feature recalibration on the rotated candidate box to obtain classification features.
[0014] Preferably, the geometric perception region expansion module obtains the aspect ratio prior information of the current rotated candidate box, calculates the height expansion factor and width expansion factor, and calculates the expanded candidate box size based on this, and adaptively expands the short side of the elongated candidate box.
[0015] Preferably, the height expansion factor satisfies The width expansion factor satisfies The aspect ratio of the current rotated candidate box. , and These represent the width and length of the rotated candidate box, respectively. For the Sigmoid function, The maximum expansion ratio, Aspect ratio symmetrical threshold, This is the transition sharpness parameter.
[0016] Preferably, after expansion, the candidate box size is .
[0017] Preferably, the geometry-guided channel calibration module acquires the extended features extracted by the geometry-aware region expansion module after spatial expansion of the rotated candidate box, and performs global average pooling on the extended features to obtain the channel descriptor.
[0018] Construct a geometric descriptor that includes information on the shape, orientation, and scale of the expanded candidate boxes;
[0019] The geometric descriptor and the channel descriptor are concatenated and input into a multilayer perceptron to generate channel weights. The channel weights are then used to recalibrate the extended features.
[0020] Preferably, the geometric descriptor satisfy,
[0021]
[0022] in, For linear projection layers, The angle is the rotation angle.
[0023] Preferably, the channel weights satisfy the following:
[0024]
[0025] in, The channel descriptor obtained by global average pooling is... Indicates splicing, For ReLU function, and These are the weight matrices for the first and second fully connected layers in a multilayer perceptron, respectively.
[0026] Preferably, the recalibration satisfies,
[0027]
[0028] in, To extend features;
[0029] Based on recalibrated features Input the classification header to predict the class probability.
[0030] A task decoupling detection system for targets with extreme aspect ratios in aerial images includes:
[0031] Memory, used to store computer programs;
[0032] A processor is used to execute the computer program to implement the task decoupling detection method for targets with extreme aspect ratios in aerial images.
[0033] This invention relates to a task-decoupled detection method and system for targets with extreme aspect ratios in aerial images. The method involves constructing a task-decoupled spatial-channel collaborative target detection model, acquiring a long-distance sample dataset, and training the model. The aerial image to be detected is input into the trained model to generate rotated candidate boxes. The target detection head is decoupled into independent localization and classification branches, which are used to locate the rotated candidate boxes and extract context-enhanced classification features. The outputs of the localization and classification branches are fused to obtain the directional target detection result. The system is implemented based on this method.
[0034] The beneficial effects of this invention are as follows:
[0035] (1) A task-decoupled spatial-channel collaborative framework (TD-SCS) is proposed. By explicitly decoupling the feature extraction of classification and localization, each task can operate under its own optimal conditions. The localization branch retains the original proposal geometry for accurate bounding box regression, avoiding feature misalignment problems.
[0036] (2) Design a task-decoupled geometric perception region extension module (TD-GREM), using aspect ratio as a geometric prior to perform anisotropic RoI extension, prioritizing the expansion of the short side of slender targets, thereby introducing richer lateral context, while avoiding including too much background along the long axis;
[0037] (3) A geometry-guided channel calibration module (GCCM) is proposed to reduce background interference introduced by spatial expansion. By incorporating geometric descriptors into a lightweight channel recalibration mechanism, feature refinement based on geometry is achieved through residual formulas, which significantly improves the detection accuracy of targets with extreme aspect ratios such as bridges and ports. Attached Figure Description
[0038] Figure 1 This is a flowchart of the method of the present invention;
[0039] Figure 2 This is a schematic diagram of the overall process of implementing the present invention;
[0040] Figure 3 This is a schematic diagram of the architecture of the space-channel cooperative target detection model for task decoupling of the present invention;
[0041] Figure 4 This is a schematic diagram of the geometric perception region extension module for task decoupling in this invention;
[0042] Figure 5 This is a schematic diagram of the geometrically guided channel calibration module in this invention. Detailed Implementation
[0043] The present invention will be further described in detail below with reference to embodiments, but the scope of protection of the present invention is not limited thereto.
[0044] This invention relates to a task decoupling detection method for targets with extreme aspect ratios in aerial images, applicable to target detection in long-distance scenarios such as UAV aerial photography and aerial photography, especially for targets with extreme aspect ratios (EAR) such as bridges, large ships, and ports.
[0045] First, it should be noted that for "extreme aspect ratio targets", when the aspect ratio of a candidate box meets the preset threshold condition, the corresponding candidate box is determined to be an extreme aspect ratio candidate box (target). The preset threshold here is a configurable parameter, preferably 3~20, and more preferably 5~15. For candidate boxes that do not reach the preset threshold, slight expansion, proportional expansion or no expansion can be performed.
[0046] The method includes the following steps:
[0047] (1) Construct a task-decoupled spatial-channel collaborative target detection model;
[0048] (2) Obtain a long-distance sample dataset and train the model;
[0049] (3) Input the aerial image to be detected into the trained model, generate rotating candidate boxes, and decouple the target detection head into independent localization and classification branches. Perform localization and context-enhanced classification feature extraction on the rotating candidate boxes respectively, and fuse the outputs of the localization and classification branches to obtain the directional target detection results.
[0050] The steps of the method are explained in detail below.
[0051] (1) Construct a task-decoupled spatial-channel syndrome (TD-SCS) target detection model;
[0052] The task-decoupled spatial-channel cooperative target detection model includes a backbone network, a feature pyramid network (FPN), a region generation network (RPN), and a target detection head. The target detection head includes a decoupled localization branch and a classification branch.
[0053] The localization branch uses the rotation region of interest alignment module to extract features of the rotation candidate boxes, and performs bounding box regression based on the rotation candidate boxes;
[0054] In this invention, specifically, the backbone network can be a ResNet-50 pre-trained on ImageNet. The input image first passes through the backbone network to extract deep features, and then passes through a feature pyramid network to form a multi-scale feature map. The directional candidate region generation network generates multiple original rotated candidate boxes based on the multi-scale feature map, each rotated candidate box being represented as... , where (x i ,y i () is the center coordinate, w i and h i Width and height, respectively, θ i The rotation angle is used; the original rotation candidate boxes are input into the positioning branch and the classification branch respectively;
[0055] The localization branch directly utilizes the Rotated RoI Align (RRoI Align) module to align the original rotated candidate bounding boxes. Extracting features at fixed resolution This branch is used for accurate bounding box regression. It preserves the geometry of the original candidate boxes, ensuring the integrity of boundary-sensitive information.
[0056] The classification branch sequentially includes a Task-Decoupled Geometry-Aware Region Expansion Module (TD-GREM) and a Geometry-Guided Channel Calibration Module (GCCM), which are used to perform geometrically aware spatial expansion and feature recalibration on the rotated candidate boxes to obtain classification features.
[0057] (1-1) Task-Decoupled Geometry-AwareRegion Expansion Module (TD-GREM)
[0058] In this invention, TD-GREM is used to adaptively expand the candidate boxes in the classification branch according to the geometry of the candidate boxes, so as to improve the context awareness capability required for the classification of slender targets.
[0059] The geometry-aware region expansion module obtains the current rotated candidate box. Based on the prior information of aspect ratio, the height expansion factor and width expansion factor are calculated, and the expanded candidate box size is calculated based on this. The short side of the long strip candidate box is adaptively expanded.
[0060] The height expansion factor satisfies The width expansion factor satisfies The aspect ratio of the current rotated candidate box. , and These represent the width and length of the rotated candidate box, respectively. For the Sigmoid function, The maximum expansion ratio, Aspect ratio symmetrical threshold, This refers to the transition sharpness parameter; it should be noted that in this invention, "aspect ratio" is a conventional usage in the art, and in practice... The formula is expressed as "width-to-length ratio," which is equivalent to the conventional "length-to-width ratio" and is used to describe the shape of an object (elongated shape); in this embodiment, , , .
[0061] After expansion, the candidate box size is .
[0062] This invention maintains the center coordinates (x) i ,y i and rotation angle θ i Without changing the bounding box, only adjust the width and / or height to obtain the expanded candidate box; when (For horizontally slender targets) Larger, primarily expanding the height (i.e., the shorter side in the current case); when (For longitudinally slender targets) For larger targets, the expansion primarily affects the width (i.e., the shorter side in the current case); for targets that are nearly square, only slight expansion is achieved in both directions; subsequently, the expanded features are extracted from the expanded candidate boxes using a rotational RoI alignment module. , The multi-scale feature maps output by the backbone network and the feature pyramid network. The rotated candidate box expanded by TD-GREM is denoted as RRoIAlign represents the rotation alignment operation for the region of interest, used to align regions from the feature map. Extracting candidate boxes The corresponding fixed-size features, where C is the number of channels and K×K is the spatial resolution of the output feature map;
[0063] In this way, TD-GREM introduces richer lateral contextual information for targets with extreme aspect ratios, while avoiding the introduction of too much irrelevant background along the long axis.
[0064] (1-2) Geometry-Guided Channel Calibration Module (GCCM)
[0065] In this invention, GCCM is used to perform geometrically guided channel recalibration on the expanded features extracted from the expanded candidate box, so as to suppress background interference introduced by spatial expansion and improve the response intensity of the channels related to slender targets.
[0066] The geometry-guided channel calibration module acquires the extended features extracted by the geometry-aware region expansion module after spatial expansion of the rotated candidate box. Global average pooling is performed on the extended features to obtain channel descriptors. , ;
[0067] Construct a geometric descriptor that includes information on the shape, orientation, and scale of the expanded candidate boxes;
[0068] Specifically, the geometry descriptor satisfy,
[0069]
[0070] in, As a linear projection layer, the above 5-dimensional vectors are mapped to a low-dimensional geometric embedding space (with dimension d). For rotation angle, The aspect ratio of the original candidate boxes is used; the dynamic range of the scale is compressed using logarithmic transformation, and the periodicity of the angles is handled using biangular trigonometric functions.
[0071] geometry descriptor With channel descriptor After splicing, a joint representation is obtained. The input is a two-layer multilayer perceptron, which generates channel weights. The extended features are then recalibrated using these channel weights.
[0072] The channel weights satisfy the following:
[0073]
[0074] in, The channel descriptor obtained by global average pooling is... Indicates splicing, For ReLU function, and These are the weight matrices of the first and second fully connected layers in a multilayer perceptron, respectively.
[0075] The calibration weights for each channel range from (0,1). This mechanism makes the channel weights dependent on the geometric properties of the candidate box, thereby producing differentiated channel responses for targets of different shapes, orientations, and scales.
[0076] The recalibration is satisfied.
[0077]
[0078] in, To extend features; Channel-level multiplication; residual form ensures scaling factor Strictly positive, this residual recalibration method can enhance the target-related channel response while reducing background channel interference introduced by the expansion and avoiding excessive suppression of features in the early stages of training;
[0079] Based on recalibrated features Input the classification header to predict the class probability.
[0080] GCCM utilizes geometric information to guide channel recalibration, highlighting target-related channels and implicitly suppressing background response, thereby improving the discriminative power of classification features when background noise is introduced by spatial expansion.
[0081] After completing the above processing, the bounding box regression results and classification results are decoupled and predicted and assembled to output the directional target detection results.
[0082] (2) Obtain a long-distance sample dataset and train the model;
[0083] This method acquires target detection image datasets for distant scenes, such as images taken by satellites or drones, but is applicable to aerial images captured by drones. The dataset contains target instances of different terrain features, scales, and orientations. Each image is accompanied by an annotation file recording the target's category and its rotated bounding box parameters. As an example, the DOTA-v1.0 and HRSC2016 public datasets can be used.
[0084] (2-1) Data Preprocessing
[0085] For the DOTA dataset, the original images are segmented into 1024×1024 image patches with a 200-pixel overlap between patches; for the HRSC2016 dataset, the original images are segmented into 800×800 image patches. The processed data are then divided into training, validation, and test sets proportionally, such as 70% training, 15% validation, and 15% test.
[0086] (2-2) Model Training
[0087] Implemented using the MMRotate toolkit. Oriented R-CNN was used as the baseline detector, with ResNet-50-FPN as the backbone network, pre-trained on ImageNet. Training followed a standard 1×schedule (12 epochs), with a batch size of 2 or 4, an initial learning rate of 0.005, and a stochastic gradient descent optimizer with a momentum of 0.9 and weight decay of 0.0001. Input images were uniformly scaled to 1024×1024 (DOTA) or 800×800 (HRSC2016), and data augmentation methods such as random horizontal flipping and rotation were employed.
[0088] (3) Input the aerial image to be detected into the trained model, generate rotating candidate boxes, and decouple the target detection head into independent localization and classification branches. Perform localization and context-enhanced classification feature extraction on the rotating candidate boxes respectively, and fuse the outputs of the localization and classification branches to obtain the directional target detection results.
[0089] The trained model is fed with an aerial image (such as a drone image) to be detected. Multi-scale features are extracted through the backbone network and FPN, and a set of rotated candidate boxes are generated by the directional RPN. Subsequently, the object detection head is decoupled into independent localization and classification branches, which respectively localize the rotated candidate boxes and extract classification features based on context enhancement. The localization branch performs bounding box regression based on the original rotated candidate boxes and outputs the accurate bounding box parameters after regression. The classification branch performs geometrically aware spatial expansion and feature recalibration on the rotated candidate boxes and outputs the class probability of each candidate box.
[0090] The regression bounding box output by the localization branch is assembled with the class probability output by the classification branch to obtain the final directional target detection result (rotated bounding box and its class label). Non-maximum suppression can be used to remove redundant detection boxes, and the threshold is set to 0.5.
[0091] This invention also relates to a task decoupling detection system for targets with extreme aspect ratios in aerial images, comprising:
[0092] Memory, used to store computer programs;
[0093] A processor is used to execute the computer program to implement the task decoupling detection method for targets with extreme aspect ratios in aerial images.
[0094] This system can be deployed in UAV ground stations, embedded devices, or cloud servers to process UAV aerial images in real time and output directional target detection results.
[0095] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0096] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0097] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0098] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0099] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0100] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A task decoupling detection method for targets with extreme aspect ratios in aerial images, characterized in that, Includes the following steps: Construct a spatial-channel collaborative target detection model with task decoupling; Obtain a dataset of samples from distant locations and train the model; The aerial image to be detected is input into the trained model to generate rotating candidate boxes. The target detection head is decoupled into independent localization and classification branches. The rotating candidate boxes are localized and the classification features are extracted based on context enhancement, respectively. The outputs of the localization and classification branches are fused to obtain the directional target detection results.
2. The task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 1, characterized in that, The task-decoupled spatial-channel collaborative target detection model includes a backbone network, a feature pyramid network, a region generation network, and a target detection head. The target detection head includes a decoupled localization branch and a classification branch. The localization branch uses the rotation region of interest alignment module to extract features of the rotation candidate boxes, and performs bounding box regression based on the rotation candidate boxes; The classification branch is sequentially equipped with a task-decoupled geometric perception region expansion module and a geometrically guided channel calibration module, which are used to perform geometric perception spatial expansion and feature recalibration on the rotated candidate box to obtain classification features.
3. The task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 2, characterized in that, The geometric perception region expansion module obtains the aspect ratio prior information of the current rotated candidate box, calculates the height expansion factor and width expansion factor, and calculates the expanded candidate box size based on this, and adaptively expands the short side of the elongated candidate box.
4. The task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 3, characterized in that, The height expansion factor satisfies The width expansion factor satisfies The aspect ratio of the current rotated candidate box. , and These represent the width and length of the rotated candidate box, respectively. For the Sigmoid function, The maximum expansion ratio, Aspect ratio symmetrical threshold, This is the transition sharpness parameter.
5. The task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 4, characterized in that, After expansion, the candidate box size is .
6. The task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 2, characterized in that, The geometry-guided channel calibration module acquires the extended features extracted by the geometry-aware region expansion module after spatial expansion of the rotated candidate box, and performs global average pooling on the extended features to obtain the channel descriptor. Construct a geometric descriptor that includes information on the shape, orientation, and scale of the expanded candidate boxes; The geometric descriptor and the channel descriptor are concatenated and input into a multilayer perceptron to generate channel weights. The channel weights are then used to recalibrate the extended features.
7. The task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 6, characterized in that, The geometry descriptor satisfy, , in, For linear projection layers, The angle is the rotation angle.
8. A task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 6, characterized in that, The channel weights satisfy the following: , in, The channel descriptor obtained by global average pooling is... Indicates splicing, For ReLU function, and These are the weight matrices for the first and second fully connected layers in a multilayer perceptron, respectively.
9. A task decoupling detection method for targets with extreme aspect ratios in aerial images according to claim 6, characterized in that, The recalibration is satisfied. , in, To extend features; Based on recalibrated features Input the classification header to predict the class probability.
10. A task decoupling detection system for targets with extreme aspect ratios in aerial images, characterized in that, include: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the task decoupling detection method for targets with extreme aspect ratios in aerial images as described in any one of claims 1 to 9.