Mixed anchor point remote sensing image target detection method based on multi-scale large kernel convolution
Through the HMLSKNet method, combined with multi-scale large-core convolution and hybrid anchor point technology, the problem of insufficient detection accuracy and speed in remote sensing image object detection is solved, and more efficient object detection effect is achieved.
Patent Information
- Application Number
- CN202510341325.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing remote sensing image object detection methods are insufficient in the detection accuracy and speed when processing remote sensing images with high resolution, small targets and rotation angles, and are difficult to effectively extract target information in complex backgrounds and diverse data.
HMLSKNet, a hybrid anchor remote sensing image object detection method based on multi-scale large-core convolution, combines multi-scale selective large-core convolution network and hybrid anchor target box, and improves target box design by introducing mixed anchor point and direction-aware convolution technology, and improves detection accuracy and speed.
The accuracy and speed of object detection are significantly improved on the remote sensing image dataset, the calculation cost is reduced, the detection ability of small and large targets is enhanced, and the accuracy and efficiency of feature extraction are improved.
Smart Images

Figure CN120279310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image target detection in the field of computer vision, and specifically relates to a method for hybrid anchor remote sensing image target detection based on multi-scale large kernel convolution. Background Art
[0002] Object detection has always been an important part of the image field. For different data sets, the corresponding algorithms used are also different. In the face of different data sets, the network needs to be changed to extract features in order to achieve better detection effects. With the continuous progress of remote sensing technology, the obtained remote sensing image data is large in quantity and wide in coverage, containing rich surface information. It is necessary to efficiently and accurately extract target information from it. Remote sensing image target detection can be applied to multiple fields such as smart city construction, environmental monitoring, and resource management, such as road traffic monitoring, land use planning, and natural disaster warning. In recent years, deep learning technology has made remarkable progress in the field of computer vision, especially in object detection tasks, providing new solutions for remote sensing image target detection. Remote sensing image target detection has characteristics such as high resolution and large scale, complex background and diversity, multi-source data and multi-modal fusion, diversity and scale change of targets, spatio-temporal characteristics, large data volume and high computational complexity, high-precision requirements, uncertainty and noise. These characteristics make remote sensing image target detection a challenging but very important research field, which requires comprehensive use of various technical means to continuously improve the accuracy, efficiency and robustness of detection algorithms to cope with complex and changing remote sensing data and application requirements. For the object detection of remote sensing images, the characteristics of its data set are that the objects are small, and because the images are taken by remote sensing satellites, the images are relatively blurred, and the object detection of remote sensing images is mostly applied to buildings such as ships and bridges, and their objects have a certain angle. If the target boxes that appear can fit the objects well, the detection effect can be improved. In view of the characteristics of complex background, variable target scale and large required computing resources of remote sensing images, the present invention proposes a hybrid anchor object detection algorithm based on multi-scale selective large kernel convolution.
[0003] High-performance remote sensing object detectors usually rely on the Region Convolutional Neural Network (RCNN) framework, which consists of a Region Proposal Network and a Region Convolutional Neural Network (CNN) detection head. The Region Proposal Network proposes high-quality Regions of Interest (RoIs) from the backbone feature map, while the Region Convolutional Neural Network detection head is responsible for object classification and bounding box regression. In recent years, several variants of the Region Convolutional Neural Network framework have been proposed. The two-stage Region of Interest (RoI) Transformer rotates the candidate horizontal anchor boxes using fully connected layers in the first stage and then extracts the features within the boxes for further regression and classification. The More Robust Detection for Small, Cluttered, and Rotated Objects (SCRDet) uses an attention mechanism to reduce background noise and improve the modeling of crowded and small objects. The Oriented Region Convolutional Neural Network for Rotated Object Detection (Oriented RCNN) and the Gliding Vertex based on horizontal bounding boxes for multi-directional object detection introduce a new box encoding system to address the instability of training losses caused by the periodicity of rotation angles. Treating remote sensing detection as a point detection task provides another approach to solving remote sensing detection problems. The single-stage detection framework does not rely on proposed anchors but directly classifies and regresses the oriented bounding boxes from the anchor points densely sampled from the grid. The Rotated Object Detection with Feature Alignment (S2ANet) network extracts robust object features through oriented feature alignment and azimuth-invariant feature extraction. On the other hand, the Dilated Residual Network (DRN) uses an attention mechanism to dynamically refine the features extracted by the backbone network to obtain more accurate predictions. Different from the Oriented Region Convolutional Neural Network for Rotated Object Detection and the Gliding Vertex based on horizontal bounding boxes for multi-directional object detection, the Rotation Object Detection with Learning Modulated Loss (RSDet) solves the discontinuity problem of regression losses by introducing a modulated loss. Localization Distillation improves the localization quality of the oriented bounding boxes through distillation. The Anchor-free Oriented Proposal Generator and the Single-stage Detection Network for Rotated Object Detection in Remote Sensing Images (R3Det) adopt a progressive regression method to refine the bounding boxes from coarse-grained to fine-grained. Besides Convolutional Neural Networks, the Arbitrary Orientation Object Detection Transformer (AO2-DETR) introduces the Transformer-based detection framework DETR into the remote sensing detection task, bringing more research diversity. Some methods provide another approach to solving remote sensing detection problems. Although these methods have achieved promising results in solving the rotation variance problem, they do not consider the strong and valuable prior information presented in aerial images.To address these problems, the present invention proposes a new hybrid anchor remote sensing image object detection method based on multi-scale large kernel convolution, which is implemented using HMLSKNet (Hybrid Anchor Multi-Scale Large Selective Kernel Network). The main features of the present invention are as follows.
[0004] 1. Extracting features using a multi-scale selective large kernel convolution network can improve accuracy. For the feature extraction backbone network part, MLSKNet (Multi-Scale Large Selective Kernel Network) is used, which utilizes large kernels and a spatial selection mechanism to better simulate this prior information without modifying the current detection framework. By leveraging the inherent features of remote sensing images, which require a more extensive and adaptable combination of the background, MLSKNet can effectively simulate the different context nuances of different object types by adjusting its larger spatial receptive field. A large number of experiments have shown that the lightweight model MLSKNet achieves the best performance on competitive remote sensing datasets.
[0005] 2. Hybrid anchors have the advantages of reducing computational costs, increasing processing speed, and relatively improving detection accuracy. The reason is that after feature extraction, a hybrid anchor rotation detector is used in the detection stage. Since object detection usually relies on two-stage or one-stage methods and typically adopts an anchor-based strategy, the number of anchors generated during training is redundant, which may lead to computationally expensive operations. In contrast, the anchor-free mechanism provides a faster processing speed, but the reduction in the number of training samples may affect detection accuracy. Hybrid anchors combine the advantages of anchor-based and anchor-free schemes for oriented object detection. By using only one preset anchor at each position on the feature map and optimizing these anchors using our direction-aware convolution technique.
[0006] A series of experiments were conducted on the DOTA-v1.0 and FAIR1M-v1.0 datasets, and HMLSKNet was compared with other state-of-the-art technologies. The results show that HMLSKNet is competitive compared to other methods in terms of object detection.
[0007] The main content of the research on this algorithm invention is as follows:
[0008] The present invention proposes a hybrid anchor remote sensing image object detection method based on multi-scale large kernel convolution: HMLSKNet (Hybrid Anchor Multi-Scale Large Selective Kernel Network). LSKNet has achieved superior results on public datasets. On this basis, the present invention adds a multi-scale attention mechanism MSAA (Multi-Scale Attention Aggregation module) to further extract features and attempts to improve the performance of the entire network. Due to the characteristics of remote sensing images, such as high resolution, small targets, and certain rotation angles, traditional object bounding box detection methods are difficult to accurately fit the target object, resulting in a decrease in detection accuracy. Currently, the advanced model LSKNet for remote sensing image object detection achieves efficient object detection by using large-scale selectable convolution kernels. LSKNet dynamically adjusts the weights of different-scale convolution kernels through a designed selection mechanism to achieve effective fusion of multi-scale feature maps, enabling the network to adaptively select the most suitable convolution kernel size, thereby improving the effect and accuracy of feature extraction. However, LSKNet mainly improves the detection effect by changing the backbone network and uses object bounding boxes with rotation angles for object detection. The research objective of this paper is to improve the design of the object bounding box by introducing hybrid anchors, enabling it to combine the advantages of anchor-based and anchor-free mechanisms, and introducing the orientation-aware convolution (O-AwareCon) technology to improve the detection quality by adapting to the shape and orientation of the target, thereby enhancing the detection effect on remote sensing image datasets. The key problems to be solved include how to improve the accuracy and detection speed of object detection, reduce the computational parameters to save computing power and increase the detection speed; how to further improve the backbone network to enhance the accuracy of the model, and how to optimize the combination of hybrid anchors and LSKNet, conduct comparative experiments on multiple public datasets, and further optimize the network structure to improve the overall detection performance of the model. Summary of the Invention
[0009] Based on a deep understanding of selective large kernel convolution and hybrid anchor object bounding boxes, the present invention proposes a hybrid anchor remote sensing image object detection method based on multi-scale large kernel convolution - HMLSKNet. This method is an improvement based on LSKNet. Further, a multi-scale attention mechanism (MSAA) is added. The selective convolution kernel mechanism is used for feature extraction, and at the same time, hybrid anchors are introduced to improve the accuracy and detection speed of the object bounding box. Specifically, by using the MLSKNet module as the backbone network to extract features, the traditional object bounding box method is improved. The hybrid anchor is used to represent the oriented bounding box, making it more suitable for object detection tasks with rotation angles and complex shapes and simultaneously increasing the detection speed.
[0010] The technical problem to be solved by the present invention is to overcome the neglect of the target detection box in the existing large kernel selective convolution network, and further improve the accuracy and detection speed of target detection, and provide a new detection method that makes improvements to the target detection box.
[0011] The present invention is a target detection algorithm based on deep learning and is implemented based on a convolutional neural network. The present invention uses a large kernel and a spatial selection mechanism to better simulate these prior information without modifying the current detection framework, and utilizes the inherent characteristics of remote sensing images: the need to more widely and adaptively combine the background, by adjusting its larger spatial receptive field and proposing to use a hybrid anchor to represent the oriented bounding box. The network proposed by the present invention consists of two components: a large kernel convolutional neural network and a hybrid anchor target box.
[0012] The technical solution of the present invention is as follows:
[0013] A method for detecting targets in remote sensing images with hybrid anchors based on multi-scale large kernel convolution is implemented by using HMLSKNet. HMLSKNet includes three parts: backbone, neck, and head. The specific steps are as follows:
[0014] Step 1: Dataset processing:
[0015] The dataset is divided into a training set, a validation set, and a test set. Both the training set and the validation set are used for training, and the test set is used for testing.
[0016] Preferably: The DOTA-v1.0 dataset crops a series of 1024×1024 graphics from the original images with a step size of 824. During the training process, we only adopt the method of random horizontal flipping to avoid overfitting. For a fair comparison with other methods, we adopt data augmentation (i.e., random rotation) in the training stage. The dataset is cropped into 1024×1024 images with a step size of 512.
[0017] Step 2: Extract features by multi-scale selective large kernel convolution (the backbone part, implemented by using MLSKNet)
[0018] The training set and validation set processed in step 1 are input into the MLSKNet backbone network. Let the size of the input feature map be [B, C, H, W], where B is the batch size, C is the number of channels, and H and W are the height and width. First, it undergoes normalization to scale different feature maps to a unified numerical range. Then the feature map enters the LSK module, undergoes non-linear combination through a fully connected layer, and then undergoes Gaussian error activation for regularization to reduce overfitting. The processed features reach the selective large kernel convolution layer. In the selective large kernel convolution layer, first, it passes through multiple parallel convolution kernel branches of different sizes. Each convolution kernel branch uses different large kernel depthwise separable convolutions to obtain multiple groups of output feature maps of different scales. Then feature concatenation is performed, and the outputs of all convolution kernel branches are concatenated along the channel dimension to obtain [B, N×C, H, W], where N is the number of branches. Then maximum average pooling is performed, global average pooling is performed on the concatenated feature map to compress spatial information to obtain [B, N×C, 1, 1]. After passing through the spatial attention module, spatial selection is performed to obtain the weights corresponding to the output feature maps of each branch. Multiply the output feature map of each branch by its corresponding weight and add them to obtain the fused feature map [B, C, H, W]. The processed feature map further passes through a feedforward neural network, which includes a fully connected layer, a depth convolution, a GELU activation function, and a second fully connected layer. The feature map processed by the feedforward neural network is input into the multi-scale attention module (MSAA), which performs feature aggregation using spatial and channel dual paths: One is the spatial refinement path: First, channel compression is performed through a 1×1 convolution, and then parallel convolution operations of multi-scale convolution kernels (such as 3×3, 5×5, 7×7) are performed and summed. Further, spatial features are aggregated through a combination operation of mean pooling and max pooling, and after a 7×7 convolution, element-wise multiplication is performed with the feature map activated by Sigmoid to obtain the spatial refinement feature map. The other is the channel aggregation path: Global average pooling is used to compress the feature dimension to C1×1×1, and then a channel attention map is generated through a 1×1 convolution and ReLU activation. The channel attention map is upsampled and fused with the spatial refinement feature map. Finally, the input feature map is merged with the fused feature map to obtain the extracted features.
[0019] Step 3: The Feature Pyramid Network performs feature fusion (Neck part)
[0020] The features obtained in step 2 are input into the Feature Pyramid Network FPN (Feature Pyramid Networks). After extracting multi-scale features through the Resnet (Residual Network), the top-down path is used to upsample the high-level features, and they are fused with the low-level features through lateral connections. Finally, through convolution processing, multi-scale prediction feature maps with both high-resolution details and rich semantic information are generated, that is, feature pyramids at different levels.
[0021] Step 4: Mixing Anchors to Achieve Object Classification and Oriented Bounding Box Regression (Head Part)
[0022] The features extracted in Step 3 are fed into the mixing anchor region proposal network. The mixing anchor region proposal network uses an anchor-free assignment strategy to generate horizontal candidate boxes for the rectangularized oriented ground truth. The horizontal candidate boxes are refined through direction-aware convolution and subsequent anchor-based detection heads to generate high-quality horizontal candidate regions. Then, the lightweight horizontal candidate regions use the RotatedShared2FCBBoxHead for the classification head of the rotated bounding box regression to perform rotation alignment processing on the horizontal candidate regions of interest, generating oriented candidate regions. Finally, these oriented candidate regions are further refined through the oriented bounding box detection head to complete the object classification and bounding box regression tasks.
[0023] Calculation formula for the loss function of the mixing anchor:
[0024]
[0025] where a x , a y , a w , a h represent horizontal anchor boxes, t x , t y , t w , t h represent the ground truth, δ = (δ x , δ y , δ w , δ h ) represents the predicted encoding value, represents the target encoding value, a′x, a′ y , a′ h , a′w represent candidate proposals (Proposals), w AF represents the loss weight for fine-tuning, Intersection(·) represents the intersection, and Union(·) represents the union.
[0026] The total loss function is: L RPN = L AF + L AB
[0027] where L AB represents the loss function calculated for the anchor target boxes.
[0028] Step 5: Model Prediction
[0029] The test set is put into the model trained in steps 2-4 for detection to obtain the target information in the image, and the performance of the model is evaluated by the mean average precision (mAP) metric.
[0030] The main features of the present invention are as follows:
[0031] Selective large convolution kernel: By designing a selective mechanism, HMLSKNet can dynamically adjust the weights of convolution kernels of different scales, so as to effectively fuse between feature maps of different scales. This mechanism allows the network to adaptively select the most appropriate convolution kernel size according to the characteristics of the input image, thereby improving the effect and accuracy of feature extraction.
[0032] Hybrid anchor target box representation method: Combining the advantages of anchor-based and anchor-free mechanisms, an anchor is used at each position on the feature map. In the anchor-based component, the orientation-aware convolution (O-AwareCon) technology is introduced to improve the detection quality by adapting to the shape and orientation of the target, thereby enhancing the detection effect on the remote sensing image dataset.
[0033] Multi-scale feature extraction: By introducing a multi-scale feature extraction module, it can effectively capture target information of different scales, enhancing the detection ability for small and large targets. While ensuring computational efficiency, it improves the resolution and receptive field of the feature map.
[0034] New loss function: A special loss function is designed to better optimize various parts of object detection, including the regression of position, scale, and polar coordinate radius. The new loss function can more effectively handle targets of different scales and shapes, improving the detection accuracy.
[0035] The beneficial effects of the present invention: The present invention uses MLSKNet as the backbone network, introduces hybrid anchors to improve the design of the target box, enabling it to combine the advantages of anchor-based and anchor-free mechanisms, and introduces the orientation-aware convolution technology to improve the detection quality by adapting to the shape and orientation of the target, thereby enhancing the detection effect on the remote sensing image dataset. Brief Description of the Drawings
[0036] Figure 1 It is the overall network structure of the present invention.
[0037] Figure 2 It is a schematic diagram of the MLSKNet module of the present invention. An MLSKNet block consists of two residual sub-blocks and a multi-scale attention module: a large kernel selection sub-block (LSK module), a feed-forward network (FFN) sub-block, and a multi-scale attention module (MSAA).
[0038] Figure 3 It is a schematic diagram of the LSK module of the present invention.
[0039] Figure 4 Schematic diagram of the MSAA structure of the present invention.
[0040] Figure 5 Hybrid anchor target box module structure of the present invention. Specific implementation manners
[0041] The specific implementation manners of the present invention will be further described below in conjunction with the accompanying drawings and technical solutions.
[0042] Figure 1 Overall structure of HMLSKNet proposed by the present invention. The remote sensing image is first fed into the backnone - MLSKNet multi-scale large kernel convolution used in the present invention for feature extraction, and the extracted remote sensing image features are fed into the iterative feature fusion sub-network. The iterative feature enhancement sub-network selects the FPN feature pyramid structure and first uses the MLSKNet convolutional neural network as its underlying feature extraction network, enabling it to effectively extract the features of the image and providing a basis for subsequent feature fusion. Then comes the bottom-up path, that is, features are gradually extracted from the input image through convolutional layers to obtain feature maps of multiple scales. These feature maps usually have different resolutions and semantic information, providing a rich feature source for subsequent feature fusion.
[0043] Figure 2 It is the backbone network part of the present invention - MLSKNet. An MLSKNet includes a large kernel selection sub-block (LSK module), a feed-forward network (FFN) sub-block, and a multi-scale attention module (MSAA); Figure 3It is a schematic diagram of the LSK module in the MLSKNet structure of the present invention. When processing images, the large kernel selection sub-block is the key to the processing process. It uses a series of depth convolutional kernels with different dilation rates to process the input feature map. The weights of these convolutional kernels are dynamically determined according to the input, and can capture context information at different scales as needed, thereby dynamically adjusting the receptive field of the network. This process not only expands the receptive field, but also effectively reduces the number of model parameters and improves the computational efficiency. Then, the spatial selection mechanism weights the convolutional feature map, and these weights are also dynamically calculated based on the input feature map, allowing the network to further adjust the receptive field according to the spatial position of the target object and focus on the important parts of the image. Subsequently, the feed-forward network sub-block performs channel mixing and feature refinement on the features processed by the large kernel selection sub-block. This sub-block contains a fully connected layer, a depth convolution, a GELU activation function, and a second fully connected layer. These components work together to improve the quality of the features and provide the necessary information for subsequent classification and detection tasks. By processing the input feature map layer by layer and stacking multiple LSK modules, MLSKNet can extract high-level features with rich context information. These feature maps are then sent to the object detection head for processing to output the final object detection results. The whole process not only efficiently captures the context information required by different targets, but also improves the detection performance for targets of different scales by dynamically adjusting the receptive field and using the spatial selection mechanism.
[0044] Figure 4 It is a schematic diagram of the MSAA structure in the MSLKNet structure of the present invention. This module first receives feature maps of different scales and further refines these features through parallel processing paths. One path focuses on feature aggregation in the spatial dimension. By performing convolutional operations on the feature map at different scales and then combining means such as average pooling and max pooling, the fusion of multi-scale spatial information is achieved. The other path focuses on feature interaction in the channel dimension. Through the attention mechanism, the correlation between different channels is captured, and the weights of each channel are adjusted accordingly. The outputs of these two paths are then combined to achieve comprehensive enhancement of spatial and channel information.
[0045] Figure 5 It is the part of the hybrid anchor object detection box. First, the remote sensing image extracts deep features through the FPN feature pyramid network. Then, the extracted features generate horizontal anchors on the rectified oriented ground truth using the anchor-free assignment strategy through the hybrid anchor. The anchors are refined through azimuth-aware convolution and the subsequent anchor-based head to generate high-quality horizontal proposal boxes. After that, the lightweight proposal box transformation network learns to generate oriented proposal boxes from the predicted horizontal proposal boxes after RoI alignment processing. Finally, these oriented proposal boxes are refined through the oriented bounding box head for classification and bounding box regression.
[0046] The present invention uses the DOTA-v1.0 dataset and the FAIR1M-v1.0 dataset. The DOTA-v1.0 dataset contains 2,806 remote sensing images. It contains 188,282 instances of 15 categories: airplane (PL), baseball diamond (BD), bridge (BR), grand track field (GTF), small vehicle (SV), large vehicle (LV), ship (SH), tennis court (TC), basketball court (BC), storage tank (ST), soccer ball field (SBF), roundabout (RA), port (HA), swimming pool (SP), and helicopter (HC). The FAIR1M-v1.0 dataset is a dataset composed of 15,266 high-resolution images, with more than 1 million instances. It contains 5 categories and 37 sub-category objects, and its technical highlight lies in its fine-grained object classification and precise annotation method. The dataset covers 37 different categories, including different models of airplanes, various ships and vehicles, as well as sports venues, road facilities, etc. The annotation uses the XML file format, and each object is represented by an OBB, ensuring the accurate depiction of the object boundary in different directions. Each XML file corresponds to an image, which contains the size information of the image and the sequence of coordinate points of each target object, forming a closed polygon that precisely outlines the target contour.
[0047] The present invention is mainly divided into two subtasks: data processing and network training. In data processing, the present invention sets both the training set and the validation set for training, and the test set for testing. A series of 1024×1024 graphics are cropped from the original images, with a step size of 824. During the training process, we adopt the method of random horizontal flipping to avoid overfitting. To make a fair comparison with other methods, we adopt data augmentation (i.e., random rotation) in the training stage. In the multi-scale experiment, we first resize the original images according to three scales (0.5, 1.0, and 1.5), and then crop them into 1024×1024 images, with a step size of 512. There are more than 1 million instances and more than 15,000 images in the FAIR1M-v1.0 dataset. All objects in the FAIR1M dataset are annotated for 5 categories and 37 sub-categories through oriented bounding boxes. The size of each image ranges from 1000×1000 to 10,000×10,000 pixels and contains objects showing various scales, orientations, and shapes.
[0048] In network training, all experiments of the present invention were implemented for training and testing using the PyTorch framework on a server with 16GB RAM and a single NVIDIA 4090Ti GPU (11GB). During training, the present invention was trained on the processed dataset. In addition, data augmentation operations such as random horizontal flipping, mirroring, and rotation were adopted by the present invention to avoid overfitting. The present invention used the stochastic gradient descent method to optimize the entire network, with a momentum value of 0.9, a batch size of 4, and a total of 40 training rounds. The initial learning rate was 5e-3. During the processing, after the DOTA-v1.0 dataset and the FAIR1M-v1.0 dataset were input into the network, a dataset processing was first performed, and the dataset trimming size was 1024×1024.
[0049] Table 1 Specific experimental results on the DOTA-v1.0 dataset
[0050]
[0051] Table 2 Specific experimental results on the FAIR1M-v1.0 dataset
[0052]
[0053] The comparison method selected in the specific implementation of the present invention is to compare with LSKNet. LSKNet uses LSKNet as the backbone network for feature extraction, uses the Oriented RPNHead for object box detection, takes into account the unique prior knowledge in the remote sensing scene, and uses selective large kernel convolution to further expand the range of the receptive field. LSKNet achieved new state-of-the-art results in standard benchmark tests (i.e., DOTA-v1.0 (81.85% mAP), FAIR1M-v1.0 (47.87% mAP)). For a fair comparison, the HMLSKNet method used its publicly available code and its recommended parameter settings, and both used the same pre-trained network and were tested on the same test set. Judging from the final experimental results, the HMLSKNet method proposed by the present invention obtained the best performance in terms of the map index in different experimental settings (i.e., DOTA-v1.0 (83.19% mAP), FAIR1M-v1.0 (66.48% mAP)). The higher the map index, the higher the accuracy of the method. The specific experimental results on the DOTA-v1.0 dataset are shown in Table 1: The specific experimental results on the FAIR1M-v1.0 dataset are shown in Table 2.
Claims
1. A hybrid anchor remote sensing image target detection method based on multi-scale large kernel convolution, characterized in that It is implemented using HMLSKNet, which consists of three parts: backbone, neck, and head. The specific steps are as follows: Step 1: Dataset processing: Divide the dataset into a training set, a validation set, and a test set. The training set and the validation set are both used for training, and the test set is used for testing; Step 2: Extract features using multi-scale selective large kernel convolution The training set and the validation set processed in Step 1 are input into the MLSKNet backbone network. Assume the input feature map size is [B, C, H, W], where B is the batch size, C is the number of channels, and H and W are the height and width. First, perform normalization to scale different feature maps to a unified numerical range. Then, the feature map enters the LSK module, undergoes non-linear combination through a fully connected layer, and then undergoes Gaussian error activation for regularization to reduce overfitting. The processed features reach the selective large kernel convolution layer. In the selective large kernel convolution layer, first pass through multiple parallel convolution kernel branches of different sizes. Each convolution kernel branch uses different large kernel depthwise separable convolutions to obtain multiple groups of output feature maps of different scales. Then, perform feature concatenation, concatenating the outputs of all convolution kernel branches along the channel dimension to obtain [B, N×C, H, W], where N is the number of branches. Then, perform max-average pooling, perform global average pooling on the concatenated feature map to compress spatial information to obtain [B, N×C, 1, 1], and then perform spatial selection after passing through the spatial attention module to obtain the weights corresponding to the output feature maps of each branch; Multiply the output feature map of each branch by its corresponding weight and add them to obtain the fused feature map [B, C, H, W]. The processed feature map further passes through a feed-forward neural network, which includes a fully connected layer, a depth convolution, a GELU activation function, and a second fully connected layer. The feature map processed by the feed-forward neural network is input into the multi-scale attention module MSAA, which performs feature aggregation using a spatial and channel dual-path: One is the spatial refinement path: First, perform channel compression through a 1×1 convolution, and then perform parallel convolution operations with multi-scale convolution kernels and sum them. Further aggregate spatial features through a combination operation of mean pooling and max pooling, and then perform element-wise multiplication with the feature map activated by Sigmoid after a 7×7 convolution to obtain the spatial refinement feature map; The other is the channel aggregation path: Use global average pooling to compress the feature dimension to C1×1×1, and then generate a channel attention map through a 1×1 convolution and ReLU activation. The channel attention map is upsampled and fused with the spatial refinement feature map; Finally, merge the input feature map with the fused feature map to obtain the extracted features; Step 3 : Feature pyramid network for feature fusion The features obtained in Step 2 are input into the feature pyramid network FPN to extract multi-scale features. Then, use a top-down path to upsample the high-level features and fuse them with the low-level features through lateral connections. Finally, generate a multi-scale prediction feature map with both high-resolution details and rich semantic information through convolution processing, that is, feature pyramids at different levels; Step 4: Implement object classification and oriented bounding box regression using hybrid anchors Feed the features extracted in Step 3 into the hybrid anchor region proposal network. The hybrid anchor region proposal network uses an anchor-free assignment strategy to generate horizontal candidate boxes for the rectified oriented ground truth. The horizontal candidate boxes are refined through orientation-aware convolution and subsequent anchor-based detection heads to generate high-quality horizontal candidate regions. Then, the lightweight horizontal candidate regions are rotationally aligned using a rotated bounding box regression classification head for the horizontal candidate regions of interest to generate oriented candidate regions. Finally, these oriented candidate regions are further refined through an oriented bounding box detection head to complete the object classification and bounding box regression tasks; Step 5: Model prediction Detect by putting the test set into the model trained in Steps 2 - 4 to obtain the object information in the image, and at the same time evaluate the model performance using the mean average precision (mAP) metric.
2. The method for detecting remote sensing image targets with hybrid anchors based on multi-scale large kernel convolution according to claim 1, wherein, In Step 4 described above, the calculation formula for the loss function of the hybrid anchor: Among them, a x , a y , a w , a h represent horizontal anchor boxes, t x , t y , t w , t h represent ground truths, δ = (δ x , δ y , δ w , δ h ) represents the predicted coding value, represents the target coding value, a′ x , a′ y , a′ h , a′ w represent candidate proposals, w AF represents the fine-tuning loss weight, Intersection(·) represents the intersection, and Union(·) represents the union; The total loss function is: L RPN = L AF + L AB Among them, L AB represents the loss function calculated based on the anchored target box.
Citation Information
Cited By
Remote sensing image target detection method based on feature pyramid and boundary perception vector
CN120451516A
Remote sensing image target detection method based on feature pyramid and boundary-aware vector
CN120451516B