Remote Sensing Target Detection Method Based on a Centralized Visual Processing Center
By adopting a centralized visual processing center method in remote sensing image object detection, combining Transformer encoder and CNN, and through multiple average pooling layers and sparse convolution detection heads, the problems of high similarity and serious occlusion of remote sensing image targets are solved, achieving higher detection accuracy and efficiency.
Patent Information
- Application Number
- CN202410983311.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-07-22
AI Technical Summary
The remote sensing image has high similarity in the similarity of the target features and serious occlusion problems, which makes it difficult for existing detection models to achieve the ideal effect of detection accuracy and efficiency.
Using a remote sensing object detection method based on a centralized vision processing center, combined with a visual processing center parallelized by Transformer encoder and CNN, the channel dependence is extracted through multiple average pooling layers and established pixel-level correlation, enhancing the representation ability of detail features, and introducing sparse convolution detection heads to improve detection accuracy and efficiency.
The detection accuracy of remote sensing targets in complex backgrounds is improved, the detection efficiency in high feature similarity and severe occlusion environments are enhanced, and the detection accuracy and efficiency are achieved.
Smart Images

Figure CN119169446B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of remote sensing and computer vision, and in particular to a remote sensing target detection method based on a centralized vision processing center. Background Art
[0002] The rapid development of satellite sensors and remote sensing technology has greatly promoted the research of remote sensing image target detection. The high efficiency of remote sensing image target detection has made it widely used in many fields, such as urban monitoring, forest early warning, resource detection, crop monitoring, etc. Remote sensing images have the characteristics of large target scale span, dense distribution, high feature similarity, and severe occlusion, which makes it difficult for existing detection models to achieve the ideal effect of balancing detection accuracy and efficiency.
[0003] Against the backdrop of the rapid development of deep learning, object detection tasks based on convolutional neural networks (CNNs) have made significant progress. However, most object detection models are designed for object detection tasks in conventional scenarios, and are trained and tested using the MSCOCO dataset as a benchmark. If these object detection models used in conventional scenarios are directly applied to remote sensing image object detection, there will be two major problems: first, the size span of remote sensing targets is too large and the distribution is dense, which leads to the occurrence of missed detection problems in conventional detection models; second, remote sensing overhead imaging, wide field of view, and complex background lead to high similarity of target features and severe occlusion, resulting in a high false detection rate in conventional detection models. CNN-based object detection models perform downsampling during feature extraction to reduce the amount of calculation, which can easily lead to feature blurring or even loss of small target features. When target features and background features have similar textures and sizes, they are difficult to distinguish. In addition, occlusion problems usually lead to incomplete feature learning, resulting in semantic ambiguity. Therefore, more refined spatial features and global context information are needed as clues for target feature learning.
[0004] CNN has the advantage of spatial position expression, but due to the locality of convolution operations, it is difficult to directly model global semantic interactions and contextual information. Researchers have made different attempts to address this issue. For example, FPN uses the idea of image pyramids and proposes a top-down inter-layer feature interaction structure that enables shallow features to obtain the semantic representation of deep features. ExFuse proposes to introduce semantic information into low-level features, while introducing shallow high-resolution details into high-level features to enhance the later feature fusion effect. Lin uses the inherent multi-scale pyramid hierarchy of deep convolutional networks and adds branches to construct feature pyramids at marginal additional costs. These methods do not directly encode global contextual information, and it is difficult to directly obtain clear global scene information from remote sensing images with complex backgrounds.
[0005] Regarding the problem of the limited receptive field in CNN, object detection methods based on vision transformers have been recently proposed. These methods first divide the input image into different image patches, and then utilize the feature interaction based on multi-head attention among the patches to obtain global long-range dependencies. Although these methods can strengthen the utilization of the limited receptive field and local context information in CNN, they also increase the computational complexity at the same time. The research on other advanced methods shows that shallow features mainly contain general feature patterns of objects, such as texture, color, and orientation, which are often not global; in contrast, deep features reflect object-specific information, which usually requires global information. Therefore, we believe that the Transformer encoder can be used in deep features. The convolutional multi-head self-attention (CMHSA) of the efficient convolutional transform block (ECTB) proposed by Tao et al. uses convolutional projection instead of positional linear projection to extract the context information of objects to improve the recognition ability of occluded objects, reducing the amount of computation without significantly improving the detection accuracy of occluded targets.
[0006] Object detection in remote sensing images is widely applied in various fields, and many people have conducted research on its characteristics. Ouyang et al. modified the YOLOv3 network structure to enable the small feature maps corresponding to smaller objects to be fused with the feature maps of larger objects, achieving the simultaneous detection of large and small objects. Lu et al. designed an end-to-end network for attention feature fusion SSD, using spatial attention and channel attention to suppress background noise and highlight key features, solving the problem of complex backgrounds in remote sensing object detection. Chen et al. used the model of natural images to propose a domain-adaptive Faster R-CNN structure for image horizontal and strength levels, solving the problem of insufficient remote sensing image data. Yang et al. used a CNN model to achieve dual detection of the position and orientation of remote sensing images. The above methods have fully utilized the advantages of CNN in spatial position to solve the problems of multi-scale objects and complex backgrounds in remote sensing images. However, the problems of high feature similarity and severe occlusion caused by complex backgrounds have not been well solved.
[0007] The attention mechanism is widely used in semantic segmentation and image detection tasks. When using attention to process tasks, the importance of different information is reflected by weights. The channel attention mechanism SENet is divided into two parts: compression and excitation, enabling it to adaptively select and emphasize important features and improve the feature expression ability. The spatial attention mechanism STN can transform various deformed data in space and automatically capture the features of important regions. The hybrid attention mechanism CBAM applies channel-level attention and spatial-level attention to adaptive feature refinement. The self-attention mechanism ASPP performs convolution operations on the input feature map using dilated convolutions with different dilation rates to obtain the context information of the image at multiple different scales.
[0008] The Transformer was first proposed for the machine translation task and outperformed previous sequence transduction models based on complex recurrence or CNNs. Subsequently, Transformer-based models have been widely applied to various fields of natural language processing. The standard Transformer block consists of multi-head self-attention (MSA), multi-layer perceptron (MLP), and layer normalization (LN). To mitigate the large computation of complex Transformer models, PVT introduced a pyramid structure into ViT to obtain multi-scale feature maps. Specifically, it flexibly controls the length of the Transformer sequence through the patch embedding layer. Although PVT reduces the consumption of computing resources to some extent, its complexity is still quadratic in the image size. Liu et al. proposed Swin Transformer based on the shifted window strategy, which confines the computation of MSA to non-overlapping windows while allowing cross-window information interaction. Swin Transformer has only linear computational complexity and achieves state-of-the-art performance in various vision tasks, including image classification, object detection, and semantic segmentation. Although Transformer models perform well in computer vision tasks and can better establish long-range dependencies / global relationships of features, they still lack the ability to capture detailed feature representations, resulting in unsatisfactory detection accuracy when dealing with small objects and complex backgrounds. Summary of the Invention
[0009] To solve the technical problems of high similarity of target features and serious occlusion problems in remote sensing images, the present invention provides a remote sensing target detection method based on a centralized vision processing center. While selecting the advantages of CNN in terms of spatial position, the Transformer method is used to propose a new framework to inherit the advantages of both and solve the problems of high feature similarity and serious occlusion; by using multiple average pooling layers to extract vertical and horizontal channel dependencies, and then multiplying them to establish pixel-level correlations of features and enhance the representation ability of detailed features; by using a Lightweight encoder to capture the global context information and long-range dependencies of the input image, and combining with the PLC module to capture fine-grained feature representations, so as to improve the remote sensing target detection accuracy in complex backgrounds.
[0010] The remote sensing target detection method based on a centralized vision processing center provided by the present invention includes the following steps:
[0011] a. Construct a centralized vision processing center (CVPC) that includes a Transformer encoder and a convolutional neural network (CNN) in parallel, where;
[0012] b. Construct a centralized feature cross-layer fusion pyramid structure and fuse it with the results of the Centralized Visual Processing Center (CVPC) to enhance the detail expression ability of features in each layer;
[0013] c. Introduce a Sparse Convolution Detection Head (CEASC) to improve detection accuracy while ensuring detection efficiency.
[0014] Preferably, the Centralized Visual Processing Center (CVPC) includes a Lightweight encoder and a Pixel level Learning Center (PLC) module. The Lightweight encoder captures global long-range dependencies, and the Pixel level Learning Center (PLC) module establishes pixel-level correlations to enhance the representation ability of detailed features.
[0015] Preferably, the Lightweight encoder module consists of two residual blocks, namely the Neighborhood Attention module and the Channel-based MLP module. The Transformer encoder of the Centralized Visual Processing Center (CVPC) uses the Lightweight encoder, which can reduce the computational amount and improve the detection efficiency.
[0016] Preferably, the Pixel level learning center (PLC) module establishes pixel-level correlations through global average pooling and matrix multiplication in the spatial direction to enhance the detection ability of occluded targets in remote sensing images.
[0017] Preferably, the centralized feature cross-layer fusion pyramid structure uses a top-down method to fuse high-level semantic information with low-level detail information to improve the accuracy of remote sensing target detection.
[0018] Preferably, the Sparse Convolution Detection Head (CEASC) adopts an adaptive multi-layer mask strategy to generate the optimal mask ratio. The Sparse Convolution Detection Head (CEASC) improves the accuracy of target detection while maintaining detection efficiency through sparse convolution operations.
[0019] Compared with related technologies, the remote sensing target detection method based on the Centralized Visual Processing Center provided by the present invention has the following beneficial effects:
[0020] The present invention provides a remote sensing target detection method based on the Centralized Visual Processing Center:
[0021] 1. A centralized visual processing center (CVPC) is proposed. This structure is a visual processing center with parallel Transformer encoder and CNN. The Lightweight encoder is used to capture global long-range dependencies, and the Pixellevel learning center (PLC) module is used to establish pixel-level correlations, enhance the representation ability of detailed features, and improve the detection accuracy of remote sensing targets in complex backgrounds;
[0022] 2. A centralized feature cross-layer fusion pyramid structure is proposed. In a top-down manner, it is fused with the results of the centralized visual processing center, enhancing the detailed expression ability of features at each layer and improving the detection accuracy of remote sensing targets at various scales;
[0023] 3. A sparse convolutional detection head (CEASC) is introduced, which ensures the detection efficiency while improving the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is the overall structure diagram of CF-YOLO;
[0025] Figure 2 is the structure diagram of CVPC;
[0026] Figure 3 is the structure diagram of the centralized feature cross-layer fusion pyramid;
[0027] Figure 4 is the structure diagram of CEASC;
[0028] Figure 5 is the visualization of the detection results of CF-YOLO in the DOTA-v1.0 dataset;
[0029] Figure 6 is the visualization of the detection results of CF-YOLO in the DIOR dataset;
[0030] Figure 7 is the visualization of the detection results of CF-YOLO in the RSOD dataset. DETAILED DESCRIPTION OF THE INVENTION
[0031] The present invention will be further described below in conjunction with the drawings and embodiments.
[0032] As Figure 1 shown, this is the overall structure diagram of CF-YOLO in the present invention. The upper part is the overall network structure, and the lower part is the details of each operation.
[0033] The remote sensing target detection method based on a centralized vision processing center proposed by the present invention includes the steps of: constructing a centralized vision processing center (CVPC); constructing a centralized feature cross-layer fusion pyramid structure; and introducing a sparse convolution detection head (CEASC).
[0034] A. Centralized vision processing center (CVPC):
[0035] We utilize the advantages of CNN in spatial position and the global context information advantages of vision transformers to propose a centralized vision processing center (CVPC). The CVPC aims to effectively improve the detection efficiency of remote sensing targets with high feature similarity and severe occlusion. As Figure 2 shown, we propose that the CVPC mainly consists of two blocks connected in parallel, where a Lightweight encoder is used to capture global long-range dependencies; a PLC module is used to establish pixel-level correlations and enhance the representation ability of strong detail features. The resulting feature maps of these two blocks are concatenated along the channel dimension and used as the input to the downstream centralized feature cross-layer fusion pyramid. The above process can be expressed as:
[0036] X = cat(LC(X in ), PLC(X in ))
[0037] where X is the output of the CVPC; cat(.) represents the operation of concatenating feature maps along the channel dimension; LC(X in ), PLC(X in ) represent the output features of the Lightweight encoder and the PLC module respectively. X in is the output feature after the SPPCSPC operation at the bottom of the Backbone. Lightweight encoder: The Lightweight encoder mainly consists of two residual blocks: a neighborhood attention module and a channel-based MLP module. The input of the channel-based MLP module is the output of the neighborhood attention module. Dropout operations are performed after these two blocks to help the network converge better, prevent the network from overfitting, and improve the feature generalization and robustness capabilities. For the multi-head neighborhood attention module, first, X in is fed into the LayerNorm layer, then the NeighborhoodAttention operation is implemented on these features, then a Dropout operation is performed, and finally, it is added to X in to implement the residual structure. The above process is expressed as:
[0038]
[0039] where is the output of the neighborhood attention module, LN(.) is the LayerNorm operation, and NA(·) is the NeighborhoodAttention operation.
[0040] For the channel-based MLP module, first, the output features of the multi-head neighborhood attention module
[0041] are fed into the LayerNorm layer, and then the channel MLP is implemented on these features. Compared with the spatial MLP, the channel MLP can effectively reduce the computational complexity. Then, the Dropout operation and connection are performed. The above process is expressed as:
[0042]
[0043] where LN(.) is the LayerNorm operation and CMLP(.) is the channel MLP. PLC: Given the feature map X ∈ R (h×w)×c1 , to reduce the computational cost, the number of channels of the feature map is reduced to c1 / 2 through the Reshape operation. Then the feature X is fed into a dilated convolutional layer with a 3×3 kernel and a dilation rate of 2 to reconstruct the structural information of the feature map through a large receptive field. Global average pooling operations are used on the vertical and horizontal spatial directions, and the formula for calculating the elements in each direction is expressed as:
[0044]
[0045] where i, j, and q are the indices of the vertical direction, horizontal direction, and channel respectively, 0 ≤ i < h, 0 ≤ j < w, 0 ≤ q < c1 / 2. BN(.) is the BatchNormalization, and GELU(.) is the GELU activation function of the dilated convolutional layer. We obtain the aggregated tensors k h and k w in the vertical and horizontal directions from the above formula, where k h ∈ R (h×1)×c1 / 2 , k w ∈ R (1×w)×c1 / 2 . k h and k w converge the pixel-level weights of the feature map spatially. Therefore, we perform matrix multiplication on the two to obtain the position-related feature map M, M ∈ R (h×w)×c1 / 2 . Finally, M is passed through a 1×1 convolutional layer and reshaped again to restore the number of channels to C1 to obtain the output feature map PLC(X in ), PLC(X in ) ∈ R (h×w)×c1 , and the calculation process is shown as follows:
[0046]
[0047] where × represents matrix multiplication, denotes a 1×1 convolutional layer with Batch Normalization and GELU.
[0048] B. Centralized Feature Cross - layer Fusion Pyramid:
[0049] The Backbone network is used for feature extraction. As the number of its network layers deepens, the receptive field of the network becomes larger and the semantic expression ability is enhanced. However, after multiple convolutional operations, many detailed features become blurred or even lost. Using the deep - layer feature maps of the Backbone for downstream object detection tasks may lead to a decline in the model's detection ability. The Centralized Visual Processing Center (CVPC) is used for centralized feature expression to enhance the representation ability of detailed features. Therefore, we propose a centralized feature cross - layer fusion pyramid structure to fuse the result feature maps of the Centralized Visual Processing Center (CVPC) with the Neck part. As Figure 3 shown, we add four branches. After upsampling the result feature maps of the Centralized Visual Processing Center (CVPC) respectively, we perform concat feature fusion with Stage1, Stage2, Stage3, and Stage4. The centralized feature cross - layer fusion pyramid effectively integrates the context information of the CVPC and the Neck detection layer, provides more comprehensive and rich visual context information for the Head, improves the detection accuracy of remote sensing targets at various scales, and makes the model more robust and accurate.
[0050] C. Sparse Convolution Detection Head (CEASC):
[0051] While multi - scale feature cross - layer fusion brings the benefit of improving accuracy, it may also bring higher computational complexity. Optimizing computational efficiency and reducing computational complexity is a problem that needs to be solved. The present invention introduces an optimized design of a detection head (CEASC) based on sparse convolution, which effectively balances detection accuracy and efficiency. CEASC proposes a new global context - enhanced adaptive sparse convolution network. The algorithm first develops a context - enhanced group normalization (CE - GN) layer, replacing the statistical data based on sparsely sampled features with global context features, and then designs an adaptive multi - layer mask strategy to generate the optimal mask ratio at different scales to achieve compact foreground coverage, improving both accuracy and efficiency. The structure of the CEASC detection head is as Figure 4 shown.
[0052] To verify the effectiveness and robustness of the proposed CF-YOLO method, we conducted sufficient experiments on three large public remote sensing datasets, namely DOTA-v1.0, DIOR, and RSOD, and compared them with advanced methods.
[0053] A.Dataset:
[0054] The DOTA dataset is divided into three versions: v1.0, v1.5, and v2.0. Here, we use the DOTA-v1.0 version. DOTA-v1.0 contains 2,806 images and 188,282 instances. The image size ranges from 800×800 to 4,000×4,000. The targets are divided into 15 categories, namely: Helicopter (HC), Harbor (HA), Ship (SH), Large Vehicle (LV), Tennis Court (TC), Plane (PL), Small Vehicle (SV), Bridge (BR), Baseball Diamond (BD), Ground Track Field (GTF), Basketball Court (BC), Roundabout (RA), Soccer Ball Field (SBF), Storage Tank (ST), and Swimming Pool (SP). The split ratio of the training, validation, and test sets of this dataset is 3:1:2.
[0055] The DIOR dataset contains 23,463 images and 192,518 instances. The image size is fixed at 800×800. The targets are divided into 20 categories, namely: Chimney (CH), Airplane (APL), Train Station (TS), Airport (APO), Basketball Court (BC), Bridge (BR), Storage Tank (STO), Dam (DAM), Highway Service Area (ESA), Baseball Field (BF), Highway Toll Station (ETS), Golf Course (GF), Harbor (HA), Overpass (OP), Ship (SH), Windmill (WM), Ground Track Field (GTF), Tennis Court (TC), Vehicle (VE), and Stadium (STA). The training and validation sets contain 11,725 images and 68,073 instances, and the test set contains 11,738 images and 124,445 instances.
[0056] The RSOD dataset contains four types: airplanes, fuel tanks, sports fields, and overpasses, and is annotated in the format of the PASCAL VOC dataset. The dataset includes 4 folders, each folder containing one type of object: the airplane dataset, with 4,993 airplane instances in 446 images; the sports field, with 191 sports field instances in 189 images; the overpass, with 180 overpass instances in 176 images; and the fuel tank, with 1,586 fuel tank instances in 165 images. The split ratio of the training, validation, and test sets of this dataset is 3:1:2.
[0057] B.Implementation details
[0058] (1) Baseline: To verify the generality of CVPC and the usability of the overall idea in remote sensing object detection, we selected YOLOv7 as the baseline, used an end-to-end training strategy, and followed the default training and inference settings of YOLOv7. Different from previous YOLO series models, YOLOV7 introduced model re-parameterization into the network architecture, and the idea of re-parameterization first appeared in REPVGG; a training method for the auxiliary head was proposed to increase the training cost and improve the accuracy without affecting the inference time because the auxiliary head only appears during the training process. In addition, the YOLOv7 label assignment strategy adopts the cross-grid search of YOLOV5 and the matching strategy of YOLOX.
[0059] (2) Training Settings: We implemented our network on the public deep learning framework PyTorch. All experiments were conducted on an A6000 GPU with the experimental parameters: epoch = 300, b atchsize = 16, imgsize = 640×640.
[0060] (3) Evaluation metrics: In the ablation experiments, we mainly followed the commonly used object detection evaluation metric - mean average precision (mAP@0.5). In addition, three parameters, Prams (M), GFLOPs (G), and Speed (ms), were used to quantify the efficiency of the model. In the comparative experiments, only the mean average precision (mAP@0.5) metric was used.
[0061] C. Ablation study
[0062] We conducted ablation experiments and comparisons on the effects of the Centralized Visual Processing Center (CVPC), the Centralized Feature Cross-layer Fusion Pyramid, and the Sparse Convolution Detection Head (CEASC) using the DOTA-V1.0 dataset.
[0063] (1) Effect of centralized visual processing center (CVPC): In remote sensing images, the features of the "large and small vehicles" category are easily similar to complex backgrounds and easily blocked by shadows and trees, making it difficult to extract and recognize its semantic features. Therefore, we specifically compared the detection effects of "large and small vehicles". As shown in Table I, in order to verify the effectiveness of Lightweight encoder and PLC, we conducted an ablation experiment on the baseline. After using the Lightweight encoder module alone, the detection accuracy of "large and small vehicles" increased by 2.01% and 2.23% respectively, and the average detection accuracy of all types increased by 1.14%. After using the PLC module alone, the detection accuracy of "large and small vehicles" increased by 2.08% and 2.30% respectively, and the average detection accuracy of all types increased by 1.08%. After using the Lightweight encoder module and the PLC module at the same time, the detection accuracy of "large and small vehicles" increased by 4.12% and 3.98% respectively, and the average detection accuracy of all types increased by 1.88%. It can be seen that the introduction of CVPC effectively reduces the negative impact of high target feature similarity and severe occlusion caused by the complex background of remote sensing images.
[0064] Table I Comparison of CVPC ablation experiments
[0065]
[0066] (2) Effect of centralized feature cross-layer fusion pyramid: In order to verify the effectiveness of centralized feature cross-layer fusion pyramid, we used four-stage feature fusion schemes for ablation experiments. In the first stage, only the results of CVPC were fused with Stage 1, and then one stage of feature fusion was added in sequence to test the effect. As shown in Table II, it can be seen that as the stage of fusion of CVPC results with the following hierarchical features increases, the detection effect becomes better. It can be seen that the introduction of centralized feature cross-layer fusion pyramid effectively increases the accuracy of remote sensing target detection, while having little effect on computational efficiency.
[0067] Table II Comparison of centralized feature cross-layer fusion pyramid ablation experiments
[0068]
[0069]
[0070] (3) Effect of sparse convolution detection head (CEASC): In order to verify the effectiveness of the CEASC detection head, we conducted ablation experiments on the baseline, CEASC and DyHead. As shown in Table III, we used CEASC and DyHead to replace the original YOLOv7 detection head respectively. It can be seen that CEASC effectively balances the detection accuracy and efficiency. While ensuring the detection accuracy, the sparse convolution is used to greatly reduce the operation parameters and improve the detection efficiency.
[0071] Table III Comparison of CEASC ablation experiments
[0072]
[0073] B. Comparisons With State-of-the-Arts
[0074] (1) Experimental results of DOTA dataset: We conducted experiments on the HBB task of the DOTA-v1.0 dataset. We patch-cropped the images in this dataset to expand the data volume and solve the problem of small category distribution. This remote sensing dataset has the problems of large target scale span, dense distribution, high feature similarity, and severe occlusion, but our proposed method can solve these challenging problems well and has achieved good results in all scale categories. At the same time, it is particularly effective for the "large and small vehicles" category that is similar to complex background features and easily obscured by shadows and trees. As can be seen from Table IV, LV and SV have the largest accuracy improvements, which are 3.46% and 3.04% respectively. Figure 5 As shown in Figure 2, we visually compare the model effects.
[0075] Table IV Comparative experiments on the DOTA-v1.0 dataset
[0076]
[0077] (2) Experimental results on the DIOR dataset: We conducted comparative experiments with several advanced horizontal box detectors using the DIOR dataset under the same conditions as when validating the DOTA dataset. The data results are shown in Table V. All experiments used the train and val sets during training and were evaluated on the test set. The experimental results show that our method also achieved good results in various scale categories of this dataset. In particular, the effect on the "vehicle" category, which is easily similar to complex background features and easily blocked by shadows and trees, is obvious. As can be seen from Table V, VE has the largest improvement in accuracy, which is 7.01%. This verifies the robustness and generalization of our method, and our CF-YOLO network is challenging for the latest horizontal detectors and has great room for growth. Figure 6 The comparison chart of the effects is shown.
[0078] Table V Comparative experiments on the DIOR dataset
[0079]
[0080]
[0081] (3) Experimental results on the RSOD dataset: We conducted comparative experiments with several advanced horizontal box detectors using the RSOD dataset under the same conditions as when validating the DOTA dataset, and their results are listed in Table VI. All experiments used the train and val sets during training and were evaluated on the test set. Figure 7 The comparison chart of the effects is shown.
[0082] Table VI Comparative experiments on the RSOD dataset
[0083]
[0084] Compared with related technologies, the remote sensing target detection method based on a centralized vision processing center provided by the present invention has the following beneficial effects:
[0085] In the present invention, we first propose a centralized visual processing center (CVPC), which is a visual processing center with the Transformer encoder module and the CNN module in parallel. The Lightweight encoder is used to capture global long-range dependencies, and the PLC module is used to establish pixel-level correlations, enhance the representation ability of detailed features, and improve the detection accuracy of remote sensing targets in complex backgrounds; the CVPC effectively improves the detection efficiency of remote sensing targets with high feature similarity and severe occlusion. Secondly, the present invention proposes a centralized feature cross-layer fusion pyramid structure, which is fused with the results of the centralized visual processing center in a top-down manner, enhances the detailed expression ability of features at each layer, and improves the detection accuracy of remote sensing targets at various scales. Finally, we introduce a sparse convolution detection head (CEASC), which ensures the detection efficiency while improving the accuracy. The present invention uses YOLOv7 as the baseline and achieves remarkable detection performance on three public remote sensing datasets, namely DOTA-v1.0, DIOR, and RSOD.
[0086] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A remote sensing target detection method based on a centralized visual processing center, characterized in that: The following steps are involved: a. Build a centralized visual processing center (CVPC) that includes a Transformer encoder and a convolutional neural network (CNN) in parallel, where; b. Construct a centralized feature cross-layer fusion pyramid structure and fuse it with the results of the centralized visual processing center (CVPC) to enhance the detail expression ability of each layer of features; c. Introducing sparse convolution detection head (CEASC) to improve detection accuracy while ensuring detection efficiency; The centralized visual processing center (CVPC) includes the Lightweight encoder and the Pixel level Learning Center (PLC) modules. The Lightweight encoder captures the global long-range dependencies, and the Pixel level Learning Center (PLC) module establishes pixel-level correlations to enhance the representation of detailed features. The Lightweight encoder consists of two residual blocks: a neighborhood attention module and a channel MLP-based module. The input of the channel MLP-based module is the output of the neighborhood attention module. Both blocks are followed by Dropout operations to help the network converge better. PLC module: Given a feature map X∈R (h×w)×c1 , the number of feature map channels is reduced to c1 / 2 through the Reshape operation, and then the feature X is fed into a 3×3 dilated convolution layer with a dilation rate of 2 to reconstruct the structural information of the feature map through a large receptive field, and a global average pooling operation is used for the vertical and horizontal spatial directions, and then a matrix multiplication is performed to obtain the position-related feature map M. Finally, M is passed through a 1×1 convolution layer and the Reshape operation is performed again, and the number of channels is restored to C1 to obtain the output feature map; The centralized feature cross-layer fusion pyramid structure uses a top-down approach to fuse high-level semantic information with low-level detail information to improve the accuracy of remote sensing target detection; The sparse convolution detection head (CEASC) adopts an adaptive multi-layer mask strategy to generate the optimal mask ratio. The sparse convolution detection head (CEASC) performs sparse convolution operations to improve the accuracy of target detection while maintaining detection efficiency.
2. The remote sensing target detection method based on a centralized visual processing center according to claim 1 is characterized in that: The pixel level learning center (PLC) module establishes pixel-level correlations through global average pooling and matrix multiplication in the spatial direction to enhance the detection capability of occluded targets in remote sensing images.
Citation Information
Patent Citations
CNN + Transform remote sensing image detection method
CN116977872A