Rotating target detection method based on remote sensing image

By using the combination method of the attitude guide feature acquisition module, the enhanced path feature pyramid network and the feature refinement module in the remote sensing image rotation object detection, the limitations of the prior art when processing large scale differences, dense and rotating targets are solved, and high-precision and robust remote sensing image rotation object detection is achieved.

CN120014436APending Publication Date: 2025-05-16BEIJING UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411991662.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing remote sensing image rotation object detection method has limitations when processing targets with large scale differences, dense targets and rotating targets, making it difficult to achieve high-precision and robust detection.

Method used

The feature extraction backbone composed of the pose-guided feature acquisition module is adopted to extract image features through convolutional serial structure, and combined with the enhanced path feature pyramid network and feature refinement module, the precise detection of the rotation target is achieved.

Benefits of technology

This method can effectively extract the fine-grained features of rotating objects, reduce feature losses, improve detection accuracy and robustness, and adapt to complex remote sensing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014436A_ABST
    Figure CN120014436A_ABST
Patent Text Reader

Abstract

The invention discloses a rotating target detection method based on a remote sensing image, and belongs to the field of computer vision. The method comprises the following steps: firstly, extracting fine-grained features of an RGB remote sensing image by using a feature extraction trunk formed by a target attitude guide feature extraction module to obtain a series of feature map sequences with different spatial resolutions and channel numbers; then, the last three layers of features of the feature map sequences are selected, context information of the features is aggregated by enhancing a path feature pyramid network, and feature loss in paths from bottom to top in the aggregation process is made up; and finally, the feature refining module performs weighted reconstruction and optimization on the spliced features by introducing an attention mechanism to obtain aggregated multi-scale fine-grained features, and sends the aggregated multi-scale fine-grained features into a classification head and a regression head to obtain a final target detection result. According to the method, the competitive precision is achieved in a rotating target detection method of a mainstream remote sensing image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of computer vision, and in particular to a rotating target detection method for remote sensing images. Background Art

[0002] Rotated target detection in remote sensing images refers to regressing the position of the target from the remote sensing image using an angled rotating rectangular box and classifying the target. This task has a wide range of application prospects in urban planning, agriculture, military operations and other fields, and is a very valuable research direction in the field of computer vision. Remote sensing target detection takes a single RGB remote sensing image as input, predicts the areas of interest in the image (such as bridges, vehicles, ships, ports, etc.), and outputs the location and category parameters of these targets. Traditional remote sensing target detection methods mainly rely on manually designed features and machine learning-based classifiers, such as texture analysis, threshold segmentation or edge detection. However, these methods have weak adaptability to complex scenes and often fail to achieve satisfactory results when faced with problems such as target scale changes, background interference and dense arrangement. Deep learning technology represented by convolutional neural networks (CNNs) can automatically learn high-level features in images, significantly improving the accuracy and robustness of target detection. The current mainstream remote sensing target detection framework usually uses a general deep learning backbone network (such as VGG, ResNet or EfficientNet) for feature extraction, and combines it with target detection algorithms (such as Faster R-CNN, YOLO, RetinaNet, etc.) to complete the detection task. However, the unique complexity of remote sensing images makes the existing technology still face many challenges in practical applications. For example, the targets in remote sensing images have significant scale differences. Small targets (such as vehicles and ships) and large targets (such as buildings and bridges) need to be accurately identified at the same time. The targets are densely arranged and the target frame rotation angle is arbitrary. However, conventional feature extraction networks have limitations in processing small targets and dense targets. In addition, the model's multi-scale sensitivity to the target directly affects the accuracy of target detection. The mainstream method to improve the target scale sensitivity mainly relies on the feature pyramid network to interactively transmit scale information. Although this method can improve the network's ability to perceive the scale of the target, the feature loss and degradation caused by the interaction process are also obvious. In summary, there are still many shortcomings in the current technical field. Summary of the invention

[0003] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a rotating target detection method based on remote sensing images.

[0004] The technical solution adopted by the present invention is:

[0005] A method for detecting a rotating target based on a remote sensing image, the method comprising the following steps:

[0006] S1. Extracting features of remote sensing images through a feature extraction backbone composed of a target attitude guided feature acquisition module: Extracting features of a single RGB remote sensing image through a feature extraction backbone composed of an attitude guided feature acquisition module Tp-GFA. The backbone uses a convolutional serial structure to extract image features, and obtains a series of feature map sequences with different spatial resolutions and number of channels.

[0007] S2. According to a series of feature graph sequences extracted in S1, the last three layers of features are selected as the input of the enhanced path feature pyramid network EPFPN; EPFPN adopts a multi-path feature pyramid design to fully aggregate features, and fuses the original input with the output through lateral connections to make up for the loss of original information in the bottom-up path during feature interaction; the lateral connection refers to fusing the original input features with the output features to make up for the loss of original information in the bottom-up path during feature interaction;

[0008] S3: After each splicing in step S2, the features are optimized by feature refinement module RFCA in the form of feature reconstruction to ensure that the features are fully integrated and make up for the feature differences between classes. The optimized features are then sent to the classification head and regression head to obtain the final target detection results.

[0009] Furthermore, the posture-guided feature acquisition module described in step S1 uses the idea of ​​deformable convolution to introduce a learnable offset so that the sampling position of the convolution kernel can be adaptively adjusted to perceive the geometric shape and posture changes of the remote sensing target, so as to adapt to the characteristics of the arbitrary direction of the target in the remote sensing image, and obtain a feature representation rich in the target morphology; then the feature rich in morphological representation is regressed through a multilayer perceptron (MLP) to obtain a direction parameter θ; the direction parameter θ is used to provide a direction indication for the rotation convolution, and further extract the fine-grained features of the target. The two jointly promote the feature extraction capability of the remote sensing target, which can be formulated as follows:

[0010] θ = σ(MLP(MaxPool(F))),

[0011] W'=RotateMatrix(W,θ),

[0012] M h (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))),

[0013] Y(final)=M h (F)⊙Conv2d(X,W'),

[0014] Among them, σ is the activation function sigmoid, MaxPool represents the maximum pooling operation, MLP is a multi-layer perceptron, which is used to regress the direction parameter θ. After the direction parameter θ and the original convolution weight W are mapped by the RotateMatrix rotation matrix, the rotated convolution kernel parameter W' is obtained. F represents the feature map output by the previous level, AvgPool represents average pooling, each channel has a different weight, and the channel weight and the channel attention map are multiplied to form the final attention weight M h (F). Finally, the attention weight M h (F) is used to adjust the importance of different channels. Conv2d represents the standard two-dimensional convolution, and ⊙ represents the element-by-element point multiplication operation. As the training progresses, the network is gradually optimized so that the attention weights are gradually concentrated in the target area, thereby improving the detection accuracy and robustness.

[0015] Furthermore, the enhanced path feature pyramid network (EPFPN) described in step S2 adopts a multi-path design, which effectively fuses high-level semantic information with low-level detail information to enhance the model's ability to recognize multi-scale targets. At the same time, the enhanced path introduced in the network fuses the original input and output to make up for the loss of the original features in the interaction process, which is formulated as follows:

[0016] Pl(top-down)=Upsample(Pl+1)+Cl

[0017] Pl(final)=(Pl(top-down)+Conv2d(Bl))⊕Pl

[0018] Among them, in the top-down path, the feature map Pl+1 of the previous layer is enlarged to the same spatial resolution as the input feature map Cl of the current layer through an upsampling operation, and Upsample(Pl+1) is obtained, and then it is summed pixel by pixel with the underlying feature Cl to obtain the aggregated feature map Pl(top-down) of the top-down path. In the bottom-up path, Bl represents the feature map from the bottom, and a new feature is generated after the convolution operation Conv2d(Bl). Pl(top-down) is added to the convolved Conv2d(Bl) to obtain the final output of the current layer, which is finally spliced ​​with the original feature, and feature aggregation is performed layer by layer in a recursive manner.

[0019] Furthermore, the feature refinement module RFCA described in step S3 adopts the idea of ​​feature reconstruction to reconstruct and refine the overall features after each fusion to generate reconstructed fusion features; the input features are first convolved with a 3x3 kernel to widen the receptive field, and then batch normalization BN and SiLu activation are performed; subsequently, the output feature map is merged with the original input through collaborative convolution; finally, a new output is generated after dimension adjustment.

[0020] After the above steps, the fine-grained features rich in high-level semantic information are regressed and classified respectively through two subsequent convolutional layers. The classification head is responsible for predicting the category probability distribution of each anchor box to determine the category of the object. The regression head also predicts the bounding box parameters of each region or anchor box through the convolutional layer to achieve the regression of the target position. After removing redundant boxes through non-maximum suppression, the final target detection result is obtained.

[0021] Compared with the previous remote sensing target detection method, the innovation of the present invention is that a remote sensing object detection network is proposed to effectively extract the features of rotating objects and reduce feature loss. The present invention combines deformable convolution and rotational convolution, and thus constructs a feature extraction backbone for targets in remote sensing images, and dynamically extracts the features of remote sensing targets in a serial manner, thereby adapting to the characteristics of remote sensing objects to achieve the purpose of fully extracting fine-grained features of remote sensing targets. Specifically, in this serial feature extraction network structure, the constructed attitude-guided feature acquisition module (Tp-GFA) runs through the entire feature extraction process, and the features are serially passed through the attitude-guided feature extraction module. The features are gradually encoded in this process, and the feature representation of the remote sensing object is gradually sufficient and rich. Under ideal conditions, as the number of layers increases, more attitude-guided feature acquisition modules will obtain more comprehensive fine-grained information. The above method can enable the model to show good detection performance when facing remote sensing targets with arbitrary directions, and can better adapt to complex real-world environments.

[0022] The beneficial effects of the present invention are:

[0023] The rotating target detection method based on remote sensing images proposed in the present invention can realize accurate detection of targets in remote sensing scenes. The proposed attitude-guided feature acquisition module can adapt to the target characteristics in remote sensing scenes, so as to stably and accurately obtain fine-grained features of multiple types of remote sensing objects. At the same time, the enhanced path feature pyramid network can effectively compensate for the information loss and degradation in the feature interaction process, thereby further improving the detection accuracy. Finally, the feature refinement module can take into account both inter-class differences and intra-class similarities, generate high-quality aggregate features, and ensure the stability and robustness of the method in multi-target scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings on the premise of creative work.

[0025] Figure 1A schematic diagram of the architecture of a remote sensing rotating target detection network according to an embodiment of the present invention;

[0026] Figure 2 A structural diagram of a feature acquisition module guided by posture according to an embodiment of the present invention;

[0027] Figure 3 Schematic diagram of visualization results according to an embodiment of the present invention.

[0028] Figure 4 Flow chart for implementing this method. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.

[0030] According to a rotating target detection method based on remote sensing images according to an embodiment of the present invention, in practical applications, Figure 1 As shown, the network structure of the present invention is deployed, including:

[0031] A serial backbone acquisition network combined with a target posture guided feature acquisition module takes a single image as input and generates a total of five different levels of features C1, C2, C3, C4, and C5. The last three layers C3, C4, and C5 are passed to EPFPN for feature aggregation.

[0032] A feature fusion network with reduced feature loss is used to enhance the detection of multi-scale objects while reducing the loss and degradation of features during interaction.

[0033] A feature refinement mechanism refines the feature representation of the integrated features as a whole through feature reconstruction, alleviates inter-class differences, generates high-quality aggregated features, and ensures stability and robustness in multi-target scenarios.

[0034] In order to facilitate understanding of the above technical solutions of the present invention, the above technical solutions of the present invention are described in detail below through actual deployment and application.

[0035] The model-based remote sensing rotation target detection task is specified as a tuple α=(x, y, w, h, θ), where x and y represent the center point position of the detection frame, w and h represent the length and width of the detection frame, and θ represents the rotation angle of the detection frame. The tuple α constitutes a five-parameter regression paradigm for remote sensing rotation target detection. Compared with the four-parameter regression paradigm b=(x, y, w, h) of remote sensing horizontal target detection, there is an additional rotation angle θ of the detection frame to achieve more accurate positioning of the remote sensing target. The paradigm α is more common in remote sensing target detection, and the present invention is also based on the five-parameter representation paradigm. For a detection task with N categories, the tuple γ=(n1, n2, ..., nN), where n1, n2, ..., nN represents the probability values ​​of the N categories of the current detection frame, so that M(α, γ) constitutes a complete representation of a single target. For an end-to-end remote sensing target detection network, the final output feature of the network is (B, C, W, H). After the classification network branch and the regression network branch, the feature dimension is converted from (B, C, W, H) to (B, γ, W, H) and (B, α, W, H), where B represents the batch size, B represents the number of network channels, and W and H represent the length and width of the feature map.

[0036] In order to design an end-to-end trainable model, the idea of ​​existing deformable convolution and rotation convolution is used to obtain fine-grained features of remote sensing targets from highly complex remote sensing images. The target posture guided feature acquisition module (Tp-GFA) plays a key role in this process. Specifically, the backbone network composed of this module gradually extracts the representation of the image from low resolution to high resolution and marks it as (C1, C2, C3, C4, C5). Each Tp-GFA module in the network has the same network structure. The input of the next stage is the output of the previous stage. As the number of layers increases, the feature information is richer and the semantic information is more advanced. The final output Ci = (B, Ci, Wi, Hi), i = 1, 2, ..., 5, where i represents the level number. The target posture guided feature acquisition module proposed in the example of the present invention can be defined as:

[0037] Y(p 0 )=∑(w(p n )X(p 0 +p n +Δp n ))

[0038] where p 0 represents the position on the output feature map, p n is the regular sampling position of the convolution kernel, Δp n is the learned offset, w(p n) is the weight of the convolution kernel, and X is the input feature map. This step is used to learn the offset and apply it to the feature map to perceive the shape and posture of remote sensing objects in any direction. Then it is necessary to provide direction guidance for the rotation convolution and obtain the corresponding rotation convolution kernel parameters, which can be expressed as:

[0039] W'=RotateMatrix(W,σ(MLP(MaxPool(Y(p 0 )))))

[0040] Among them, σ is the activation function sigmoid, MaxPool represents the maximum pooling operation, MLP is a fully connected layer, which is used to regress the direction parameters. After the reverse parameters and the original convolution parameters are rotated by the RotateMatrix matrix, the rotated convolution kernel parameters W' are obtained. Then, in order to highlight the most important feature areas, the channel attention optimization layer is introduced to try to learn the weights between different channels to distinguish the importance of different channels. Each channel represents a different image feature, which enhances the network's attention to important channels. The whole process can be expressed as follows:

[0041] M h (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))

[0042] Where F represents the feature map output by the previous level, AvgPool represents average pooling, each channel has a different weight, and the weight is multiplied by the channel attention map to form the final attention weight M h .

[0043] Y(final)=M h (F)⊙Conv2d(X,W')

[0044] Finally, the attention weight M h (F) is used to adjust the importance of different channels. Conv2d represents the standard two-dimensional convolution, and ⊙ represents the element-by-element point multiplication operation. As the training progresses, the network is gradually optimized so that the attention weights are gradually concentrated in the target area, thereby improving the detection accuracy and robustness.

[0045] Although the backbone feature extraction network mentioned above is sufficient and robust for the feature extraction of remote sensing targets with arbitrary directions, focusing only on the feature representation of remote sensing targets is not enough to achieve satisfactory predictions, because although these features contain fine-grained high-level semantic information, they lack the scale information of the target. At the same time, it is also unable to solve the prediction errors caused by inter-costal differences and intra-class similarities, thus creating an accuracy bottleneck for remote sensing target detection.

[0046] Therefore, the present invention designs an enhanced path feature pyramid network to fuse features of different scales and levels, enhance the expression ability of the model, and perceive the scale differences and changes of the target. Figure 1 As shown, the present invention proposes a feature aggregation network with an enhancement path for fusing multi-scale features in an image. The network contains multiple convolutional layers, through which low-level and high-level features of the image are extracted. In order to better fuse these features, the present invention introduces an enhancement path, which enables the network to effectively integrate feature information at different levels through a series of jump connections and fusion modules. Each jump connection corresponds to the original feature of a level, which not only retains the detailed information of the image, but also enhances the network's adaptability to scale changes. The aggregation process can be expressed as:

[0047] Pl(top-down)=Upsample(Pl+1)+Cl

[0048] In the top-down path, the feature map Pl+1 of the previous layer is enlarged to the same spatial resolution as the input feature map Cl of the current layer through the upsampling operation, and Upsample(Pl+1) is obtained. Then, it is summed pixel by pixel with the underlying feature Cl to obtain the aggregated feature map Pl(top-down) of the top-down path. Next, in the bottom-up path, Bl represents the feature map from the bottom, which generates new features after the convolution operation Conv(Bl). Pl(top-down) is summed with the convolutional Conv(Bl) to obtain the final feature map Pl(min) of the current layer. The formula can be expressed as:

[0049] Pl(min)=Pl(top-down)+Conv2d(Bl)

[0050] In addition, if Figure 1 As shown by the red line, the present invention designs an enhancement path strategy to assist feature fusion and compensate for the loss and degradation of features during the interaction process. The defined enhancement path feature aggregation method can be expressed as:

[0051] Pl(final)=(Pl(top-down)+Conv2d(Bl))⊕Pl

[0052] Ultimately, through the flow of information in this bidirectional path, feature aggregation is performed layer by layer in a recursive manner, which includes both the top-down fusion information from the previous layer and the bottom-up enhancement information from the lower layer, ensuring the effective expression of multi-scale features and compensating for the loss and degradation of the original features during the interaction process, outputting more accurate target detection results.

[0053] For the above aggregation network, each time the feature is spliced, it will be refined by the feature refinement module to achieve the purpose of repeatedly aggregating context information and focusing on spatial features. In the feature refinement module proposed in the present invention, the input feature is first convolved with a 3x3 kernel to widen the receptive field, and then batch normalization BN and SiLu activation are performed. Subsequently, the output feature map is merged with the original input through collaborative convolution to generate a new output Y. The whole process can be expressed as follows:

[0054] Ymid=RFCAConv(SiLu(BN(Convi×i(X)))),i=3

[0055] Y=SiLu(BN(Convi×i(Ymid))⊕X),i=1

[0056] Where ⊕ represents a connection and Y represents the output feature. BN and SiLu represent batch normalization and sigmoid weighted linear unit SiLU activation operations respectively. RFCAConv can be expressed as follows:

[0057] Arf=Softmax(Ad(gi×i(AvgPool((X))))

[0058] Frf=Ad(ReLU(Norm(gk×k(X))))

[0059] F=Arf×Frf

[0060] Among them, X represents the input feature, F represents the output feature, gi×i represents the group convolution of size i×i, k represents the size of the convolution kernel, Norm represents normalization, and Ad represents the reshape operation, as shown below:

[0061] Y=Ad(X)

[0062] Where X represents the input feature and Y represents the reconstructed output feature. For the input feature dimension C×H×W, the input feature X undergoes group convolution with a convolution kernel size of K×K to obtain a feature with a dimension of CK2×H×W. After dimensionality adjustment, the dimension of X is converted from CK2×H×W to C×KH×KW. This process assigns different weights to feature maps and implements feature reconstruction to refine features, thereby helping the network filter and retain key information, ensuring that context information fully interacts and reducing feature differences between classes, significantly improving the accuracy and robustness of target detection.

[0063] After the above steps, the fine-grained features rich in high-level semantic information are regressed and classified respectively through two subsequent convolutional layers. First, the feature map passes through the classification head, and the convolutional layer is used to predict the category probability distribution of each region or anchor box to determine the category of the object. The regression head also predicts the bounding box parameters of each region or anchor box through the convolutional layer to achieve the regression of the target position. Finally, the classification and regression results are combined, and the redundant boxes are removed through non-maximum suppression to obtain the final target detection result.

[0064] In order to better illustrate the performance of the rotating target detection algorithm based on remote sensing images provided by the present invention, a comparative experiment is conducted on the public remote sensing dataset DOTA with respect to the existing network model method. The experimental results are shown in Table 1.

[0065] The DOTA dataset is an open dataset designed for the task of rotating object detection in remote sensing images. The dataset contains 2,806 images, 188,282 instances, and 15 different categories. The objects in the images show various scales, directions, and shapes;

[0066] Five representative categories (SBF-football field, RA-roundabout, HA-port, SP-swimming pool, HC-helicopter) with large aspect ratio or densely arranged features were selected from 22 different methods and their average precision (AP) and overall mean average precision (mAP) were presented.

[0067] Table 1 Comparison results of different models on the DOTA dataset

[0068]

[0069]

[0070] In summary, with the help of the above technical solution. The present invention proposes a rotating target detection network based on remote sensing images. The basic idea is to use the target posture to guide feature acquisition, enrich the fine-grained feature representation of the target, and improve the adaptability of the network in the remote sensing scene, so as to fully extract the remote sensing target features with arbitrary direction characteristics. In addition, in order to alleviate the common problem of information loss and degradation in the feature aggregation process, the present invention designs an enhanced path feature golden network. While ensuring the full fusion of multi-scale information, it makes up for the feature dilution caused by upsampling and the interpolation operation of downsampling, which will cause the loss and degradation of semantic information, prompting the network to repeatedly aggregate and retain high-level semantic information. Finally, the present invention proposes a feature refinement module, which uses the idea of ​​feature reconstruction to ensure that the context information is fully interactive and make up for the feature differences between classes. This network structure can fully aggregate features and reduce feature loss, significantly improving the accuracy and robustness of target detection. Overall, in most cases, the method proposed in the present invention has outstanding performance in the detection results of targets in high-resolution and wide-coverage remote sensing images, and can accurately obtain the position and category of objects in remote sensing images.

[0071] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are included in the protection scope of the present invention.

Claims

1. A rotating target detection method based on remote sensing images, characterized in that: The method comprises the following steps: S1. Extracting features of remote sensing images through a feature extraction backbone composed of a target attitude guided feature acquisition module: Extracting features of a single RGB remote sensing image through a feature extraction backbone composed of an attitude guided feature acquisition module Tp-GFA. The backbone uses a convolutional serial structure to extract image features, and obtains a series of feature map sequences with different spatial resolutions and number of channels. S2. According to a series of feature graph sequences extracted in S1, the last three layers of features are selected as the input of the enhanced path feature pyramid network EPFPN; EPFPN adopts a multi-path feature pyramid design to fully aggregate features, and fuses the original input with the output through lateral connections to make up for the loss of original information in the bottom-up path during feature interaction; the lateral connection refers to fusing the original input features with the output features to make up for the loss of original information in the bottom-up path during feature interaction; S3: The features after each splicing in step S2 are weighted reconstructed and optimized through the feature refinement module RFCA that introduces the attention mechanism to ensure that the features are fully integrated and make up for the feature differences between classes. The optimized features are then sent to the classification head and regression head to obtain the final target detection results.

2. The method for detecting rotating targets based on remote sensing images according to claim 1, characterized in that: The posture-guided feature acquisition module described in step S1 uses the idea of ​​deformable convolution to introduce a learnable offset so that the sampling position of the convolution kernel can be adaptively adjusted to perceive the geometric shape and posture changes of the target in the remote sensing image, and obtain a feature representation rich in the target morphology; the feature rich in the target morphology is regressed through a multi-layer perceptron MLP to obtain a direction parameter; the direction parameter is used to provide a direction indication for the rotation convolution RCN, and further extract the fine-grained features of the target in the remote sensing image, which is formally expressed as: θ = σ(MLP(MaxPool(F))), W'=RotateMatrix(W,θ), M h (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))), Y(final)=M h (F)⊙Conv2d(X,W'), Among them, σ is the activation function sigmoid, MaxPool represents the maximum pooling operation, MLP is a multi-layer perceptron, which is used to regress the direction parameter θ. After the direction parameter θ and the original convolution weight W are mapped by the RotateMatrix rotation matrix, the rotated convolution kernel parameter w' is obtained; F represents the feature map output by the previous level, AvgPool represents average pooling, each channel has a different weight, and the channel weight and the channel attention map are multiplied to form the final attention weight M h (F); Attention weight M h (F) is used to adjust the importance of different channels. Conv2d represents the standard two-dimensional convolution, and ⊙ represents the element-by-element point multiplication operation. As the training progresses, the enhanced path feature pyramid network is gradually optimized, so that the attention weights are gradually concentrated in the target area, thereby improving the detection accuracy and robustness.

3. The rotating target detection method based on remote sensing images according to claim 1 is characterized in that: The enhanced path feature pyramid network described in S2 adopts a multi-path design and effectively integrates high-level semantic information with low-level detail information to promote the enhanced path feature pyramid network's ability to recognize multi-scale targets; the enhanced path introduced in the enhanced path feature pyramid network fuses the original input with the input to compensate for the loss of original information of the original features during the interaction process, which is formulated as follows: Pl(top-down)=Upsample(Pl+1)+Cl Among them, in the top-down path, the feature map Pl+1 of the previous layer is enlarged to the same spatial resolution as the input feature map Cl of the current layer through an upsampling operation, and Upsample(Pl+1) is obtained, and then it is summed pixel by pixel with the underlying feature Cl to obtain the aggregated feature map Pl(top-down) of the top-down path. In the bottom-up path, Bl represents the feature map from the bottom, and a new feature is generated after the convolution operation Conv2d(Bl). Pl(top-down) is added to the convolved Conv2d(Bl) to obtain the final output of the current layer, which is finally spliced ​​with the original feature, and feature aggregation is performed layer by layer in a recursive manner.

4. The method for detecting rotating targets based on remote sensing images according to claim 1, characterized in that: The feature refinement module RFCA described in step S3 reconstructs and refines the overall features after each remote sensing image splicing using the idea of ​​feature reconstruction to generate reconstructed fusion features; the input features are first convolved with a 3x3 kernel to widen the receptive field, and then batch normalization BN and SiLu activation are performed; then, the output feature map is merged with the original input through collaborative convolution; finally, a new remote sensing image output is generated after dimension adjustment; After the above step S3, the fine-grained features rich in high-level semantic information are regressed and classified respectively through two subsequent convolutional layers; the classification head is responsible for predicting the category probability distribution of each anchor box, thereby determining the category of the object in the remote sensing image; The regression head also predicts the bounding box parameters of each region or anchor box through the convolution layer to achieve the regression of the target position; after removing redundant boxes through non-maximum suppression, the target detection result of the final remote sensing image is obtained.

Citation Information

Cited By

  • Unmanned aerial vehicle target detection method based on feature fusion DINO

    CN120544081A

  • Feature enhancement method and system for sea target detection and related equipment

    CN121391645A