A method for omnidirectional object detection in remote sensing images based on multi-layer feature interactive pyramid and lightweight enhanced detection head

The omnidirectional target detection method for remote sensing images based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head solves the problem of imbalance between accuracy and complexity in rotation target detection in remote sensing images, and achieves efficient and accurate rotation target detection.

CN120510357BActive Publication Date: 2025-09-30SHIJIAZHUANG TIEDAO UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510616306.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-30
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Existing remote sensing image rotation target detection methods have difficulty in achieving accurate detection when processing rotating targets and have high computational complexity, making it difficult to balance detection accuracy and model complexity.

Method used

An omnidirectional target detection method for remote sensing images is proposed, which adopts a multi-layer feature interactive pyramid and a lightweight enhanced detection head. By aggregating and fusion multi-layer feature maps, combined with lightweight enhanced convolution and Lamp pruning methods, the model complexity is reduced, and the KLD divergence loss function is introduced to improve the detection accuracy.

Benefits of technology

The proposed method achieves efficient and accurate detection of rotating targets in remote sensing images, reduces model complexity and computational complexity, and improves the feature extraction capability and detection accuracy of rotating targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510357B_ABST
    Figure CN120510357B_ABST
Patent Text Reader

Abstract

The present invention discloses an omnidirectional target detection algorithm for remote sensing images based on a multi-layer feature interaction pyramid and a lightweight enhanced detection head, which belongs to the field of computer vision. The method comprises the following steps: 1. Preprocessing a remote sensing image dataset. 2. Building a remote sensing image omnidirectional target detection model: designing a multi-layer feature interaction pyramid, obtaining intermediate feature maps and fused feature maps by aggregating multi-layer feature maps, realizing cross-layer feature fusion, avoiding information interaction being limited between adjacent layers, and generating rotation-invariant feature representations to enhance the feature information of rotated targets; constructing a lightweight enhanced detection head, adopting a shared enhanced convolution strategy to reduce the number of model parameters, and extracting feature information through central difference convolution, so that the model can capture the detailed features of rotated targets; adopting the Lamp pruning method to reduce the complexity of the model without losing accuracy. 3. Constructing a loss function for the model, introducing the KLD divergence loss function to solve the problem of periodic angle changes in rotated target detection. 4. Iteratively training the model until the model reaches convergence and obtains the optimal weight. 5. The obtained optimal weight is tested on the test set to obtain an evaluation result. The method reduces the number of model parameters and computational complexity while ensuring the accuracy of rotating target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote sensing image omnidirectional target detection method based on a multi-layer feature interaction pyramid and a lightweight enhanced detection head, and belongs to the field of computer vision. Background Art

[0002] With the rapid development of deep learning, many excellent horizontal bounding box object detection algorithms have emerged. These methods have achieved promising results in object detection tasks. These methods use horizontal bounding boxes to delineate and classify objects. However, these methods struggle to effectively model directional changes when dealing with arbitrarily oriented objects. This results in horizontal bounding box-based methods being unable to provide accurate detection results when handling rotated objects.

[0003] To this end, many researchers have moved beyond simple horizontal bounding boxes and begun exploring methods for detecting rotated objects in remote sensing images. These methods not only provide more precise object localization but also effectively reduce background interference, improving detection accuracy. Dong et al. combined a shape-aware label assignment strategy with a fine-grained feature alignment module to detect objects in remote sensing images with arbitrary orientations. The fine-grained feature alignment module adaptively aligns features by refining anchor points to adjust the sampling positions of convolutional kernels. Zheng et al. proposed an object-level rotation-invariant semantic representation framework that synergistically integrates the exploration of latent supervision, rotation-invariant learning, and guided attention mechanisms into a unified network to improve the detection performance of rotated objects in remote sensing images. Xie et al. proposed a dual-focus detector that simultaneously focuses on the exploration of contextual knowledge and the mitigation of angle sensitivity. Luo et al. learned the scale and orientation of objects from three different views and designed the SSC loss to enhance the network's ability to perceive object scale. Furthermore, they introduced a scale-guided DS matching strategy to improve the accuracy of object angle prediction. Dai et al. proposed a Transformer-based object detection framework, AO2-DETR, specifically for detecting objects in arbitrary orientations. It improves detection efficiency and accuracy through a directional proposal generation mechanism, an adaptive proposal refinement module, and a rotation-aware matching loss. Hou et al. proposed a new flexible shape adaptive selection and shape adaptive measurement strategy for target detection, including the SA-S strategy for sample selection and the SA-M strategy for positive sample quality estimation. However, these methods improve the accuracy of rotated target detection at the expense of increasing the computational complexity of the model. At the same time, existing models have difficulty extracting rotation-invariant features of targets in different directions. There is an urgent need to design an algorithm that strikes a good balance between detection accuracy and model computational complexity to efficiently meet the needs of rotated target detection in remote sensing images. Summary of the Invention

[0004] To this end, the present invention provides a remote sensing image omnidirectional target detection method based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head.

[0005] The present invention is implemented by the following scheme:

[0006] Step 1: Preprocess the remote sensing image dataset;

[0007] Step 2: Establish an omnidirectional target detection model for remote sensing images: Design a multi-layer feature interaction pyramid to aggregate multi-layer feature maps to obtain intermediate feature maps and fused feature maps, and perform path interaction through upsampling and downsampling to achieve cross-layer feature fusion and generate rotation-invariant feature representations; construct a lightweight enhanced detection head, adopt a shared enhanced convolution strategy to reduce the number of model parameters, and extract feature information through central difference convolution, so that the model can capture the detailed features of rotated targets; use the lamp pruning method to prune the weights with the lowest lamp score in each layer to reduce the complexity of the model;

[0008] Step 3: Construct a loss function to update the model weights. The KLD divergence loss is used as the regression loss function to measure the difference in rotation angle between the predicted box and the true box, thereby improving the rotation intersection-over-union ratio between the predicted box and the true box.

[0009] Step 4: Iterate the model training until the model reaches convergence and obtains the optimal weight;

[0010] Step 5: Use the obtained optimal weight to detect the test set images and obtain the final test results.

[0011] Furthermore, in step 1, the DOTA remote sensing dataset was cropped into 1024×1024 pixel blocks, with a 200-pixel overlap between adjacent blocks to minimize information loss. If the original image is smaller than 1024×1024, it is padded with zeros. DOTA covers 15 common object categories, including airplanes (PL), baseball fields (BD), bridges (BR), track and field fields (GTF), small vehicles (SV), large vehicles (LV), ships (SH), tennis courts (TC), basketball courts (BC), storage tanks (ST), football fields (SBF), roundabouts (RA), ports (HA), swimming pools (SP), and helicopters (HC).

[0012] Furthermore, the step 2 establishes a remote sensing image target detection model: a multi-layer feature interaction pyramid is used to aggregate multi-layer feature maps as an intermediate hub for feature interaction, thereby realizing cross-layer feature fusion, avoiding information interaction being limited between adjacent layers, and generating rotation-invariant feature representations; a lightweight enhanced detection head is used to reduce the number of parameters and computational complexity of the model while enhancing feature expression capabilities; and a Lamp pruning method is used to remove unimportant connections, thereby reducing the complexity of the model without sacrificing accuracy.

[0013] Furthermore, the multi-layer feature interaction pyramid takes the C3, C4, and C5 feature maps generated by the backbone network as input. First, the feature map C3 is downsampled to obtain the F1' feature map of the same size as the C4 feature map. The feature map C5 is upsampled to obtain the F3' feature map of the same size as the C4 feature map. The feature maps F1', C4, and F3' are concatenated to obtain the intermediate feature map F m , serving as the intermediate hub for subsequent feature interactions. This design enables cross-layer feature fusion, retaining shallow detail information while integrating deep semantic information, enhancing the expressive power of features. The process is defined as follows:

[0014] F1'=Downsample(C3)

[0015] F3'=Upsample(C5)

[0016] F m =Concat(F1',C4,F3')

[0017] Furthermore, the intermediate feature map F m Downsampling is performed and feature concatenation is performed with the C5 feature map to obtain feature map f2, which enhances the richness of deep semantic expression. Simultaneously, the new feature map obtained by upsampling the intermediate feature map is concatenated with the C3 feature map to obtain feature map f1, enhancing the representation of detailed information. This interaction between upsampling and downsampling paths helps generate high-quality feature representations. The process definition is as follows:

[0018] f1=Concat(Upsample(F m ),C3)

[0019] f2=Concat(Downsample(F m ),C5)

[0020] Furthermore, F m , f1 and f2 are spliced ​​to generate a fusion feature map f m The fused feature map contains feature information at different scales. The intermediate and fused feature maps are then upsampled and concatenated with the generated f1 feature map to obtain the feature map f1', enhancing spatial resolution and detail expression. At the same time, the intermediate and fused feature maps are downsampled and concatenated with the generated f2 feature map to obtain the feature map f2', highlighting the abstraction capability of semantic information.

[0021] Furthermore, the output feature map fully utilizes the details and semantic information of the feature maps of each layer. This module implements the interaction and fusion of feature maps of different resolutions to enhance the global semantic expression capability while preserving detailed information, breaking the limitations of feature interaction and enriching the information flow path. The process definition is as follows:

[0022] f m =Concat(Downsample(f1),F m ,Upsample(f2))

[0023] f1'=Concat(Upsample(F m ),f1,Upsample(f m ))

[0024] f2'=Concat(Downsample(F m ),f2,Downsample(f m ))

[0025] Furthermore, the lightweight enhanced detection head applies 1×1 convolution to the output of the multi-layer feature interaction pyramid to adjust the number of channels. Then, a weight-sharing strategy is used to perform enhanced convolution operations. This not only effectively learns image features but also reduces the number of parameters and complexity of the model. The enhanced convolution consists of two parallel convolution layers, including standard convolution and center difference convolution. The calculation process of the enhanced convolution is as follows:

[0026] F out =EConv(F in )=F in *(K VC +K CDC )

[0027] Among them, F in is the given input feature, EConv() represents the enhanced convolution operation, K VC represents the convolution kernel of the standard convolution, K CDC The kernel represents the center difference convolution. The center difference convolution consists of two steps: sampling and aggregation. Sampling involves sampling a pixel block of the same size as the convolution kernel from the input feature map. Aggregation involves subtracting the center element from each element in the sampled pixel block, and then performing a standard convolution to obtain the output. Finally, the shared convolutional features are used to predict the bounding box position and the probability of the object class, respectively, to generate the final detection result.

[0028] Furthermore, the Lamp pruning treats each neural network layer as an operator and determines the impact of pruning on the model output. The square of the target connection weight is calculated and normalized by the sum of the squares of all surviving weights in the same layer. Lamp prunes the weight with the lowest Lamp score in each layer until the pruning criteria are met and the required global sparsity constraint is achieved. This method ensures that at least one weight is retained in each layer while retaining the most important weights. The specific scoring formula of Lamp is as follows:

[0029]

[0030] Where W[u] represents the weight tensor W mapped by index u. (W[u]) 2 Indicates the size of the weight. ∑ v≥u (W[v]) 2 The cumulative sum of all squares of weights is used to normalize the weights. u and v are sorted in ascending order to represent the index mapping corresponding to the weights.

[0031] Furthermore, step 3 constructs a rotation target detection loss function: KLD divergence loss is used as a regression loss function for rotation target detection to measure the difference in rotation angle between the predicted box and the true box, thereby improving the rotation intersection-over-union ratio between the predicted box and the true box. KLD divergence loss converts the rotated bounding box into a two-dimensional Gaussian distribution, and uses the rotation matrix and scale factor to map the rotated box parameters to the mean and covariance matrix of the Gaussian distribution. By calculating the Gaussian distance between the predicted box and the true box, KLD divergence loss can more stably measure the difference between the true value and the predicted value, solve the angle mutation problem of rotation target detection, and achieve more accurate target detection.

[0032] Furthermore, the step 4: iteratively trains the model until the model reaches convergence, thereby obtaining the optimal weight.

[0033] Furthermore, the step 5: using the obtained optimal weight to detect the test set images to obtain the final test results.

[0034] Beneficial effects of the present invention:

[0035] The present invention proposes a method for omnidirectional target detection in remote sensing images based on a multi-layer feature interaction pyramid and a lightweight enhanced detection head; the present invention designs a multi-layer feature interaction pyramid, obtains intermediate feature maps and fused feature maps by aggregating multi-layer feature maps, realizes cross-layer feature fusion, avoids information interaction being limited between adjacent layers, generates rotation-invariant feature representation, and enhances the feature information of rotated targets; the present invention constructs a lightweight enhanced detection head, adopts shared convolution to reduce the number of model parameters, and extracts feature information through central difference convolution, so that the model can capture detailed features, thereby improving the feature extraction capability of rotated targets; the present invention adopts the Lamp pruning method to reduce the complexity of the model while maintaining the original performance; the present invention introduces the KLD divergence loss function to solve the problem of angle periodicity in rotating target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0037] Figure 1 This is a flowchart of the overall implementation of an embodiment of the present invention;

[0038] Figure 2 A framework diagram of the overall network model in an embodiment of the present invention;

[0039] Figure 3 Schematic diagram of a multi-layer feature interaction pyramid structure in an embodiment of the present invention;

[0040] Figure 4 This is a schematic diagram of the structure of a lightweight enhanced detection head in an embodiment of the present invention;

[0041] Figure 5 Schematic diagram of the Lamp pruning process in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The technical solutions in the embodiments of the present invention will be further described below with reference to the accompanying drawings in the embodiments of the present invention, but are not intended to limit the present invention.

[0043] like Figure 1 As shown, the present invention provides a method for omnidirectional target detection in remote sensing images based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head, the steps of which are as follows:

[0044] Step 1: Preprocess the remote sensing image dataset;

[0045] Step 2: Establish an omnidirectional target detection model for remote sensing images: Design a multi-layer feature interaction pyramid to aggregate multi-layer feature maps to obtain intermediate feature maps and fused feature maps, and perform path interaction through upsampling and downsampling to achieve cross-layer feature fusion and generate rotation-invariant feature representations; construct a lightweight enhanced detection head, adopt a shared enhanced convolution strategy to reduce the number of model parameters, and extract feature information through central difference convolution, so that the model can capture the detailed features of rotated targets; use the lamp pruning method to prune the weights with the lowest lamp score in each layer to reduce the complexity of the model;

[0046] Step 3: Construct a loss function to update the model weights. KLD divergence loss is used as the regression loss function for rotated object detection. It is used to measure the difference in rotation angle between the predicted box and the true box, and to improve the rotation intersection-over-union ratio between the predicted box and the true box.

[0047] Step 4: Iterate the model training until the model reaches convergence and obtains the optimal weight;

[0048] Step 5: Use the obtained optimal weight to detect the test set images and obtain the final test results.

[0049] Furthermore, in step 1, the DOTA remote sensing dataset was cropped into 1024×1024 pixel blocks, with a 200-pixel overlap between adjacent blocks to minimize information loss. If the original image is smaller than 1024×1024, it is padded with zeros. DOTA covers 15 common object categories, including airplanes (PL), baseball fields (BD), bridges (BR), track and field fields (GTF), small vehicles (SV), large vehicles (LV), ships (SH), tennis courts (TC), basketball courts (BC), storage tanks (ST), football fields (SBF), roundabouts (RA), ports (HA), swimming pools (SP), and helicopters (HC).

[0050] Furthermore, the step 2 establishes a remote sensing image target detection model: a multi-layer feature interaction pyramid is used to aggregate multi-layer feature maps as an intermediate hub for feature interaction, thereby realizing cross-layer feature fusion, avoiding information interaction being limited between adjacent layers, and generating rotation-invariant feature representations; a lightweight enhanced detection head is used to reduce the number of parameters and computational complexity of the model while enhancing feature expression capabilities; and a Lamp pruning method is used to remove unimportant connections, thereby reducing the complexity of the model without significantly sacrificing accuracy.

[0051] like Figure 2As shown in the figure, the remote sensing image rotation target detection model: first, the input image is extracted through the backbone network to generate multi-scale feature maps C1 to C5. Secondly, the multi-layer feature interaction feature pyramid performs cross-layer feature fusion, combines deep semantic information with shallow detail information to achieve information interaction, and generates rotation-invariant feature representations. Then, a lightweight enhancement detection head is used to predict the position and category of the target, and redundant bounding boxes are removed by non-maximum suppression. At the same time, the KLD divergence loss function is introduced to solve the problem of periodic changes in angles and achieve a stable training process. Finally, the Lamp pruning method is used to optimize the model, reducing the network parameters while maintaining the original performance and reducing the network volume and computational complexity.

[0052] like Figure 3 As shown in FIG, the multi-layer feature interaction pyramid takes the C3, C4, and C5 feature maps generated by the backbone network as input. The feature map C3 is downsampled to obtain the F1' feature map of the same size as the C4 feature map. The feature map C5 is upsampled to obtain the F3' feature map of the same size as the C4 feature map. The feature maps F1', C4, and F3' are concatenated to obtain the intermediate feature map F m , serving as the intermediate hub for subsequent feature interactions. This design enables cross-layer feature fusion, retaining shallow detail information while integrating deep semantic information, enhancing the expressive power of features. The process is defined as follows:

[0053] F1'=Downsample(C3)

[0054] F3'=Upsample(C5)

[0055] F m =Concat(F1',C4,F3')

[0056] The intermediate feature map F m Downsampling is performed and feature concatenation is performed with the C5 feature map to obtain feature map f2, which enhances the richness of deep semantic expression. Simultaneously, the new feature map obtained by upsampling the intermediate feature map is concatenated with the C3 feature map to obtain feature map f1, enhancing the representation of detailed information. This interaction between upsampling and downsampling paths helps generate high-quality feature representations. The process definition is as follows:

[0057] f1=Concat(Upsample(F m ),C3)

[0058] f2=Concat(Downsample(F m ),C5)

[0059] F m, f1 and f2 are spliced ​​to generate a fusion feature map f m The fused feature map contains feature information at different scales. The intermediate and fused feature maps are then upsampled and concatenated with the generated f1 feature map to obtain the feature map f1', enhancing spatial resolution and detail expression. At the same time, the intermediate and fused feature maps are downsampled and concatenated with the generated f2 feature map to obtain the feature map f2', highlighting the abstraction capability of semantic information.

[0060] The output feature map fully utilizes the details and semantic information of each layer's feature map. This module implements the interaction and fusion of feature maps of different resolutions to enhance the global semantic expression capability while preserving detailed information, breaking the limitations of feature interaction and enriching the information flow path. The process definition is as follows:

[0061] f m =Concat(Downsample(f1),F m ,Upsample(f2))

[0062] f1'=Concat(Upsample(F m ),f1,Upsample(f m ))

[0063] f2'=Concat(Downsample(F m ),f2,Downsample(f m ))

[0064] like Figure 4 As shown in the figure, the lightweight enhanced detection head uses 1×1 convolution processing on the output of the neck to adjust the number of channels; then, a shared weight strategy is used to perform enhanced convolution operations; this not only effectively learns the features of the image, but also reduces the number of parameters and complexity of the model. The enhanced convolution consists of two parallel convolution layers, which include standard convolution and center difference convolution. The center difference convolution first samples a pixel block with the same size as the convolution kernel in the input feature map. This process is called sampling. Then, the center element is subtracted from each element in the sampled pixel block, and the output is obtained. This process is called aggregation. The features learned by the two convolution layers are added to obtain the output of the enhanced convolution. The calculation process of the enhanced convolution is as follows:

[0065] F out =EConv(F in )=F in *(K VC +K CDC )

[0066] Among them, F inis the given input feature, EConv() represents the enhanced convolution operation, K VC represents the convolution kernel of the standard convolution, K CDC The kernel represents the center difference convolution. The center difference convolution consists of two steps: sampling and aggregation. Sampling involves sampling a pixel block of the same size as the convolution kernel from the input feature map. Aggregation involves subtracting the center element from each element in the sampled pixel block, and then performing a standard convolution to obtain the output. Finally, the shared convolutional features are used to predict the bounding box position and the probability of the object class, respectively, to generate the final detection result.

[0067] like Figure 5 As shown in Figure 1, the Lamp pruning algorithm treats each neural network layer as an operator and determines the impact of pruning on the model output. This is done by calculating the square of the target connection weight and normalizing it by the sum of the squares of all "surviving weights" in the same layer. The specific scoring formula for Lamp is as follows:

[0068]

[0069] Where W[u] represents the weight tensor W mapped by index u. (W[u]) 2 Indicates the size of the weight. ∑ v≥u (W[v]) 2 The cumulative sum of all squared magnitudes is used to normalize the weights. u and v are sorted in ascending order to represent the index mapping corresponding to the weights. Lamp prunes the weights with the lowest Lamp score in each layer until the pruning criteria are met and the required global sparsity constraint is achieved. This approach ensures that at least one weight is retained in each layer while preserving the most important weights. Figure 5 -a sorts neurons by the absolute value of their weights, Figure 5 -b is to calculate the Lamp score, Figure 5 -c removes neurons with scores less than the Lamp score threshold to meet pruning requirements. This scoring mechanism achieves effective pruning by identifying and removing less important connections while retaining important connections, thereby reducing model complexity without significantly sacrificing accuracy.

[0070] Furthermore, in step 3, a loss function is constructed to update the model weights: the KLD divergence loss converts the rotated bounding box into a two-dimensional Gaussian distribution, and uses the rotation matrix and scale factor to map the rotated box parameters to the mean and covariance matrix of the Gaussian distribution; by calculating the Gaussian distance between the predicted box and the true box, the KLD divergence loss can more stably measure the difference between the true value and the predicted value, solve the angle mutation problem of rotated target detection, and achieve more accurate target detection.

[0071] All experiments were conducted using the PyTorch framework. The experimental environment included a GeForce RTX 3060 graphics card, an i5-12600KF 3.70GHz processor, and a Windows operating system. During training, the initial learning rate, momentum, weight decay, and batch size were set to 0.01, 0.937, 0.0005, and 4, respectively.

[0072] The model is evaluated using the average precision (AP) and mean average precision (mAP). AP is calculated as follows:

[0073]

[0074] Where P refers to the number of samples predicted as positive samples that are actually positive samples. R refers to the number of samples that are correctly identified among the true positive samples. mAP is the average AP of all categories. mAP is calculated as follows:

[0075]

[0076] Where n represents the total number of categories in the dataset. i represents the accuracy of the i-th category.

[0077] Table 1 shows the comparative experimental results of different models on the DOTA dataset. The values ​​in bold represent the best results in that category. As shown in the table, the proposed method achieves the best overall mAP, reaching 79.54%. The proposed algorithm achieves mAPs of 95.64%, 60.84%, 89.95%, 91.83%, 95.31%, 70.92%, 69.55%, 87.04%, and 75.06% for aircraft, bridges, large vehicles, ships, tennis courts, soccer fields, roundabouts, ports, and swimming pools, respectively. The target accuracy for these categories outperforms the comparative methods, achieving optimal detection results.

[0078] Table 1 Comparative experiments on the DOTA dataset (unit: %)

[0079]

[0080] To reduce model complexity, we performed Lamp pruning on the model, which incorporated a multi-layer feature interaction pyramid, a lightweight enhanced detection head, and the KLD loss function. This reduced computational complexity from 7.6 GB to 5.1 GB, a 32.9% reduction; and the number of parameters from 2.8 MB to 2.4 MB, a 14.2% reduction. This made the model more efficient overall.

[0081] The above are specific embodiments of the present invention. It should be noted that the present invention is not limited to the above specific embodiments. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for omnidirectional target detection in remote sensing images based on a multi-layer feature interaction pyramid and a lightweight enhanced detection head, characterized in that: The method comprises: Step 1: Preprocess the remote sensing image dataset; Step 2: Build a remote sensing image omnidirectional target detection model: Design a multi-layer feature interaction pyramid to aggregate multi-layer feature maps to obtain intermediate feature maps and fused feature maps, and perform path interaction through upsampling and downsampling to achieve cross-layer feature fusion and generate rotation-invariant feature representation; The multi-layer feature interaction pyramid first takes the C3, C4 and C5 feature maps generated by the backbone network as input, performs feature splicing on these three layers of input, and obtains the intermediate feature map F m , as the intermediate hub for subsequent feature interactions; secondly, the intermediate feature map F m After downsampling and C5 feature map feature splicing, the feature map is obtained f 1 , improve the richness of deep semantic expression, and at the same time, the new feature map obtained by upsampling the intermediate feature map is concatenated with the C3 feature map to obtain the feature map f 2 , enhance the representation ability of detail information; then, F m 、 f 1 and f 2 These three layers of feature maps are spliced ​​to generate a fusion feature map f m , the fused feature map contains feature information of different scales; finally, the intermediate feature map and the fused feature map are upsampled separately and combined with the generated f 1 The feature maps are spliced ​​to enhance the spatial resolution and detail expression. At the same time, the intermediate feature maps and the fused feature maps are downsampled and combined with the generated f 2 The feature maps are spliced ​​to highlight the abstract expression of semantic information. The generated output feature map fully utilizes the details and semantic information of each layer of feature maps and generates a rotation-invariant feature representation. A lightweight enhanced detection head is constructed, which uses a shared enhanced convolution strategy to reduce the number of model parameters and extracts feature information through central difference convolution, allowing the model to capture the detailed features of the rotating target. Lamp pruning method is used to prune the weight with the lowest Lamp score in each layer to reduce the complexity of the model; Step 3: Construct a loss function to update the model weights, where the KLD divergence loss is used as the regression loss for rotated object detection to measure the difference in rotation angle between the predicted box and the true box; Step 4: Iterate the model training until the model reaches convergence and obtains the optimal weight; Step 5: Use the obtained optimal weight to detect the test set images and obtain the final test results.

2. The method for omnidirectional target detection in remote sensing images based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head according to claim 1, characterized in that: The remote sensing image dataset preprocessing is to crop the DOTA remote sensing image dataset into 1024×1024 pixel blocks, and set an overlapping area of ​​200 pixels between adjacent pixel blocks; if the original image is smaller than 1024×1024, it is processed by zero padding.

3. The method for omnidirectional target detection in remote sensing images based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head according to claim 1, characterized in that: The lightweight enhanced detection head first applies 1×1 convolution to the output of the multi-layer feature interaction pyramid to adjust the number of channels. Then, a weight-sharing strategy is used to perform enhanced convolution operations, which not only effectively learns image features but also reduces the number of parameters and complexity of the model. Finally, the features after shared convolution are used to predict the position of the bounding box and the probability of the target category, respectively. Non-maximum suppression is used to remove redundant bounding boxes to generate the final detection results.

4. The method for omnidirectional target detection in remote sensing images based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head according to claim 3, characterized in that: The enhanced convolution consists of two parallel convolution layers, including a standard convolution layer and a center difference convolution layer.

5. The method for omnidirectional target detection in remote sensing images based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head according to claim 1, characterized in that: The KLD divergence loss converts the rotated bounding box into a two-dimensional Gaussian distribution, and uses the rotation matrix and scale factor to map the rotated box parameters to the mean and covariance matrix of the Gaussian distribution. By calculating the Gaussian distance between the predicted box and the true box, the KLD divergence loss can more stably measure the difference between the true value and the predicted value, and can solve the problem of angle mutation in rotated object detection.

6. The method for omnidirectional target detection in remote sensing images based on a multi-layer feature interactive pyramid and a lightweight enhanced detection head according to claim 1, characterized in that: The lamp pruning is performed by calculating the square of the target connection weight and normalizing it according to the sum of the squares of all surviving weights in the same layer.