A ship detection method and system based on synthetic aperture radar

CN122780904APending Publication Date: 2026-09-18NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611264341.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]然而,尽管现有方法取得了一定进展,但在实际应用中仍面临着多方面的核心挑战,制约了检测性能的进一步提升

Benefits of technology

1.通过并行与串联增强路径,实现了空间域特征与频率域特征在不同层级上的深度协同融合。相较于传统单一依赖卷积的模块,本发明能够更全面地捕获SAR图像的本质信息,有效地区分真实船舶目标与海杂波、散斑噪声及近岸复杂地物等干扰,从而提升了船舶检测模型在复杂场景下的检测精度和鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780904A_ABST
    Figure CN122780904A_ABST
Patent Text Reader

Abstract

The application discloses a ship detection method and system based on a synthetic aperture radar, and belongs to the technical field of image recognition. The method comprises the following steps: firstly, an initial ship detection model containing a feature extraction network and a group of initial prediction heads is constructed based on a multi-path multi-level space-frequency cooperative enhancement module, so that the representation ability of the synthetic aperture radar image signal is improved. Secondly, an adaptive head pruning strategy is used to optimize the network structure, so that the optimal matching between the ship detection model structure and the prediction task is realized. Finally, the optimized network is trained by using a scene-adaptive label allocation strategy, so that the detection robustness of the ship detection model in a complex and changeable environment is improved. The application uses the multi-path multi-level space-frequency cooperative enhancement module, the adaptive head pruning strategy and the scene-adaptive label allocation strategy, so that the capability, efficiency and scene adaptability of the synthetic aperture radar ship detection are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image or video recognition or understanding technology, and in particular to a ship detection method and system based on synthetic aperture radar. Background Technology

[0002] Synthetic Aperture Radar (SAR) is a remote sensing technology that utilizes the reflection characteristics of microwave signals and achieves high-resolution imaging through the principle of synthetic aperture. Unlike optical sensors, SAR can operate in all weather conditions and is unaffected by clouds, fog, or day / night cycles. Therefore, SAR exhibits unique advantages in maritime surveillance, and SAR-based ship detection technology has become an important means of maritime monitoring. This technology aims to accurately identify and locate ship targets from complex marine backgrounds and has wide applications in various scenarios. For example, in maritime traffic safety, it can monitor waterways in real time to prevent collisions; in maritime law enforcement, it helps combat illegal fishing and maintain order; and in emergency rescue, it can quickly locate distressed vessels and improve search and rescue efficiency. These advantages make SAR ship detection an important tool for ensuring maritime safety and sustainable development.

[0003] Many SAR image-based ship detection models typically consist of three main parts: a backbone network, a neck network, and a head network. The backbone is responsible for extracting multi-scale basic features from the input SAR image, generally employing a deep convolutional neural network (CNN) structure to generate multi-scale feature maps with rich semantic information. The neck uses techniques such as cross-layer connections and multi-scale fusion to further enhance the expressive power of features, thereby improving the model's adaptability to complex scenes. In the final stage of the detection model, the head often employs a prediction head design, outputting classification results and regression predictions at different scales. Furthermore, during model training, the label assignment strategy is one of the key factors affecting model performance. For example, task-aligned label assignment strategies calculate the alignment score between classification and regression tasks and select the top N candidate samples with the highest scores as positive samples, enriching the model's supervision signal and enabling the model to more effectively learn the target's features and location information.

[0004] However, despite the progress made by existing methods, several core challenges remain in practical applications, hindering further improvements in detection performance. First, at the feature extraction level, current models rely excessively on convolutional operations to extract spatial features, underutilizing the rich frequency domain and texture information in SAR images that effectively distinguishes targets from the background. This becomes a bottleneck for performance improvement under strong interference from speckle noise, sea clutter, and complex nearshore features. Second, at the network structure level, the multiple prediction heads designed for detecting multi-scale ships generally lack targeted optimization, often resulting in redundant functions and wasted computational resources for some prediction heads, or a mismatch between the target scale range they are responsible for and the actual data distribution, leading to prediction bias. Finally, at the model training level, existing task alignment label allocation strategies generally adopt a "one-size-fits-all" fixed Top-K method, which cannot adapt to the significant differences between nearshore (complex background) and offshore (simple background) scenarios in SAR images, limiting the model's learning potential and generalization ability in diverse scenarios. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a ship detection method and system based on synthetic aperture radar. By utilizing a multi-path, multi-level space-frequency collaborative enhancement module, an adaptive head pruning strategy, and a scene-adaptive label allocation strategy, the aim is to obtain a ship detection model with stronger feature representation capabilities, a more streamlined network structure, and better scene adaptability, thereby comprehensively improving the ship detection capability and robustness in complex SAR scenarios.

[0006] This invention is achieved through the following technical solution: A ship detection method based on synthetic aperture radar includes the following steps: S1: Construct an initial ship detection model that includes a feature extraction network and a set of initial prediction heads. The feature extraction network is constructed using a multi-path, multi-level space-frequency co-enhancement module as the basic computing unit. Select any synthetic aperture radar image from the dataset and input it into the initial ship detection model. By applying the multi-path, multi-level space-frequency co-enhancement module at multiple scale levels, a set of multi-scale feature maps that are more representative of the characteristics of synthetic aperture radar image signals are obtained and fed into the initial prediction head. S2: The configuration of the initial prediction head is optimized through an adaptive head pruning strategy, and the structure of the feature extraction network is adjusted synchronously according to the pruning results to obtain the optimized ship detection model; S3: The optimized ship detection model is trained on the synthetic aperture radar images in the training set using a scene-adaptive label allocation strategy. The scene-adaptive label allocation strategy dynamically assigns different numbers of positive sample candidate boxes to real targets based on whether the scene type of the input synthetic aperture radar image is nearshore or offshore. After training, the final ship detection model is obtained. S4: Input the synthetic aperture radar image to be detected into the final ship detection model, and output the classification and regression results of ship detection.

[0007] Furthermore, the processing procedure of the multi-path, multi-level space-frequency collaborative enhancement module in step S1 is as follows: a) The input synthetic aperture radar image is segmented into multiple parts and then distributed to at least two different types of processing paths; b) The first type of processing path adopts a direct parallel space-frequency structure, which includes a convolutional sub-branch for extracting spatial domain features and a stationary wavelet transform sub-branch for extracting frequency domain features. The convolutional sub-branch and the stationary wavelet transform sub-branch output parallel space-frequency features. c) The second type of processing path adopts a multi-level serial enhancement structure, which includes multiple serial feature deepening layers, and sets up a parallel space frequency structure as described in the first type of processing path at the output of at least one of the feature deepening layers. d) Integrate the parallel spatial frequency features output from all processing paths to obtain the enhanced feature representation at the current scale.

[0008] Furthermore, the specific implementation of the multi-path, multi-level space-frequency collaborative enhancement module is as follows: S111: Perform initial convolution on the input synthetic aperture radar image, and divide the processed initial features equally along the channel dimension. One main branch, It is an even number greater than or equal to 0; S112: Before Each main branch is processed as the first type of processing path, with convolutional sub-branches and stationary wavelet transform sub-branches set in parallel for each main branch; then... Each main branch is processed as a processing path of the second type, and a method containing [a specific function] is applied to each main branch. The enhancement structure consists of a series of convolutional layers, and within the enhancement structure, a convolutional sub-branch and a stationary wavelet transform sub-branch are set in parallel at the output of each series of convolutional layers. S113: Perform final splicing and convolution integration on all parallel spatial-frequency features generated by all processing paths in step S112 to achieve lossless fusion of spatial and frequency domain information at the same resolution.

[0009] Furthermore, the construction process of the feature extraction network in step S1 is as follows: S121: Construct a backbone network by downsampling at each level, and use the multi-path multi-level space-frequency collaborative enhancement module to enhance features at each downsampling level; S122: A feature fusion neck is constructed by upsampling step by step and fusing it with the features of the corresponding layer of the backbone network. The fusion process after upsampling also adopts the multi-path multi-level space-frequency collaborative enhancement module.

[0010] The optimized initial prediction head constructed in S1 includes a regression branch and a classification branch, and the regression branch and the classification branch adopt an asymmetric structure.

[0011] Furthermore, the method for optimizing the initial prediction head configuration in S2 using an adaptive head pruning strategy is as follows: S211: Starting from the high-level prediction head responsible for detecting large targets, execute a top-down redundant prediction head elimination strategy, calculate the contribution of each prediction head in each preset target scale interval, and then identify and eliminate prediction heads whose contribution is not the highest in all scale intervals. S212: Starting from the low-level prediction head responsible for detecting small targets, execute a bottom-up mismatch prediction head removal strategy. Calculate the proportion of predicted targets in each preset target scale interval and the proportion of real targets in the training data. When the difference between the proportion of predicted targets and the proportion of real targets exceeds a preset threshold, identify and remove the prediction head that contributes the most predictions in the current scale interval.

[0012] Furthermore, in step S3, in order to allocate the optimal number of positive candidate boxes to the real target based on the scene type during training, the method for pre-determining these optimal numbers is as follows: S311: First, fix the value of the number of positive candidate boxes for nearshore scenes, train on the training set, and find the value of the number of positive candidate boxes for offshore scenes that makes the optimized ship detection model most accurate on the validation set. Use this value as the optimal value of the number of positive candidate boxes for offshore scenes. S312: Fix the value of the optimal number of positive candidate boxes for offshore scenes in S311, train on the training set, and find the value of the number of positive candidate boxes for nearshore scenes that makes the optimized ship detection model most accurate on the validation set, and use it as the value of the optimal number of positive candidate boxes for nearshore scenes.

[0013] A ship detection system based on synthetic aperture radar is used to execute a ship detection method based on synthetic aperture radar as described above, comprising a multi-path space-frequency cooperative enhancement network construction unit, a predictor head adaptive pruning optimization unit, a scene adaptive training unit, and an output unit. The multi-path space-frequency cooperative enhancement network construction unit is used to perform step S1 to construct an initial ship detection model based on the multi-path multi-level space-frequency cooperative enhancement module; The predictive head adaptive pruning optimization unit is used to execute step S2, which optimizes the initial predictive head configuration through the adaptive head pruning strategy. The scene-adaptive training unit is used to execute step S3, which uses the scene-adaptive label allocation strategy to train the optimized ship detection model. The output unit is used to execute step S4, inputting the synthetic aperture radar image to be detected into the final ship detection model, and outputting the classification and regression results of ship detection.

[0014] Beneficial effects of the invention: The present invention provides a ship detection method and system based on synthetic aperture radar, which has the following advantages: 1. By enhancing the path through parallel and serial connections, deep collaborative fusion of spatial domain features and frequency domain features at different levels is achieved. Compared to traditional modules that rely solely on convolution, this invention can capture the essential information of SAR images more comprehensively, effectively distinguishing real ship targets from interference such as sea clutter, speckle noise, and complex nearshore features, thereby improving the detection accuracy and robustness of the ship detection model in complex scenarios.

[0015] 2. By eliminating redundancy from the top down, prediction heads with minimal performance contribution were simplified, significantly reducing computational overhead and resource waste. Simultaneously, through bottom-up mismatch correction, the prediction preferences of the ship detection model for targets at various scales were calibrated, making it more consistent with the distribution of real data. This two-way optimization mechanism ensures a highly streamlined and efficient ship detection model structure, enabling the model to obtain better feature representations at each scale, achieving simultaneous improvement in efficiency and accuracy.

[0016] 3. The optimal number of positive samples is dynamically allocated to each image scene. For complex nearshore scenes, more stringent and higher-quality supervision signals are provided to avoid interference; for simple offshore scenes, richer supervision signals are provided to promote learning. This personalized training approach enables the ship detection model to obtain better training guidance in diverse scenarios, enhancing its environmental adaptability and generalization performance. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the process of this invention.

[0018] Figure 2 This is a schematic diagram of the multi-path, multi-level space-frequency collaborative enhancement module structure of the present invention.

[0019] Figure 3(a) is a schematic diagram of the optimized synthetic aperture radar ship detection model being trained when the number of Top-K positive samples in a fixed near-shore scenario is equal to 10.

[0020] Figure 3(b) is a schematic diagram of the optimized synthetic aperture radar ship detection model being trained when the number of Top-K positive samples in the fixed offshore scenario of the present invention is equal to 5.

[0021] Figure 4 This is a schematic diagram of the system structure of the present invention. Detailed Implementation

[0022] A ship detection method based on synthetic aperture radar includes the following steps, the flowchart of which is shown below. Figure 1 As shown: S1: Construct an initial ship detection model that includes a feature extraction network and a set of initial prediction heads. The feature extraction network is constructed using a multi-path, multi-level space-frequency co-enhancement module as the basic computing unit. Select any synthetic aperture radar image from the dataset and input it into the initial ship detection model. By applying the multi-path, multi-level space-frequency co-enhancement module at multiple scale levels, a set of multi-scale feature maps that are more representative of the characteristics of synthetic aperture radar image signals are obtained and fed into the initial prediction head.

[0023] Specifically, the processing procedure of the multi-path, multi-level space-frequency collaborative enhancement module is as follows: a) The input synthetic aperture radar image is segmented into multiple parts and then distributed to at least two different types of processing paths; b) The first type of processing path adopts a direct parallel space-frequency structure, which includes a convolutional sub-branch for extracting spatial domain features and a stationary wavelet transform sub-branch for extracting frequency domain features. The convolutional sub-branch and the stationary wavelet transform sub-branch output parallel space-frequency features. c) The second type of processing path adopts a multi-level serial enhancement structure, which includes multiple serial feature deepening layers, and sets up a parallel space frequency structure as described in the first type of processing path at the output of at least one of the feature deepening layers. d) Integrate the parallel spatial frequency features output from all processing paths to obtain the enhanced feature representation at the current scale.

[0024] The specific implementation process of the multi-path, multi-level space-frequency collaborative enhancement module is as follows: S111: Perform initial convolution on the input synthetic aperture radar image, and divide the processed initial features equally along the channel dimension. One main branch, It is an even number greater than or equal to 0; The synthetic aperture radar image is initially convolved, and the specific processing formula is shown in equation (1): (1); in: Indicates initial features, This represents the features of the input synthetic aperture radar image. This represents the convolution operation.

[0025] Subsequently, the processed features are divided into n main branches along the channel dimension, and the features corresponding to each main branch are... ; S112: Before Each main branch is processed as the first type of processing path, with convolutional sub-branches and stationary wavelet transform sub-branches set in parallel for each main branch; then... Each main branch is processed as a processing path of the second type, and a method containing [a specific function] is applied to each main branch. The enhancement structure consists of a series of convolutional layers, and within the enhancement structure, a convolutional sub-branch and a stationary wavelet transform sub-branch are set in parallel at the output of each series of convolutional layers. Specifically, convolutional sub-branches can be set in parallel for each main branch according to equation (2), and stationary wavelet transform sub-branches can be set in parallel for each main branch according to equation (3): (2); (3); in: , Indicates the first Initial characteristics of each main branch Indicates the first Local spatial detail features output by each main branch Indicates the first Global frequency domain features of the output of each main branch This indicates a stationary wavelet transform operation.

[0026] For the later In each main branch, within the cascaded convolutional layer enhancement structure corresponding to each main branch, the output of each cascaded convolutional layer is similar to equations (2) and (3), with a convolutional sub-branch and a stationary wavelet transform sub-branch set in parallel to extract spatial and frequency domain features. For the first main branch, The first main branch corresponds to the cascaded convolutional layer enhancement structure. Layered convolution features The spatial frequency features extracted through the convolutional sub-branch and the stationary wavelet transform sub-branch can be expressed as: ; S113: Perform final splicing and convolution integration on all parallel spatial-frequency features generated by all processing paths in step S112 to achieve lossless fusion of spatial and frequency domain information at the same resolution.

[0027] The specific calculation formula for this step is formula (4): (4); in: This indicates the characteristics after channel splicing. This indicates that the splicing operation is performed along the channel dimension. This indicates the output of the multi-path, multi-level space-frequency collaborative enhancement module.

[0028] The feature extraction network obtains a set of multi-scale feature maps that are more representative of the characteristics of synthetic aperture radar image signals by applying the multi-path, multi-level, space-frequency collaborative enhancement module at multiple scale levels.

[0029] Optimized, allowing selection of the number of main branches. This allows the model to achieve good feature extraction capabilities while maintaining its lightweight nature. The structure diagram of the multi-path, multi-level space-frequency collaborative enhancement module is shown below. Figure 2 As shown, the processing procedure of the multi-path, multi-level space-frequency collaborative enhancement module is as follows: S111a: Perform initial convolution on the input synthetic aperture radar image, and divide the processed initial features into two main branches along the channel dimension; S112a: The first main branch is processed as the first type of processing path, and a convolutional sub-branch and a stationary wavelet transform sub-branch are set in parallel; the second main branch is processed as the second type of processing path, and an enhancement structure containing a single cascaded convolutional layer is applied in parallel to the output of the cascaded convolutional layer, and a convolutional sub-branch and a stationary wavelet transform sub-branch are set in parallel. S113a: Perform final splicing and convolution integration on all parallel spatial-frequency features generated by the two main branches in step S112a to achieve lossless fusion of spatial and frequency domain information at the same resolution.

[0030] Specifically, the construction process of the feature extraction network is as follows: S121: Construct a backbone network by progressive downsampling, and use the multi-path multi-level space-frequency collaborative enhancement module to enhance features at each downsampling level; The specific process of building the Backbone component is as follows: The synthetic aperture radar images in the training set are processed through a single convolutional layer to obtain a 2x downsampled feature map; The 2x downsampled feature map is processed by a combination of a convolutional layer and a multi-path, multi-level space-frequency collaborative enhancement module to obtain a 4x downsampled feature map. The 4x downsampled feature map is processed by a combination of a convolutional layer and a multi-path, multi-level space-frequency collaborative enhancement module to obtain an 8x downsampled feature map. The 8x downsampled feature map is processed by a combination of a convolutional layer and a multi-path, multi-level space-frequency collaborative enhancement module to obtain a 16x downsampled feature map. The 16x downsampled feature map is processed by combining a convolutional layer with a multi-path, multi-level space-frequency collaborative enhancement module to obtain a 32x downsampled feature map.

[0031] S122: A feature fusion neck is constructed by upsampling step by step and fusing it with the features of the corresponding layer of the backbone network. The fusion process after upsampling also adopts the multi-path multi-level space-frequency collaborative enhancement module.

[0032] The process of building the Neck section is as follows: The 32x downsampled feature map is upsampled to obtain a new 16x downsampled feature map. The new 16x downsampled feature map is concatenated with the 16x downsampled feature map obtained in the previous step along the channel dimension, and then fused through the multi-path multi-level space-frequency collaborative enhancement module to obtain the final 16x downsampled feature map. The final 16x downsampled feature map is then upsampled to obtain a new 8x downsampled feature map. The new 8x downsampled feature map is concatenated with the 8x downsampled feature map obtained in the previous step along the channel dimension, and then fused through the multi-path multi-level space-frequency collaborative enhancement module to obtain the final 8x downsampled feature map. The final 8x downsampled feature map is then upsampled to obtain a new 4x downsampled feature map. The new 4x downsampled feature map is concatenated with the 4x downsampled feature map obtained in the previous step along the channel dimension, and then fused through the multi-path multi-level space-frequency collaborative enhancement module to obtain the final 4x downsampled feature map. The final 4x downsampled feature map is then upsampled to obtain a new 2x downsampled feature map. The new 2x downsampled feature map is concatenated with the 2x downsampled feature map obtained in the previous step along the channel dimension, and then fused through the multi-path multi-level space-frequency collaborative enhancement module to obtain the final 2x downsampled feature map.

[0033] Then, the constructed feature extraction network is connected to a set of prediction heads. The specific implementation process is as follows: The final 2x downsampled feature map, the final 4x downsampled feature map, the final 8x downsampled feature map, the final 16x downsampled feature map obtained from the Neck part, and the 32x downsampled feature map obtained from the Backbone part are concatenated to the corresponding prediction heads.

[0034] The optimized initial prediction head includes a regression branch and a classification branch, and the regression branch and the classification branch adopt an asymmetric structure.

[0035] The regression branch and the classification branch adopt an asymmetric structure, which can be adapted to the different characteristics of regression tasks and classification tasks respectively.

[0036] In traditional default configurations, prediction heads for different scales are designed symmetrically, meaning that the classification and regression branches use the same number of convolutional layers across all scales, typically three. This approach fails to consider the unique properties of features at different resolutions. Specifically, high-resolution features excel at capturing fine spatial details but may lack sufficient semantic information, posing a challenge for classification tasks; while low-resolution features, although rich in semantic cues, suffer from lower spatial localization accuracy, impacting regression performance.

[0037] To address this, this application employs an asymmetric design, specifically enhancing the classification branch of the high-resolution prediction head and strengthening the regression branch of the low-resolution prediction head. By adding an additional depthwise separable convolutional layer to these two key branches, not only is the model's ability to handle features at different scales improved, but the overall performance of the head part is also optimized. This asymmetric design allows the model to more effectively utilize the advantages of features at each scale, thereby achieving more accurate feature representations at different scales. Ultimately, this approach improves the performance of multi-scale object detection, demonstrating its potential for handling diverse targets in complex scenes.

[0038] S2: The configuration of the initial prediction head is optimized through an adaptive head pruning strategy, and the structure of the feature extraction network is adjusted synchronously according to the pruning results to obtain the optimized ship detection model; The method for optimizing the initial prediction head configuration based on multi-scale feature maps using an adaptive head pruning strategy is as follows: S211: Starting with the high-level prediction head responsible for detecting large targets, a top-down redundant prediction head removal strategy is implemented. The contribution of each prediction head in each preset target scale interval is calculated, and then prediction heads whose contribution is not the highest in all scale intervals are identified and removed. This aims to identify and remove redundant prediction heads that have a low contribution to the overall task.

[0039] The contribution here is defined as the ratio of the number of predictions made by a certain predictor in a certain interval to the total number of predictions made by all predictors in that interval.

[0040] The specific implementation process is as follows: First, the area of ​​the ship targets in the training set is divided into multiple preset scale intervals.

[0041] Then, to quantify the contribution of each predictor, for any given scale interval, its contribution is determined by calculating the ratio of the number of targets predicted by a specific predictor within that interval to the total number of targets predicted by all predictors within the same interval. Specifically, this can be calculated using equation (5): (5); in: Indicates the first The first prediction head Contribution across scale intervals, Indicates the first The first prediction head Number of predicted targets in each scale interval This represents the total number of prediction heads. It is calculated... It can quantify the contribution ratio of each prediction head within a specific scale range.

[0042] Based on this assessment, the "dominant predictor" with the highest contribution is identified for each scale interval. After completing the assessment of all scale intervals, if a predictor fails to become the "dominant predictor" in any scale interval, it is considered a redundant predictor and is removed.

[0043] In this embodiment, the number of each predictor of the synthetic aperture radar ship detection model distributed across each scale interval is shown in Table 1, and the contribution of each predictor of the synthetic aperture radar ship detection model in each scale interval calculated according to Table 1 is shown in Table 2. Table 1

[0044] Table 2

[0045] As shown in Table 2, the contribution of the 32x downsampled head is not the highest across all scale ranges, indicating that this prediction head is redundant. Therefore, this prediction head is removed from the Head portion of the constructed synthetic aperture radar ship detection model. Thus, 2x, 4x, 8x, and 16x downsampled heads are selected.

[0046] S212: Starting with the low-level prediction head responsible for detecting small targets, a bottom-up mismatch prediction head removal strategy is implemented. The proportion of predicted targets in each preset target scale interval is calculated, along with the proportion of real targets in the training data. When the difference between the predicted target proportion and the real target proportion exceeds a preset threshold, the prediction head contributing the most predictions in the current scale interval is identified and removed. This aims to correct the prediction attention bias of the ship detection model in different scale intervals, ensuring that the prediction distribution of the ship detection model matches the scale distribution of the real data.

[0047] First, for each preset target scale interval, the proportion of real targets in the training data is statistically analyzed (i.e., the number of real targets in that interval / the total number of real targets). This represents the "expected scale distribution" of the data itself.

[0048] Secondly, for each scale interval, the percentage of the total number of predictions output by the ship detection model in that interval is statistically analyzed (i.e., the sum of the predictions of all prediction heads in that interval / the sum of the total predictions of all prediction heads across all intervals). This represents the current "actual prediction attention distribution" of the ship detection model.

[0049] Then, the "expected scale distribution" is compared with the "actual predicted attention distribution". If the difference between the predicted proportion and the actual proportion exceeds a preset threshold in a certain scale interval, it indicates that the ship detection model has an over-concentration problem in the prediction of that scale interval. In order to correct this bias, the prediction head that contributed the most predictions (i.e., the one with the most predictions in that interval) will be found in the mismatched scale interval and removed, thereby reducing the ship detection model's tendency to over-predict at that scale.

[0050] Specifically, the scale ranges of ship targets in the synthetic aperture radar images in the training set and the number of samples in each scale range are shown in Table 3: Table 3

[0051] Table 3 identifies the scale regions with high ship target density as 163-655 and 655-2621, and these regions should receive more attention in ship detection model processing.

[0052] The distribution of the total number of predicted targets for all prediction heads in different intervals is shown in Table 4. Table 4

[0053] As shown in Table 4, the predictions overemphasize extremely small targets in the scale ranges of 0-163 and 163-655, especially in the 0-163 range, where the difference between the predicted target ratio and the actual target ratio is 0.21, exceeding the threshold of 0.18. Therefore, it is necessary to remove the 2x downsampling head, resulting in a new prediction distribution as shown in Table 5. Table 5

[0054] As can be seen from Table 5, the new predicted distribution obtained after removing the 2x downsampling head is more consistent with the actual distribution shown in Table 3.

[0055] During the initial prediction head configuration optimization, the feature extraction network structure is adjusted synchronously. In S211 and S212, for each prediction head removed, the feature extraction network described in S1 is adjusted synchronously, removing the feature extraction layer corresponding to the removed prediction head (such as the corresponding stage in the Backbone or the corresponding fusion path in the Neck) to achieve overall simplification of the ship detection model structure.

[0056] S3: The optimized ship detection model is trained using a scene-adaptive label allocation strategy. The scene-adaptive label allocation strategy dynamically assigns different numbers of positive candidate boxes (Top-K) to the real target based on whether the scene type of the input synthetic aperture radar image is near-shore or offshore. After training, the final ship detection model is obtained for synthetic aperture radar ship detection.

[0057] Specifically, in order to assign the optimal number of positive candidate boxes to the real target based on the scene type during training, the method for pre-determining these optimal numbers is as follows: S311: First, fix the value of the number of positive candidate boxes for nearshore scenes, train on the training set, and find the value of the number of positive candidate boxes for offshore scenes that makes the optimized ship detection model most accurate on the validation set. Use this value as the optimal value of the number of positive candidate boxes for offshore scenes. S312: Fix the value of the optimal number of positive candidate boxes for offshore scenes in S311, train on the training set, and find the value of the number of positive candidate boxes for nearshore scenes that makes the optimized ship detection model most accurate on the validation set, and use it as the value of the optimal number of positive candidate boxes for nearshore scenes.

[0058] In relatively simple offshore environments, ship targets are clearly distinguishable and easy to differentiate, so a small number of Top-K positive samples are sufficient for training. However, in cluttered and complex nearshore scenarios, the number of Top-K positive samples needs to be increased to improve the robustness of the model in order to cope with the effects of noise and interference.

[0059] By adopting the above-mentioned scenario-adaptive label allocation strategy, the number of Top-K positive samples is dynamically adjusted according to the complexity of the scenario. This can better adapt to different detection environments and provide differentiated supervision signals for samples in different scenarios. This flexible adjustment mechanism helps to improve the detection performance of ship detection models in diverse scenarios.

[0060] Specifically, the detection accuracy diagrams for training the synthetic aperture radar (SAR) ship detection model using a scene-adaptive label allocation strategy on the training set are shown in Figures 3(a) and 3(b). In step S311, the number of Top-K positive samples in the nearshore scene is fixed at 10. The optimized SAR ship detection model is trained on the training set using a scene-adaptive label allocation strategy, as shown in Figure 3(a). In step S312, the number of Top-K positive samples in the offshore scene is fixed at 5. The optimized SAR ship detection model is trained on the training set using a scene-adaptive label allocation strategy, as shown in Figure 3(b). In the figures, AP50 represents the accuracy of the ship detection model when the overlap between the predicted and actual results is 50%, and AP75 represents the accuracy of the ship detection model when the overlap between the predicted and actual results is 75%. This represents the number of Top-K positive samples in the offshore scenario. This represents the number of Top-K positive samples in nearshore scenarios.

[0061] As can be seen from Figure 3(a), When AP50 and AP75 are equal to 5, they both reach their highest values. Therefore, the optimal number of Top-K positive samples in the offshore scenario is 5, as shown in Figure 3(b). When AP50 and AP75 are equal to 10, they both reach their highest values. Therefore, the optimal number of Top-K positive samples in the nearshore scenario is 10.

[0062] S4: Input the synthetic aperture radar image to be detected into the final ship detection model, and output the classification and regression results of ship detection.

[0063] A ship detection system based on synthetic aperture radar (SAR) is used to execute a ship detection method based on SAR as described in any one of the above embodiments, and its system diagram is shown below. Figure 4 As shown, it includes a multi-path space-frequency collaborative enhancement network construction unit, a prediction head adaptive pruning optimization unit, a scene adaptive training unit, and an output unit; The multi-path space-frequency cooperative enhancement network construction unit is used to execute step S1 to construct an initial ship detection model based on the multi-path multi-level space-frequency cooperative enhancement module. The predictor head adaptive pruning optimization unit is used to execute step S2, which optimizes the initial predictor head configuration through the adaptive head pruning strategy. The scene-adaptive training unit is used to execute step S3, which uses the scene-adaptive label allocation strategy to train the optimized ship detection model. The output unit is used to execute step S4, inputting the synthetic aperture radar image to be detected into the final ship detection model, and outputting the classification and regression results of ship detection.

[0064] In summary, the present invention provides a ship detection method and system based on synthetic aperture radar. By utilizing a multi-path, multi-level space-frequency collaborative enhancement module, an adaptive head pruning strategy, and a scene-adaptive label allocation strategy, it aims to obtain a ship detection model with stronger feature representation capabilities, a more streamlined network structure, and better scene adaptability, thereby comprehensively improving the ship detection capability and robustness in complex SAR scenarios.

[0065] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A ship detection method based on synthetic aperture radar, characterized in that: Includes the following steps: S1: Construct an initial ship detection model that includes a feature extraction network and a set of initial prediction heads. The feature extraction network is constructed using a multi-path, multi-level space-frequency co-enhancement module as the basic computing unit. Select any synthetic aperture radar image from the dataset and input it into the initial ship detection model. By applying the multi-path, multi-level space-frequency co-enhancement module at multiple scale levels, a set of multi-scale feature maps that are more representative of the characteristics of synthetic aperture radar image signals are obtained and fed into the initial prediction head. S2: The configuration of the initial prediction head is optimized through an adaptive head pruning strategy, and the structure of the feature extraction network is adjusted synchronously according to the pruning results to obtain the optimized ship detection model; S3: The optimized ship detection model is trained on the synthetic aperture radar images in the training set using a scene-adaptive label allocation strategy. The scene-adaptive label allocation strategy dynamically assigns different numbers of positive sample candidate boxes to real targets based on whether the scene type of the input synthetic aperture radar image is nearshore or offshore. After training, the final ship detection model is obtained. S4: Input the synthetic aperture radar image to be detected into the final ship detection model, and output the classification and regression results of ship detection.

2. The ship detection method based on synthetic aperture radar according to claim 1, characterized in that: The processing procedure of the multi-path, multi-level space-frequency collaborative enhancement module in step S1 is as follows: a) The input synthetic aperture radar image is segmented into multiple parts and then distributed to at least two different types of processing paths; b) The first type of processing path adopts a direct parallel space-frequency structure, which includes a convolutional sub-branch for extracting spatial domain features and a stationary wavelet transform sub-branch for extracting frequency domain features. The convolutional sub-branch and the stationary wavelet transform sub-branch output parallel space-frequency features. c) The second type of processing path adopts a multi-level serial enhancement structure, which includes multiple serial feature deepening layers, and sets up a parallel space frequency structure as described in the first type of processing path at the output of at least one of the feature deepening layers. d) Integrate the parallel spatial frequency features output from all processing paths to obtain the enhanced feature representation at the current scale.

3. The ship detection method based on synthetic aperture radar according to claim 2, characterized in that: The specific implementation of the multi-path, multi-level space-frequency coordinated enhancement module is as follows: S111: Perform initial convolution on the input synthetic aperture radar image, and divide the processed initial features equally along the channel dimension. One main branch, It is an even number greater than or equal to 0; S112: Before Each main branch is processed as the first type of processing path, with convolutional sub-branches and stationary wavelet transform sub-branches set in parallel for each main branch; then... Each main branch is processed as a processing path of the second type, and a method containing [a specific function] is applied to each main branch. The enhancement structure consists of a series of convolutional layers, and within the enhancement structure, a convolutional sub-branch and a stationary wavelet transform sub-branch are set in parallel at the output of each series of convolutional layers. S113: Perform final splicing and convolution integration on all parallel spatial-frequency features generated by all processing paths in step S112 to achieve lossless fusion of spatial and frequency domain information at the same resolution.

4. The ship detection method based on synthetic aperture radar according to claim 1, characterized in that: The construction process of the feature extraction network in step S1 is as follows: S121: Construct a backbone network by downsampling at each level, and use the multi-path multi-level space-frequency collaborative enhancement module to enhance features at each downsampling level; S122: A feature fusion neck is constructed by upsampling step by step and fusing it with the features of the corresponding layer of the backbone network. The fusion process after upsampling also adopts the multi-path multi-level space-frequency collaborative enhancement module.

5. The ship detection method based on synthetic aperture radar according to claim 1, characterized in that: The initial prediction head constructed in step S1 includes a regression branch and a classification branch, and the regression branch and the classification branch adopt an asymmetric structure.

6. A ship detection method based on synthetic aperture radar according to claim 5, characterized in that: The method for optimizing the initial prediction head configuration in S2 using an adaptive head pruning strategy is as follows: S211: Starting from the high-level prediction head responsible for detecting large targets, execute a top-down redundant prediction head elimination strategy, calculate the contribution of each prediction head in each preset target scale interval, and then identify and eliminate prediction heads whose contribution is not the highest in all scale intervals. S212: Starting from the low-level prediction head responsible for detecting small targets, execute a bottom-up mismatch prediction head removal strategy. Calculate the proportion of predicted targets in each preset target scale interval and the proportion of real targets in the training data. When the difference between the proportion of predicted targets and the proportion of real targets exceeds a preset threshold, identify and remove the prediction head that contributes the most predictions in the current scale interval.

7. The ship detection method based on synthetic aperture radar according to claim 1, characterized in that: In step S3, to allocate the optimal number of positive candidate boxes to the real target based on the scene type during training, the method for pre-determining these optimal numbers is as follows: S311: First, fix the value of the number of positive candidate boxes for nearshore scenes, train on the training set, and find the value of the number of positive candidate boxes for offshore scenes that makes the optimized ship detection model most accurate on the validation set. Use this value as the optimal value of the number of positive candidate boxes for offshore scenes. S312: Fix the value of the optimal number of positive candidate boxes for offshore scenes in S311, train on the training set, and find the value of the number of positive candidate boxes for nearshore scenes that makes the optimized ship detection model most accurate on the validation set, and use it as the value of the optimal number of positive candidate boxes for nearshore scenes.

8. A ship detection system based on synthetic aperture radar, used to execute a ship detection method based on synthetic aperture radar as described in any one of claims 1 to 7, characterized in that, It includes a multi-path space-frequency collaborative enhancement network construction unit, a prediction head adaptive pruning optimization unit, a scene adaptive training unit, and an output unit; The multi-path space-frequency cooperative enhancement network construction unit is used to execute step S1 to construct an initial ship detection model based on the multi-path multi-level space-frequency cooperative enhancement module. The predictor head adaptive pruning optimization unit is used to execute step S2, which optimizes the initial predictor head configuration through the adaptive head pruning strategy. The scene-adaptive training unit is used to execute step S3, which uses the scene-adaptive label allocation strategy to train the optimized ship detection model. The output unit is used to execute step S4, inputting the synthetic aperture radar image to be detected into the final ship detection model, and outputting the classification and regression results of ship detection.