A SAR image ship target segmentation method based on diffusion model

By using a diffusion model-based approach, the SAR ship instance segmentation task is represented as a denoising and reconstruction process. By combining a feature pyramid network and the cross-union loss function, the problem of insufficient ship segmentation accuracy and time-consuming hyperparameter adjustment in existing technologies is solved, and high-precision ship instance segmentation is achieved.

CN119741496BActive Publication Date: 2026-04-28UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2024-12-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing deep learning-based SAR ship instance segmentation techniques suffer from the problem that pre-defined target priors cannot perfectly match the shape and size of ship targets, and hyperparameter adjustment is time-consuming and laborious, resulting in insufficient segmentation accuracy.

Method used

A diffusion-based approach is adopted to represent the SAR ship instance segmentation task as a denoising process from noisy boxes to target boxes and a reconstruction process from target boxes to ship instances. By combining a spatial context-based joint enhancement feature pyramid network, a focused cross-union loss function, and an instance-aware mask representation, SAR ship instances are adaptively acquired.

Benefits of technology

It achieves the best average accuracy in near-shore and near-shore scenarios, avoiding the difficulties of setting target prior hyperparameters and the time-consuming and laborious problems of hyperparameter adjustment in traditional segmentation methods, thus improving segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741496B_ABST
    Figure CN119741496B_ABST
Patent Text Reader

Abstract

The application discloses a SAR image ship target segmentation method based on a diffusion model, and solves the problems of insufficient precision and time and labor consuming of existing segmentation technologies. The diffusion model is introduced in original SAR ship instance segmentation, and the ship instance segmentation task is processed from a generative perspective for the first time. In addition, various effective improvements such as a spatial context joint enhancement feature pyramid network, a focal IoU loss function and a mask representation based on instance perception are provided to ensure superior instance segmentation precision. Compared with the traditional method, the application has moderate model calculation complexity and the highest spatial complexity, and in actual application scenarios, the slightly complex model can bring higher segmentation precision and other advantages. Therefore, the application can be widely applied to the SAR image ship target instance segmentation task, and the model complexity can be further improved in the future; and the best average precision of the SAR image ship target segmentation in the near sea and near shore scenes is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of Synthetic Aperture Radar (SAR) image interpretation technology, and relates to a SAR image ship target segmentation method based on a diffusion model. Background Technology

[0002] Synthetic Aperture Radar (SAR) is an active microwave sensor capable of operating around the clock and in all weather conditions. Compared to optical sensors, SAR can penetrate clouds and fog, enabling observation tasks to be completed even under adverse weather conditions. SAR has become an important means of Earth observation today, with increasingly widespread applications in national economic fields such as terrain image generation, target detection and reconnaissance, land resource exploration, and natural disaster monitoring. In particular, SAR also has extensive applications in the marine field, especially in ship detection using SAR images, which is of great significance in the field of remote sensing. For details, see the literature "Zhang Qingjun, Han Xiaolei, Liu Jie. Progress and Development Trend of Spaceborne Synthetic Aperture Radar Remote Sensing Technology [J]. Spacecraft Engineering, 2017, 26(06):1-8."

[0003] With the massive acquisition of SAR ship data and the rapid development of SAR ship intelligent interpretation technology, SAR ship instance segmentation has become an advanced target interpretation method. In particular, SAR ship target instance segmentation, as a key component of ship surveillance, can effectively achieve pixel-level ship target detection and refined marine information perception. See the reference "Gao, F., Huo, Y., Wang, J., Hussain, A., Zhou, H., 2021. Anchor-free SAR ship instancesegmentation with centroid-distance based loss. IEEE J.Sel.Topics Appl.EarthObserv.Remote Sens.14, 11352–11371."

[0004] Currently, deep learning-based SAR ship instance segmentation technology has matured. Fully automated, efficient, and high-precision network models can be designed to achieve data-driven SAR ship instance segmentation. However, existing deep learning-based SAR ship instance segmentation techniques often approach the problem from a regression optimization perspective, continuously tuning prior hyperparameters and iteratively optimizing the network model. This leaves room for further improvement in SAR ship instance segmentation accuracy, primarily due to two challenges: the difficulty in perfectly matching the shape and size of ship targets with pre-defined target priors, and the time-consuming and labor-intensive nature of setting prior parameters.

[0005] Therefore, to address the aforementioned problems, this invention proposes a SAR image ship target segmentation method based on a diffusion model. This method proposes a novel diffusion model, representing the SAR ship instance segmentation task as a denoising process from noise boxes to target boxes and a reconstruction process from target boxes to ship instances. Innovatively, it avoids the hyperparameter difficulties of traditional segmentation methods due to pre-defined target priors and the time-consuming and laborious hyperparameter adjustment process, while ensuring optimal average segmentation accuracy. Summary of the Invention

[0006] This invention discloses a SAR image ship target segmentation method based on a diffusion model, addressing the problems of insufficient accuracy and time-consuming nature of existing segmentation techniques. Based on diffusion model theory, this method represents the SAR ship instance segmentation task as a denoising process from noisy boxes to target boxes and a reconstruction process from target boxes to ship instances. It provides a spatial context-jointly enhanced feature pyramid network to enhance spatial location and multi-level contextual information, promoting multi-scale ship feature extraction; a focused intersection-over-union (IoU) loss function to alleviate the positive-negative sample imbalance problem; and an instance-aware mask representation to adaptively extract SAR ship instances from denoised target boxes. This invention achieves the best average accuracy for SAR image ship target segmentation in near-shore and offshore scenes.

[0007] To facilitate the description of the present invention, the following terms are defined first:

[0008] Definition 1: SAR image

[0009] In remote sensing, SAR (Synthetic Aperture Radar) acquires surface images using airborne or spaceborne platforms. This process involves transmitting phase-encoded pulses along a direction nearly perpendicular to the sensor's motion vector, and receiving and recording the echoes reflected from the surface. To form an image, intensity measurements are required along two mutually orthogonal axes. For details on SAR images, see the literature "Principles of Synthetic Aperture Radar Imaging," edited by Pi Yiming et al., published by University of Electronic Science and Technology of China Press.

[0010] Definition 2: SSDD dataset

[0011] The SSDD dataset refers to the SAR ship detection dataset. This dataset is the first open dataset widely used for researching advanced ship detection techniques based on deep learning (DL) synthetic aperture radar (SAR) imagery. The official version of SSDD was released on December 1, 2017, at the BIGSARDATA conference. The official version of SSDD includes three types: BBox-SSDD, RBox-SSDD, and PSeg-SSDD, containing 1160 SAR images, each approximately 480*330 pixels in size, facilitating use by researchers according to different task requirements. Detailed information about the SSDD dataset can be found in the reference “Zhang,T.,Zhang,X.,Li,J.,Xu,X.,Wang,B.,Zhan,X.,Xu,Y.,Ke,X.,Zeng,T.,Su,H.,Ahmad,I.,Pan,D.,Liu,C.,Zhou,Y.,Shi,J.,Wei,S.,2021.SAR ship detection dataset(SSDD):Official release and comprehensive data analysis.Remote Sens.13(18),3690.”

[0012] Definition 3: HRSID dataset

[0013] The HRSID dataset is a high-resolution SAR image dataset. It was created to address the problems of small SAR image size, limited training samples, and inappropriate annotation in existing SAR ship datasets. HRSID enables not only object detection but also instance segmentation. The dataset uses 136 panoramic SAR images with resolutions ranging from 1m to 5m, cropped to 800*800 pixels with a 25% overlap. HRSID contains 5604 cropped SAR images and 16951 ship instances. Detailed information about the HRSID dataset can be found in the reference “Wei, S., Zeng, X., Qu, Q., Wang, M., Su, H., Shi, J., 2020. HRSID: A high-resolution SAR images dataset for ship detection and instance segmentation. IEEE Access 8, 120234–120254.”

[0014] Definition 4: Classic ResNet-101 Network

[0015] ResNet-101 is a residual network with powerful feature extraction capabilities and moderate model complexity, used to map raw images to a high-level deep feature space for subsequent decoder use. A variant of the ResNet architecture, ResNet-101 has 101 layers and is designed to learn abstract features at different levels. It has proven to have good generalization and robustness in various SAR intelligent interpretation tasks. For details on the ResNet-101 model, please refer to the paper "He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 770–778."

[0016] Definition 5: Traditional deformable convolution methods

[0017] Deformable convolutional methods enhance the transformation modeling capabilities of CNNs by introducing two modules: deformable convolution and deformable region-of-interest pooling, into the classic convolutional neural network. Spatial sampling locations within the module are enhanced through additional offsets, which are learned from the target task without additional supervision. These new modules can easily replace ordinary modules in existing CNNs and can be easily trained end-to-end using standard backpropagation, resulting in deformable convolutional networks. For details on traditional deformable convolutional methods, please refer to the literature “Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., Wei, Y., 2017. Deformable convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 764–773.”

[0018] Definition 6: Classical Convolutional Neural Network

[0019] Classical convolutional neural networks (CNNs) refer to a class of feedforward neural networks that incorporate convolutional computations and have a deep structure. CNNs are constructed by mimicking the visual perception mechanisms of biological systems, enabling both supervised and unsupervised learning. The shared parameters of the convolutional kernels within their hidden layers and the sparsity of inter-layer connections allow CNNs to extract features with relatively low computational cost. CNNs have made rapid progress in fields such as computer vision, natural language processing, and speech recognition. For a detailed description of classic CNN methods, please refer to the literature "Zhang Suofei, Feng Ye, Wu Xiaofu. Progress in object detection algorithms based on deep convolutional neural networks [J / OL]. Journal of Nanjing University of Posts and Telecommunications (Natural Science Edition), 2019(05):1-9."

[0020] Definition 7: Classic CNN Feature Extraction Method

[0021] Classical CNN feature extraction involves using a CNN to extract features from the original input image. In short, the original input image is transformed into a series of feature maps through convolutional operations on different features. In a CNN, the convolutional kernels in the convolutional layers continuously slide across the image for computation. Simultaneously, the max-pooling layer is responsible for taking the maximum value of each local block in the inner product result. Therefore, CNNs implement image feature extraction methods through convolutional layers and max-pooling layers. For a detailed explanation of classic CNN feature extraction, please refer to the website "https: / / blog.csdn.net / qq_30815237 / article / details / 86703620".

[0022] Definition 8: Classical Convolutional Layer

[0023] A convolutional layer consists of several convolutional units, each with parameters optimized using the backpropagation algorithm. The purpose of convolution is to extract different features from the input. The first convolutional layer may only extract low-level features such as edges, lines, and corners, while more layers can iteratively extract more complex features from these low-level features. For a detailed explanation of classic convolutional layers, please refer to the website "https: / / www.zhihu.com / question / 49376084".

[0024] Definition 9: Traditional methods for generating recombinant nuclei

[0025] The Reassembly Kernel is a convolutional kernel generation method for dynamic upsampling, primarily applied to content-aware feature reconstruction. This method encodes the content of the input feature map and dynamically generates convolutional kernels suitable for different locations based on the context, upsampling the feature map to a higher resolution. First, a 1×1 convolutional layer is used to reduce the number of channels in the input feature map, thereby reducing computational complexity. Then, a convolution operation is performed on the compressed feature map to generate dynamic convolutional kernels. Each generated convolutional kernel is adjusted according to the contextual features at a specific location. The Softmax function is used to normalize the convolutional kernels, ensuring that the sum of the kernel weights is 1, thus achieving smooth upsampling similar to interpolation methods. For details on traditional recombinant kernel generation methods, please refer to the literature "Wang, J., Chen, K., Xu, R., Liu, Z., Loy, CC, Lin, D., 2019. CARAFE: Content-aware reassembly of features. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 3007–3016."

[0026] Definition 10: Classic Feature Pyramid Network

[0027] Feature Pyramid Networks (FPNs) are a top-down architecture with lateral connections used to construct high-level semantic feature maps at all scales. FPNs primarily address the multi-scale problem in object detection, significantly improving the detection performance of small objects by simply changing network connections without substantially increasing the computational cost of the original model. For details on Feature Pyramid Networks, see the literature “T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan and S. Belongie, “Feature Pyramid Networks for Object Detection,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 936-944, doi:10.1109 / CVPR.2017.106.” Definition 11: Traditional Content-Aware Feature Reconstruction Methods

[0028] Traditional content-aware feature reconstruction (CARAFE) methods possess large receptive field information aggregation and adaptive content recognition capabilities, comprising a kernel prediction module and a content-aware reconstruction module. Specifically, CARAFE can generate a lower-level feature layer P1 from the original bottom-level feature layer P2, thereby providing rich spatial location information and improving the feature representation capability of small-scale targets. Using CARAFE to generate the additional feature layer P1 can be described as follows:

[0029] P1 = CARAFE ×2 (P2)

[0030] Among them, CARAFE ×2 This indicates a 2x upsampling operation performed using CARAFE. When using the CARAFE module, an input feature map and an upsampling ratio must be provided first, and CARAFE will generate an output feature map. First, the kernel prediction module predicts and reassembles an adaptive quadratic upsampling kernel K of size n×n, representing an n×n pixel neighborhood at location I. Convolutional layers are then used for channel compression, generating n channels. 2 ×2 2 The reassembly kernel is then used. The content-aware reassembly module provides the reassembly feature map generated by the convolutional layer, resulting in the upsampled feature map P1. For details on the traditional content-aware feature reassembly (CARAFE) method, please refer to the literature "Wang, J., Chen, K., Xu, R., Liu, Z., Loy, CC, Lin, D., 2019. CARAFE: Content-aware reassembly of features. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 3007–3016.".

[0031] Definition 12: Classic Two-Stage ACmix

[0032] ACmix is ​​a fusion method of convolution and self-attention mechanisms, which can be used to refine the contextual representation of different feature levels of integrated feature maps, ultimately improving multi-scale feature representation. ACmix breaks down the barriers between the two attention mechanisms by exploring the similarities between convolution and self-attention in 1×1 convolution operations and fusing them together. ACmix consists of two stages. The final output is a weighted sum of the convolutional path and the self-attention path.

[0033] F out =αF conv +βF self

[0034] Among them, F outIt is the output feature, F conv It is the output feature of the convolution path, F self α and β are the output features of the self-attention path, and α and β are the learnable weight coefficients involved in network training. For details on the classic two-stage ACmix, please refer to the literature "Pan, X., Ge, C., Lu, R., Song, S., Chen, G., Huang, Z., Huang, G., 2022. On the integration of self-attention and convolution. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 805-815.".

[0035] Definition 13: Traditional intermediate feature set generation methods

[0036] In the first stage of ACmix, the input features need to be projected through three 1×1 convolutions and reshaped into N blocks, resulting in an intermediate feature set consisting of 3 x N feature maps. The core idea of ​​using 1x1 convolutions for projection is to use the 1x1 convolution kernel to linearly transform or adjust the number of channels (dimensions) of the input feature map without changing the spatial dimensions (width and height). For traditional intermediate feature set generation methods, please refer to the literature "Pan, X., Ge, C., Lu, R., Song, S., Chen, G., Huang, Z., Huang, G., 2022. On the integration of self-attention and convolution. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 805-815."

[0037] Definition 14: Traditional Convolutional Path Method

[0038] In the second stage of ACmix, one path is called the convolutional path. This involves generating a feature map using a lightweight fully connected layer, then performing convolution on the generated feature map through shift and summation operations, and extracting the information contained within from the perspective of the local receptive field. For details on traditional convolutional path methods, see the literature "Pan, X., Ge, C., Lu, R., Song, S., Chen, G., Huang, Z., Huang, G., 2022. On the integration of self-attention and convolution. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 805-815." Definition 15: Traditional Self-Attention Path Method

[0039] In the second stage of ACmix, another path is called the self-attention path. This involves dividing the intermediate features generated in the first stage into N groups, each containing three feature maps. These three feature maps are then processed using multi-head self-attention through attention and aggregation operations, serving as the query, key, and value, respectively. For details on traditional self-attention path methods, please refer to the literature "Pan, X., Ge, C., Lu, R., Song, S., Chen, G., Huang, Z., Huang, G., 2022. On the integration of self-attention and convolution. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 805-815."

[0040] Definition 16: Traditional Cascaded Neural Network Method

[0041] Traditional cascaded neural network methods involve connecting multiple neural network models according to specific rules. A cascaded neural network typically consists of multiple independent subnetworks, each responsible for a specific task or extracting specific features. By linking multiple subnetworks together, cascaded neural networks can progressively extract and refine features from the input data, thereby improving the overall model performance. For details on traditional cascaded neural network methods, please see the website "https: / / www.kepuchina.cn / article / articleinfo?business_type=100&ar_id=249206".

[0042] Definition 17: Traditional Cascaded Random Box Filling Strategy

[0043] Cascaded random box filling strategy refers to generating noisy random boxes step by step, starting from the ground truth boxes of the data. Each level of generated random boxes adds different levels of Gaussian noise to the original ground truth boxes. In each cascade, the intensity of the Gaussian noise gradually increases, perturbing the position, size, or shape of the boxes. Traditional cascaded random box filling strategies can be found in the literature "Song, J., Meng, C., & Ermon, S., 2020. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502."

[0044] Definition 18: Traditional diffusion model method

[0045] The diffusion model is a likelihood-based latent variable model that maps data to a latent space using Markov chains. The diffusion model consists of two steps: a forward diffusion process and a backward denoising process. The former adds noise to the data progressively at each time step until completely random noise is obtained; the latter denoises the data progressively at each sampling step until the ground truth is obtained. Specifically, given an input image distribution x0~q(x0), the detailed implementation formula for the forward diffusion process is defined as:

[0046]

[0047] Where, x t This represents a noise sample with a time step of t∈{0,1,…,T}; β t This represents the variance table as Gaussian noise is gradually added to x0, β t Set as a linearly increasing variable from 0.0004 to 0.02, with an upper limit of time step T of 1000.

[0048] Therefore, given an original input x0, by gradually adding Gaussian noise... Until a completely random noise sample is obtained (when t = T), as follows:

[0049]

[0050] in,

[0051] During training, network f θ (x t ,t) will attempt to minimize the l2 norm loss from the noisy sample x tFit the original target sample x0:

[0052]

[0053] During the reasoning process, network f θ (x t ,t) Try to get from noisy sample x t The original target sample x0 is gradually reconstructed, and there is an update rule for each sample step, namely X. T →X T-Δ →X0.

[0054] In ship instance segmentation, the diffusion model is described as a denoising process from noisy boxes to target boxes and a reconstruction process from target ship boxes to instances.

[0055] Traditional diffusion model methods can be obtained from the literature "Song, J., Meng, C., & Ermon, S., 2020. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502."

[0056] Definition 19: Classical Cross-Union Ratio Loss Function

[0057] The Intersection over Union (IoU) ratio is defined as the ratio of the intersection of the predicted bounding box and the union of the ground truth bounding box. The formula is as follows:

[0058]

[0059] Here, A represents the predicted bounding box, B represents the ground truth bounding box, ∩ represents the intersection operation, and ∪ represents the union operation. Typically, during model training, the IoU value is used to determine whether the result is good or bad.

[0060] When the intersection-union ratio (CIU) is used as the loss function, the commonly used calculation formula is:

[0061]

[0062] The Intersection over Union (IoU) loss function better reflects the degree of overlap and is scale-invariant; however, when the predicted bounding box and the ground truth bounding box have no intersection, the IoU loss cannot be calculated, resulting in a gradient of 0 and preventing optimization. Detailed information about the classic IoU loss function can be found at "https: / / blog.csdn.net / weixin_43482623 / article / details / 119574564".

[0063] Definition 20: Traditional method for calculating total training loss

[0064] During training, the total training loss consists of detection loss and segmentation loss:

[0065]

[0066] in, Indicates detection loss, Let λ represent the segmentation loss. bal This represents the hyperparameter balancing detection and segmentation. Specifically, and It can be represented as:

[0067]

[0068] Where, λ cls λ represents the classification weight. reg λ represents the regression weights. IoU Indicates the intersection-union ratio weights; Represents the classification loss, α t p represents the balance factor, γ represents the modulation factor, and p represents the balance factor. t Indicates the predicted category probability; Let y represent the regression loss, n represent the number of samples, and y represent the regression loss. i Let f(X) represent the predicted value of the i-th sample. i ) represents the true value of the i-th sample; This indicates the focus on the intersection and comparison loss; m represents the predicted ship mask, m gt This indicates the actual ship mask.

[0069] Definition 21: Traditional Denoising Diffusion Implicit Model Method

[0070] Denoising diffusion implicit models (DDIM) are proposed based on denoising diffusion probabilistic models. The training process is the same as that of denoising diffusion probabilistic models, generating samples by simulating Markov chains. The difference lies in that DDIM is a more efficient iterative implicit probabilistic model, which constructs a class of non-Markov diffusion processes. Training the reverse process with the same objective allows for faster sampling. For details on traditional denoising diffusion implicit model methods, please refer to the literature Song, J., Meng, C., & Ermon, S., 2020. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502.

[0071] Definition 22: Reconstructing Ship Instance Mask

[0072] During inference, starting with noisy boxes with a Gaussian distribution, ground truth boxes conforming to the target distribution are gradually generated through a denoising process. A denoising model is used to estimate the detection boxes for the next sampling step, and a box update strategy is employed to better align inference with training. Box update refers to filtering out low-scoring boxes after each sampling step, then concatenating the remaining boxes with newly sampled random boxes from a Gaussian distribution. Instance-specific filters are then established, ultimately reconstructing the ship instance mask m.

[0073] b0 = f(...(f(b) T-s ,Ts))),s={0,...,T}

[0074] θ0=η(b0)

[0075] m=φ(F mask ;θ0)

[0076] Where b0 is a true bounding box that conforms to the target distribution, generated step by step through the denoising process; b T-s η is a noise box that conforms to a Gaussian distribution, f(·) is a denoising network, η(·) is a fully connected layer, and θ0 is an instance-specific filter.

[0077] Definition 23: Traditional polarization self-attention module method

[0078] Self-attention (SA), also known as internal attention, is an attention mechanism that associates different positions within a single sequence to compute a representation of the same sequence. The Polarimetric Self-Attention Module (PSA) consists of two branches: a channel self-attention branch and a spatial self-attention branch. Through polarimetric filtering and enhancement, it achieves high-quality pixel regression, ensuring effective information integration. Specifically, given an input feature map... The process can be described as follows:

[0079] F ‘’ =f psa (F′)

[0080] F′=f re (concat(conv 1×3 (F),conv 3×3 (F),conv 3×1 (F)))

[0081] Where, conv 1×3 conv represents a 1×3 convolution operation. 3×3 conv represents a 3×3 convolution operation. 3×1 This represents a 3×1 convolution operation; f psa(·) represents the PSA operator; concat(·) represents the concatenation operator; f re (·) denotes a feature refinement operation, such as a 1×1 convolutional layer. For details on traditional polarized self-attention module methods, see the literature "Liu, H., Liu, F., Fan, X., Huang, D., 2021. Polarized self-attention: Towards high-quality pixel-wise regression. arXiv preprint arXiv:2107.00782.". Definition 24: Conditional convolution methods for traditional instance segmentation.

[0082] Conditional convolutions (CondInst) for instance segmentation employ a dynamic instance-aware network conditioned on instances. Instance segmentation is solved by a fully convolutional network, eliminating the need for ROI cropping and feature alignment. Furthermore, the dynamically generated conditional convolutions significantly increase capacity, allowing for very compact masking and significantly accelerating inference speed. For details on traditional conditional convolutions for instance segmentation, please refer to the literature "Tian, ​​Z., Shen, C., Chen, H., 2020. Conditional convolutions for instancesegmentation. In: Proceedings of the IEEE European Conference on ComputerVision. pp. 282–298."

[0083] Definition 25: Classic AdamW Optimization Method

[0084] The classic Adam optimization method is a first-order optimization algorithm that can replace the traditional stochastic gradient descent process. It iteratively updates the weights of the neural network based on training data. Adam designs independent adaptive learning rates for different parameters by calculating the first and second moment estimates of the gradient. AdamW, through a simple modification, decouples the weight decay from the optimization steps of the loss function, thereby restoring the original formula for weight decay regularization and greatly improving the generalization performance of Adam. For details on the classic Adam optimization method, please refer to the literature "Loshchilov, I., Hutter, F., 2019. Decoupled weight decay regularization. In: Proceedings of the IEEE International Conference on Learning Representations. pp. 1–19.".

[0085] Definition 26: Standard Network Testing Methods

[0086] Standard network testing methods refer to performing final testing on the model on a test set to obtain the model's test results on the test set. For details on standard network testing methods, please refer to the literature "C. Lu, and W. Li, "Ship Classification in High-Resolution SAR Images via Transfer Learning with Small Training Dataset," Sensors, vol. 19, no. 1, pp. 63, 2018."

[0087] Definition 27: Classical COCO metric

[0088] The COCO metric is used as the evaluation standard. The core criterion of the COCO metric is the average precision (AP), which is the average precision across 10 IoU thresholds, ranging from 0.50 to 0.95 with intervals of 0.05. 50 This represents the average precision when the IoU threshold is 0.50; AP 75 This represents the average accuracy when the IoU threshold is 0.75; AP S This represents the average precision of small pixel targets (ships <32). 2 (pixels); AP M The average precision (32) represents the target with medium pixel count. 2 <Ships<96 2 (pixels); AP L This represents the average precision of large pixel targets (ships > 96). 2 (pixels). In ship instance segmentation, the average precision AP value is represented by mask average precision (maskAP). In addition, the number of parameters (#Para) is used to measure the computational complexity of the algorithm, and the number of floating-point operations (FLOPs) is used to measure the space complexity of the algorithm. The smaller the values ​​of both, the lighter the network is. For the standard COCO metric, please refer to the literature "Zhang,T.,Zhang,X.,Ke,X.,Zhan,X.,Shi,J.,Wei,S.,Pan,D.,Li,J.,Su,H.,Zhou,Y.,Kumar,D.,2020.LS-SSDD-v1.0:Adeep learning dataset dedicated to small ship detection from large-scale sentinel-1SAR images.Remote Sens.12(18),2997.".

[0089] Definition 28: Nine conventional methods with similar functions to this invention

[0090] They are:

[0091] 1. Traditional Mask Object Instance Segmentation Framework Method (Mask R-CNN), see the literature "He, K., Gkioxari, G., Dollár, P., & Girshick, R., 2017. Mask R-CNN. In: Proceedings of the IEEE International Conference on Computer Vision.pp.2980–2988.";

[0092] 2. Traditional multi-stage object detector method (Cascade Mask R-CNN), see the literature "Cai, Z., Vasconcelos, N., 2019. Cascade R-CNN: High quality object detection and instancesegmentation. IEEE Trans. Pattern Anal. Mach. Intell. 43(5), 1483–1498.";

[0093] 3. The traditional real-time instance segmentation model (YOLACT) method, see the literature "Bolya, D., Zhou, C., Xiao, F., Lee, Y".

[0094] J., 2019. YOLACT: Real-time instance segmentation. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 9156–9165.";

[0095] 4. The traditional conditional convolution method (CondInst) for instance segmentation is detailed in the literature "Tian, ​​Z., Shen, C., Chen, H., 2020. Conditional convolutions for instance segmentation. In: Proceedings of the IEEE European Conference on Computer Vision. pp. 282–298.";

[0096] 5. Traditional efficient real-time target detector methods (RTMDet), see the literature "Lyu, C., Zhang, W., Huang, H.,

[0097] Zhou,Y.,Wang,Y.,Liu,Y.,Zhang,S.,Chen,K.,2022.RTMDet: An empirical study of designing realtime object detectors.arXiv preprint arXiv:2212.07784.";

[0098] 6. Traditional instance segmentation model method based on rotated bounding boxes (SRNet), see the literature "Yang, X., Zhang, Q.,

[0099] Dong,Q.,Han,Z.,Luo,X.,Wei,D.,2023.Ship Instance Segmentation BasedonRotated Bounding Boxes for SAR Images.Remote Sens.15(5),pp.1324.";

[0100] 7. Traditional instance segmentation model method based on Gaussian heatmap (BFSS-Inst), see the literature "Gao, F., Zhong, F., Sun,

[0101] J.,Hussain,A.,Zhou,H.,2024.BBox-free SAR ship instance segmentationmethodbased on gaussian heatmap.IEEE Trans.Geosci.Remote Sens.62,1–18.";

[0102] 8. Traditional high-quality instance segmentation network method for remote sensing images (HQ-ISNet), see the literature "Su, H., Wei, S., Liu, S.,

[0103] Liang,J.,Wang,C.,Shi,J.,Zhang,X.,2020.HQ-ISNet: High-quality instance segmentation for remote sensing imagery.Remote Sens.12(6),989.";

[0104] 9. Traditional instance segmentation model method based on centroid distance loss (AFSS-Inst), see the literature "Gao, F., Huo, Y.,

[0105] Wang,J.,Hussain,A.,Zhou,H.,2021.Anchor-free SAR ship instancesegmentationwith centroid-distance based loss.IEEE J.Sel.Topics Appl.EarthObserv.Remote Sens.

[0106] 14,11352–11371.

[0107] This invention provides a SAR image ship target segmentation method based on a diffusion model, which includes the following steps:

[0108] Step 1: Prepare the dataset

[0109] For the SSDD dataset provided in Definition 2, the ratio of the training set to the test set is set to 8:2, and the training set is denoted as SSDD. train The test set is denoted as SSDD. test ;

[0110] For the HRSID dataset provided in Definition 3, the ratio of the training set to the test set is set to 13:7, and the training set is denoted as HRSID. train The test set is denoted as HRSID test .

[0111] Step 2: Encoder Construction

[0112] Step 2.1: Backbone Network Construction

[0113] For the classic ResNet-101 network model in Definition 4, the traditional deformable convolution method in Definition 5 is used to process it, resulting in the backbone classic convolutional neural network of the SAR image ship target segmentation network in Definition 6, denoted as Backbone.

[0114] Using the classic CNN feature extraction method in Definition 7, the training set SSDD obtained in step 1 is processed. train or HRSID train Any SAR image in the dataset is processed on the classical convolutional neural network backbone obtained in the previous step, and the feature output is denoted as F. in ;

[0115] Step 2.2: Construction of Spatial Context Jointly Enhanced Feature Pyramid Network

[0116] The sampling kernel size n is set to 5, and the traditional recombination kernel generation method in Definition 9 is used to process the classic convolutional layer in Definition 8, generating 5 channels. 2 ×22 The recombinant nucleus;

[0117] The feature output F obtained in step 2.1 in As input, the traditional content-aware feature reorganization method in Definition 11 is used to process the classic feature pyramid network in Definition 10, generating additional low-level feature maps to obtain the spatial information enhancement module, denoted as SIE. The output feature of the SIE module is denoted as F. CARAFE ;

[0118] For the classic two-stage ACmix in Definition 12, in the first stage, the feature output F obtained above is... CARAFE As input, the traditional intermediate feature set generation method in Definition 13 is used to obtain the intermediate feature set, denoted as F. m ;

[0119] In the second stage, the intermediate feature set F obtained in the first stage is... m As input, the traditional convolutional path method in Definition 14 and the traditional self-attention path method in Definition 15 are used to obtain the semantic information enhancement module, denoted as CIE. The weighted sum of the features of the two paths output by the CIE module is denoted as F. out .

[0120] Thus, the SIE and CIE modules obtained in the first two steps are cascaded using the traditional cascaded neural network method in Definition 16 to obtain the Spatial Context Joint Enhanced Feature Pyramid Network, denoted as SCJE-FPN.

[0121] The feature output F obtained in step 2.1 in The input is fed into the Spatial Context Joint Enhanced Feature Pyramid Network (SCJE-FPN), and the output of the Spatial Context Joint Enhanced Feature Pyramid Network (SCJE-FPN) is the deep feature F after spatial context joint enhancement. out The encoder part of the SAR image ship target segmentation network based on the diffusion model has been completed.

[0122] Step 3: Construction of Ship Target Segmentation Network

[0123] During training, the deep features F in step 2 are... out Using the traditional cascaded random box filling strategy method in Definition 17, the noise box B is obtained. t ;

[0124] For the noise box B obtained in the previous step t The traditional diffusion model method defined in Definition 18 is used for processing to obtain an instance-specific filter θ0 = η(f(B)). t ,t)), where f(·) is the denoising network and η(·) is the fully connected layer;

[0125] For the classic cross-union ratio (CUI) loss function in Definition 19, by adding a new focusing factor α, we obtain the focused CUI loss function, denoted as FIoU:

[0126] Specifically, given a predicted bounding box B and a ground truth bounding box B gt The FIoU loss function is defined as follows:

[0127]

[0128] Where p1 is the coordinate of the point at the top left of the prediction box, and p2 is the coordinate of the point at the bottom right of the prediction box. The coordinates of the top-left point of the true bounding box. Let be the coordinates of the bottom right point of the ground truth bounding box, (·) be the Euclidean distance, and c be the diagonal length of the minimum bounding box between the predicted and ground truth bounding boxes. It is the focus crossover ratio, α is the focus factor with a value of 0.95, and IoU is the crossover ratio;

[0129] The focused intersection-union ratio (FIoU) loss function obtained in the previous step is processed using the traditional training total loss calculation method in Definition 20 to obtain the total loss function during training. in Indicates detection loss, Let λ represent the segmentation loss. bal This represents the balance hyperparameter between detection and segmentation, with a value of 5.

[0130] During the inference process, the noisy box B from the training process... t After training the network using the traditional denoising diffusion implicit model method in Definition 21, the reconstructed ship instance mask m in Definition 22 is obtained.

[0131] At this point, the decoder module of the SAR image ship target segmentation network based on the diffusion model has been completed.

[0132] The feature output F from step 2 out As input, the traditional polarization self-attention module method in Definition 23 is used for information integration to obtain a multiple asymmetric convolutional self-attention module, denoted as MACSA.

[0133] The MACSA module obtained in the previous step is processed using the conditional convolution method for traditional instance segmentation in Definition 24 to obtain the instance-aware mask representation module, denoted as IAMR.

[0134] This completes the construction of the SAR image ship target segmentation network based on the diffusion model, denoted as DiffSARShipInst.

[0135] Step 4: Establish a ship segmentation model

[0136] The training set SSDD obtained in step 1 train or HRSID train As input, the classic AdamW optimization method in Definition 25 is used to optimize and train the SAR image ship target segmentation network DiffSARShipInst based on the diffusion model completed in step 3. After training, the ship segmentation model is obtained, denoted as DiffSARShipInst_Model.

[0137] Step 5: Test the ship segmentation model

[0138] Using the test set SSDD obtained in step 1 test or HRSID test The ship segmentation model DiffSARShipInst_Model obtained in step 4 is tested using the standard network testing method in definition 26, and the test results of the test set on DiffSARShipInst_Model are obtained, denoted as Result.

[0139] Step 6: Evaluate the ship segmentation model

[0140] Using the test result Result of the ship segmentation model obtained in step 5 as input, the classic COCO metric method in definition 27 is used to obtain the mask average accuracy maskAP.

[0141] This concludes the entire method.

[0142] The innovation of this invention lies in introducing a diffusion model into the original SAR ship instance segmentation, thus addressing the ship instance segmentation task from a generative perspective for the first time. It also provides several effective improvements, including a spatial context-jointly enhanced feature pyramid network, a focused intersection-over-union loss function, and instance-aware mask representation, to ensure superior instance segmentation accuracy. Experimental results on the SAR Ship Detection Dataset (SSDD) and the High Resolution SAR Image Dataset (HRSID) show that this invention achieves the best average mask accuracy in nearshore and offshore scenes.

[0143] The advantage of this invention is that it represents the SAR ship instance segmentation task as a denoising process from a noise box to a target box and a reconstruction process from a target box to a ship instance, thus avoiding the difficulties of setting target prior hyperparameters in traditional segmentation methods and the time-consuming and laborious hyperparameter adjustment process. Attached Figure Description

[0144] Figure 1 This is a flowchart illustrating the SAR image ship target segmentation method of the present invention.

[0145] Figure 2 This table compares the average accuracy of the SAR image ship target segmentation method in this invention with that of different ship instance segmentation methods on the SSDD dataset.

[0146] Where: AP represents average accuracy, the best model is marked in bold, and the second best model is marked in underline.

[0147] Figure 3 This table compares the average accuracy of the SAR image ship target segmentation method in this invention with that of different ship instance segmentation methods on the HRSID dataset.

[0148] Where: AP represents average accuracy, the best model is marked in bold, and the second best model is marked in underline.

[0149] Figure 4 This table compares the model complexity of the SAR image ship target segmentation method in this invention with different ship instance segmentation methods.

[0150] in: This represents the FLOP calculated based on SSDD; Represents the FLOP calculated based on HRSID; #Para represents the complexity evaluation index calculated based on SSDD / HRSID. The best complexity evaluation index value is marked in bold, and the second best complexity evaluation index value is marked with a subscript. Detailed Implementation

[0151] The following is in conjunction with the appendix Figure 1 The present invention will be described in further detail below.

[0152] Step 1: Prepare the dataset

[0153] For the SSDD dataset provided in Definition 2, the ratio of the training set to the test set is set to 8:2, and the training set is denoted as SSDD. train The test set is denoted as SSDD. test ;

[0154] For the HRSID dataset provided in Definition 3, the ratio of the training set to the test set is set to 13:7, and the training set is denoted as HRSID. train The test set is denoted as HRSID test .

[0155] Step 2: Encoder Construction

[0156] Step 2.1: Backbone Network Construction

[0157] For the Resnet-101 network model in Definition 4, the deformable convolution method in Definition 5 is used to obtain the backbone convolutional neural network of the SAR image ship target segmentation network in Definition 6, denoted as Backbone;

[0158] Using the classic CNN feature extraction method in Definition 7, the training set SSDD obtained in step 1 is processed. train or HRSID train A SAR image is processed on the backbone convolutional neural network obtained in the previous step, and the feature output is denoted as F. in ;

[0159] Step 2.2: Construction of Spatial Context Jointly Enhanced Feature Pyramid Network

[0160] For the classic convolutional layer in Definition 8, the sampling kernel size n is set to 5, and the traditional recombination kernel generation method in Definition 9 is used to generate 5 channels. 2 ×2 2 The recombinant nucleus;

[0161] For the classic feature pyramid network in Definition 10, the feature output F obtained in step 2.1 is... in As input, the content-aware feature reorganization technique in Definition 11 is used to generate additional low-level feature maps, resulting in a spatial information enhancement module, denoted as SIE. The output feature of the SIE module is denoted as F. CARAFE ;

[0162] The feature output F from the previous step CARAFE As input, in the first stage of the classic two-stage ACmix in Definition 12, the traditional intermediate feature set generation method in Definition 13 is used for processing to obtain the intermediate feature set, denoted as F. m ;

[0163] The intermediate feature set F obtained in the first stage m As input, in the second stage of the classic two-stage ACmix in Definition 12, the traditional convolutional path method in Definition 14 and the traditional self-attention path method in Definition 15 are used to obtain the semantic information enhancement module, denoted as CIE. The weighted sum of the features of the two paths output by the CIE module is denoted as F. out .

[0164] Thus, the SIE and CIE modules obtained in the first two steps are cascaded using the cascaded neural network in Definition 16 to obtain the Spatial Context Joint Enhanced Feature Pyramid Network, denoted as SCJE-FPN.

[0165] The feature output F obtained in step 2.1 inAs input to the spatial context joint augmentation feature pyramid network SCJE-FPN, the feature is strong, and the output is the deep feature F after spatial context joint augmentation. out The encoder part of the SAR image ship target segmentation network based on the diffusion model has been completed.

[0166] Step 3: Construction of Ship Target Segmentation Network

[0167] During training, the feature output F in step 2 is... out Using the cascaded random box filling strategy in Definition 17, noise box B is obtained. t ;

[0168] For the noise box B obtained in the previous step t The denoising network in the diffusion model defined in Definition 18 is used for processing to obtain an instance-specific filter θ0 = η(f(B)). t ,t)), where f(·) is the denoising network and η(·) is the fully connected layer;

[0169] For the cross-union ratio (CU) loss function in Definition 19, a focused CU loss function, denoted as FIoU, is obtained by adding a new focusing factor α. Specifically, given a predicted bounding box B and a ground truth bounding box B... gt The FIoU loss is defined as follows:

[0170]

[0171] Where p1 is the coordinate of the point at the top left of the prediction box, and p2 is the coordinate of the point at the bottom right of the prediction box. The coordinates of the top-left point of the true bounding box. Let be the coordinates of the bottom right point of the ground truth bounding box, (·) be the Euclidean distance, and c be the diagonal length of the minimum bounding box between the predicted and ground truth bounding boxes. It is the focus crossover ratio, α is the focus factor with a value of 0.95, and IoU is the crossover ratio;

[0172] Using the focus intersection-over-union (CIU) loss function obtained in the previous step, the total training loss calculation method in Definition 20 is used to obtain the total loss function during the training process. in Indicates detection loss, Let λ represent the segmentation loss. bal This represents the balance hyperparameter between detection and segmentation, with a value of 5.

[0173] During the inference process, the noisy box B from the training process... t After training the network using the denoising diffusion implicit model in Definition 21, the reconstructed ship instance mask m in Definition 22 is obtained.

[0174] At this point, the decoder module of the SAR image ship target segmentation network based on the diffusion model has been completed.

[0175] The feature output F from step 2 out As input, the polarization self-attention module in Definition 23 is used for information integration to obtain the multiple asymmetric convolutional self-attention module, denoted as MACSA.

[0176] For the MACSA module obtained in the previous step, the conditional convolution method for instance segmentation in Definition 24 is used to obtain the instance-aware mask representation module, denoted as IAMR.

[0177] This completes the construction of the SAR image ship target segmentation network based on the diffusion model, denoted as DiffSARShipInst.

[0178] Step 4: Establish a ship segmentation model

[0179] The training set SSDD obtained in step 1 train or HRSID train As input, the classic AdamW algorithm in Definition 25 is used to optimize and train the SAR image ship target segmentation network DiffSARShipInst based on the diffusion model completed in step 3. After training, the ship segmentation model is obtained, denoted as DiffSARShipInst_Model.

[0180] Step 5: Test the ship segmentation model

[0181] Using the test set SSDD obtained in step 1 test or HRSID test The ship segmentation model DiffSARShipInst_Model obtained in step 4 is tested using the standard network testing method in definition 26, and the test results of the test set on DiffSARShipInst_Model are obtained, denoted as Result.

[0182] Step 6: Evaluate the ship segmentation model

[0183] Using the test result Result of the ship segmentation model obtained in step 5 as input, the COCO metric method in definition 27 is used to obtain the mask average accuracy maskAP.

[0184] This concludes the entire method.

[0185] As can be seen from the above specific embodiments, the present invention can effectively segment ship target instances in SAR images. Figure 2As shown, the AP value achieved by this invention on the SSDD dataset is 70.6% for offshore scenarios and 56.2% for nearshore scenarios; Figure 3 As shown, the AP value achieved by this invention on the HRSID dataset is 70.9% for offshore scenarios and 42.6% for nearshore scenarios, both of which are the best values ​​among all comparison methods. Figure 4 As shown, compared with nine traditional methods with similar functions, this invention exhibits moderate model computational complexity and the highest space complexity. In practical applications, it offers advantages such as higher segmentation accuracy for slightly more complex models. Therefore, this invention can be widely applied to SAR image ship target instance segmentation tasks, and its model complexity can be further improved in the future.

Claims

1. A SAR image ship target segmentation method based on a diffusion model, characterized by its... Includes the following steps: Step 1: Prepare the dataset For the SSDD dataset, the ratio of training set to test set is set to 8:2, and the training set is denoted as SSDD. train The test set is denoted as SSDD. test ; For the HRSID dataset, the ratio of the training set to the test set is set to 13:7, and the training set is denoted as HRSID. train The test set is denoted as HRSID test ; Step 2: Encoder Construction Step 2.1: Backbone Network Construction For the classic ResNet-101 network model, the traditional deformable convolution method is used to process it, and the backbone classical convolutional neural network of the SAR image ship target segmentation network is obtained, denoted as Backbone. The classic CNN feature extraction method is used to extract features from the training set SSDD obtained in step 1. train or HRSID train Any SAR image in the dataset is processed on the classical convolutional neural network backbone obtained in the previous step, and the feature output is denoted as F. in ; Step 2.2: Construction of Spatial Context Jointly Enhanced Feature Pyramid Network The sampling kernel size n is set to 5, and the classic convolutional layer is processed using the traditional kernel reconstruction method, generating 5 channels. 2 ×2 2 The recombinant nucleus; The feature output F obtained in step 2.1 in As input, the classic feature pyramid network is processed using a traditional content-aware feature reconstruction method to generate additional low-level feature maps, resulting in a spatial information enhancement module, denoted as SIE. The output feature of the SIE module is denoted as F. CARAFE ; For the classic two-stage ACmix, in the first stage, the feature output F obtained above is... CARAFE As input, a traditional intermediate feature set generation method is used to obtain an intermediate feature set, denoted as F. m ; In the second stage, the intermediate feature set F obtained in the first stage is... m As input, the input is processed using traditional convolutional path methods and traditional self-attention path methods to obtain a semantic information enhancement module, denoted as CIE. The weighted sum of the features from the two paths output by the CIE module is denoted as F. out ; Thus, the SIE and CIE modules obtained above are cascaded using the traditional cascaded neural network method to obtain the Spatial Context Joint Enhanced Feature Pyramid Network, denoted as SCJE-FPN; The feature output F obtained in step 2.1 in The input is fed into the Spatial Context Joint Enhancement Feature Pyramid Network (SCJE-FPN), and the output of the Spatial Context Joint Enhancement Feature Pyramid Network (SCJE-FPN) is the deep feature F after spatial context joint enhancement of the features. out The encoder part of the SAR image ship target segmentation network based on the diffusion model has been completed; Step 3: Construction of Ship Target Segmentation Network During training, the deep features F in step 2 are... out Using the traditional cascaded random box filling strategy, the noise box B is obtained. t ; For the noise box B obtained in the previous step t The traditional diffusion model method is used to process the data, resulting in an instance-specific filter θ0 = η(f(B)). t ,t)), where f(·) is the denoising network and η(·) is the fully connected layer; For the classic cross-union ratio (CUR) loss function, by adding a new focusing factor α, we obtain the focused CUR loss function, denoted as FIoU: Specifically, given a predicted bounding box B and a ground truth bounding box B gt The FIoU loss function is defined as follows: Where p1 is the coordinate of the point at the top left of the prediction box, and p2 is the coordinate of the point at the bottom right of the prediction box. The coordinates of the top-left point of the true bounding box. Let ρ be the coordinates of the bottom right point of the ground truth bounding box, ρ(·) be the Euclidean distance, and c be the diagonal length of the minimum bounding box between the predicted and ground truth bounding boxes. It is the focus crossover ratio, α is the focus factor with a value of 0.95, and IoU is the crossover ratio; The Focused Intersection over Union (FIoU) loss function obtained in the previous step is processed using the traditional method for calculating the total training loss to obtain the total loss function during the training process. in Indicates detection loss, Let λ represent the segmentation loss. bal This represents the balance hyperparameter between detection and segmentation, with a value of 5; During the inference process, the noisy box B from the training process... t After training the network using the traditional denoising diffusion implicit model method, the reconstructed ship instance mask m is obtained; At this point, the decoder module of the SAR image ship target segmentation network based on the diffusion model has been completed; The feature output F from step 2 out As input, information is integrated using the traditional polarization self-attention module method to obtain a multiple asymmetric convolutional self-attention module, denoted as MACSA. The MACSA module obtained in the previous step is processed using the traditional conditional convolution method for instance segmentation to obtain the instance-aware mask representation module, denoted as IAMR. This completes the construction of the SAR image ship target segmentation network based on the diffusion model, denoted as DiffSARShipInst; Step 4: Establish a ship segmentation model The training set SSDD obtained in step 1 train or HRSID train As input, the classic AdamW optimization method is used to optimize and train the SAR image ship target segmentation network DiffSARShipInst based on the diffusion model completed in step 3. After training, the ship segmentation model is obtained, denoted as DiffSARShipInst_Model. Step 5: Test the ship segmentation model Using the test set SSDD obtained in step 1 test or HRSID test The ship segmentation model DiffSARShipInst_Model obtained in step 4 is tested using standard network testing methods to obtain the test results of the test set on DiffSARShipInst_Model, denoted as Result; Step 6: Evaluate the ship segmentation model Using the test result Result of the ship segmentation model obtained in step 5 as input, the mask average accuracy maskAP is obtained by using the classic COCO metric method. This concludes the entire method.

Citation Information

Patent Citations

  • Improvement in step-spindles

    US101109A

  • Diffusion model-based radar image aircraft target detection method

    CN117788956A

  • Diffusion model-based copy movement tampering detection method

    CN118038090A