Semi-supervised cross-modal SAR image ship detection method and system
By constructing a semi-supervised cross-modal SAR image ship detection method, using the teacher network to acquire pseudo-labels and introduce background feature alignment and target enhancement modules, the problem of low detection accuracy of optical image to SAR image is solved, and higher detection accuracy and generalization performance are achieved.
Patent Information
- Application Number
- CN202510340498.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, when the model trained with optical image is used for SAR image ship detection, there is a problem of low accuracy. It is mainly due to the significant differences in imaging mechanism, texture characteristics and noise distribution of optical images and SAR images, resulting in unsatisfactory cross-modal migration task.
A semi-supervised cross-modal SAR image ship detection method is constructed, and a teacher network and a student network are adopted to obtain pseudo-labeled SAR images without labels through the pseudo-label banking strategy, and a background feature alignment module and a target enhancement module are introduced to calculate the background feature alignment loss and target feature alignment loss, and combined with the pseudo-label classification positioning loss and target-background separation confrontation loss, the model training process is optimized.
It improves the accuracy of SAR image ship detection, enhances the feature expression ability and generalization performance of the model, reduces false alarms of pseudo-labels, and improves the detection effect in the target domain.
Smart Images

Figure CN120279309A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar image interpretation, and in particular to a semi-supervised cross-modal SAR image ship detection method and system. Background Art
[0002] As a maritime transportation carrier and an important military target, realizing accurate and real-time detection of ships and predicting their behaviors provides important data support for activities such as military deployment, maritime traffic management, marine rescue, and maritime supervision, and also provides a reliable basis for decision-making. Deep learning technology has become the mainstream technology for ship target detection, but its detection accuracy requires obtaining sufficient labeled data. Due to the complexity of SAR image imaging, compared with optical remote sensing images, labeling ship targets in SAR images requires a lot of domain knowledge, so it requires a high labor cost.
[0003] To solve this problem, existing research has used models trained with optical remote sensing images for SAR image ship detection. However, the distribution differences between optical images and SAR images make the models trained with optical images unable to achieve satisfactory results when used for SAR images. To solve the impact of this domain shift problem on the target detection accuracy, existing methods can be divided into three categories:
[0004] (1) Methods based on image-to-image conversion: For ship detection in SAR images, methods based on image-to-image conversion can help solve problems in cross-modal learning. By converting the ship features in optical images to the SAR image domain, the model can use the dataset of optical images during training to generate target images similar to SAR images and train in SAR images, thereby improving the performance of the model. This method is especially suitable for situations where data is scarce, reducing the need for a large amount of labeled data. For example, the commonly used generative adversarial networks (GANs) is a common image-to-image conversion method, which includes a generator and a discriminator. The generator tries to deceive the discriminator by generating images, while the discriminator evaluates the authenticity of the images.
[0005] (2) Methods based on feature alignment: The goal of feature alignment is to map the features from different modalities to the same space, so that the features of these two modalities can be effectively compared and fused. In this way, even if the original information of the images (such as texture, color, or structure) is quite different, through feature alignment, the model can still extract effective and shared features from different modalities, thereby improving the performance of target detection.
[0006] (3) Pseudo-label self-training method: As one of the most intuitive methods to solve the domain adaptation problem, pseudo-label self-training usually relies on pseudo-labels in the target domain to train the network. However, the pseudo-labels may be inaccurate, and direct use may lead to error accumulation. Therefore, filtering strategies are often adopted, such as removing low-confidence pseudo-labels, to reduce the impact of noise. Compared with feature alignment by minimizing the upper bound of the prediction error, the pseudo-label self-training method is simple and efficient because it directly minimizes the prediction error in the target domain.
[0007] However, most of the above methods focus on the application of domain adaptation in natural images or images of the same modality. The source domain and the target domain usually have a consistent imaging mechanism, and there are only differences in image styles. However, due to the significant differences in imaging mechanisms, texture features, and noise distributions between optical images and SAR images, when these methods are directly applied to the cross-modal transfer task from optical images to SAR images, it is often difficult to achieve ideal results, reducing the accuracy of ship detection using SAR images. Summary of the Invention
[0008] To this end, the technical problem to be solved by the present invention is to overcome the problem of low accuracy of ship detection using SAR images in the prior art.
[0009] To solve the above technical problem, the present invention provides a semi-supervised cross-modal ship detection method for SAR images, including:
[0010] Construct a ship detection model, including a teacher network and a student network; both the teacher network and the student network include a backbone network, a feature pyramid network, and a detector; the student network also includes a background feature alignment module and a target enhancement module;
[0011] Train the ship detection model with labeled optical images, labeled SAR images, and unlabeled SAR images, including:
[0012] Input the unlabeled SAR images into the teacher network to obtain pseudo-labels and construct a pseudo-label bank;
[0013] After passing the labeled optical images, labeled SAR images, and unlabeled SAR images through the backbone network and the feature pyramid network of the student network, obtain the multi-scale features of each image, and then input them into the background feature alignment module to obtain the modal category probability maps of the features at each scale; calculate the feature alignment loss of the features at each scale, and add them up to obtain the background feature alignment loss;
[0014] Input the multi-scale features of the labeled SAR images and labeled optical images into the target enhancement module, and intercept the target features in the corresponding scale features according to the size of the true bounding boxes in each image; calculate the target feature alignment loss according to the modal category probabilities of all target features;
[0015] The unlabeled SAR image is passed through the student network to obtain the predicted localization and the first-class probability of each candidate box; if the unlabeled SAR image exists in the pseudo-label bank, the pseudo-label classification localization loss is calculated.
[0016] The total loss is constructed from the background feature alignment loss, the target feature alignment loss, the pseudo-label classification localization loss, and the supervised classification localization loss; the parameters of the student network are updated with the total loss, and the parameters of the teacher network are assigned with the parameters of the student network.
[0017] The teacher network of the trained ship detection model is used to detect ships in the SAR image.
[0018] Preferably, the detector uses Oriented R-CNN, which includes a region proposal network and an ROI detection head.
[0019] The ROI detection head includes an original regression branch, an original classification branch, and a background difference branch; the background difference branch is used to magnify the candidate box by a preset multiple, and then intercept the target feature from the corresponding scale feature map output by the feature pyramid network according to the size of the magnified candidate box. The target feature is flattened, passed through a fully connected layer, and then passed through a softmax activation layer to obtain the second-class probability.
[0020] Preferably, the unlabeled SAR image is input into the teacher network to obtain pseudo-labels, and a pseudo-label bank is constructed, including:
[0021] The unlabeled SAR image is input into the teacher network, and the first-class probability and the second-class probability of each candidate box are respectively output by the original classification branch and the background difference branch.
[0022] After the first-class probability and the second-class probability of each candidate box are weighted and summed, the target-class probability of each candidate box is obtained; the maximum value in the target-class probability is selected as the confidence score of the candidate box.
[0023] If the confidence score of the candidate box is greater than the first threshold, the candidate box is used as a pseudo-label.
[0024] The confidence score of the unlabeled SAR image is calculated according to the confidence scores of all candidate boxes, and the formula is:
[0025]
[0026] where c i represents the confidence score of the i-th unlabeled SAR image, c i,m represents the confidence score of the m-th candidate box of the i-th unlabeled SAR image, M represents the total number of candidate boxes, and θ ins represents the first threshold. It means that the confidence score of the m-th candidate box in the i-th unlabeled SAR image is greater than the first threshold, and the value is 1;
[0027] If the confidence score of the unlabeled SAR image is greater than the second threshold, the unlabeled SAR image and its pseudo-label are stored in the pseudo-label bank.
[0028] Preferably, the total loss further includes a background difference branch loss, including:
[0029] The labeled SAR image is input into the student network, and the second category probability of each candidate box in the labeled SAR image is output by the background difference branch, and the background difference branch loss is calculated according to the label.
[0030] Preferably, the calculation method of the supervised classification and localization loss includes:
[0031] The labeled SAR image is input into the student network, and the predicted localization and the first category probability of each candidate box in the labeled SAR image are output by the original regression branch and the original classification branch of the detector respectively, and the supervised classification and localization loss of the SAR image is calculated according to the label;
[0032] The labeled optical image is input into the student network, and the predicted localization and the first category probability of each candidate box in the labeled optical image are output by the original regression branch and the original classification branch of the detector respectively, and the supervised classification and localization loss of the optical image is calculated according to the label;
[0033] The supervised classification and localization loss of the SAR image and the supervised classification and localization loss of the optical image are added together to obtain the supervised classification and localization loss.
[0034] Preferably, the background feature alignment module includes a number of non-local discriminant networks, and the number of non-local discriminant networks is the same as the number of levels of the feature pyramid network; each non-local discriminant network includes a first convolutional block, a second convolutional block, a non-local block, a third convolutional block, and a sigmoid activation layer connected in sequence; each convolutional block includes a convolutional layer, a normalization layer, and an activation layer connected in sequence.
[0035] Preferably, the labeled optical image, the labeled SAR image, and the unlabeled SAR image pass through the backbone network and the feature pyramid network of the student network to obtain the multi-scale features of each image, and then are input into the background feature alignment module to obtain the modal category probability maps of the features at each scale; the feature alignment loss of the features at each scale is calculated, and the sum is used as the background feature alignment loss, including:
[0036] The labeled optical image, the labeled SAR image, and the unlabeled SAR image pass through the backbone network and the feature pyramid network of the student network to obtain the corresponding multi-scale features of each image;
[0037] Input the multi-scale features of all images into the corresponding non-local discriminant network according to the size to obtain the modal class probability maps of the optical images and SAR images for each scale feature; among them, the modal class probability map of the SAR image includes the modal class probability maps of the labeled SAR image and the unlabeled SAR image.
[0038] Calculate the feature alignment loss for each scale feature according to the modal class probability map of each scale feature, and add them up to obtain the background feature alignment loss. The formula is:
[0039]
[0040] Among them, L bfa represents the background feature alignment loss, n represents the total number of non-local discriminant networks, represents the modal class probability map of the i-th scale feature of the SAR image, represents the modal class probability map of the i-th scale feature of the optical image; when the image is an optical image, d = 0; when the image is a SAR image, d = 1.
[0041] Preferably, calculate the target feature alignment loss according to the modal class probabilities of all target features, including:
[0042] After the target feature passes through the reverse gradient layer, flattening layer, fully connected layer, and softmax activation layer connected in sequence, the modal class probability of the target feature is obtained.
[0043] Calculate the target feature alignment loss according to the modal class probabilities of all target features. The formula is:
[0044]
[0045] Among them, L ofe represents the target feature alignment loss, N represents the total number of true bounding boxes in the labeled SAR image and the labeled optical image, represents the modal class probability of the j-th target feature of the labeled SAR image, represents the modal class probability of the j-th target feature of the labeled optical image; when the image is an optical image, d = 0; when the image is a SAR image, d = 1.
[0046] Preferably, the total loss further includes a multi-level feature transfer loss, including:
[0047] The unlabeled SAR image passes through the backbone network and feature pyramid network of the teacher network to obtain multi-scale features
[0048] The unlabeled SAR image passes through the backbone network and feature pyramid network of the student network to obtain multi-scale features
[0049] According to multi-scale features and calculate the multi-level feature transfer loss, and the formula is:
[0050]
[0051] where, L mfkt represents the multi-level feature transfer loss, n represents the number of levels of the feature pyramid network, represents the k-th scale feature obtained by inputting the unlabeled SAR image into the teacher network, represents the k-th scale feature obtained by inputting the unlabeled SAR image into the student network.
[0052] The present invention also provides a semi-supervised cross-modal SAR image ship detection system, including:
[0053] A model construction module for constructing a ship detection model, including a teacher network and a student network; both the teacher network and the student network include a backbone network, a feature pyramid network, and a detector; the student network also includes a background feature alignment module and a target enhancement module;
[0054] A training module for training the ship detection model with labeled optical images, labeled SAR images, and unlabeled SAR images, including:
[0055] A pseudo-label acquisition unit for inputting the unlabeled SAR image into the teacher network to obtain pseudo-labels and constructing a pseudo-label bank;
[0056] A background feature alignment loss calculation unit for obtaining the multi-scale features of each image after passing the labeled optical image, labeled SAR image, and unlabeled SAR image through the backbone network and feature pyramid network of the student network, and then inputting them into the background feature alignment module to obtain the modal class probability map of each scale feature; calculating the feature alignment loss of each scale feature and adding them up to obtain the background feature alignment loss;
[0057] A target feature alignment loss calculation unit for inputting the multi-scale features of the labeled SAR image and the labeled optical image into the target enhancement module, intercepting target features in the corresponding scale features according to the size of the true bounding box in each image; calculating the target feature alignment loss according to the modal class probability of all target features;
[0058] A target-background separation adversarial loss calculation unit for constructing a target-background separation adversarial loss with the background feature alignment loss and the target feature alignment loss;
[0059] The pseudo-label classification and localization loss calculation unit is used to obtain the predicted localization and the first-class probability of each candidate box for the unlabeled SAR image through the student network; if the unlabeled SAR image exists in the pseudo-label bank, the pseudo-label classification and localization loss is calculated.
[0060] The parameter update unit is used to construct the total loss with the background feature alignment loss, the target feature alignment loss, the pseudo-label classification and localization loss, and the supervised classification and localization loss; update the parameters of the student network with the total loss, and assign the parameters of the student network to the parameters of the teacher network.
[0061] The detection module is used to detect ships in the SAR image by using the teacher network of the trained ship detection model.
[0062] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0063] A semi-supervised cross-modal SAR image ship detection method according to the present invention uses the mean teacher network as the basic framework. During training, the teacher network adopts the pseudo-label bank strategy to obtain the pseudo-labels of the unlabeled SAR images, and calculates the pseudo-label classification and localization loss to achieve semi-supervised cross-modal detection; further introduces the target-background separation adversarial module, uses the optical image and the SAR image to achieve background feature alignment, and uses the labeled optical image and the labeled SAR image to extract the ship target features on the feature map for feature alignment, so as to avoid the problem of insufficient discriminability of the target features caused by the alignment of the ship target and the background, and calculates the target-background separation adversarial loss according to the background feature alignment loss and the target feature alignment loss; finally, the pseudo-label classification and localization loss and the target-background separation adversarial loss are introduced into the total loss, improving the accuracy of ship detection.
[0064] Furthermore, the present invention introduces the multi-level feature knowledge transfer loss. By comparing the differences between the teacher network and the student network at the feature level, it guides the student network to learn the intermediate layer feature representation of the teacher network, enhances the feature expression ability of the model, and improves its generalization performance in the target domain, thereby improving the accuracy of ship detection.
[0065] Furthermore, due to the interference of onshore ground object targets in the nearshore scene, false alarms occur when the teacher network generates pseudo-labels, thus reducing the reliability of the pseudo-labels. The present invention introduces a background difference branch in the ROI detection head. By expanding the candidate boxes to obtain more background region features as inputs, it can effectively capture the features of the background region. Finally, the prediction results of the original classification branch and the background difference branch are fused to screen and optimize the pseudo-labels, thereby improving the reliability of the pseudo-labels. Description of the Drawings
[0066] To make the content of the present invention easier to be clearly understood, the following further describes the present invention in detail according to specific embodiments of the present invention and in conjunction with the accompanying drawings, where:
[0067] Figure 1 is a flowchart of a semi-supervised cross-modal SAR image ship detection method of the present invention;
[0068] Figure 2 is a structural diagram of a ship detection model;
[0069] Figure 3 is a structural diagram of a target-background separation adversarial network;
[0070] Figure 4 is a structural diagram of a background feature alignment module;
[0071] Figure 5 is a structural diagram of a background-aware collaborative classification module;
[0072] Figure 6 is a schematic diagram of screening pseudo-labels by a pseudo-label bank strategy;
[0073] Figure 7 is a precision-recall curve graph of the method of the present invention and other methods;
[0074] Figure 8 is an example graph of visualized detection results of the method of the present invention and other methods on SSDD data, where Figure 8 (a1) and (a2) in it are the detection results of Unbiased Teacher for near-shore and far-shore images respectively, Figure 8 (b1) and (b2) in it are the detection results of Soft Teacher for near-shore and far-shore images respectively, Figure 8 (c1) and (c2) in it are the detection results of DualTeacher(SSCMOD) for near-shore and far-shore images respectively, Figure 8 (d1) and (d2) in it are the detection results of DualTeacher(SSOD) for near-shore and far-shore images respectively, Figure 8 (e1) and (e2) in it are the detection results of the present invention for near-shore and far-shore images respectively, Figure 8 (f1) and (f2) in it are the true annotations for near-shore and far-shore images respectively.
[0075] Figure 9 is the mAP 50 index graph of the full-supervised method with different proportions of labeled data and the method of the present invention. Detailed implementation manners
[0076] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0077] Embodiment 1
[0078] The existing cross-modal ship detection work is mainly based on the horizontal box detection framework, and the rotation box detection method, which is more applicable to cross-modal ship target detection in SAR images, is still in a blank stage at present. Therefore, the present invention provides a semi-supervised cross-modal SAR image ship detection method, including:
[0079] S1: Construct a ship detection model, the structure of which is referred to Figure 2 as shown, including a teacher network and a student network. Among them, both the teacher network and the student network include a backbone network, a Feature Pyramid Network (FPN), and a detector.
[0080] After the Feature Pyramid Network of the student network, there is also an Object-Background Separation Adversarial Network (OBSA), the structure of which is referred to Figure 3 as shown; the Object-Background Separation Adversarial Network is connected to the Feature Pyramid Network of the student network through a reverse gradient layer, and includes a Background Feature Alignment Module (BFA) and an Object Enhancement Module (OFE).
[0081] The Background Feature Alignment Module includes a number of non-local discriminant networks, and the number of non-local discriminant networks is the same as the number of levels of the Feature Pyramid Network; referring to Figure 4 as shown, each non-local discriminant network includes a first convolutional block, a second convolutional block, a non-local block, a third convolutional block, and a sigmoid activation layer connected in sequence; each convolutional block includes a convolutional layer, a normalization layer, and an activation layer connected in sequence. In this embodiment, the number of levels of the Feature Pyramid Network is 5.
[0082] The Object Enhancement Module is used to intercept object features in the corresponding scale features according to the size of the true bounding box in each image, and output the modal class probability of the object features through a reverse gradient layer, a flattening layer, a fully connected layer, and a softmax activation layer connected in sequence.
[0083] The detector uses Oriented R-CNN for rotation box detection, including a Region Proposal Network (RPN) and an ROI detection head.
[0084] Preferably, the ROI detection head includes an Original Regression Branch (ORB), an Original Classification Branch (OCB), and a Background Difference Branch (BDB). The Original Regression Branch is used to output the predicted location of each candidate box; the Original Classification Branch is used to output the first category probability [p0, p1,..., p c, where p0 represents the first probability of the background, p c represents the first probability of class c, and c is the total number of classes.
[0085] Refer to Figure 5 As shown, the original classification branch and the background difference branch form a background-aware collaborative classification module (BACC). The background difference branch includes a flattening layer, a fully connected layer, and a softmax activation layer connected in sequence. After magnifying the candidate box by a preset multiple, the target feature is intercepted from the corresponding scale feature map output by the feature pyramid network according to the size of the magnified candidate box, and the second-class probability [p'0, p'1, …, p' c is obtained, where p'0 represents the second probability of the background, p' c represents the second probability of class c.
[0086] S2: Train the ship detection model with labeled optical images, labeled SAR images, and unlabeled SAR images.
[0087] S21: Select the ship dataset of the optical remote sensing image with full labels and the ship dataset of the partially labeled SAR image as the training data. As a preferred solution of the present invention, it mainly includes the following steps:
[0088] S211: Scale the optical image and the SAR image to 1024 pixels by the long side, and keep the aspect ratio of the image unchanged during the scaling process;
[0089] S212: Generate each batch of training data, including 1 labeled optical image, 1 labeled SAR image, and 1 unlabeled SAR image;
[0090] S213: Perform weak data augmentation operations on the labeled optical image and the labeled SAR image; the weak data augmentation operations include horizontal flipping, vertical flipping, and diagonal flipping;
[0091] S214: Perform strong data augmentation and weak data augmentation operations on the unlabeled SAR image; the strong data augmentation includes horizontal flipping, vertical flipping, diagonal flipping, contrast enhancement, brightness enhancement, sharpness enhancement, and histogram equalization; the weak data augmentation includes horizontal flipping, vertical flipping, and diagonal flipping.
[0092] S22: Refer to Figure 6 As shown, input the unlabeled SAR image into the teacher network to obtain pseudo-labels, and construct a pseudo-label bank, including:
[0093] Input the unlabeled SAR image into the teacher network, and after passing through the backbone network and the feature pyramid network of the teacher network, obtain the multi-scale features of the unlabeled SAR image
[0094] The multi-scale features of the unlabeled SAR image Input the detector of the teacher network. The candidate boxes of the unlabeled SAR image are output by the RPN, and then the target features are intercepted from the corresponding scale feature maps according to the sizes of the candidate boxes. The target features pass through the original classification branch to respectively output the first-class probabilities of each candidate box in the unlabeled SAR image;
[0095] After magnifying the candidate boxes by a preset multiple, the target features are intercepted from the corresponding scale feature maps according to the sizes of the magnified candidate boxes. The target features pass through the flattening layer, fully connected layer, and softmax activation layer connected in sequence in the background difference branch to output the second-class probabilities of each candidate box in the unlabeled SAR image;
[0096] After weighted summation of the first-class probabilities and the second-class probabilities of each candidate box, the target-class probabilities of each candidate box are obtained; the maximum value in the target-class probabilities is selected as the confidence score of the candidate box. In this embodiment, the weighted summation adopts a ratio of 1:1;
[0097] If the confidence score of the candidate box is greater than the first threshold, the candidate box is used as a pseudo-label to complete the instance-level pseudo-label screening;
[0098] After instance-level pseudo-label screening, image-level confidence is introduced to evaluate the overall quality of the pseudo-labels in this image. Calculate the confidence score of the unlabeled SAR image according to the confidence scores of all candidate boxes. The formula is:
[0099]
[0100] where c i represents the confidence score of the i-th unlabeled SAR image, c i,m represents the confidence score of the m-th candidate box in the i-th unlabeled SAR image, M represents the total number of candidate boxes, θ ins represents the first threshold, represents that the confidence score of the m-th candidate box in the i-th unlabeled SAR image is greater than the first threshold, and the value is 1;
[0101] If the confidence score of the unlabeled SAR image is greater than the second threshold, the unlabeled SAR image and its pseudo-labels are stored in the pseudo-label bank.
[0102] Aiming at the problem that semi-supervised cross-modal detection methods often need to generate pseudo-labels for unlabeled images, and it is difficult to obtain high-quality pseudo-labels, this embodiment adopts the pseudo-label bank strategy to screen the pseudo-labels.
[0103] To solve the problem that in the near - shore scenario, the interference of on - shore ground object targets leads to false alarms when the teacher network generates pseudo - labels, thereby reducing the reliability of the pseudo - labels, the present invention proposes a background - aware collaborative classification network. The background - difference branch is used to obtain more background - region features as input by expanding the candidate boxes, which can effectively capture the feature information of the background region. Finally, the prediction results of the original classification branch and the background - difference branch are fused to screen and optimize the pseudo - labels, thereby improving the reliability of the pseudo - labels.
[0104] S23: Input the labeled optical image, labeled SAR image, and unlabeled SAR image into the student network to construct the total loss function. The specific steps are as follows:
[0105] S231: After passing the labeled optical image, labeled SAR image, and unlabeled SAR image through the backbone network and feature pyramid network of the student network, multi - scale features of each image are obtained.
[0106] S232: Calculate the supervised classification and localization loss.
[0107] The multi - scale features of the labeled SAR image are input into the detector of the student network. The RPN outputs the candidate boxes of the labeled SAR image, and then the target features are intercepted from the corresponding scale feature maps output by the feature pyramid network according to the size of the candidate boxes. The target features pass through the original regression branch and the original classification branch respectively to output the predicted localization and the first - class probability of each candidate box in the labeled SAR image, and the supervised classification and localization loss of the SAR image is calculated according to the labels;
[0108] The multi - scale features of the labeled optical image are input into the detector of the student network. The RPN outputs the candidate boxes of the labeled optical image, and then the target features are intercepted from the corresponding scale feature maps output by the feature pyramid network according to the size of the candidate boxes. The target features pass through the original regression branch and the original classification branch respectively to output the predicted localization and the first - class probability of each candidate box in the labeled optical image, and the supervised classification and localization loss of the optical image is calculated according to the labels;
[0109] The supervised classification and localization loss of the SAR image and the supervised classification and localization loss of the optical image are added together to obtain the supervised classification and localization loss L det-l 。
[0110] S233: Calculate the pseudo - label classification and localization loss.
[0111] The multi - scale features of the unlabeled SAR image in the student network The detector inputting the student network outputs candidate boxes of the unannotated SAR image by the RPN, and then intercepts the target features from the corresponding scale feature maps output by the feature pyramid network according to the sizes of the candidate boxes; the target features pass through the original regression branch and the original classification branch to respectively output the predicted localization and the first category probability of each candidate box in the unannotated SAR image; query whether the unannotated SAR image exists in the pseudo-label bank; if it exists, calculate the pseudo-label classification localization loss of the unannotated SAR image; if it does not exist, do not calculate the pseudo-label classification localization loss L of the unannotated SAR image det-ul 。
[0112] S234: Calculate the background difference branch loss.
[0113] The multi-scale features of the annotated SAR image are input into the detector of the student network, and the RPN outputs the candidate boxes of the annotated SAR image; after magnifying the candidate boxes by a preset multiple, intercept the target features from the corresponding scale feature maps output by the feature pyramid network according to the sizes of the magnified candidate boxes; the target features pass through the flattening layer, the fully connected layer and the softmax activation layer connected in sequence in the background difference branch, and output the second category probability of each candidate box in the annotated SAR image; calculate the background difference branch loss L according to the label bdb 。
[0114] S235: Calculate the multi-level feature transfer loss.
[0115] The multi-scale features of the unannotated SAR image obtained in the teacher network and the multi-scale features of the unannotated SAR image obtained in the backbone network of the student network Calculate the multi-level feature transfer loss, and the formula is:
[0116]
[0117] where, L mfkt represents the multi-level feature transfer loss, n represents the number of levels of the feature pyramid network, represents the k-th scale feature obtained by inputting the unannotated SAR image into the teacher network, represents the k-th scale feature obtained by inputting the unannotated SAR image into the student network.
[0118] The present invention introduces the multi-level feature knowledge transfer loss. By comparing the differences between the teacher network and the student network at the feature level, it guides the student network to learn the intermediate layer feature representations of the teacher network, so as to enhance the feature expression ability of the model and improve its generalization performance in the target domain.
[0119] S236: Input the multi-scale features of the labeled optical images, labeled SAR images, and unlabeled SAR images into the target-background separation adversarial network, and calculate the target-background separation adversarial loss.
[0120] Considering the problem that the current feature alignment method directly aligning the overall feature distribution may lead to the weakening of the feature expression of ship targets, thus affecting the target detection performance, the present invention proposes a target-background separation adversarial network, which extracts ship target features on the feature map using all optical images and labeled SAR images and performs feature alignment to avoid the problem of insufficient discriminability of target features caused by the alignment of ship targets with the background. The specific steps include:
[0121] S236-1: Calculate the background feature alignment loss.
[0122] Input the multi-scale features of the labeled optical images, labeled SAR images, and unlabeled SAR images into the background feature alignment module, and input the corresponding non-local discriminant network according to the size to obtain the modal class probability maps of the optical images and SAR images for each scale feature; among them, the modal class probability maps of the SAR images include the modal class probability maps of the labeled SAR images and unlabeled SAR images, specifically including:
[0123] Obtain the multi-scale feature F of the SAR image S and the multi-scale feature F of the optical image O , where F S represents the multi-scale features of the unlabeled SAR image and the labeled SAR image; input the features F S and F O of each scale feature through a non-local discriminant network respectively. First, downsample these features through 2 convolutional blocks to obtain the features D S and D O :
[0124] D S =(LrakyReLU(BN(Conv(LeakyReLU(BN(Conv(F S )))))))
[0125] D O =(LrakyReLU(BN(Conv(LeakyReLU(BN(Conv(F O )))))))
[0126] Then input D S and D O into the non-local block to obtain the non-local feature maps N S and N O , and then obtain the modal class probability map P through the convolutional block and the Sigmoid activationS and P O :
[0127] P S = Sigmoid(LeakyReLU(BN(Conv(N S ))))
[0128] P O = Sigmoid(LeakyReLU(BN(Conv(N O ))))
[0129] Calculate the feature alignment loss of each scale feature according to the modal class probability map of each scale feature, and add them to obtain the background feature alignment loss. The formula is:
[0130]
[0131] Among them, L bfa represents the background feature alignment loss, n represents the total number of non-local discriminant networks, represents the modal class probability map of the i-th scale feature of the SAR image, represents the modal class probability map of the i-th scale feature of the optical image; when the image is an optical image, d = 0; when the image is a SAR image, d = 1.
[0132] S236-2: Calculate the target feature alignment loss.
[0133] Input the multi-scale features of the labeled SAR image and the labeled optical image into the target enhancement module, and intercept the target features in the corresponding scale features according to the size of the true bounding box in each image; each target feature passes through a reverse gradient layer, a flattening layer, a fully connected layer, and a softmax activation layer connected in sequence to obtain the modal class probability of the target feature.
[0134] In this embodiment, since the number of levels of the feature pyramid network is 5, after discarding the topmost feature, that is, the scale feature with the smallest size, according to the size of the true bounding box in the SAR image and the optical image, the targets are divided into small targets, medium targets, medium-large targets, and large targets, and the target features are intercepted from the high-level, middle-high-level, middle-level, and low-level feature maps of the multi-scale feature map corresponding to the sizes of different targets.
[0135] Calculate the target feature alignment loss according to the modal class probabilities of all target features. The formula is:
[0136]
[0137] Among them, L ofedenotes the target feature alignment loss, and N denotes the total number of ground truth bounding boxes in the annotated SAR images and the annotated optical images. denotes the modal class probability of the j-th target feature in the annotated SAR image. denotes the modal class probability of the j-th target feature in the annotated optical image; when the image is an optical image, d = 0; when the image is a SAR image, d = 1.
[0138] S236-3: Construct the target-background separation adversarial loss L using the background feature alignment loss and the target feature alignment loss obsa , and the formula is:
[0139] L obsa = L bfa + L ofe
[0140] S237: Construct the total loss L, and the formula is:
[0141] L = L det-l + λ1L det-ul + λ2L obsa + λ3L mfkt + λ4L bdb
[0142] where λ1, λ2, λ3, λ4 are hyperparameters.
[0143] S24: Use the stochastic gradient descent method to calculate the derivatives of the parameters of the student network according to the total loss and update the parameters of the student network.
[0144] S25: The parameters of the teacher network are assigned using the strategy of moving exponential average based on the parameters of the student network.
[0145] S3: Use the teacher network of the trained ship detection model to detect ships in SAR images.
[0146] Input the SAR image to be detected into the trained teacher network, and pass through the backbone network, feature pyramid network and detector in the teacher network in sequence. The ROI detection head in the detector outputs the predicted localization and class probability of each candidate box through the original regression branch and the original classification branch respectively.
[0147] In summary, for the semi-supervised cross-modal SAR image ship detection method described in the present invention, with the mean teacher network as the basic framework, during training, the teacher network adopts the pseudo-label bank strategy to obtain pseudo-labels of unlabeled SAR images, and calculates the pseudo-label classification and localization loss to achieve semi-supervised cross-modal detection; further, a target-background separation adversarial module is introduced to align the background features using optical images and SAR images, and to align the ship target features extracted from the labeled optical images and labeled SAR images on the feature map to avoid the problem of insufficient discriminability of target features caused by the alignment of ship targets and the background, and calculate the target-background separation adversarial loss according to the background feature alignment loss and the target feature alignment loss; finally, the pseudo-label classification and localization loss and the target-background separation adversarial loss are introduced into the total loss, improving the accuracy of ship detection.
[0148] Furthermore, the present invention introduces a multi-level feature knowledge transfer loss. By comparing the differences between the teacher network and the student network at the feature level, it guides the student network to learn the intermediate layer feature representations of the teacher network, enhancing the feature expression ability of the model and improving its generalization performance in the target domain, thereby improving the accuracy of ship detection.
[0149] Furthermore, due to the interference of onshore object targets in the nearshore scene, false alarms occur when the teacher network generates pseudo-labels, reducing the reliability of the pseudo-labels. The present invention introduces a background difference branch in the ROI detection head. By expanding the candidate boxes to obtain more background region features as input, it can effectively capture the features of the background region. Finally, the prediction results of the original classification branch and the background difference branch are fused to screen and optimize the pseudo-labels, thereby improving the reliability of the pseudo-labels.
[0150] Embodiment 2
[0151] To verify the effectiveness of the semi-supervised cross-modal deep learning SAR image rotated ship detection method proposed in the present invention, in this embodiment, the DOTA dataset is selected as the optical image training data, and the SSDD data is used as the SAR dataset.
[0152] 1. Experimental data
[0153] DOTA is a currently mainstream optical remote sensing image dataset, and the image sources are Google Earth and two Chinese satellites (Gaofen-2 and Jilin-1). This dataset contains a total of 2,806 remote sensing images, with image sizes ranging from 800×800 to 4,000×4,000, and a total of 188,282 instances, divided into 15 categories such as airplanes and ships, using rotated annotation. First, 31,879 images with a size of 1,024×1,024 are obtained by slicing, and 1,869 images containing ship targets are selected as the training set.
[0154] SSDD is the first available SAR image target detection dataset. The images come from three spaceborne SAR imaging platforms, namely Radarsat-2, TerraSAR-X, and Sentinel-1. The polarization methods used by the sensors include HH, VV, VH, and HV. It consists of 1,160 SAR images with an average size of 500×500, containing a total of 2,456 ships, with two annotation methods: horizontal bounding boxes and rotated bounding boxes. The resolution of the images ranges from 1 to 15 m, covering ship targets of different sizes and materials in both inshore and offshore scenarios. In the experiment, 928 training set images and 232 test set images are divided without repetition at a ratio of 8:2. The training set is further divided at a ratio of 1:9 to obtain 93 annotated training images and 835 unannotated images.
[0155] 2. Experimental Setup
[0156] This embodiment is based on the MMRotate environment and is carried out on an NVIDIA RTX 3090 GPU. The two-stage object detection network OrientedR-CNN with ResNet50-FPN as the backbone is selected as both the teacher network and the student network. The initial learning rate for training is 0.0025, the weight decay is 0.0001, and the momentum is 0.9. A total of 30,000 iterations of training are performed. The learning rate is reduced to 0.1 times the previous value at the 20,000th and 27,500th iterations respectively. The hyperparameters λ1, λ2, λ3, λ4 are set to 2.0, 0.1, 0.5, and 1.0 respectively. During the training and testing processes, detection boxes with an intersection over union (IoU) greater than 0.5 between the candidate box and the ground truth bounding box are considered correct detections. mAP50 (the average precision when the IoU threshold for determining true targets is 0.5) is selected as the evaluation metric.
[0157] 3. Experimental Results
[0158] Table 1 shows the comparative experimental results with other advanced methods, including the semi-supervised methods Unbiased Teacher, SoftTeacher, DualTeacher (SSOD), and the semi-supervised cross-domain method Dual Teacher (SSCMOD). OrientedR-CNN (10%) in the table indicates that the network is trained using 10% of the SAR annotation data. It can be seen from the results that the average precision mAP50 of the present invention is improved by 5.4% compared to the base model and is superior to the current mainstream semi-supervised object methods (Unbiased Teacher, Soft Teacher, DualTeacher (SSOD)) and semi-supervised cross-modal method (DualTeacher (SSCMOD)).
[0159] Table 1 Experimental results of different methods in the DOTA→SSDD task
[0160]
[0161] Figure 7 shows the precision-recall (PR) curves of the method proposed by the present invention and other advanced methods. The area enclosed by the PR curve and the horizontal axis (recall) reflects the overall detection accuracy of the model and can intuitively measure the detection performance of different methods. In the evaluation of the PR curve, the closer the curve is to the upper right of the coordinate axis, the higher the precision the model can maintain at different recall levels, that is, the better the detection performance. It can be observed from the figure that the PR curve of the method proposed by the present invention is generally above all the comparison methods, indicating that the present invention can maintain a high precision within different recall ranges and is significantly better than other methods, fully verifying the effectiveness and superiority of the present invention.
[0162] Figure 8 shows the visual detection results of the method proposed by the present invention and other advanced methods on the SSDD data, including two scenarios of inshore and offshore. Among them Figure 8 (a1) and (a2) in it are the detection results of UnbiasedTeacher for inshore and offshore images respectively, Figure 8 (b1) and (b2) in it are the detection results of SoftTeacher for inshore and offshore images respectively, Figure 8 (c1) and (c2) in it are the detection results of DualTeacher (SSCMOD) for inshore and offshore images respectively, Figure 8 (d1) and (d2) in it are the detection results of DualTeacher (SSOD) for inshore and offshore images respectively, Figure 8 (e1) and (e2) in it are the detection results of the present invention for inshore and offshore images respectively, Figure 8(f1) and (f2) in it are the true annotations for the nearshore and offshore images respectively. As can be seen from the figure, the method proposed in the present invention has achieved the best detection effect in both scenarios. In the nearshore scenario, UnbiasedTeacher, SoftTeacher, DualTeacher (SSCMOD), and the present invention can all correctly detect all ship targets. However, due to the existence of a large number of areas similar to ship targets on land, the first three methods have a large number of false alarms. In contrast, the number of false alarms of the present invention is significantly reduced. In the offshore scenario, due to the influence of sea clutter, all methods have a certain degree of false alarms. Among them, UnbiasedTeacher, SoftTeacher, and DualTeacher (SSCMOD) generate more false alarms. Although DualTeacher (SSOD) has fewer false alarms, there are missed detections. In contrast, the present invention can effectively suppress false alarms while ensuring the detected targets, further verifying its effectiveness.
[0163] Figure 9 The mAP50 metrics of the fully supervised method with different proportions of labeled data and the proposed semi-supervised cross-modal object detection method are shown. The red dots represent the mAP50 of the fully supervised method with different proportions of labeled data, which are 10%, 20%, 30%, 50%, 70%, 90%, and 100% of the labeled data used during training from left to right respectively. The straight lines represent the mAP50 of the semi-supervised cross-modal method with different proportions of labeled data. The three red, green, and orange straight lines represent the use of 10%, 30%, and 50% of the labeled data during training respectively. It can be seen that in the case of 10% SAR labeled data, the method proposed in the present invention has a significant improvement compared with the fully supervised method and is close to the accuracy of the fully supervised method when 20% of the SAR labeled data is used, indicating that the present invention can save nearly half of the labeled data when the labeled data is less.
[0164] Embodiment III
[0165] Based on the semi-supervised cross-modal SAR image ship detection method described in Embodiment I, this embodiment provides a semi-supervised cross-modal SAR image ship detection system, including:
[0166] A model construction module for constructing a ship detection model, including a teacher network and a student network; both the teacher network and the student network include a backbone network, a feature pyramid network, and a detector; the student network also includes a background feature alignment module and a target enhancement module;
[0167] A training module for training the ship detection model with labeled optical images, labeled SAR images, and unlabeled SAR images, including:
[0168] The pseudo-label acquisition unit is used to input the unlabeled SAR image into the teacher network to obtain pseudo-labels and construct a pseudo-label bank;
[0169] The background feature alignment loss calculation unit is used to input the labeled optical image, the labeled SAR image, and the unlabeled SAR image into the backbone network and the feature pyramid network of the student network to obtain the multi-scale features of each image, and then input them into the background feature alignment module to obtain the modal category probability maps of the features at each scale; calculate the feature alignment losses of the features at each scale, and add them up to obtain the background feature alignment loss;
[0170] The target feature alignment loss calculation unit is used to input the multi-scale features of the labeled SAR image and the labeled optical image into the target enhancement module, intercept the target features in the corresponding scale features according to the size of the true bounding box in each image; calculate the target feature alignment loss according to the modal category probabilities of all target features;
[0171] The target-background separation adversarial loss calculation unit is used to construct the target-background separation adversarial loss with the background feature alignment loss and the target feature alignment loss;
[0172] The pseudo-label classification and localization loss calculation unit is used to input the unlabeled SAR image into the student network to obtain the predicted localization and the first category probability of each candidate box; if the unlabeled SAR image exists in the pseudo-label bank, calculate the pseudo-label classification and localization loss;
[0173] The parameter update unit is used to construct the total loss with the background feature alignment loss, the target feature alignment loss, the pseudo-label classification and localization loss, and the supervised classification and localization loss; update the parameters of the student network with the total loss, and assign the parameters of the student network to the parameters of the teacher network;
[0174] The detection module is used to detect ships in the SAR image by using the teacher network of the trained ship detection model.
[0175] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0176] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0179] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A semi-supervised cross-modal ship detection method for SAR images, characterized in that Including: Construct a ship detection model, including a teacher network and a student network; Both the teacher network and the student network include a backbone network, a feature pyramid network, and a detector; the student network also includes a background feature alignment module and a target enhancement module; Train the ship detection model with labeled optical images, labeled SAR images, and unlabeled SAR images, Including: Input the unlabeled SAR image into the teacher network to obtain pseudo-labels and construct a pseudo-label bank; After passing the labeled optical images, labeled SAR images, and unlabeled SAR images through the backbone network and the feature pyramid network of the student network, obtain the multi-scale features of each image, and then input them into the background feature alignment module to obtain the modal category probability maps of the features at each scale; calculate the feature alignment loss of the features at each scale, and sum them up to obtain the background feature alignment loss; Input the multi-scale features of the labeled SAR images and labeled optical images into the target enhancement module, Intercept the target features from the corresponding scale feature maps output by the feature pyramid network according to the size of the true bounding box in each image; Calculate the target feature alignment loss according to the modal category probabilities of all target features; Input the unlabeled SAR image through the student network to obtain the predicted localization and the first category probability of each candidate box; if the unlabeled SAR image exists in the pseudo-label bank, calculate the pseudo-label classification localization loss; Construct the total loss with the background feature alignment loss, target feature alignment loss, pseudo-label classification localization loss, and supervised classification localization loss; update the parameters of the student network with the total loss, and assign the parameters of the student network to the parameters of the teacher network; Use the teacher network of the trained ship detection model to detect ships in SAR images.
2. A semi-supervised cross-modal SAR image ship detection method according to claim 1, characterized in that, The detector uses Oriented R-CNN, including a region proposal network and an ROI detection head; The ROI detection head includes an original regression branch, an original classification branch, and a background difference branch; The background difference branch is used to magnify the candidate box by a preset multiple, and then intercept the target features from the corresponding scale feature map output by the feature pyramid network according to the size of the magnified candidate box. The target features pass through a flattening layer, a fully connected layer, and a softmax activation layer to obtain the second category probability.
3. A semi - supervised cross - modal SAR image ship detection method according to claim 2, characterized in that, Input the unlabeled SAR image into the teacher network to obtain pseudo-labels and construct a pseudo-label bank, including: Input the unlabeled SAR image into the teacher network, and the original classification branch and the background difference branch respectively output the first category probability and the second category probability of each candidate box; After weighted summation of the first category probability and the second category probability of each candidate box, obtain the target category probability of each candidate box; select the maximum value in the target category probabilities as the confidence score of the candidate box; If the confidence score of the candidate box is greater than the first threshold, then use the candidate box as a pseudo-label; Calculate the confidence score of the unlabeled SAR image according to the confidence scores of all candidate boxes. The formula is: Among them, c i represents the confidence score of the i-th unlabeled SAR image, c i,m represents the confidence score of the m-th candidate box of the i-th unlabeled SAR image, M represents the total number of candidate boxes, and θ ins represents the first threshold, represents that the confidence score of the m-th candidate box of the i-th unlabeled SAR image is greater than the first threshold, and the value is 1; If the confidence score of the unlabeled SAR image is greater than the second threshold, then store the unlabeled SAR image and its pseudo-labels in the pseudo-label bank.
4. A semi-supervised cross-modal SAR image ship detection method according to claim 2, characterized in that, The total loss also includes the background difference branch loss, including: The labeled SAR image is input into the student network. The background difference branch outputs the second-class probability of each candidate box in the labeled SAR image, and the background difference branch loss is calculated according to the label.
5. A semi-supervised cross-modal SAR image ship detection method according to claim 1, characterized in that, The calculation method of the supervised classification and localization loss includes: The labeled SAR image is input into the student network. The original regression branch and the original classification branch of the detector respectively output the predicted localization and the first-class probability of each candidate box in the labeled SAR image, and the supervised classification and localization loss of the SAR image is calculated according to the label. The labeled optical image is input into the student network. The original regression branch and the original classification branch of the detector respectively output the predicted localization and the first-class probability of each candidate box in the labeled optical image, and the supervised classification and localization loss of the optical image is calculated according to the label. The supervised classification and localization loss of the SAR image and the supervised classification and localization loss of the optical image are added to obtain the supervised classification and localization loss.
6. A semi-supervised cross-modal SAR image ship detection method according to claim 1, characterized in that, The background feature alignment module includes several non-local discriminant networks, and the number of non-local discriminant networks is the same as the number of levels of the feature pyramid network; each non-local discriminant network includes a first convolutional block, a second convolutional block, a non-local block, a third convolutional block, and a sigmoid activation layer connected in sequence; each convolutional block includes a convolutional layer, a normalization layer, and an activation layer connected in sequence.
7. A semi-supervised cross-modal SAR image ship detection method according to claim 6, wherein, The labeled optical image, the labeled SAR image, and the unlabeled SAR image are passed through the backbone network and the feature pyramid network of the student network to obtain the multi-scale features of each image, and then input into the background feature alignment module to obtain the modal class probability maps of the features at each scale. Calculate the feature alignment loss of the features at each scale, and add them up to obtain the background feature alignment loss, including: The labeled optical image, the labeled SAR image, and the unlabeled SAR image are passed through the backbone network and the feature pyramid network of the student network to obtain the corresponding multi-scale features of each image. The multi-scale features of all images are input into the corresponding non-local discriminant networks according to the size to obtain the modal class probability maps of the optical image and the SAR image of the features at each scale; among them, the modal class probability map of the SAR image includes the modal class probability maps of the labeled SAR image and the unlabeled SAR image. Calculate the feature alignment loss of the features at each scale according to the modal class probability maps of the features at each scale. Add them up to obtain the background feature alignment loss, and the formula is: Among them, L bfa represents the background feature alignment loss, n represents the total number of non-local discriminant networks, represents the modal class probability map of the i-th scale feature of the SAR image, represents the modal class probability map of the i-th scale feature of the optical image; when the image is an optical image, d = 0; when the image is a SAR image, d = 1.
8. A semi-supervised cross-modal SAR image ship detection method according to claim 1, characterized in that, Calculate the target feature alignment loss according to the modal class probabilities of all target features, including: The target feature passes through a reverse gradient layer, a flattening layer, a fully connected layer, and a softmax activation layer connected in sequence to obtain the modal class probability of the target feature. Calculate the target feature alignment loss according to the modal class probabilities of all target features, and the formula is: Among them, L ofe represents the target feature alignment loss, N represents the total number of ground truth bounding boxes in the labeled SAR images and the labeled optical images, represents the modal class probability of the j-th target feature in the labeled SAR image, represents the modal class probability of the j-th target feature in the labeled optical image; When the image is an optical image, d = 0; when the image is a SAR image, d = 1.
9. A semi-supervised cross-modal SAR image ship detection method according to claim 1, characterized in that The total loss also includes a multi-level feature transfer loss, including: The unannotated SAR image passes through the backbone network and the feature pyramid network of the teacher network to obtain multi-scale features The unannotated SAR image passes through the backbone network and the feature pyramid network of the student network to obtain multi-scale features According to multi-scale features and calculate the multi-level feature transfer loss, and the formula is as follows: Among them, L mfkt represents the multi-level feature transfer loss, n represents the number of levels of the feature pyramid network, represents the k-th scale feature obtained by inputting the unlabeled SAR image into the teacher network, represents the k-th scale feature obtained by inputting the unlabeled SAR image into the student network.
10. A semi-supervised cross-modal SAR image ship detection system, characterized in that, including: A model construction module for constructing a ship detection model, including a teacher network and a student network. Both the teacher network and the student network include a backbone network, a feature pyramid network, and a detector; the student network also includes a background feature alignment module and a target enhancement module. A training module for training a ship detection model with annotated optical images, annotated SAR images, and unannotated SAR images, including: A pseudo-label acquisition unit for inputting unannotated SAR images into a teacher network to obtain pseudo-labels, Constructing a pseudo-label bank; A background feature alignment loss calculation unit for obtaining multi-scale features of each image after passing the annotated optical images, annotated SAR images, and unannotated SAR images through the backbone network and feature pyramid network of the student network, and then inputting them into the background feature alignment module to obtain the modal category probability maps of the features at each scale; calculating the feature alignment losses of the features at each scale, and adding them up to obtain the background feature alignment loss; An object feature alignment loss calculation unit for inputting the multi-scale features of the annotated SAR images and annotated optical images into an object enhancement module, and intercepting object features in the corresponding scale features according to the sizes of the true bounding boxes in each image; calculating the object feature alignment loss according to the modal category probabilities of all object features; An object-background separation adversarial loss calculation unit for constructing an object-background separation adversarial loss with the background feature alignment loss and the object feature alignment loss; A pseudo-label classification and localization loss calculation unit for obtaining the predicted localization and the first category probability of each candidate box by passing the unannotated SAR image through the student network; if the unannotated SAR image exists in the pseudo-label bank, calculating the pseudo-label classification and localization loss; A parameter update unit for constructing a total loss with the background feature alignment loss, the object feature alignment loss, the pseudo-label classification and localization loss, and the supervised classification and localization loss; updating the parameters of the student network with the total loss, and assigning the parameters of the student network to the parameters of the teacher network; A detection module for detecting ships in SAR images using the teacher network of the trained ship detection model.