Downstream visual detection model driven SAR-optical image conversion method and system, and storage medium

Through the downstream visual detection model-driven SAR-optical image conversion method, the problem of insufficient structural integrity and semantic consistency in the prior art is solved, the image conversion quality and the accuracy of downstream detection tasks are improved, and the rapid training and application of the model are achieved.

CN120388208APending Publication Date: 2025-07-29NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510381383.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing SAR-optical image conversion technology lacks structural integrity and semantic consistency, resulting in low accuracy of downstream visual detection tasks and lack of mature application implementation, resulting in difficulty in reproducing algorithms and high learning costs.

Method used

Using a downstream visual detection model-driven method, a closed-loop optimization system is built by pre-training semantic segmentation and object detection model, combining weighted loss functions of countermeasures, supervision losses and detection losses, and a closed-loop optimization system is built, and the SAR-optical image conversion model is trained, and pseudo-optical images with optical characteristics are generated using the image conversion model, and quality evaluation is performed.

Benefits of technology

It improves the quality and fidelity of the converted images, improves the recognition accuracy of downstream visual detection tasks, realizes the rapid convergence and application implementation of the model, and reduces learning costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388208A_ABST
    Figure CN120388208A_ABST
Patent Text Reader

Abstract

The invention provides an SAR-optical image conversion method and system driven by a downstream visual inspection model and a storage medium. The method comprises the following steps: acquiring SAR-optical image data, marking the data, and preprocessing the data; pre-training of the SAR and the optical image domain is completed by using the visual detection model; selecting an image conversion model according to the characteristics of the data set, constructing a conversion model driven by downstream visual detection, detecting, evaluating and feeding back the generated pseudo optical image and the real optical image by the model in combination with a detection task branch, and determining the real optical image by designing a weighted loss function including confrontation loss, supervision loss and detection loss. Carrying out model training optimization; and finally, performing downstream visual inspection task evaluation by using the generated pseudo-optical image. According to the method, semantic information of the SAR and the optical image is fully utilized, and the quality and fidelity of the converted image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology, computer vision, and software development, and particularly relates to a SAR-optical image conversion method, system, and storage medium driven by a downstream visual detection model. Background Art

[0002] Optical remote sensing based on the visible light band and synthetic aperture radar (SAR) systems based on the principle of electromagnetic wave scattering constitute the two core observation systems of remote sensing for earth observation technology.

[0003] With its active detection method and all-weather imaging ability, SAR imaging technology breaks through the limitations of optical imaging restricted by cloud cover and lighting conditions, providing reliable data support for key fields such as disaster monitoring and target detection. However, due to the special side-looking imaging mechanism and complex backscattering characteristics of the SAR system, the generated grayscale images pose significant challenges in terms of visual interpretation, which directly affects the large-scale application and popularization of this technology. In contrast, optical remote sensing images are widely recognized for their intuitive color representation and delicate ground object texture features, but the defect that their imaging quality is easily interfered by atmospheric conditions has always been difficult to completely overcome. It is worth noting that SAR and optical images exhibit significant complementary characteristics in terms of spatial resolution, spectral information, penetration ability, etc. However, the essential differences in the electromagnetic wave action mechanisms (active emission and passive reception) and signal manifestation forms (phase information and radiance) between the two lead to a huge gap in their data characteristics.

[0004] With the booming development of deep learning and image conversion technology, SAR-optical image translation (S2OIT) technology centered on generative adversarial networks (GANs) has opened up new research paths. Researchers have systematically analyzed the non-linear mapping rules between the two modal data by constructing deep neural network models, and successfully achieved the conversion of SAR images into images with optical characteristics.

[0005] Currently, the research on SAR-optical image conversion technology generally focuses on optimizing pixel-level quantization parameters such as SSIM and PSNR. These metrics only measure the numerical similarity between the generated image and the real image by pixel-by-pixel comparison, ignoring the structural correlation at the semantic level of the image, resulting in defects in the structural integrity and semantic consistency of the generated image. For example, the blurred contour of the ground object or the absence of key targets seriously affects the detection efficiency and accuracy of downstream visual detection tasks, and directly leads to the failure of the conversion result in practical application scenarios such as object detection and land cover classification. In addition, most of the existing SAR-optical image conversion technologies are theoretical algorithms, and there is no corresponding implementation in terms of software application and hardware deployment, resulting in difficulties in algorithm reproduction, high learning costs, and low work efficiency for technicians using relevant algorithms.

[0006] In view of the problems in the above-mentioned related technologies, that is, the lack of structural integrity and semantic consistency in the SAR-optical image conversion technology leads to the failure of downstream visual detection tasks, and the lack of complete and mature application implementation, no effective solution has been proposed yet. Summary of the Invention

[0007] In order to overcome the deficiencies of the prior art, the present invention provides a SAR-optical image conversion method, system and storage medium driven by a downstream visual detection model. The method includes: obtaining SAR-optical image data, performing labeling and data preprocessing; using a visual detection model to complete pre-training in the SAR and optical image domains; selecting an image conversion model according to the characteristics of the data set, constructing a conversion model driven by downstream visual detection, which combines a detection task branch to detect, evaluate and feedback the generated pseudo-optical image and the real optical image, and through designing a weighted loss function including adversarial loss, supervised loss and detection loss, performing model training optimization; finally using the generated pseudo-optical image to evaluate the downstream visual detection task. The present invention makes full use of the semantic information of SAR and optical images, can improve the quality and fidelity of the converted image, and solves the problem that the recognition accuracy of the converted image in the downstream visual detection task is low due to the defects in structural integrity and semantic consistency in the prior art.

[0008] A SAR-optical image conversion method driven by a downstream visual detection model is characterized by the following steps:

[0009] Step 1: Obtain remote sensing image data and perform annotation. The remote sensing image data includes SAR images and optical images of the corresponding scene, and the annotation is semantic or category label annotation;

[0010] Step 2: Perform preprocessing on the image data obtained in Step 1, including normalization of the image spatial resolution, normalization of pixel values, and image enhancement processing;

[0011] Step 3: Use the image data processed in Step 2 to pre-train the downstream visual detection model, where the downstream visual detection model includes a semantic segmentation task model and an object detection task model;

[0012] Step 4: Select an image conversion model. Specifically, for the paired dataset of SAR images and optical images, a supervised image conversion model is adopted; for the unpaired dataset of SAR images and optical images, an unsupervised image conversion model is adopted;

[0013] Step 5: Construct a SAR-optical image conversion model driven by the downstream visual detection model, including a forward branch and a feedback branch. The forward branch uses the visual detection model pre-trained on SAR images to detect the input SAR image, and after cascading the detection results with the SAR image at the channel level, inputs them into the image conversion model for training. The feedback branch uses the visual detection model pre-trained on optical images to detect the generated pseudo-optical image and the real optical image respectively, and evaluates the quality of the detection results, and feeds the evaluation results back to the image conversion model to guide its training;

[0014] Step 6: Design a loss function, which includes three parts: adversarial loss, supervised loss, and detection loss of the downstream visual detection model. Specifically, for the unsupervised conversion model, the detection loss of the downstream visual detection model includes CIoU detection box regression loss, confidence prediction loss using cross-entropy loss, and class prediction loss. For the supervised conversion model, the weighted sum of cross-entropy loss, feature matching loss, and semantic segmentation-guided L1 loss is used as the detection loss of the downstream visual detection model;

[0015] Step 7: Perform joint optimization training on the SAR-optical image conversion model. Specifically: Select a suitable training strategy, and continuously optimize the parameters of the SAR-optical image conversion model through repeated iterative training processes. In each iteration, adjust the parameters of the image conversion model according to the conversion results of the image conversion model and the information feedback of the downstream visual detection model to reduce the difference between the generated image and the real image. The training strategy includes the weights of the losses, the number of rounds of iterative training, and the step size of training;

[0016] Step 8: Use the trained SAR-optical image conversion model to convert the input SAR image into a pseudo-optical image with optical image features, completing the SAR-optical image conversion;

[0017] Step 9: Use the downstream visual detection model to detect the pseudo-optical image generated by the SAR-optical image conversion model, and evaluate the quality of the detection results. The downstream visual detection model and quality evaluation metrics used are consistent with those in the SAR-optical image conversion model trained in Step 7.

[0018] Furthermore, the acquisition of the remote sensing image dataset described in step 1 refers to obtaining or self-collecting and constructing a remote sensing image dataset from a public dataset or a satellite image database; the annotation is to perform semantic or class label annotation on the image using an interactive segmentation annotation tool.

[0019] Furthermore, the image enhancement processing described in step 2 includes operations such as random cropping, rotation, and contrast enhancement of the image.

[0020] Furthermore, the semantic segmentation task model described in step 3 uses the U-Net network, and the object detection task model uses the YOLOv5 network.

[0021] Furthermore, the specific process of the model pre-training in step 3 is as follows: First, using SAR images and corresponding annotations as inputs, train the semantic segmentation task model and the object detection task model to obtain the semantic segmentation task model and the object detection task model for SAR images; then, using optical images and corresponding annotations as inputs, train the semantic segmentation task model and the object detection task model to obtain the semantic segmentation task model and the object detection task model for optical images; during training, for the semantic segmentation task model, use the cross-entropy loss function, and for the object detection task model, use the detection box regression loss, confidence prediction loss, and class prediction loss; during training, optimize the network parameters through backpropagation.

[0022] Furthermore, in the SAR-optical image conversion model driven by the downstream visual detection model constructed in step 5, for the semantic segmentation model, use the pix2pix supervised framework; for the object detection model, use the CycleGAN unsupervised framework.

[0023] Further, the adversarial loss mentioned in step 6 refers to the adversarial training of optimizing the generator and discriminator using the least squares loss; the supervision loss refers to introducing the pixel-level L1 loss under the condition of the registration dataset; the CIoU consists of the overlapping region IoU, the center point distance penalty term, and the aspect ratio penalty term. The IoU calculates the intersection over union of the predicted box and the ground truth box to measure the overlapping degree between the two. The center point distance penalty term constrains the position offset between the two through the Euclidean distance between the center points of the predicted box and the ground truth box. The aspect ratio penalty term introduces a consistency measure of the aspect ratio to ensure the shape similarity between the predicted box and the ground truth box. The cross-entropy loss refers to using the semantic segmentation task model pre-trained on the optical image to segment the optical image generated by the conversion model, and calculating the cross-entropy loss between the segmentation result and the semantic segmentation label of the SAR image of the image as the final cross-entropy loss. The feature matching loss refers to using the downsampling network of the semantic segmentation task model pre-trained on the optical image to downsample the input real optical image and the pseudo-optical image generated by the conversion model respectively, calculating the L1 loss for the features obtained by downsampling at each level, and calculating the average value of the total L1 loss as the final feature matching loss. The semantic segmentation-guided L1 loss refers to using the semantic segmentation task model pre-trained on the optical image to perform binary classification segmentation on the input optical image, multiplying the 01 matrix of the segmentation result element-wise with the input real optical image and the pseudo-optical image generated by the conversion model respectively, calculating the L1 loss of the two images obtained after multiplication, and multiplying by a weight as the final semantic segmentation-guided enhanced L1 loss, where the weight value ranges from 10 to 20.

[0024] Further, in step 7, an alternating optimization strategy is adopted to perform end-to-end training on the model. The specific process is as follows: First, fix the parameters of the conversion model, and use the visual detection model on the SAR image to obtain the corresponding model output; then, splice the model output and the original SAR image and input them into the generator to generate a pseudo-optical image; next, use the visual detection model on the SAR image to extract the multi-level features of the generated image and the real sample, and calculate the semantic consistency loss; finally, synchronously optimize the image conversion ability of the generator and the discrimination ability of the discriminator through backpropagation; an adaptive learning rate strategy is adopted during the training process, and the training is terminated when the segmentation accuracy of the validation set reaches a stable threshold.

[0025] A SAR-optical image conversion system driven by a downstream visual detection model includes an acquisition module, a transmission device, and a server device; among them, the acquisition module is used to obtain the video stream data of SAR and optical images, and is connected to the server device through the transmission device; the server device is used to execute the SAR-optical image conversion method driven by the downstream visual detection model, generate conversion results and evaluation metrics, and display them in real time; the transmission device is used to transmit the video stream data to the server device.

[0026] A storage medium stores a computer program which, when run, is used to execute the steps of a SAR-optical image conversion method driven by a downstream visual detection model disclosed in the present invention.

[0027] The beneficial effects of the present invention are as follows: due to the use of image data preprocessing including image spatial resolution normalization, pixel value normalization, and random cropping / rotation / contrast enhancement, it can effectively eliminate the modality differences between SAR and optical images, enhance the consistency of data distribution, and thus alleviate the overfitting problem that may exist in subsequent model training; due to the use of U-Net and YOLOv5 for independent pre-training in the SAR and optical image domains respectively, a cross-modal feature decoupling representation ability is constructed, which can effectively improve the feature extraction accuracy of the semantic segmentation model, and thus provide more accurate semantic information for the target detection task; due to the use of pix2pix (supervised) or CycleGAN (unsupervised) as the basic conversion framework, it effectively solves the model generalization bottleneck of existing methods when the data alignment degree is uncertain; due to the use of channel concatenation input in the forward branch and CIoU quality evaluation in the feedback branch, a closed-loop optimization system is constructed, which can effectively suppress the common texture distortion phenomenon in traditional methods; due to the use of an innovatively designed weighted loss system that combines adversarial loss, pixel-level supervision loss, and task-driven detection loss, it can effectively improve the mAP index of the downstream target detection task after model training, significantly superior to the single adversarial loss scheme; due to the use of an alternating optimization training and dynamic learning rate adjustment mechanism, and controlling the training termination timing through the validation set accuracy threshold, it can accelerate the model convergence speed and avoid the risk of overfitting. Description of the Drawings

[0028] Figure 1 is a flowchart of the SAR-optical image conversion method driven by the downstream visual detection model of the present invention;

[0029] Figure 2 is a framework diagram of the SAR-optical image conversion system driven by the downstream visual detection model of the present invention. Detailed Embodiments

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] An embodiment of the SAR-optical image conversion method driven by a downstream vision detection model of the present invention mainly solves the defects in structural integrity and semantic consistency in the related prior art, resulting in a low recognition accuracy of the converted image in the downstream vision detection task, and realizes SAR-optical image conversion and downstream task detection.

[0032] As Figure 1 shown, this embodiment includes the following steps:

[0033] Step 1: Remote sensing data acquisition and annotation. Based on the application scenario requirements, a data set containing SAR images and corresponding optical images is established. The data sources preferably include public data sets, satellite image databases, and self-collected remote sensing image data, and preferably cover multi-scene samples of different seasons, terrains, and target types.

[0034] For the registered data set, the SAR image and the optical image have been spatially aligned; for the unregistered data set, a weak association relationship can be established through image retrieval. During the annotation process, an interactive segmentation annotation tool is used to annotate semantic or category labels for the images, and the SAR and optical images of the registered data set can share the annotation results.

[0035] The embodiment of the present invention extracts typical samples of the road data set from the SEN1-2 data set as the data set driven by the semantic segmentation model, including four-column diagrams of spring and summer scenes, respectively presenting optical images, SAR images, binary segmentation masks, and visualization annotation effects. The embodiment of the present invention extracts SAR image data containing ship targets from TerraSAR-X satellite images in the HRSID data set, extracts optical image data containing ship targets from the FAIR1M data set, and combines the two as the data set driven by the target detection model.

[0036] Step 2: Data preprocessing. The embodiment of the present invention performs standardization processing on the original remote sensing data, specifically including: (1) Spatial resolution normalization processing, using the interpolation method to normalize the image size; (2) Dynamic range adjustment, normalizing the intensity values of the SAR image and the multi-spectral channels of the optical image respectively; (3) Data augmentation strategy, applying methods such as random cropping, rotation, and contrast enhancement to expand the training samples and improve the generalization ability of the model.

[0037] Step 3: Pretraining of the downstream visual detection task model. The downstream visual detection model includes a semantic segmentation task model and an object detection task model. In the embodiment of the present invention, the U-Net network architecture is selected as the benchmark model for the downstream semantic segmentation task, and YOLOv5 is selected as the benchmark model for the downstream object detection task. In the embodiment of the present invention, two independent training processes are respectively established: the first training process takes the SAR image and the corresponding annotation as the input, and trains to obtain a semantic segmentation model or an object detection model of the SAR image; the second training process takes the optical image and the same annotation as the input, and trains to obtain a segmentation model or an object detection model of the optical image. For the semantic segmentation model, the cross-entropy loss function is adopted in the embodiment of the present invention; for the object detection model, the detection box regression loss, the confidence prediction loss, and the class prediction loss are adopted in the embodiment of the present invention, and both optimize the network parameters through backpropagation to provide prior knowledge support for the subsequent conversion model.

[0038] Step 4: Selection of the image conversion model. For the paired dataset and the unpaired dataset, image conversion models suitable for the characteristics of the dataset are respectively adopted, that is, for the paired dataset and the unpaired dataset, supervised and unsupervised image conversion models are respectively adopted.

[0039] Step 5: Construction of the SAR-optical image conversion model driven by the downstream visual detection model. For the conversion framework driven by the downstream visual detection model, the embodiment of the present invention establishes an improved SAR-optical image conversion model including a dual-branch structure:

[0040] 1) The forward branch uses the semantic segmentation task model or the object detection task model pre-trained on the SAR image to process the input SAR image, and cascades the SAR image and its semantic segmentation result or object detection result at the channel level and inputs them into the image conversion model for training;

[0041] 2) The feedback branch uses the semantic segmentation task model or the object detection task model pre-trained on the optical image to detect the generated pseudo-optical image and the real optical image respectively, and evaluates the detection results, and feeds the evaluation results back to the image conversion model to guide its training.

[0042] Among them, for the semantic segmentation model, the pix2pix supervised framework is adopted in the embodiment of the present invention; for the object detection model, the CycleGAN unsupervised framework is adopted in the embodiment of the present invention.

[0043] Step 6: Loss function design. The loss function system of the embodiments of the present invention includes three parts: (1) Adversarial loss, using least squares loss to optimize the adversarial training of the generator and discriminator; (2) Supervision loss, introducing pixel-level L1 loss under the condition of the registration dataset; (3) Detection loss of the downstream visual detection model. For the unsupervised conversion framework, the embodiments of the present invention use CIoU as the detection box regression loss and cross-entropy loss as the confidence prediction loss and class prediction loss; for the supervised conversion framework, the embodiments of the present invention use the weighted sum of cross-entropy loss, feature matching loss, and semantic segmentation-guided L1 loss as the detection loss of the downstream visual detection model to strengthen the target area constraint. The specific design is as follows:

[0044] CIoU: It consists of the overlapping area (IoU), the center point distance penalty term, and the aspect ratio penalty term. IoU calculates the intersection over union of the predicted box and the ground truth box to measure the overlapping degree between the two; the center point distance penalty term constrains the position offset between the two through the Euclidean distance between the center points of the predicted box and the ground truth box; the aspect ratio penalty term introduces a consistency measure of the aspect ratio to ensure the shape similarity between the predicted box and the ground truth box.

[0045] Cross-entropy loss: Use the semantic segmentation model pre-trained on the real optical image to segment the optical image generated by the conversion model, and calculate the cross-entropy loss between the segmentation result and the semantic segmentation label of the SAR image of the image as the final cross-entropy loss.

[0046] Feature matching loss: Use the downsampling network of the semantic segmentation model pre-trained on the real optical image to downsample the input real optical image and the pseudo-optical image generated by the conversion model respectively, calculate the L1 loss for the features obtained by downsampling at each level, and calculate the average value of the total L1 loss as the final feature matching loss.

[0047] Semantic segmentation-guided L1 loss: Use the semantic segmentation model pre-trained on the real optical image to perform binary classification segmentation on the input optical image, multiply the 01 matrix of the segmentation result element-wise with the input real optical image and the pseudo-optical image generated by the conversion model respectively, calculate the L1 loss of the two images obtained after multiplication, and multiply by a weight as the final semantic segmentation-guided enhanced L1 loss, where the weight value ranges from 10 to 20.

[0048] Step 7: Joint optimization training. That is, select a suitable training strategy, and continuously optimize the parameters of the SAR-optical image conversion model in steps 5-6 through repeated iterative training processes. At each iteration, adjust the parameters of the image conversion model according to the conversion result of the image conversion model and the information feedback of the downstream visual detection model to reduce the difference between the generated image and the real image; the training strategy includes the weight of the loss, the number of rounds of iterative training, and the step size of training.

[0049] The embodiments of the present invention adopt an alternating optimization strategy for end-to-end training: First, the parameters of the conversion model are fixed, and the corresponding model output is obtained by using the SAR-semantic segmentation model or the SAR-object detection model; Subsequently, the model output and the original SAR image are spliced and input into the generator to generate a pseudo-optical image; Then, the SAR-semantic segmentation model or the SAR-object detection model is used to extract multi-level features of the generated image and the real sample, and the semantic consistency loss is calculated; Finally, the image conversion ability of the generator and the discrimination ability of the discriminator are synchronously optimized through backpropagation. In the training process of the embodiments of the present invention, an adaptive learning rate strategy is adopted, and the training is terminated when the segmentation accuracy of the validation set reaches a stable threshold.

[0050] Step 8: SAR-optical image conversion. Use the trained SAR-optical image conversion model to convert the input SAR image into a pseudo-optical image with optical image features, and complete the SAR-optical image conversion.

[0051] Step 9: Evaluation of downstream vision detection tasks. In the embodiments of the present invention, for the pseudo-optical image generated by the SAR-optical image conversion model driven by the constructed downstream vision detection model, the corresponding downstream vision detection tasks are tested, that is, the downstream vision detection model is used to detect the generated pseudo-optical image, and the detection results are evaluated for quality. Among them, the downstream vision detection model and the quality evaluation index used are the same as those in the SAR-optical image conversion model trained in Step 7. In the embodiments of the present invention, for the semantic segmentation task, the class accuracy of pixels is calculated, and then the average value is accumulated; for the object detection task, the mean average precision mAP with confidence levels of 0.5 and 0.95 is calculated.

[0052] The conversion result generated by the SAR-optical image conversion model driven by the trained downstream vision detection model is a characterization result with optical features, and this conversion result can be directly input into the pre-trained optical semantic segmentation model or object detection model to improve the performance of downstream tasks such as semantic segmentation and object detection.

[0053] A SAR-optical image conversion system driven by a downstream vision detection model, as Figure 2 shown, includes an acquisition module, a transmission device, and a server device. Among them, the acquisition module is used to obtain the video stream data of SAR and optical images, and is connected to the server device through the transmission device; The server device is used to execute the SAR-optical image conversion method driven by the downstream vision detection model, generate conversion results and evaluation indicators, and display them in real time; The transmission device is used to transmit the video stream data to the server device.

[0054] A storage medium stores a computer program which, when run, is used to execute the steps of a SAR-optical image conversion method driven by a downstream vision detection model disclosed by the present invention. The memory may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories.

[0055] Among them, the development process of the computer program can be carried out according to the following steps:

[0056] Step 1: Selection of development framework. Build a desktop application based on a cross-platform framework, develop using a high-level programming language, and make full use of the flexibility and cross-platform characteristics of the framework.

[0057] Step 2: Design of the main interface functions of the application. It includes the following core modules:

[0058] (1) Data upload function: Support the upload of image and text data in multiple formats, and have intelligent verification functions (such as resolution, number of channels, etc.).

[0059] (2) Cache cleaning function: Record the temporary file log and support cleaning the cache by type.

[0060] (3) Function navigation area: Provide entrances to each task module to facilitate switching to different function interfaces.

[0061] Step 3: Design of the downstream vision detection task interface. It includes the following core modules:

[0062] (1) Model training: Build in multiple model architectures, support parameter customization and visualization of the training process.

[0063] (2) Model testing: Support batch processing and result display, and provide model inference acceleration and metric visualization.

[0064] (3) Function navigation: An operation interface connecting the training and testing functions.

[0065] Step 4: Design of the SAR-optical image conversion interface. It includes the following core modules:

[0066] (1) Model training: Support multiple model types, provide parameter configuration and visualization of the training process.

[0067] (2) Model testing: Support batch processing, and provide display of conversion results and metrics.

[0068] (3) Function navigation: Connect to the relevant training and testing function interfaces.

[0069] Step 5: Design of the model training interface. It includes a dataset selection function and a model architecture selection function.

[0070] Step 6: Design of the model test interface. It includes functions such as data selection, model configuration, metric calculation, and result visualization.

Claims

1. A SAR-optical image conversion method driven by a downstream visual detection model, characterized in that The steps are as follows: Step 1: Obtain remote sensing image data and perform annotation. Among them, the remote sensing image data includes SAR images and optical images of the corresponding scenes, and the annotation is semantic or class label annotation; Step 2: Preprocess the image data obtained in Step 1, including normalization processing of the image spatial resolution, pixel value normalization processing, and image enhancement processing; Step 3: Use the image data processed in Step 2 to pre-train the downstream visual detection model. The downstream visual detection model includes a semantic segmentation task model and an object detection task model; Step 4: Select an image conversion model. Among them, for the paired dataset of SAR images and optical images, a supervised image conversion model is adopted, and for the unpaired dataset of SAR images and optical images, an unsupervised image conversion model is adopted; Step 5: Construct a SAR-optical image conversion model driven by the downstream visual detection model, including a forward branch and a feedback branch. Among them, the forward branch uses the visual detection model pre-trained on the SAR image to detect the input SAR image, and after cascading the detection result with the SAR image at the channel level, it is input into the image conversion model for training; the feedback branch uses the visual detection model pre-trained on the optical image to detect the generated pseudo-optical image and the real optical image respectively, and performs quality evaluation on the detection results, and feeds back the evaluation results into the image conversion model to guide its training; Step 6: Design a loss function, including three parts: adversarial loss, supervised loss, and detection loss of the downstream visual detection model. Among them, for the unsupervised conversion model, the detection loss of the downstream visual detection model includes CIoU detection box regression loss, and confidence prediction loss and class prediction loss using cross-entropy loss; for the supervised conversion model, the weighted sum of cross-entropy loss, feature matching loss, and semantic segmentation-guided L1 loss is used as the detection loss of the downstream visual detection model; Step 7: Perform joint optimization training on the SAR-optical image conversion model. Specifically: select a suitable training strategy, and continuously optimize the parameters of the SAR-optical image conversion model through repeated iterative training processes. At each iteration, adjust the parameters of the image conversion model according to the conversion result of the image conversion model and the information feedback of the downstream visual detection model to reduce the difference between the generated image and the real image; the training strategy includes the weight of the loss, the number of rounds of iterative training, and the step size of training; Step 8: Use the trained SAR-optical image conversion model to convert the input SAR image into a pseudo-optical image with optical image characteristics to complete the SAR-optical image conversion; Step 9: Use the downstream visual detection model to detect the pseudo-optical image generated by the SAR-optical image conversion model and perform quality evaluation on the detection results. Among them, the downstream visual detection model and quality evaluation index used are consistent with those in the SAR-optical image conversion model trained in Step 7.

2. The downstream visual detection model-driven SAR-to-optical image conversion method according to claim 1, characterized in that: The acquisition of the remote sensing image dataset described in step 1 refers to obtaining or self-collecting and constructing a remote sensing image dataset from a public dataset or a satellite image database; the annotation refers to using an interactive segmentation annotation tool to perform semantic or class label annotation on the image.

3. The SAR-optical image conversion method driven by a downstream vision detection model according to claim 1, wherein: The image enhancement processing described in step 2 includes operations such as random cropping, rotation, and contrast enhancement of the image.

4. The SAR-optical image conversion method driven by a downstream vision detection model according to claim 1, wherein: The semantic segmentation task model described in step 3 uses the U-Net network, and the object detection task model uses the YOLOv5 network.

5. A SAR-optical image conversion method driven by a downstream vision detection model according to claim 1, characterized in that: The specific process of the model pre-training described in step 3 is as follows: First, using SAR images and corresponding annotations as inputs, train the semantic segmentation task model and the object detection task model to obtain the semantic segmentation task model and the object detection task model for SAR images; then, using optical images and corresponding annotations as inputs, train the semantic segmentation task model and the object detection task model to obtain the semantic segmentation task model and the object detection task model for optical images; during training, for the semantic segmentation task model, use the cross-entropy loss function, and for the object detection task model, use the detection box regression loss, confidence prediction loss, and class prediction loss; during training, optimize the network parameters through backpropagation.

6. The downstream visual detection model-driven SAR-to-optical image conversion method according to claim 1, characterized in that: In the SAR-optical image conversion model driven by the downstream visual detection model constructed in step 5, for the semantic segmentation model, use the pix2pix supervised framework; for the object detection model, use the CycleGAN unsupervised framework.

7. The downstream visual detection model-driven SAR-to-optical image conversion method according to claim 1, characterized in that: The adversarial loss described in step 6 refers to using the least squares loss to optimize the adversarial training of the generator and the discriminator; the supervised loss refers to introducing the pixel-level L1 loss under the condition of the registration dataset; the CIoU consists of the overlap region IoU, the center point distance penalty term, and the aspect ratio penalty term. The IoU calculates the intersection over union of the predicted box and the ground truth box to measure the overlap degree between the two. The center point distance penalty term constrains the position offset between the two through the Euclidean distance between the center points of the predicted box and the ground truth box; the aspect ratio penalty term introduces a consistency measure of the aspect ratio to ensure the shape similarity between the predicted box and the ground truth box; the cross-entropy loss refers to using the semantic segmentation task model pre-trained on the optical image to segment the optical image generated by the conversion model, and calculating the cross-entropy loss between the segmentation result and the semantic segmentation label of the SAR image of the image as the final cross-entropy loss; the feature matching loss refers to using the downsampling network of the semantic segmentation task model pre-trained on the optical image to downsample the input real optical image and the pseudo-optical image generated by the conversion model respectively, calculating the L1 loss for the features obtained by downsampling at each level, and calculating the average value of the total L1 loss as the final feature matching loss; The semantic segmentation-guided L1 loss refers to using the semantic segmentation task model pre-trained on the optical image to perform binary classification segmentation on the input optical image, multiplying the 01 matrix of the segmentation result by the input real optical image and the pseudo optical image generated by the conversion model element by element, calculating the L1 loss of the two images obtained after the multiplication, and multiplying it by a weight as the final semantic segmentation-guided enhanced L1 loss, where the weight value is 10 to 20.

8. A SAR-optical image conversion method driven by a downstream visual detection model according to claim 1, characterized in that: In step 7, an alternating optimization strategy is used to train the model end-to-end. The specific process is as follows: first, the conversion model parameters are fixed, and the corresponding model output is obtained by using the visual detection model on the SAR image; then, the model output and the original SAR image are spliced and input into the generator to produce a pseudo-optical image; then, the visual detection model on the SAR image is used to extract the multi-level features of the generated image and the real sample, and the semantic consistency loss is calculated; finally, the image conversion ability of the generator and the discrimination ability of the discriminator are simultaneously optimized through backpropagation; an adaptive learning rate strategy is used during the training process, and the training is terminated when the segmentation accuracy of the validation set reaches a stable threshold.

9. A SAR-optical image conversion system driven by a downstream visual detection model, characterized in that include: An acquisition module, a transmission device, and a server device; wherein the acquisition module is used to obtain video stream data of SAR and optical images and connect to the server device through the transmission device; the server device is used to execute the SAR-optical image conversion method driven by the downstream visual detection model, generate conversion results and evaluation indicators, and display them in real time; the transmission device is used to transmit the video stream data to the server device.

10. A storage medium, characterized in that: A computer program is stored thereon, and when the computer program is run, it is used to execute the steps of the method according to any one of claims 1 to 7.