A high-efficiency single-step diffusion SAR image colorization generation method

CN122820875APending Publication Date: 2026-09-25XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610765346.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

生成对抗网络(GAN)方法推理速度快,但因其对图像底层物理特征的建模能力有限,在处理SAR图像特有的相干斑噪声与强散射点时,易产生结构畸变及颜色失真

Benefits of technology

本发明通过在预训练的潜空间扩散模型中加入LoRA层构建生成器模型,在预训练的DINO模型中加入适配器模块构建对比特征提取器模型,并在联合训练过程中冻结两个预训练模型的全部原始参数,仅对新增的LoRA层及适配器模块等微量参数进行更新。一方面极大地降低了模型在SAR图像彩色化任务上的训练算力成本与显存占用;另一方面,由于完整继承了预训练模型强大的生成先验与语义表征能力,使得生成器模型仅需单步前向推理即可输出高质量彩色化结果,从根本上解决了传统扩散模型依赖多步迭代、推理效率低下的问题。同时,在联合训练框架下,生成器模型与对比特征提取器模型协同优化,对比特征提取器能够提取输入SAR图像与生成图像的多层语义特征,为生成器提供结构一致性的监督信号,从而迫使生成图像精准保留原始SAR图像的空间布局与地物轮廓信息,克服了现有生成对抗网络易产生的结构畸变问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820875A_ABST
    Figure CN122820875A_ABST
Patent Text Reader

Abstract

The application discloses a kind of efficient single-step diffusion SAR image colorization generation methods, comprising: constructing dataset;Based on latent space diffusion model and LoRA layer generator is constructed;Based on DINO model and adapter module contrast feature extractor is constructed;Discriminator is constructed;And the joint training of three is carried out.The application freezes all original parameters of pre-training model when joint training, only updates LoRA layer and micro-parameters such as adapter module, greatly reduces training cost;While inheriting the generation prior of pre-training model and semantic representation ability, the generator only needs single-step inference to output high-quality colorization result, solves the problem of low efficiency of traditional diffusion model multi-step iteration.And generator and contrast feature extractor are optimized in training, can extract multi-layer semantic features to provide structural consistency supervision, force generated image to retain the spatial layout and feature contour of SAR image, overcome the structural distortion problem easily produced by generative adversarial network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of SAR image colorization, specifically relating to an efficient method for generating colorized single-step diffusion SAR images. Background Technology

[0002] Synthetic Aperture Radar (SAR) possesses the advantage of all-weather, all-day imaging and plays a vital role in many fields. However, the unique imaging mechanism of SAR images results in significant differences between their visual representation and natural optical images, limiting the efficient interpretation of the data. Therefore, converting grayscale SAR data into colorized natural optical images has become a key technological approach to improve its recognition and understanding efficiency.

[0003] In recent years, deep learning-based SAR image colorization methods have become a research hotspot, with generative adversarial networks (GANs) and diffusion models being the mainstream technical approaches. While GANs offer fast inference speeds, their limited ability to model the underlying physical features of images makes them prone to structural distortion and color distortion when dealing with speckle noise and strong scattering points unique to SAR images. Although diffusion models generate high-quality images, they typically rely on hundreds or even thousands of iterative denoising steps, resulting in high training costs and slow inference speeds, making it difficult to meet the demands of efficient, real-time processing in practical applications.

[0004] Therefore, how to effectively maintain the semantic structure of SAR images and improve the visual quality of colorization results while reducing model training and inference costs and improving generation efficiency is a technical problem that needs to be solved by existing technologies. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides an efficient method for single-step diffusion SAR image colorization generation. An efficient method for colorization generation of single-step spread SAR images includes: Construct a dataset that includes multiple SAR images and multiple visible optical remote sensing images; A generator model is constructed based on a preset latent space diffusion model. The generator model includes multiple LoRA layers. One of the LoRA layers is used to perform low-rank adaptation on the input features of a convolutional layer or a fully connected layer in the original module of the latent space diffusion model. The original module includes an encoder, a U-net network, and a decoder. A contrastive feature extractor model is constructed based on the preset DINO model. The contrastive feature extractor model includes multiple adapter modules. One of the adapter modules is used to perform semantic space mapping on the output features of the nested tensor module of the DINO model. Based on the dataset and the preset training framework, the generator model, the contrastive feature extractor model and the preset discriminator model are jointly trained to obtain the fully trained generator model. During the joint training process, the original modules in the latent space diffusion model and the original parameters of the DINO model are frozen. The SAR image to be colorized is input into the fully trained generator model for processing to obtain a colorized SAR image.

[0006] In one embodiment of the present invention, constructing the generator model includes: The LoRA layer is embedded in each convolutional layer and fully connected layer in the original module. The LoRA layer performs low-rank adaptation on the input features of the first network layer it is connected to, and generates the output features of the first network layer based on the result of the low-rank adaptation. The first network layer is a convolutional layer or a fully connected layer in the original module. The LoRA layer consists of a preset lower projection matrix and an upper projection matrix. The dimension of the LoRA layer of the encoder and the decoder is a preset first dimension, and the dimension of the LoRA layer of the U-net network is a preset second dimension.

[0007] In one embodiment of the present invention, the generation of the output features of the first network layer based on the result of low-rank adaptation can be expressed as: ; in, The output features of the first network layer, These are the input features of the first network layer. The weight matrix is ​​a matrix. The lower projection matrix is ​​used to project the input features. Mapping to a lower-dimensional space, matrix The up-projection matrix is ​​used to map low-dimensional features back to the original feature dimensions.

[0008] In one embodiment of the present invention, constructing the generator model includes: In the latent space diffusion model, multiple residual connection layers are set. The residual connection layers are used to perform convolution processing on the output feature map of the encoder through a preset zero convolution layer, so that the spatial size of the output feature map matches the feature map of the corresponding layer of the decoder.

[0009] In one embodiment of the present invention, the step of setting multiple residual connection layers in the latent space diffusion model includes: A total of four zero-convolutional layers are configured; The input of the first zero convolutional layer is the output of the third downsampling encoding module of the encoder. The output features are added to the output of the first decoder convolutional layer of the decoder and then used as the input of the first upsampling decoding module of the decoder. The input of the second zero convolutional layer is the output of the second downsampling encoding module of the encoder. The output features are added to the output of the first upsampling decoding module and then used as the input of the second upsampling decoding module of the decoder. The input of the third zero convolutional layer is the output of the first downsampling encoding module of the encoder. The output features are added to the output of the second upsampling decoding module and then used as the input of the third upsampling decoding module of the decoder. The input to the fourth zero convolutional layer is the output of the first encoder convolutional layer of the encoder. The output features are added to the output of the third upsampling decoding module and then used as the input to the fourth upsampling decoding module of the decoder.

[0010] In one embodiment of the present invention, the contrast feature extractor model includes a first adapter module, a second adapter module, a third adapter module, and a fourth adapter module; The DINO model includes a first DINO convolutional layer and twelve stacked nested tensor modules; The first adapter module, the second adapter module, the third adapter module, and the fourth adapter module perform semantic space mapping on the output features of the third nested tensor module, the sixth nested tensor module, the ninth nested tensor module, and the twelfth nested tensor module of the DINO model, respectively, to obtain the first semantic feature, the second semantic feature, the third semantic feature, and the fourth semantic feature; The output features of the contrast feature extractor model are generated based on the first semantic feature, the second semantic feature, the third semantic feature, and the fourth semantic feature.

[0011] In one embodiment of the present invention, the joint training of the generator model, the contrastive feature extractor model, and the preset discriminator model based on the dataset and a preset training framework includes: Normalize the images in the dataset; A training sample pair is formed based on a SAR image and a visible optical remote sensing image from the dataset. The model parameters of the generator model, the contrastive feature extractor model, and the discriminator model are updated based on a set of training sample pairs. Multiple sets of training sample pairs are generated, and the model parameters of the generator model, the contrast feature extractor model, and the discriminator model are updated based on the multiple sets of training sample pairs until a preset number of training iterations are reached, thereby obtaining the fully trained generator model.

[0012] In one embodiment of the present invention, updating the model parameters of the generator model, the contrastive feature extractor model, and the discriminator model based on a set of training sample pairs includes: The original SAR image is input into the generator model to generate a fake optical image, wherein the original SAR image is the SAR image in the training sample pair; The fake optical image is input into the discriminator model, the adversarial loss of the generator model is calculated to obtain the first adversarial loss, and the parameter gradient of the generator model is calculated based on the first adversarial loss to obtain the first parameter gradient. The model parameters of the generator model are updated according to the first parameter gradient. The original optical image is input into the discriminator model, the adversarial loss of the discriminator model is calculated to obtain the second adversarial loss, and the parameter gradient of the discriminator model is calculated based on the second adversarial loss to obtain the second parameter gradient. The model parameters of the discriminator model are updated according to the second parameter gradient. The original optical image is the visible optical remote sensing image in the training sample pair. The original optical image is input into the generator model to generate an identity mapping image, and the Manhattan distance between the identity mapping image and the original optical image is calculated, and the Manhattan distance is used as the identity mapping loss. The parameter gradient of the generator model is calculated based on the identity mapping loss to obtain the third parameter gradient, and the model parameters of the generator model are updated based on the third parameter gradient. The contrast loss is calculated based on the original SAR image and the fake optical image. The contrast loss is then used to calculate the parameter gradients of the generator model and the contrast feature extractor model, respectively, to obtain the fourth parameter gradient and the fifth parameter gradient. The model parameters of the generator model and the contrast feature extractor model are then updated using the fourth parameter gradient and the fifth parameter gradient, respectively.

[0013] In one embodiment of the present invention, the step of calculating the contrast loss based on the original SAR image and the false optical image includes: The original SAR image and the fake optical image are respectively input into the contrast feature extractor model to obtain the first feature map and the second feature map; Flatten the first feature map and the second feature map along the height and width directions of the feature map respectively to obtain a first feature matrix and a second feature matrix containing the same number of feature points; Feature vectors are extracted from multiple first positions of the first feature matrix and the second feature matrix respectively to obtain multiple first feature vectors and multiple second feature vectors. The first feature vectors and the second feature vectors are feature vectors extracted from the first feature matrix and the second feature matrix respectively. Each first position corresponds to one first feature vector and one second feature vector. The contrast loss is obtained by calculating the contrast loss based on multiple first feature vectors and multiple second feature vectors.

[0014] The beneficial effects of this invention are: This invention constructs a generator model by adding a LoRA layer to a pre-trained latent space diffusion model and a contrastive feature extractor model by adding an adapter module to a pre-trained DINO model. During joint training, all original parameters of the two pre-trained models are frozen, with only minor parameters such as the newly added LoRA layer and adapter module being updated. This significantly reduces the computational cost and memory usage for training the model in SAR image colorization tasks. Furthermore, by fully inheriting the powerful generative priors and semantic representation capabilities of the pre-trained models, the generator model can output high-quality colorization results with only a single forward inference step, fundamentally solving the problem of low inference efficiency and reliance on multi-step iterations in traditional diffusion models. Simultaneously, within the joint training framework, the generator model and the contrastive feature extractor model are co-optimized. The contrastive feature extractor can extract multi-layer semantic features from the input SAR image and the generated image, providing the generator with a supervisory signal for structural consistency. This forces the generated image to accurately retain the spatial layout and feature outline information of the original SAR image, overcoming the structural distortion problem easily generated by existing generative adversarial networks.

[0015] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating an efficient single-step diffusion SAR image colorization generation method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the framework of an efficient single-step diffusion SAR image colorization generation method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the network structure of the generator model provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the network structure of the contrast feature extractor model provided in an embodiment of the present invention; Figure 5 This is a schematic diagram illustrating an example image of the dataset provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the comparison results provided in the embodiments of the present invention. Detailed Implementation

[0017] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0018] Example 1 It should be noted that SAR image colorization falls under the category of image style transfer tasks, the core objective of which is to establish a semantically consistent mapping relationship from one image domain to another. Currently, image style transfer mainly employs two main technical approaches: Generative Adversarial Networks (GANs) and diffusion models. The core principle of GANs lies in constructing an adversarial game mechanism between a generative model and a discriminative model. The generative model learns a nonlinear mapping from the source domain to the target domain to synthesize forged images, while the discriminative model identifies the authenticity of input samples by mining the distribution characteristics of the target. This adversarial training process essentially seeks a Nash equilibrium point in the function space, making the generated result statistically close to the real image. Its significant technical advantage lies in the fact that cross-modal feature conversion can be completed in a single forward computation during the inference stage, exhibiting high generation efficiency and real-time response capabilities. Diffusion models, on the other hand, are based on a random diffusion process, achieving image reconstruction through two stages: forward diffusion and inverse denoising. In the forward process, the model gradually introduces Gaussian noise into the real image through a predefined variance scheduling scheme until it evolves into a purely random noise distribution. In the reverse process, the model learns how to accurately estimate and eliminate noise components in the image from random noise states using progressive iterations of Markov chains, gradually restoring the image's deep semantic features and fine texture information. This probability diffusion-based generative logic enables the model to effectively model the high-dimensional probability distribution of ground features in complex scenes, demonstrating significant advantages in achieving image structure preservation, detail reconstruction, and visual realism.

[0019] While existing image transfer methods have achieved significant results in style transfer, their direct application to SAR image colorization tasks still faces insurmountable technical bottlenecks. On one hand, generative adversarial networks (GANs), lacking the ability to finely model the underlying physical features of images, may exhibit structural distortion and color distortion when processing the unique speckle noise and strong scattering points of SAR images. This results in generated color images that fail to meet the requirements for detailed interpretation in terms of detail fidelity. Furthermore, the unstable training process of GANs is prone to pattern collapse, limiting their generalization ability in complex remote sensing scenarios. On the other hand, while diffusion models generate high-quality images, they face enormous computational resource challenges in practical applications. Because their core algorithm relies on hundreds or even thousands of iterations of denoising, the model not only has extremely high demands for high-performance GPU memory during training but also exhibits very limited generation efficiency during inference. The high training and inference costs make diffusion models unsuitable for real-time processing scenarios in remote sensing monitoring, which have extremely high timeliness requirements, limiting the model transfer and customized development for specific terrain scenes.

[0020] To balance the generation quality and computational efficiency of SAR image colorization, this embodiment proposes a SAR image colorization method based on comparative basic model fine-tuning. The core idea is to guide the single-step diffusion model through comparative learning objectives to achieve efficient and high-fidelity cross-modal conversion.

[0021] Combination Figure 1 As shown, an efficient single-step diffusion SAR image colorization generation method includes: S1. Construct a dataset, which includes multiple SAR images and multiple visible optical remote sensing images; S2. Construct a generator model based on a preset latent space diffusion model. The generator model includes multiple LoRA layers. One of the LoRA layers is used to perform low-rank adaptation on the input features of a convolutional layer or a fully connected layer in the original module of the latent space diffusion model. The original module includes an encoder, a U-net network, and a decoder. S3. Construct a contrastive feature extractor model based on the preset DINO model. The contrastive feature extractor model includes multiple adapter modules. One of the adapter modules is used to perform semantic space mapping on the output features of the nested tensor module of the DINO model. S4. Based on the dataset and the preset training framework, the generator model, the contrast feature extractor model and the preset discriminator model are jointly trained to obtain the fully trained generator model. During the joint training process, the original modules in the latent space diffusion model and the original parameters of the DINO model are frozen. S5. Input the SAR image to be colored into the fully trained generator model for processing to obtain a colored SAR image.

[0022] Understandably, in this embodiment, the variational autoencoder, U-Net network, and decoder of the pre-trained stable diffusion model are deeply integrated to construct an end-to-end generative architecture. A global residual connection mechanism is innovatively introduced between the encoder and decoder to enhance the cross-layer transfer of underlying geometric features. In terms of optimization strategy, the parameter-efficient Low-Rank Adaptation (LoRA) technique is adopted. While freezing the main weights of the pre-trained model, parameters are updated only for newly added fine-tuning layers, significantly reducing the computational cost during training. By constructing a joint optimization objective of adversarial learning and contrastive learning, this embodiment reduces the overall difference between the generated image and the real optical image while using contrastive constraints to guide the model to accurately retain the structural information of the input SAR image, ultimately achieving high-quality SAR image colorization generation with single-step inference capabilities. The specific steps are as follows: Specifically, in step S1, this embodiment collects SAR images and visible optical remote sensing images to construct a training dataset, wherein the SAR images serve as the source domain dataset. Visible optical remote sensing images as target domain datasets .

[0023] Furthermore, in combination Figure 2 and Figure 3 As shown, step S2 specifically includes: S21. A generator model is constructed by loading a pre-trained Latent Diffusion Model (LDM) as the architecture, and the weights of the "sd-turbo" pre-trained model are loaded from the open-source diffuses library as the generator baseline model. Its network architecture is as follows: Figure 3 As shown, it specifically consists of three parts: a variational autoencoder (VAE) encoder, a U-net network, and a VAE decoder. The VAE encoder maps the input high-dimensional SAR image to a low-dimensional latent space representation to reduce the computational complexity of the diffusion model and improve generation efficiency. Its network structure includes a first encoder convolutional layer, a first downsampling coding module, a second downsampling coding module, a third downsampling coding module, a fourth downsampling coding module, a first encoder intermediate module, and a second encoder convolutional layer.

[0024] Based on the latent space representation output by the VAE encoder, the U-net network performs distributional transfer on it, mapping it to a latent space representation corresponding to the optical image style. Its network structure includes a first U-net convolutional layer, a first U-net downsampling module, a second U-net downsampling module, a third U-net downsampling module, a fourth U-net downsampling module, a first U-net intermediate module, a first U-net upsampling module, a second U-net upsampling module, a third U-net upsampling module, a fourth U-net upsampling module, and a second U-net convolutional layer.

[0025] Finally, the VAE decoder decodes and recovers the latent space representation of the optical image style, outputting the final colorized image. Its network structure includes a first decoder convolutional layer, a first upsampling decoding module, a second upsampling decoding module, a third upsampling decoding module, a fourth upsampling decoding module, a first decoder intermediate module, and a second decoder convolutional layer.

[0026] It is understood that the basic network structure of the LDM model is well known to those skilled in the art, and the specific module details will not be described in this embodiment.

[0027] S22. In the latent space diffusion model, multiple residual connection layers are set. The residual connection layers are used to perform convolution processing on the output feature map of the encoder through a preset zero convolution layer, so that the spatial size of the output feature map matches the feature map of the corresponding layer of the decoder.

[0028] Specifically, a total of four zero-convolutional layers are set; The input of the first zero convolutional layer is the output of the third downsampling encoding module of the encoder. The output features are added to the output of the first decoder convolutional layer of the decoder and then used as the input of the first upsampling decoding module of the decoder. The input of the second zero convolutional layer is the output of the second downsampling encoding module of the encoder. The output features are added to the output of the first upsampling decoding module and then used as the input of the second upsampling decoding module of the decoder. The input of the third zero convolutional layer is the output of the first downsampling encoding module of the encoder. The output features are added to the output of the second upsampling decoding module and then used as the input of the third upsampling decoding module of the decoder. The input to the fourth zero convolutional layer is the output of the first encoder convolutional layer of the encoder. The output features are added to the output of the third upsampling decoding module and then used as the input to the fourth upsampling decoding module of the decoder.

[0029] In this embodiment, because the LDM model downsamples the input image multiple times, detailed information in the original SAR image is easily lost, resulting in the generated image not conforming to the semantic structure of the original image. Therefore, this invention inserts multiple residual connections between the encoder and decoder of the LDM. To ensure network training stability, the residual connections use zero-convolutional layers to process the encoder features, and their convolutional kernel weights are initialized to a fixed value. The structure will be adjusted as the network is optimized. A total of four zero-convolutional layers are configured. Specifically, the first zero-convolutional layer takes the output of the third downsampling encoding module as input, and its output features are added to the output of the first decoder convolutional layer to serve as the input of the first upsampling decoding module. The second zero-convolutional layer takes the output of the second downsampling encoding module as input, and its output features are added to the output of the first upsampling decoding module to serve as the input of the second upsampling decoding module. The third zero-convolutional layer takes the output of the first downsampling encoding module as input, and its output features are added to the output of the second upsampling decoding module to serve as the input of the third upsampling decoding module. The fourth zero-convolutional layer takes the output of the first encoder convolutional layer as input, and its output features are added to the output of the third upsampling decoding module to serve as the input of the fourth upsampling decoding module.

[0030] Understandably, to achieve efficient low-dimensional latent space diffusion, the LDM model's VAE encoder performs multiple downsamplings on the input image. While this significantly reduces computational complexity, it also leads to the gradual loss of fine-grained structural information from the original SAR image in the encoder's deep feature map. When the decoder relies solely on the latent space features output by the U-Net for reconstruction, it struggles to recover these lost high-frequency details, resulting in structural distortions or semantic inconsistencies.

[0031] Furthermore, the speckle noise and strong scattering points unique to SAR images make it easier for the encoder to confuse signal with noise during compression, further exacerbating the loss of structural information. Therefore, it is necessary to establish one or more additional information transmission channels between the encoder and decoder to directly transfer the relatively complete structural features preserved in the intermediate layers of the encoder to the corresponding layers of the decoder, in order to compensate for the information loss caused by downsampling.

[0032] Based on the above considerations, this embodiment adds a residual connection layer, the purpose of which is: Maintaining the structural integrity of the original SAR image: The high-resolution feature map extracted from the shallow layer of the encoder is passed to the decoder to help the decoder more accurately restore the geometric contours and spatial layout of ground features when generating color images.

[0033] Improved image quality: By preserving structural details, the generated optical images are significantly improved in edge sharpness, texture fidelity, and semantic consistency, while reducing color distortion and artifacts.

[0034] It is worth noting that the residual connections in this invention are not ordinary convolutional layers, but rather zero-convolutional layers as connection paths. The kernel weights of the zero-convolutional layer are initialized to minimum values, and the biases are initialized to zero. Since the weights are close to zero, the output feature map of the zero-convolutional layer in the initial state is also close to zero. Therefore, the residual connections pass almost no signal to the decoder in the early stages of training. At this time, the behavior of the entire generator is equivalent to the original, unmodified pre-trained LDM model. This avoids the initial output deviating drastically from the pre-training distribution due to random initialization of residual branches, thus protecting the good generative ability already learned by the pre-trained model.

[0035] Meanwhile, as training progresses, gradient descent gradually increases the weights of the zero-convolutional layers, and the residual branches gradually pass more and more encoder features to the decoder. The network can adaptively learn the optimal information transfer strength from the encoder to the decoder. For samples requiring strong structure preservation, the weights will increase; for samples with simple structures or severe noise, the weights can remain small or even zero. This adaptability further stabilizes the training process.

[0036] This invention employs four layers of residual connections at different scales. This multi-scale design ensures the effective transfer of structural information at different levels (from local texture to global shape). Each residual branch is independent and does not interfere with the others, making the overall optimization significantly less difficult than with a single deep skip connection.

[0037] S23. The LoRA layer is embedded in each convolutional layer and fully connected layer in the original module. The LoRA layer performs low-rank adaptation on the input features of the first network layer it is connected to, and generates the output features of the first network layer based on the result of the low-rank adaptation. The first network layer is a convolutional layer or a fully connected layer in the original module. The LoRA layer is composed of a preset lower projection matrix and an upper projection matrix. The dimension of the LoRA layer of the encoder and the decoder is a preset first dimension, and the dimension of the LoRA layer of the U-net network is a preset second dimension.

[0038] Specifically, this invention employs the LoRA method to perform lightweight fine-tuning of the pre-trained LDM, achieving efficient adaptation for SAR image colorization tasks without altering the original parameters and network structure of the basic model. The calculation process of the original network layers is denoted as follows:

[0039] in, As input features, For output features, This is the weight matrix. Input the feature dimension to this layer. This layer outputs the feature dimension. LoRA adds a downward projection matrix to the original network layer. and projection matrix The residual connection branch composed of two matrices, after modification, results in the following network layer calculation process:

[0040] Among them, matrix Responsible for mapping input features to a low-dimensional space, matrix Responsible for mapping low-dimensional features back to the original feature dimensions. It is a low-rank dimension. Through residual connection branches and low-rank adaptation matrices, LoRA can achieve fine-tuning and optimization of the model feature space with a very small number of low-rank parameters.

[0041] For all convolutional and fully connected layers in the encoder, U-net network, and decoder of the pre-trained LDM model, a separate LoRA layer is embedded in each layer. Specifically, the LoRA layers in the VAE encoder and VAE decoder are set to have low-rank dimensions. Set a low-rank dimension for the LoRA layer of the U-net network. Matrix of all LoRA layers Initialize using Kaiming Uniform, matrix Initialize with all zeros.

[0042] Thus, the generator model of this invention is completed and marked as follows. Its input is the original SAR image, and its output is the corresponding colorized image. During training, all original parameters of the encoder, U-net network, and decoder in the pre-trained LDM model are frozen, and only the model weights of the residual connection layer and LoRA layer are trained.

[0043] Understandably, the pre-trained LDM model was originally designed for natural image generation, and its parameters have been sufficiently trained on large-scale optical datasets. If this model is directly applied to SAR image colorization tasks, it needs to be retrained or fine-tuned to adapt to the cross-modal mapping of SAR. However, completely fine-tuning all parameters presents the following problems: First, computational costs are high. LDM models have a huge number of parameters, and full parameter training requires a significant amount of GPU memory and time. Second, there is a risk of overfitting. SAR image datasets are relatively limited in size, and full parameter fine-tuning can easily lead to model overfitting, reducing generalization ability. Furthermore, pre-training knowledge may be forgotten, and significant parameter updates may destroy the model's original good generative ability, resulting in unnecessary feature drift.

[0044] Therefore, this embodiment proposes a parameter-efficient fine-tuning method that effectively adapts to SAR image colorization tasks while minimizing training resource consumption. Low-rank adaptation (LoRA) layers are embedded in each convolutional and fully connected layer of the generator to reduce the number of trainable parameters. Through low-rank decomposition, the trainable parameters consist only of small-sized low-rank matrices A and B, significantly reducing the number of parameters. Simultaneously, pre-training capability is maintained; the original weights are completely frozen, preserving the model's fundamental generative ability, and task-specific offsets are learned only through low-rank branches. Furthermore, efficient transfer learning is achieved, quickly adapting to SAR image distributions with limited data, avoiding overfitting, while retaining the high-quality generative prior of the original LDM.

[0045] This approach significantly improves training efficiency—requiring only a small number of training parameters, reducing memory usage, and shortening training time—while also ensuring generation quality. The low-rank constraint of LoRA prevents the model from deviating drastically from the pre-trained feature space during fine-tuning. Simultaneously, inference efficiency remains unchanged; during inference, the computational burden of adding new branches can be mitigated by merging weights.

[0046] Furthermore, the contrast feature extractor model includes a first adapter module, a second adapter module, a third adapter module, and a fourth adapter module; The DINO model includes a first DINO convolutional layer and twelve stacked nested tensor modules; The first adapter module, the second adapter module, the third adapter module, and the fourth adapter module perform semantic space mapping on the output features of the third nested tensor module, the sixth nested tensor module, the ninth nested tensor module, and the twelfth nested tensor module of the DINO model, respectively, to obtain the first semantic feature, the second semantic feature, the third semantic feature, and the fourth semantic feature; The output features of the contrast feature extractor model are generated based on the first semantic feature, the second semantic feature, the third semantic feature, and the fourth semantic feature.

[0047] Specifically, in combination Figure 2 and Figure 4 As shown, the specific steps in step S3 of constructing the contrastive feature extractor model include: S31. Load the pre-trained DINO model. The purpose of contrastive learning is to force the generated image to retain the spatial semantic information of the original image and avoid pattern collapse by constructing correlation constraints between different image patches of the original image and the generated image. In this embodiment, the pre-trained DINOv2-ViT-S / 14 model is first loaded, and its powerful semantic representation capability is used to construct a contrastive feature space. Specifically, the structure of the DINO model includes a first DINO convolutional layer and 12 stacked nested tensor modules. Given an input image, this embodiment obtains the output features of the multi-layer nested tensor modules of the pre-trained DINO model. ,in Represents the selected layer number. For the dimension corresponding to the feature, This represents the height and width of the corresponding feature map. Each feature point corresponds to an image patch at the corresponding location in the original image.

[0048] S32. Constructing a multi-layer adapter module. Contrastive learning is used to preserve the original image structure and avoid information loss; therefore, the feature space is required to have a high representational ability of the local spatial structure of the image. Since the feature space of the original DINO network is constructed using a self-supervised method, it may not be the optimal structural representation space for optical-SAR images. To further construct a contrastive feature learning space, this embodiment constructs a multi-layer adapter module to further map the DINO features into a semantic space. For each layer of features... Set up a separate adapter To achieve feature space mapping, each adapter consists of two fully connected network layers, where the first fully connected layer maps the input feature dimensions to the original ones. The second fully connected layer restores the feature dimension to the original input feature dimension. In this embodiment, four adapter modules are set up to perform semantic space mapping on the output features of the third, sixth, ninth, and twelfth nested tensor modules of the pre-trained DINO model, respectively. The calculation process can be represented as follows: .

[0049] It is understandable that the obtained mapped features This feature resides in a new feature space and can be directly used to calculate the contrastive learning loss. Since the adapter parameters are trainable, the new feature space will also dynamically adjust as the network trains, improving the representation of the image's spatial structure.

[0050] Thus, the contrastive feature extractor model of this embodiment is completed, and is marked as follows. Its input is either a real image or a generated image, and its output is four layers of semantic features. The specific structural diagram is as follows: Figure 3As shown. During training, all original parameters of the pre-trained DINO model are frozen, and only the adapter module parameters are trained.

[0051] It should be noted that in this embodiment, the discriminator model adopts the CLIP model architecture from the official Vision-Aided GAN library, denoted as... Its input is either a real image or a generated image, and its output is the predicted probability that the corresponding image belongs to the real image. During training, all weights of the discriminator model participate in gradient updates.

[0052] Furthermore, this embodiment employs a training method combining adversarial learning and contrastive learning to optimize the generator model. Specifically, the generator is used... Discriminator Contrast Feature Extractor Three networks, with loss functions including adversarial loss. Comparative loss and identity mapping loss Three parts.

[0053] Specifically, step S4 includes: S41. Normalize the images in the dataset; S42. A training sample pair is formed based on a SAR image and a visible optical remote sensing image from the dataset. In this embodiment, data initialization is first performed by randomly loading an optical image from the optical and SAR image dataset. With a random SAR image This forms training sample pairs. The original image is normalized, and its pixel amplitude is linearly transformed to... Within the specified range, it facilitates network training.

[0054] S43. Update the model parameters of the generator model, the contrastive feature extractor model, and the discriminator model based on a set of training sample pairs; Specifically, step S43 includes: S431. Input the original SAR image into the generator model to generate a fake optical image, wherein the original SAR image is the SAR image in the training sample pair; S432. Input the fake optical image into the discriminator model, calculate the adversarial loss of the generator model to obtain the first adversarial loss, calculate the parameter gradient of the generator model based on the first adversarial loss to obtain the first parameter gradient, and update the model parameters of the generator model according to the first parameter gradient. In this embodiment, the original SAR image is first... Input generator To obtain false optical images ; Then, the generator adversarial loss is calculated, including: using fake optical images Input discriminator Calculate the adversarial loss of the generator , can be represented as:

[0055] According to the loss of the confrontation Calculate the parameter gradients of the generator model and update the generator model parameters using the AdamW optimizer.

[0056] S433. Input the original optical image into the discriminator model, calculate the adversarial loss of the discriminator model to obtain the second adversarial loss, and calculate the parameter gradient of the discriminator model based on the second adversarial loss to obtain the second parameter gradient. Update the model parameters of the discriminator model according to the second parameter gradient. The original optical image is the visible optical remote sensing image in the training sample pair. In this embodiment, calculating the discriminator adversarial loss includes: processing the original optical image... Input discriminator Calculation discriminator The losses of the confrontation :

[0057] According to the loss of the confrontation Calculate the parameter gradients of the discriminator model and update the discriminator model parameters using the AdamW optimizer.

[0058] S434. Input the original optical image into the generator model to generate an identity mapping image, calculate the Manhattan distance between the identity mapping image and the original optical image, and use the Manhattan distance as the identity mapping loss; In this embodiment, calculating the identity mapping loss includes: converting the original optical image... Input generator To obtain the identity mapping image Calculate the L1 distance (Manhattan distance) between it and the original optical image, as the identity mapping loss. :

[0059] According to the identity mapping loss Calculate the parameter gradients of the generator model and update the generator model parameters using the AdamW optimizer.

[0060] S435. Calculate the parameter gradient of the generator model based on the identity mapping loss to obtain the third parameter gradient, and update the model parameters of the generator model based on the third parameter gradient. S436. Calculate the contrast loss based on the original SAR image and the fake optical image to obtain the contrast loss, and calculate the parameter gradients of the generator model and the contrast feature extractor model based on the contrast loss to obtain the fourth parameter gradient and the fifth parameter gradient, and update the model parameters of the generator model and the contrast feature extractor model through the fourth parameter gradient and the fifth parameter gradient, respectively.

[0061] Specifically, step S436 includes: S4361. Input the original SAR image and the fake optical image into the contrast feature extractor model respectively to obtain the first feature map and the second feature map; S4362. Flatten the first feature map and the second feature map along the height and width directions of the feature map respectively to obtain a first feature matrix and a second feature matrix containing the same number of feature points; S4363. Extract feature vectors from multiple first positions of the first feature matrix and the second feature matrix respectively to obtain multiple first feature vectors and multiple second feature vectors. The first feature vectors and the second feature vectors are feature vectors extracted from the first feature matrix and the second feature matrix respectively. Each first position corresponds to one first feature vector and one second feature vector. S4364. Calculate the contrast loss based on multiple first feature vectors and multiple second feature vectors to obtain the contrast loss.

[0062] In this embodiment, the comparison loss is calculated, including: The purpose of contrastive learning is to construct a correspondence between image patches in the original input SAR image and the generated fake optical images, ensuring that image patches at the same locations have semantically similar structures, thereby improving the generator's ability to preserve the semantics of the original input image. The specific calculation process is as follows: Original SAR image and fake optical images Input into the contrast feature extractor respectively , to obtain output features and ,in .

[0063] Flattening is performed on all output features along the height and width of the feature map, effectively transforming the dimensions of the feature tensor. After flattening, the feature dimensions decrease from 3 to 2, which can be viewed as a feature matrix. The feature dimensions are then transformed from... Change to ,in This represents the number of feature points. Then, randomly select from the flattened result of the original SAR image. Each feature point is used as a patch feature, that is, randomly selected from the flattened feature matrix. Each feature vector is then selected from the flattened result of the spoofed optical image and compared with the result of the flattened original SAR image. Selecting the same position of each patch feature 1 feature point, to obtain the corresponding feature of the fake optical image Each patch feature, i.e. Each feature vector is used. It's necessary to ensure that the randomly selected location is identical in the flattened results of the original SAR image and the fake optical image; only then will the contrastive loss be meaningful. Thus, SAR images are obtained respectively. Patch features and fake optical images Patch features ,in .

[0064] Calculate the contrastive loss based on the patch features. :

[0065] in This is the temperature control coefficient.

[0066] In this embodiment, the contrast loss is calculated. Then, the parameter gradients of the generator and contrastive feature extractor models are calculated based on the contrastive loss, and the parameters of the generator and contrastive feature extractor are updated accordingly using the AdamW optimizer based on the obtained parameter gradients.

[0067] Understandably, existing SAR image colorization methods primarily rely on either generative adversarial training or pure diffusion model training. While adversarial loss alone can generate visually realistic images, it often suffers from structural distortion and color misalignment because it focuses mainly on the consistency of global distribution and lacks constraints on local structural semantics. On the other hand, using a diffusion model alone results in excessively high training and inference costs, making it impractical.

[0068] Therefore, in order to fully utilize the semantic modeling capabilities of the pre-trained diffusion model and efficiently adapt it to the characteristics of optical SAR images, this embodiment designs a joint optimization framework that can simultaneously constrain global style consistency, local structure preservation, and color fidelity. Specifically: This embodiment designs a joint training framework combining adversarial loss, contrast loss, and identity mapping loss. The adversarial loss drives the generated image to closely approximate the real optical image in overall visual style and texture distribution, ensuring the naturalness and realism of colorization. The contrast loss forces high similarity between the image patch features of corresponding spatial locations in the SAR image and the generated optical image, thereby accurately preserving the geometric structure and ground feature boundaries of the original SAR image and effectively suppressing structural distortion. The identity mapping loss constrains the generator to output an approximate original image when the input is an optical image, preventing the generator from meaninglessly altering the image content while enhancing color fidelity and avoiding color drift. The synergistic effect of these three loss mechanisms ensures that the generator does not lose structural information or produce color distortion during style transfer.

[0069] Meanwhile, freezing the original weights of the pre-trained LDM and the original weights of the DINO model during training, and training only the LoRA layer, residual connection layer and adapter module, can reduce training computation costs and avoid the high overhead of full parameter fine-tuning; maintain the rich visual priors of the pre-trained model to prevent overfitting on small-scale SAR datasets; and achieve task-specific lightweight adaptation to improve generalization ability.

[0070] S44. Generate multiple sets of training sample pairs, and update the model parameters of the generator model, the contrast feature extractor model, and the discriminator model based on the multiple sets of training sample pairs until a preset number of training iterations are reached to obtain the fully trained generator model.

[0071] After obtaining the fully trained generator model, the input is normalized to... The generator model can output the colorized generated result from the original SAR image within the range.

[0072] It is understood that this embodiment addresses the shortcomings of existing SAR image colorization methods based on generative adversarial networks and diffusion models, namely insufficient structure preservation, high training costs, and low inference efficiency. It proposes an efficient single-step diffusion SAR image colorization generation method. Compared to existing technologies, this embodiment introduces a low-rank adaptation strategy to efficiently fine-tune the pre-trained diffusion model while freezing the backbone model parameters, significantly reducing training resource consumption. Simultaneously, by combining adversarial and contrastive losses, it effectively improves the structural consistency and semantic fidelity of the generated images. Through efficient transfer of the pre-trained base model, this invention achieves single-step SAR image colorization generation that balances high quality and high efficiency.

[0073] This embodiment also includes an explanation of the simulation experiment setup and results, as follows: 1. Experimental conditions: The experimental data used were SAR and optical images from the publicly available dataset SAR2OPT. To facilitate network training, the original 600×600 pixel images were divided into 256×256 pixel images using a grid. 3200 pairs of optical-SAR images were randomly selected for training, and 800 pairs were used for testing. During training, the order of the optical-SAR images was randomly shuffled to achieve unpaired image training conditions. Some experimental data are shown below. Figure 5 As shown.

[0074] The experimental hardware platform used an RTX 4090 GPU, and the software environment was based on PyTorch 2.4.0 and the CUDA 12.1 framework. During training, the batch size was 1, and the total number of iterations was [number missing]. The optimizer used was the AdamW optimizer with a learning rate set to 0.00001, and the optimizer parameters were... , The weight decay is set to 0.01. The low-rank dimension of the generator is... The feature extractor selects the output feature layers of the DINO model, specifically layers 3, 6, 9, and 12. Comparison of temperature control coefficients for loss .

[0075] To verify the effectiveness of the generated fake optical images in this invention, image quality evaluation metrics LPIPS (Learned Perceptual Image Patch Similarity), FID (Fréchet Inception Distance), KID (Kernel Inception Distance), and DINO-Struct were used to quantitatively measure the quality of the generated images. LPIPS assesses the visual similarity between the generated and real images by measuring the perceptual differences between image patches in the feature space of a deep network. FID calculates the Fréchet distance between the generated and real data distributions based on the mean and covariance of the Inception feature distribution, thus measuring the overall generation quality. KID uses a kernel method to compare the differences between the generated and real sample distributions in the feature space. DINO-Struct evaluates the consistency and fidelity of the generated image with the original input image in terms of semantic structure and spatial layout based on the structural features extracted by the self-supervised model DINO. The smaller the values ​​of the four metrics, the higher the quality of the generated image.

[0076] 2. Experimental Results: 2.1 Quantitative Results: Based on the above dataset, experiments were conducted on optical-SAR image conversion using different image conversion methods, and the results are shown in Table 1. It can be seen that the proposed method achieves the best results across all metrics. It not only enables efficient single-step inference but also generates images with smaller representational differences from real optical images, achieving better results in terms of structural detail preservation and style transfer.

[0077] Table 1. Comparison of results from different image conversion methods.

[0078] 2.2 Qualitative Results: like Figure 6 As shown in the visualization comparison experiments under different terrain scenes, the mainstream models exhibit significant performance differences when handling cross-modal conversion tasks. For high-frequency detail areas such as urban buildings and industrial zones, while early methods like CycleGAN and CRAN achieved initial style transfer, they had significant shortcomings in structure preservation, with generated building outlines often accompanied by severe geometric distortion and texture disorder. Furthermore, when processing low-contrast scenes such as large areas of water and coastlines, FG-GAN and SCC-GAN are prone to uneven color distribution and artifact accumulation, failing to accurately restore the tonal consistency of optical images. In contrast, the proposed method demonstrates better stability across multiple dimensions, accurately extracting scattering features from SAR images while maintaining terrain structure consistency, achieving better SAR image colorization generation results.

[0079] Example 2 This invention also provides an efficient single-step diffusion SAR image colorization generation system, comprising: A dataset construction module is used to construct a dataset, which includes multiple SAR images and multiple visible optical remote sensing images; A generator building module is used to build a generator model based on a preset latent space diffusion model. The generator model also includes multiple LoRA layers. One of the LoRA layers is used to perform low-rank adaptation on the input features of a convolutional layer or a fully connected layer in the original module of the latent space diffusion model. The original module includes an encoder, a U-net network, and a decoder. A contrast feature extractor construction module is used to construct a contrast feature extractor model based on a preset DINO model. The contrast feature extractor model includes multiple adapter modules. One of the adapter modules is used to perform semantic space mapping on the output features of a nested tensor module of the DINO model. The model training module is used to jointly train the generator model, the contrastive feature extractor model, and the preset discriminator model based on the dataset and the preset training framework to obtain the fully trained generator model. During the joint training process, the original modules in the latent space diffusion model and the original parameters of the DINO model are frozen. The first processing module is used to input the SAR image to be colorized into the fully trained generator model for processing, so as to obtain a colorized SAR image.

[0080] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0081] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0082] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0083] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A highly efficient method for colorizing single-step diffusion SAR images, characterized in that, include: Construct a dataset that includes multiple SAR images and multiple visible optical remote sensing images; A generator model is constructed based on a preset latent space diffusion model. The generator model includes multiple LoRA layers. One of the LoRA layers is used to perform low-rank adaptation on the input features of a convolutional layer or a fully connected layer in the original module of the latent space diffusion model. The original module includes an encoder, a U-net network, and a decoder. A contrastive feature extractor model is constructed based on the preset DINO model. The contrastive feature extractor model includes multiple adapter modules. One of the adapter modules is used to perform semantic space mapping on the output features of the nested tensor module of the DINO model. Based on the dataset and the preset training framework, the generator model, the contrastive feature extractor model and the preset discriminator model are jointly trained to obtain the fully trained generator model. During the joint training process, the original modules in the latent space diffusion model and the original parameters of the DINO model are frozen. The SAR image to be colorized is input into the fully trained generator model for processing to obtain a colorized SAR image.

2. The efficient single-step diffusion SAR image colorization generation method according to claim 1, characterized in that, The construction of the generator model includes: The LoRA layer is embedded in each convolutional layer and fully connected layer in the original module. The LoRA layer performs low-rank adaptation on the input features of the first network layer it is connected to, and generates the output features of the first network layer based on the result of the low-rank adaptation. The first network layer is a convolutional layer or a fully connected layer in the original module. The LoRA layer consists of a preset lower projection matrix and an upper projection matrix. The dimension of the LoRA layer of the encoder and the decoder is a preset first dimension, and the dimension of the LoRA layer of the U-net network is a preset second dimension.

3. The efficient single-step diffusion SAR image colorization generation method according to claim 2, characterized in that, The generation of the output features of the first network layer based on the low-rank adaptation result can be expressed as: ; in, The output features of the first network layer, These are the input features of the first network layer. The weight matrix is ​​a matrix. The lower projection matrix is ​​used to project the input features. Mapping to a lower-dimensional space, matrix The up-projection matrix is ​​used to map low-dimensional features back to the original feature dimensions.

4. The efficient single-step diffusion SAR image colorization generation method according to claim 1, characterized in that, The construction of the generator model includes: In the latent space diffusion model, multiple residual connection layers are set. The residual connection layers are used to perform convolution processing on the output feature map of the encoder through a preset zero convolution layer, so that the spatial size of the output feature map matches the feature map of the corresponding layer of the decoder.

5. The efficient single-step diffusion SAR image colorization generation method according to claim 4, characterized in that, The provision of multiple residual connection layers in the latent space diffusion model includes: A total of four zero-convolutional layers are configured; The input of the first zero convolutional layer is the output of the third downsampling encoding module of the encoder. The output features are added to the output of the first decoder convolutional layer of the decoder and then used as the input of the first upsampling decoding module of the decoder. The input of the second zero convolutional layer is the output of the second downsampling encoding module of the encoder. The output features are added to the output of the first upsampling decoding module and then used as the input of the second upsampling decoding module of the decoder. The input of the third zero convolutional layer is the output of the first downsampling encoding module of the encoder. The output features are added to the output of the second upsampling decoding module and then used as the input of the third upsampling decoding module of the decoder. The input to the fourth zero convolutional layer is the output of the first encoder convolutional layer of the encoder. The output features are added to the output of the third upsampling decoding module and then used as the input to the fourth upsampling decoding module of the decoder.

6. The efficient single-step diffusion SAR image colorization generation method according to claim 1, characterized in that, The contrast feature extractor model includes a first adapter module, a second adapter module, a third adapter module, and a fourth adapter module; The DINO model includes a first DINO convolutional layer and twelve stacked nested tensor modules; The first adapter module, the second adapter module, the third adapter module, and the fourth adapter module perform semantic space mapping on the output features of the third nested tensor module, the sixth nested tensor module, the ninth nested tensor module, and the twelfth nested tensor module of the DINO model, respectively, to obtain the first semantic feature, the second semantic feature, the third semantic feature, and the fourth semantic feature; The output features of the contrast feature extractor model are generated based on the first semantic feature, the second semantic feature, the third semantic feature, and the fourth semantic feature.

7. The efficient single-step diffusion SAR image colorization generation method according to claim 1, characterized in that, The joint training of the generator model, the contrastive feature extractor model, and the preset discriminator model based on the dataset and the preset training framework includes: Normalize the images in the dataset; A training sample pair is formed based on a SAR image and a visible optical remote sensing image from the dataset. The model parameters of the generator model, the contrastive feature extractor model, and the discriminator model are updated based on a set of training sample pairs. Multiple sets of training sample pairs are generated, and the model parameters of the generator model, the contrast feature extractor model, and the discriminator model are updated based on the multiple sets of training sample pairs until a preset number of training iterations are reached, thereby obtaining the fully trained generator model.

8. The efficient single-step diffusion SAR image colorization generation method according to claim 1, characterized in that, The step of updating the model parameters of the generator model, the contrastive feature extractor model, and the discriminator model based on a set of training sample pairs includes: The original SAR image is input into the generator model to generate a fake optical image, wherein the original SAR image is the SAR image in the training sample pair; The fake optical image is input into the discriminator model, the adversarial loss of the generator model is calculated to obtain the first adversarial loss, and the parameter gradient of the generator model is calculated based on the first adversarial loss to obtain the first parameter gradient. The model parameters of the generator model are updated according to the first parameter gradient. The original optical image is input into the discriminator model, the adversarial loss of the discriminator model is calculated to obtain the second adversarial loss, and the parameter gradient of the discriminator model is calculated based on the second adversarial loss to obtain the second parameter gradient. The model parameters of the discriminator model are updated according to the second parameter gradient. The original optical image is the visible optical remote sensing image in the training sample pair. The original optical image is input into the generator model to generate an identity mapping image, and the Manhattan distance between the identity mapping image and the original optical image is calculated, and the Manhattan distance is used as the identity mapping loss. The parameter gradient of the generator model is calculated based on the identity mapping loss to obtain the third parameter gradient, and the model parameters of the generator model are updated based on the third parameter gradient. The contrast loss is calculated based on the original SAR image and the fake optical image. The contrast loss is then used to calculate the parameter gradients of the generator model and the contrast feature extractor model, respectively, to obtain the fourth parameter gradient and the fifth parameter gradient. The model parameters of the generator model and the contrast feature extractor model are then updated using the fourth parameter gradient and the fifth parameter gradient, respectively.

9. The efficient single-step diffusion SAR image colorization generation method according to claim 8, characterized in that, The calculation of contrast loss based on the original SAR image and the fake optical image includes: The original SAR image and the fake optical image are respectively input into the contrast feature extractor model to obtain the first feature map and the second feature map; Flatten the first feature map and the second feature map along the height and width directions of the feature map respectively to obtain a first feature matrix and a second feature matrix containing the same number of feature points; Feature vectors are extracted from multiple first positions of the first feature matrix and the second feature matrix respectively to obtain multiple first feature vectors and multiple second feature vectors. The first feature vectors and the second feature vectors are feature vectors extracted from the first feature matrix and the second feature matrix respectively. Each first position corresponds to one first feature vector and one second feature vector. The contrast loss is obtained by calculating the contrast loss based on multiple first feature vectors and multiple second feature vectors.

10. A high-efficiency single-step diffusion SAR image colorization generation system, characterized in that, include: A dataset construction module is used to construct a dataset, which includes multiple SAR images and multiple visible optical remote sensing images; A generator building module is used to build a generator model based on a preset latent space diffusion model. The generator model also includes multiple LoRA layers. One of the LoRA layers is used to perform low-rank adaptation on the input features of a convolutional layer or a fully connected layer in the original module of the latent space diffusion model. The original module includes an encoder, a U-net network, and a decoder. A contrast feature extractor construction module is used to construct a contrast feature extractor model based on a preset DINO model. The contrast feature extractor model includes multiple adapter modules. One of the adapter modules is used to perform semantic space mapping on the output features of a nested tensor module of the DINO model. The model training module is used to jointly train the generator model, the contrastive feature extractor model, and the preset discriminator model based on the dataset and the preset training framework to obtain the fully trained generator model. During the joint training process, the original modules in the latent space diffusion model and the original parameters of the DINO model are frozen. The first processing module is used to input the SAR image to be colorized into the fully trained generator model for processing, so as to obtain a colorized SAR image.