Synthetic aperture radar (SAR) image colorization system

By introducing a cross-attention mechanism and a diffusion model of multi-objective loss function, the mapping errors and training instability problems in SAR image colorization are solved, and more accurate and detailed optical images are generated, improving the generation efficiency and quality.

CN120472026AActive Publication Date: 2025-08-12GUANGDONG UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510543529.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The prior art has problems in the colorization of SAR images, including content mapping errors, color shifts, spatial information loss and training instability, resulting in obvious differences between the generated images and the real optical images, and the calculation complexity is high, which limits its application promotion.

Method used

A synthetic aperture radar (SAR) image colorization system based on diffusion model is used to train the noise prediction model through a cross-attention mechanism network, combining multi-objective loss function and three-step gradient denoising method to generate more accurate, rich in detail and clear structure.

Benefits of technology

It significantly improves the consistency between the generated image and the target semantics, reduces the computational complexity, and improves the inference efficiency. The generated optical images are more realistic and natural in terms of texture, edges and color transitions, and have good practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472026A_ABST
    Figure CN120472026A_ABST
Patent Text Reader

Abstract

The invention discloses a synthetic aperture radar (SAR) image colorization system, which comprises a data acquisition module, a model training module and a colorized image generation module, the data acquisition module is used for acquiring a multi-modal data set; the model training module is used for training a noise prediction model by using the multi-modal data set, and the noise prediction module is obtained by constructing a cross attention mechanism network; and the colorized image generation module is used for inputting a to-be-predicted image into the noise prediction model to obtain a color optical image. The method is suitable for disaster monitoring, landform investigation and other scenes needing high-precision SAR image analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of remote sensing images, and in particular relates to a synthetic aperture radar (SAR) image colorization system. Background Art

[0002] Given the visually unfriendly nature of SAR imagery compared to other optical imagery, SAR image colorization has become a crucial research topic for enhancing SAR image interpretation. Traditional SAR colorization methods rely on machine learning and image processing techniques. These methods require ground object classification and the construction of a classification translation knowledge base, as well as the application of feature extraction algorithms to transform features and obtain mapping relationships between various objects. These methods are costly to implement, knowledge-dependent, and difficult to generalize to complex and diverse ground scenes and detailed visual interpretation. The advancement of deep learning, particularly generative adversarial networks (GANs), has significantly boosted the development of grayscale image colorization methods. Compared to traditional feature mapping-based methods, GANs can reduce manual intervention and achieve automatic colorization through end-to-end learning. However, due to the fundamental differences in the imaging mechanisms between SAR and optical images, GAN-based colorization methods still suffer from issues such as content mapping errors, color shift, and missing spatial information. Furthermore, GAN training is prone to mode collapse and training instability, making model optimization challenging and limiting their widespread application in SAR image colorization research.

[0003] In this context, the diffusion model has gradually become a more ideal solution for SAR image colorization due to its stronger image generation capability, more stable training process and better detail restoration capability. Studies have shown that in many computer vision tasks, the image quality of the diffusion model is better than that of GAN, and it can more stably learn the mapping relationship between SAR and optical images. However, directly applying the diffusion model to SAR colorization still faces challenges. Because the backscattering characteristics of SAR images are different from the spectral reflectance characteristics of optical images, when existing diffusion models are directly applied to SAR image colorization, deviations will occur in content mapping and color mapping, resulting in significant differences between the color distribution and content consistency of the colored image and the real optical image. In addition, the traditional diffusion model uses a step-by-step denoising method for image generation, which has high computational complexity, resulting in a long inference time, limiting its practicality. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a synthetic aperture radar (SAR) image colorization system, which is suitable for scenarios requiring high-precision SAR image analysis, such as disaster monitoring and landform survey.

[0005] The present invention provides a synthetic aperture radar (SAR) image colorization system, comprising: a data acquisition module, a model training module and a colorized image generation module;

[0006] The data acquisition module is used to obtain a multimodal data set;

[0007] The model training module is used to train a noise prediction model using the multimodal dataset, wherein the noise prediction module is constructed by a cross-attention mechanism network;

[0008] The colorized image generation module is used to input the image to be predicted into the noise prediction model to obtain a color optical image.

[0009] Optionally, obtaining a multimodal dataset includes:

[0010] Acquire SAR images and optical images;

[0011] Processing the optical image to obtain a text vector and a noisy image;

[0012] The multimodal dataset is acquired based on the SAR image, the text vector, and the noisy image.

[0013] Optionally, processing the optical image to obtain a text vector includes:

[0014] extracting a preliminary text description of the optical image using a BLIP model;

[0015] fusing the preliminary text description with parameterized metadata of the remote sensing image to obtain a complete text description;

[0016] The complete text description is encoded to obtain the text vector.

[0017] Optionally, processing the optical image to obtain a noisy image includes:

[0018] Noise processing is performed on the optical image to obtain the noisy image.

[0019] Optionally, the cross-attention mechanism network is used to obtain fusion features.

[0020] Optionally, the method for obtaining the fusion feature includes:

[0021] The SAR image is used as conditional information and is spliced with the noisy image in the channel dimension to obtain joint features, and the joint features are used as the query vector Q;

[0022] The text vector is used as an initial key K and an initial value V, and a fusion feature is obtained by matching the query vector Q and the initial key K and combining the initial value V.

[0023] Optionally, constructing the noise prediction model further includes: training the noise prediction model using a multi-objective loss function, wherein the multi-objective loss function is:

[0024]

[0025] in, is a multi-objective loss function, is the image reconstruction loss, is the perceptual loss, is the semantic consistency loss, and λ1, λ2, and λ3 are loss weights.

[0026] Optionally, inputting the image to be predicted into the noise prediction model to obtain the color optical image includes:

[0027] Initialize a Gaussian noise image that matches the preset size of the SAR image and set the initial time step;

[0028] Based on the initial time step, the Gaussian noise image is denoised using a three-step gradient to obtain the color optical image.

[0029] Optionally, denoising the Gaussian noise image using a three-step gradient to obtain the color optical image includes:

[0030] Based on the initial time step, calculate the first gradient at time step t, and advance half the time step to the moment Get the second gradient and advance three-quarters of the time step to Get the third gradient;

[0031] Based on the first gradient, the second gradient, and the third gradient, updating the denoised Gaussian noise image to obtain a Gaussian noise image at time t;

[0032] Starting from the initial time step, the time step t is gradually reduced until the time step t is 0, and the color optical image is acquired.

[0033] Compared with the prior art, the present invention has the following advantages and technical effects:

[0034] The present invention can generate more accurate, detailed, and clearly structured optical images based on SAR images. By introducing semantic conditions into the training process, the consistency between the generated image and the target semantics is significantly improved, effectively avoiding the semantic bias problems that occur in traditional methods under conditions of limited image quality. For example, desert areas in some low-quality SAR images are easily misclassified as ocean. However, the present invention, by leveraging a semantic prior guidance mechanism, effectively suppresses such misclassifications and significantly enhances the semantic accuracy of the generated image.

[0035] Furthermore, based on the diffusion model, a high-performance image generation framework, this method boasts strong detail restoration capabilities, enabling the generation of more realistic and natural optical images with respect to texture, edges, and color transitions. Furthermore, a three-step gradient acceleration method is employed for feature estimation, effectively reducing computational complexity. While maintaining only a slight decrease in image quality, this method significantly improves inference efficiency, making it highly practical and promising for widespread application. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0037] Figure 1 This is a flow chart of a synthetic aperture radar (SAR) image colorization system according to an embodiment of the present invention;

[0038] Figure 2 This is a flow chart of the training process of the cross-attention mechanism network noise prediction model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0040] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0041] This embodiment proposes a synthetic aperture radar (SAR) image colorization system, including: a data acquisition module, a model training module and a colorized image generation module;

[0042] Data acquisition module, used to obtain multimodal datasets;

[0043] A model training module is used to train a noise prediction model using a multimodal dataset. The noise prediction module is constructed through a cross-attention mechanism network.

[0044] The colorized image generation module is used to input the image to be predicted into the noise prediction model to obtain a color optical image.

[0045] Specifically, the system, based on the Elucidating Diffusion Models (EDM) diffusion model, consists of two main parts: training and generation. The training phase first uses a pre-trained BLIP model based on remote sensing data to extract scene text descriptions from the optical images in the dataset. Parameters such as satellite model and polarization type are then incorporated into the text descriptions. The generated text data is used to constrain the consistency of scene object types after SAR image colorization, ensuring that the colorization results match the actual object categories. Next, the pre-trained CLIP model is used to encode the text descriptions into corresponding text vectors. The noisy optical image, time step (noise intensity), and scene information text vectors are then fed into a noise prediction network for training. To further improve colorization quality and object consistency, the noise prediction network incorporates a text-guided cross-attention mechanism to strengthen the constraints imposed by text information on the generated results. Furthermore, the system constructs a multi-objective loss function, including image reconstruction loss, color consistency loss, and semantic consistency loss. Model parameters are optimized through backpropagation to ultimately obtain the colorization model. In the generation phase of the system model, a Gaussian noise image matching the SAR image size is initialized, and an initial time step T is set. Subsequently, a fully trained SAR colorization model is used, with the SAR image serving as a guide for colorization. During the denoising process, the Bogacki-Shampine method is employed to dynamically adjust the denoising direction at different time steps t by calculating three intermediate gradients. A weighted averaging strategy is then used to update the image. This reduces the number of generation steps and accelerates the denoising process, providing more accurate denoising estimates. Ultimately, after multiple iterations, a colorization result is generated that conforms to the SAR image structure and text description.

[0046] Furthermore, obtaining a multimodal dataset includes:

[0047] Acquire SAR images and optical images;

[0048] Process the optical image to obtain text vectors and noisy images;

[0049] Based on SAR images, text vectors and noisy images, a multimodal dataset is obtained.

[0050] Specifically, the BLIP-RemoteGLM model trained on remote sensing data is used to generate text descriptions of optical image scene types. This is combined with the imaging parameters of the remote sensing image (such as the optical satellite model, SAR polarization mode, SAR imaging mode, etc.) to ensure that the colorization results maintain consistency in the object category, while enabling the system to generate images for specific satellite types. Specifically, the BLIP-RemoteGLM model first extracts a preliminary text description of the optical image, and then integrates the parameterized metadata of the remote sensing image to form a complete text description. Ultimately, the dataset is expanded from SAR images and optical images to SAR images, optical images, and their corresponding text descriptions.

[0051] Furthermore, processing the optical image to obtain the text vector includes:

[0052] Extract preliminary text description of optical images using BLIP model;

[0053] The preliminary text description is fused with the parameterized metadata of the remote sensing image to obtain a complete text description;

[0054] Encode the complete text description to obtain a text vector.

[0055] Furthermore, processing the optical image to obtain a noisy image includes:

[0056] Noise processing is performed on the optical image to obtain a noisy image.

[0057] Furthermore, the noise prediction model includes: a residual module, an upsampling module, a downsampling module and an attention module;

[0058] The residual module enhances the ability to model nonlinear features, improving training stability and model depth.

[0059] The upsampling module improves the resolution through deconvolution, interpolation, etc. for image restoration.

[0060] The downsampling module reduces the resolution through convolution or pooling to extract higher-level semantic information.

[0061] The attention module models the global dependencies within the image and improves the representation ability of image features.

[0062] The noise prediction model is a Unet-Shaped network. Similar to Unet, its main structure consists of a downsampling branch, a bottleneck, and an upsampling branch. The downsampling branch has five layers, and the upsampling branch also has five layers.

[0063] The first two layers of the downsampling branch consist of a residual module and a downsampling module, and the last three layers consist of a residual module, a self-attention module, and a downsampling module.

[0064] The bottleneck layer in the middle consists of a residual module, a cross-attention module, and a residual module. The cross-attention module is used to further interactively fuse the features of text and image, guiding the semantic consistency of the generated image.

[0065] The first three layers of the upsampling branch consist of a residual module, a self-attention module, and an upsampling module. The last two layers consist of a residual module and an upsampling module.

[0066] In addition, the input of each layer of the upsampling branch is spliced by the output of the previous layer and the output of the corresponding layer of the downsampling branch, so that the final output can better integrate the information of the input image.

[0067] Specifically, the "noise prediction network" mentioned in this embodiment is a deep network with a structure similar to U-Net, including an encoder part and a decoder part. It also includes a residual block, an upsampling layer, a downsampling layer, and an attention block. The attention block includes a self-attention network layer and a cross-attention network layer. Skip connections can fuse the shallow detailed features of the encoder with the deep semantic features of the decoder through splicing, allowing the network to better utilize contextual information.

[0068] The “text-guided cross-attention mechanism” mentioned in this embodiment is a submodule of the BottleNeck module integrated in the noise prediction network.

[0069] Furthermore, a cross-attention mechanism network is used to obtain fusion features.

[0070] Furthermore, the method for obtaining fusion features includes:

[0071] The SAR image is used as conditional information and is spliced with the noisy image in the channel dimension to obtain joint features, which are used as the query vector Q;

[0072] The text vector is used as the initial key K and initial value V. The fusion feature is obtained by matching the query vector Q with the initial key K and combining the initial value V.

[0073] Specifically, the design and implementation of a text-guided cross-attention mechanism is characterized by its interaction with Query (Q), Key (K), and Value (V), enabling more accurate matching of different regions between the SAR image and the colorized optical image, thereby improving the consistency of the scene before and after colorization. Query (Q) is influenced by the joint features of the current optical and SAR images. Key (K) incorporates text features and defines the targets to be matched between different regions. Value (V), primarily composed of text features, provides scene information from the optical image, guiding the colorization result to conform to the semantic characteristics of the target scene. Specifically, the joint features of the optical and SAR images are first obtained from the encoder output of the noise prediction network as Q. Next, the text encoding results from step 2 are used as the initial K and V. Finally, by matching regions in Q with K and combining them with the scene information provided by V, cross-attention is calculated to adjust the image generation process so that the final optical image conforms to the target scene consistency requirements.

[0074] Furthermore, constructing the noise prediction model further includes: training the noise prediction model using a multi-objective loss function, wherein the multi-objective loss function is:

[0075]

[0076] in, is a multi-objective loss function, is the image reconstruction loss, is the perceptual loss, is the semantic consistency loss, and λ1, λ2, and λ3 are loss weights.

[0077] Specifically, the step of constructing a multi-objective loss function is characterized in that the constructed loss function includes image reconstruction loss, color consistency loss and semantic consistency loss. Among them, image reconstruction uses L2 loss to measure the pixel-level difference between the generated image and the target optical image, ensure the consistency of structural details, and is used to optimize the overall structure and texture. Color consistency loss extracts features through a pre-trained VGG network to measure the feature differences between the generated image and the target optical image at different levels, constrains color consistency from the feature space, and reduces color cast. Semantic consistency loss uses CLIP or other semantic matching models to calculate the semantic similarity between the generated image and the target text description to improve the alignment between the image and the text.

[0078] Furthermore, the image to be predicted is input into the noise prediction model to obtain the color optical image, including:

[0079] Initialize a Gaussian noise image that matches the preset size of the SAR image and set the initial time step;

[0080] Based on the initial time step, the Gaussian noise image is denoised using a three-step gradient to obtain a color optical image.

[0081] Specifically, the denoising calculation step in the generation process is characterized by using the Bogacki-Shampine method to calculate three intermediate gradients and accelerating the denoising process through weighted averaging. This approach avoids local deviations associated with single-step predictions by exploring future states in stages (half-step and three-quarter-step), and can directly replace traditional solvers without the need for additional training. This simplifies the generation process and improves computational efficiency.

[0082] Furthermore, the Gaussian noise image is denoised using a three-step gradient to obtain a color optical image including:

[0083] Based on the initial time step, calculate the first gradient at time step t, and advance half the time step to the moment Get the second gradient and advance three-quarters of the time step to Get the third gradient;

[0084] Based on the first gradient, the second gradient, and the third gradient, updating the denoised Gaussian noise image to obtain a Gaussian noise image at time t;

[0085] Starting from the initial time step, the time step t is gradually reduced until the time step t is 0, and a color optical image is obtained.

[0086] Specifically, in the inference stage of the designed diffusion model, firstly, the Gaussian noise Starting from the given input time step T, denoised SAR image, and parameter text information as guidance, the Gaussian noise x is predicted based on the fully trained noise prediction model. T Denoising is performed, and the Bogacki-Shampine method is introduced to accelerate the inverse generation process of the diffusion model and reduce the number of iterations. Finally, prediction is performed using the strategy from the training phase, and this process is repeated until t reaches 0, outputting a colored image.

[0087] The present embodiment will be described in detail below with reference to the accompanying drawings:

[0088] This embodiment provides a synthetic aperture radar (SAR) image colorization system. Figure 1-2 As shown, the specific steps are as follows:

[0089] Step 1: During the training phase of the designed diffusion model, the RemoteGLM model trained based on remote sensing data is first used to generate a text description of the optical image. Parameters such as the satellite model, polarization mode, and imaging mode are then integrated into the text description. Commas are used to connect texts. The generated text description will serve as a consistency constraint for the scene and object type after the SAR image is colorized, preventing the problem of failing to generate a scene that matches the SAR image after colorization.

[0090] Specifically, read the optical image I optical , use the BLIP-RemoteGLM model fine-tuned based on remote sensing domain knowledge to understand the scene of the optical image and generate the corresponding preliminary text description Caption , that is, text Caption =BLIP(I optical ). Then, the satellite model, polarization type, imaging model and other parameter information are formed into text param , fused into the preliminary text description to improve the adaptability to SAR imaging characteristics. Finally, based on text=text Caption +text param , obtain the complete scene description text to improve the parameter information text.

[0091] Step 2: The text description is then encoded into a 512-dimensional vector using the CLIP model with frozen parameters. A small neural network is then built to map and transform the 512-dimensional text vector output by the CLIP model. By performing a linear transformation on these text features, the network adjusts them to a dimension that matches the diffusion model feature space, ensuring that the text vector can be effectively fused with image features (such as SAR images or generated optical images).

[0092] Specifically, the CLIP model with frozen parameters is composed of c Text =CLIP text (Text) Generate a 512-dimensional text vector c Text As a subsequent conditional guide, we then build a neural network with an MLP-Relu-MLP structure to enhance features and perform dimensionality transformation.

[0093] Step 3: Add noise using the time proportion noise adding strategy. The noise adding strategy of the traditional EDM model is x t =s(t)x0+σ(t)ε, where s(t) and σ(t) are functions of time step t, respectively controlling the degree of influence of the original image x0 and Gaussian noise ε. Here, s(t) = 1, σ(t) = t, and the noise addition formula is x t=x0+tε, at this time, the noise intensity increases linearly with the time step, and the time step can directly reflect the noise intensity. According to this formula, the input samples N with different noise levels are generated. optical To carry out the noise prediction model U in the next step θ training.

[0094] Step 4: The noisy optical image, time step (noise intensity), and scene information text vector are fed into the noise prediction network for training. The SAR image, serving as conditioning information, is concatenated with the noisy optical image in the channel dimension to form a joint feature. This joint feature learns the noise information of the optical image and the geometric and texture features of the SAR image. This is then used to generate the query vector (Q) in the cross-attention mechanism. Meanwhile, the scene text description is encoded and mapped to a feature space that matches the time step (noise intensity) and is used to regulate the generative dynamics during the diffusion process. The text description, after encoding and dimension transformation, is embedded into the time step (noise intensity) to regulate the generative dynamics during the diffusion process. Simultaneously, the text features are transformed to generate the key K and value V required for the cross-attention mechanism, enhancing the text guidance capability. In the cross-attention mechanism, by matching regions in Q (SAR-optical joint feature) and K (text feature), and based on the scene information provided by V, the denoising network is guided to generate an optical image that conforms to the SAR structure and text semantics. This ensures that the generated optical image matches the SAR image in visual structure and conforms to the given scene description.

[0095] Specifically, in order to improve the consistency of scene information during the colorization process of SAR images, a text-guided cross-attention mechanism is introduced into the noise prediction network, so that the model can use text information to guide the generation of colorized images. θ The input includes: the optical image N after adding noise optical SAR image I sar , text vector c Text , time step (noise intensity) t. Based on Q = W Q ·(f([I sar ,N optical ])), the query Q is calculated by stitching the SAR and noisy optical images, where f represents the feature extraction part of the denoising network. The key (K) and value (V) are calculated from the scene text K = W K c text , V=W V c text , where W Q , W K , W Vis a learnable projection matrix used to map the input features to the same feature space to calculate the attention weights. The above variables are calculated as follows after the text-guided cross-attention mechanism The final fused features contain SAR-optical information and semantic scene information. output =U θ (N optical ,I SAR ,c text ,t)Finally output the predicted image.

[0096] Step 5: Design of multi-objective loss function, constructing a joint loss function consisting of image reconstruction loss, color consistency loss and semantic consistency loss. Image reconstruction loss is Color consistency loss is composed of perceptual loss. It extracts features through the pre-trained VGG network to measure the feature differences between the generated image and the target optical image at different levels, constrains color consistency from the feature space, and reduces color cast. It is defined as Semantic consistency uses CLIP or other semantic matching models to calculate the semantic similarity between the generated image and the target text description to improve the alignment between the image and the text. The definition is as follows: Where ψ(·) is a pre-trained CLIP2 model that can map both images and text into the same feature space. Based on this, a multi-objective loss function is constructed. where λ1 = 1.0, λ2 = 0.5, and λ3 = 0.1. Finally, the model parameters are updated through backpropagation and the optimizer, making the generated image more consistent with the target optical image and text description in terms of details, color, and semantics.

[0097] Step 6: If Figure 1 As shown, the denoised SAR image I SAR With text vector text param Input the trained noise prediction network In the formula, based on the generation formula The Bogacki-Shampine method is used to accelerate the inverse process. Specifically, first, at time step T, the Gaussian noise is initialized. Then from x T Start denoising step by step, and the state of each step is recorded as x t The denoising acceleration process is realized by three-step gradient calculation. First, the gradient at time t is calculated. Then advance half a step to Predict the intermediate state and re-evaluate the gradient to get the new gradient Next, advance three-quarters of the time step to Further correct the gradient direction to get the gradient Finally, the gradient results of the three times are combined, based on Update the denoised state to get the image at the next moment, repeat this process, starting from T, gradually reduce t, and decrease it by Δt each time until t = 0, and finally get the denoised optical image x0, that is, the SAR image I SAR With text vector text param Guided color optical image I color .

[0098] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A synthetic aperture radar (SAR) image colorization system, characterized in that: include: Data acquisition module, model training module and colorized image generation module; The data acquisition module is used to obtain a multimodal data set; The model training module is used to train a noise prediction model using the multimodal dataset, wherein the noise prediction module is constructed by a cross-attention mechanism network; The colorized image generation module is used to input the image to be predicted into the noise prediction model to obtain a color optical image.

2. A synthetic aperture radar (SAR) image colorization system according to claim 1, characterized in that: Acquiring multimodal datasets includes: Acquire SAR images and optical images; Processing the optical image to obtain a text vector and a noisy image; The multimodal dataset is acquired based on the SAR image, the text vector, and the noisy image.

3. The synthetic aperture radar (SAR) image colorization system according to claim 2, characterized in that: Processing the optical image to obtain a text vector includes: extracting a preliminary text description of the optical image using a BLIP model; fusing the preliminary text description with parameterized metadata of the remote sensing image to obtain a complete text description; The complete text description is encoded to obtain the text vector.

4. The synthetic aperture radar (SAR) image colorization system according to claim 2, characterized in that: Processing the optical image to obtain a noisy image includes: Noise processing is performed on the optical image to obtain the noisy image.

5. The synthetic aperture radar (SAR) image colorization system according to claim 2, characterized in that: The cross attention mechanism network is used to obtain fusion features.

6. A synthetic aperture radar (SAR) image colorization system according to claim 5, characterized in that: Acquiring the fusion feature includes: The SAR image is used as conditional information and is spliced with the noisy image in the channel dimension to obtain joint features, and the joint features are used as the query vector Q; The text vector is used as an initial key K and an initial value V, and a fusion feature is obtained by matching the query vector Q and the initial key K and combining the initial value V.

7. The synthetic aperture radar (SAR) image colorization system according to claim 1, characterized in that: Constructing the noise prediction model further includes: training the noise prediction model using a multi-objective loss function, wherein the multi-objective loss function is: in, is a multi-objective loss function, is the image reconstruction loss, is the perceptual loss, is the semantic consistency loss, and λ1, λ2, and λ3 are loss weights.

8. The synthetic aperture radar (SAR) image colorization system according to claim 1, characterized in that: Inputting the image to be predicted into the noise prediction model to obtain a color optical image includes: Initialize a Gaussian noise image with the same size as the preset SAR image and set the initial time step; Based on the initial time step, the Gaussian noise image is denoised using a three-step gradient to obtain the color optical image.

9. The synthetic aperture radar (SAR) image colorization system according to claim 8, characterized in that: Denoising the Gaussian noise image using a three-step gradient to obtain the color optical image includes: Based on the initial time step, calculate the first gradient at time step t, and advance half the time step to the moment Get the second gradient and advance three-quarters of the time step to Get the third gradient; Based on the first gradient, the second gradient, and the third gradient, updating the denoised Gaussian noise image to obtain a Gaussian noise image at time t; Starting from the initial time step, the time step t is gradually reduced until the time step t is 0, and the color optical image is acquired.

Citation Information

Patent Citations

  • Image coding and decoding method, system, equipment and medium

    CN114882133A

  • SAR-optical image fusion method based on hybrid model

    CN117314811A

  • Method for generating multi-category special vehicle SAR image based on denoising diffusion model

    CN117541906A

  • SAR image generation method based on de-noising diffusion probability model

    CN118230191A

  • Typical ground object target segmentation method based on multi-source data attention feature fusion

    CN119131374A