Method and apparatus for generating optical images

By processing dual-temporal datasets using a conditional generative adversarial network model, features from radar and optical images are extracted and fused to generate high-quality optical images. This solves the problem of poor generation results in complex scenes and enables the generation of high-resolution optical images in complex scenarios.

CN116342902BActive Publication Date: 2026-01-23AEROSPACE INFORMATION RES INST CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310333411.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2026-01-23
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively fuse synthetic aperture radar (SAR) images and multispectral images to generate high-quality optical images in complex scenarios, especially in suburban and urban areas where there are many small-scale features and diverse terrain features, resulting in poor image generation.

Method used

A conditional generative adversarial network model is used to process the dual-temporal dataset. Change features are extracted and fused through temporal mutual attention. Optical images are generated using a multi-scale generator and discriminator. The model is optimized by combining a weight balancing module and an optimizer to generate target temporal optical images.

Benefits of technology

It significantly narrows the quality gap between optical images generated in simple and complex scenes, and the generated images can maintain good results even in complex scenes, providing support for continuous high-resolution optical remote sensing observation data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342902B_ABST
    Figure CN116342902B_ABST
Patent Text Reader

Abstract

The application provides a method and device for generating optical images, and relates to the technical field of optical images. The method for generating optical images comprises: processing synthetic aperture radar image data and multispectral image data to obtain a dual-time-phase data set; processing the dual-time-phase data set to generate optical images; wherein processing the dual-time-phase data set comprises: extracting a change feature of the dual-time-phase data set and fusing the change feature to obtain a fused feature; extracting a multi-scale feature of the fused feature to obtain a simulated target time-phase optical image; and generating a target time-phase optical image according to a discrimination result of the simulated target time-phase optical image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical imaging technology, and more specifically, to a method for generating optical images. Background Technology

[0002] Stable and continuous temporal optical images have great value in remote sensing applications such as target identification, disaster monitoring and crop monitoring. However, many regions are affected by clouds and rain, making it difficult to construct stable time series of optical images.

[0003] Currently, constructing continuous and stable temporal optical remote sensing data mainly relies on increasing the frequency of satellite observations and utilizing the complementarity of high and low spatial resolution optical data in terms of temporal and spatial resolution to compensate for missing data. However, these two methods still cannot solve the problem of insufficient optical remote sensing data in areas with continuous cloud and rain. With the development of SAR (Synthetic Aperture Radar) technology, methods have emerged to fuse SAR images and optical images to generate simulated optical images. The advantages of this method are obvious: SAR images are unaffected by clouds and rain, while visible light images have richer spectral and textural details than SAR images. However, due to the different sensing mechanisms of SAR and optical images, the characteristics of the two types of data are very different, making it difficult for traditional fusion methods to fuse the two types of data well. Although in recent years, deep learning models based on Conditional Generative Adversarial Networks (CGAN) have made significant progress in the generation of SAR to optical remote sensing images, the quality of the generated images in complex scenes is significantly different from that in simple scenes. For example, in suburban and urban scenes, there are many types of ground features, many small-scale features, and complex image transformation relationships, resulting in a generation effect that is far inferior to that in simple scenes such as farmland and forests. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] To address the aforementioned problems, this invention provides a method for generating optical images. It utilizes a conditional generative adversarial network model to process optical images generated from a dual-temporal dataset formed by synthetic aperture radar image data and multispectral image data, thus solving the technical problem of significant quality differences in optical images generated under simple and complex scenarios.

[0006] (II) Technical Solution

[0007] The first aspect of this invention provides a method for generating optical images, comprising: processing synthetic aperture radar image data and multispectral image data to obtain a dual-temporal dataset; processing the dual-temporal dataset to generate an optical image; wherein, processing the dual-temporal dataset includes: extracting and fusing the variation features of the dual-temporal dataset to obtain fused features; extracting multi-scale features of the fused features to obtain a simulated target temporal optical image; and generating a target temporal optical image based on the discrimination result of the simulated target temporal optical image.

[0008] In one embodiment of the present invention, a conditional generative adversarial network (GAN) model is used to process a dual-temporal dataset to generate an optical image. The GAN model includes a temporal mutual attention head, a multi-scale generator, and a multi-scale discriminator. The temporal mutual attention head is used to extract and fuse the variation features of the dual-temporal dataset to obtain fused features. The multi-scale generator is used to extract the multi-scale features of the fused features to obtain a simulated target temporal optical image. The multi-scale discriminator is used to generate a target temporal optical image based on the discrimination result of the simulated target temporal optical image.

[0009] In one embodiment of the present invention, the variation features of the dual-temporal dataset are extracted using a conditional generative adversarial network model and fused to obtain fused features. This includes: extracting variation features of the dual-temporal dataset using a conditional generative adversarial network model, fusing the variation features with source temporal optical image features to obtain pre-fused features; and fusing the pre-fused features with target temporal synthetic aperture radar image features to obtain fused features.

[0010] In one embodiment of the present invention, a conditional generative adversarial network model is used to extract and fuse the variation features of a dual-temporal dataset. The calculation method for the fused features is as follows:

[0011]

[0012] O = Norm(γX + MV)

[0013] S = O + Y

[0014] Where M represents variation features, O represents pre-fused features, S represents fused features, γ represents trainable parameters, X represents source temporal optical image features, Y represents target temporal synthetic aperture radar image features, Q represents query tensor, K represents encoding tensor, and V represents value tensor.

[0015] In one embodiment of the present invention, the multi-scale generator includes a first generator and a second generator; wherein, the first generator is used to extract large-scale features; the second generator is used to extract small-scale features; and the fused features input to the first generator are twice the fused features input to the second generator.

[0016] In one embodiment of the present invention, the conditional generative adversarial network model further includes a weight balancing module, wherein the weights of the source temporal synthetic aperture radar (SAR) image features, the target temporal SAR image features, and the source temporal optical image features are balanced by the weight balancing module; the weights of the source temporal SAR image features and the target temporal SAR image features are twice the weights of the source temporal optical image features.

[0017] In one embodiment of the present invention, the weight balancing module is calculated as follows:

[0018]

[0019] Where, m t σ represents the mean value of the target after linear transformation of synthetic aperture radar image data and multispectral image data. t This represents the standard deviation of synthetic aperture radar image data and multispectral image data after linear transformation.

[0020] In one embodiment of the present invention, the conditional generative adversarial network model further includes an optimizer, wherein the optimizer optimizes the conditional generative adversarial network model by utilizing adversarial loss and perceptual loss; the optimization order is: first optimize the multi-scale discriminator, and then optimize the multi-scale generator.

[0021] In one embodiment of the present invention, processing synthetic aperture radar (SAR) image data and multispectral image data to obtain a dual-temporal dataset includes: acquiring polarimetric SAR (PSAR) images and multispectral images based on Sentinel satellites; performing sub-strip separation, orbital correction, and radiometric correction on the PSAR images to obtain source temporal SAR image data and target temporal SAR image data; performing band normalization on the multispectral images to obtain source temporal optical image data; and performing sample differentiation on the source temporal SAR image data, target temporal SAR image data, and source temporal optical image data to obtain a simple scene dataset and a complex scene dataset.

[0022] A second aspect of this invention provides an optical image generation apparatus, comprising: a first processing module for processing synthetic aperture radar image data and multispectral image data to obtain a dual-temporal dataset; a second processing module for processing the dual-temporal dataset to generate an optical image; wherein the second processing module includes: a first extraction module for extracting and fusing variation features of the dual-temporal dataset to obtain fused features; a second extraction module for extracting multi-scale features of the fused features to obtain a simulated target temporal optical image; and a generation module for obtaining a target temporal optical image based on the discrimination result of the simulated target temporal optical image.

[0023] (III) Beneficial Effects

[0024] The method for generating optical images provided in this embodiment of the invention has at least the following beneficial effects:

[0025] (1) The optical image generation method provided in this embodiment of the invention uses a conditional generative adversarial network model to process the dual-temporal dataset formed by synthetic aperture radar image data and multispectral image data to generate optical images, which can significantly reduce the quality gap between optical images generated in simple and complex scenes.

[0026] (2) The optical image generation method provided in this embodiment of the invention fully utilizes the optical image features of the source phase and the synthetic aperture radar image features of the dual phase through a weight balancing strategy. This ensures that the generated target phase optical image has both real optical image texture and the sensitivity of the optical image during the change process from the source phase to the target phase. Even in complex scenarios, it can maintain good generation results and provide data support for various applications that require continuous high-resolution optical remote sensing observation data. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0028] Figure 1 The diagram illustrates a flowchart of a first embodiment of the optical image generation method provided by the present invention.

[0029] Figure 2 The diagram illustrates a second embodiment of the optical image generation method provided by the present invention.

[0030] Figure 3 The diagram illustrates the structure of the conditional generative adversarial network model provided in an embodiment of the present invention.

[0031] Figure 4 The diagram illustrates the structure of the temporal mutual attention head in the conditional generative adversarial network model provided in an embodiment of the present invention.

[0032] Figure 5 The diagram illustrates the structure of a multi-scale generator in a conditional generative adversarial network model provided in an embodiment of the present invention.

[0033] Figure 6The schematic diagram illustrates the internal structure of the first or second generator in the conditional generative adversarial network model provided in the embodiments of the present invention.

[0034] Figure 7 The diagram illustrates the structure of the multi-scale discriminator in the conditional generative adversarial network model provided in an embodiment of the present invention.

[0035] Figure 8 The illustration shows the effect generated by the optical image generation method provided in the embodiments of the present invention.

[0036] Figure 9 The illustration shows a comparison of the generation results of the optical image generation method provided in the embodiments of the present invention in a simple scene.

[0037] Figure 10 The illustration shows a comparison of the generation results of the optical image generation method provided in the embodiments of the present invention in complex scenes.

[0038] [Attached image labels]

[0039] 1-Time mutual attention;

[0040] 2-Multi-scale generator; 21-First generator; 22-Second generator;

[0041] 3- Multi-scale discriminator. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. The terminology used herein is merely for describing specific embodiments and is not intended to limit the invention. The terms "comprising," "including," etc., used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0043] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0044] In the description of this invention, it should be understood that the terms "longitudinal", "length", "circumferential", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the subsystem or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0045] Throughout the accompanying drawings, identical elements are represented by the same or similar reference numerals. Conventional structures or configurations may be omitted where they might cause confusion in understanding the invention. Furthermore, the shapes, sizes, and positional relationships of the components in the drawings do not reflect actual size, scale, or actual positional relationships. Additionally, any reference numerals placed between parentheses in the claims should not be construed as limiting the claims.

[0046] Similarly, to simplify the invention and aid in understanding one or more aspects of the invention, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. The use of terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples" indicates that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0047] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0048] Figure 1 The diagram illustrates a flowchart of a first embodiment of the optical image generation method provided by the present invention.

[0049] like Figure 1 As shown, the optical image generation method provided in the first embodiment of the present invention may include:

[0050] S1 processes synthetic aperture radar image data and multispectral image data to obtain a dual-temporal dataset.

[0051] The dual-temporal dataset obtained by processing synthetic aperture radar image data and multispectral image data includes:

[0052] S100 acquires polarimetric synthetic aperture radar and multispectral images based on Sentinel satellites.

[0053] S101 performs sub-strip separation processing, fine orbit correction processing, and radiation correction processing on the polarimetric synthetic aperture radar (PSAR) image to obtain source temporal PAR image data and target temporal PAR image data.

[0054] S102, band normalization processing is performed on the multispectral image to obtain the source temporal optical image data of the multispectral image.

[0055] S103, sample differentiation is performed on source temporal synthetic aperture radar image data, target temporal synthetic aperture radar image data and source temporal optical image data to obtain simple scene datasets and complex scene datasets.

[0056] Processing the two-temporal dataset to generate optical images, wherein processing the two-temporal dataset includes:

[0057] S2, extract the variation features of the dual-temporal dataset and fuse them to obtain the fused features.

[0058] S3 extracts multi-scale features of the fusion features to obtain a simulated temporal optical image of the target.

[0059] S4. Generate a target temporal optical image based on the discrimination results of the simulated target temporal optical image.

[0060] The optical image generation method provided in this embodiment first forms a dual-temporal dataset based on synthetic aperture radar image data and multispectral image data, and then performs feature extraction on the dual-temporal dataset to generate optical images, which greatly reduces the quality gap between optical images generated in simple and complex scenes.

[0061] Figure 2 The diagram illustrates a second embodiment of the optical image generation method provided by the present invention.

[0062] like Figure 2 As shown, the optical image generation method provided in the second embodiment of the present invention may include:

[0063] S1 processes synthetic aperture radar image data and multispectral image data to obtain a dual-temporal dataset.

[0064] The dual-temporal dataset obtained by processing synthetic aperture radar image data and multispectral image data includes:

[0065] S100 acquires polarimetric synthetic aperture radar imagery based on Sentinel Satellite 1 and multispectral imagery based on Sentinel Satellite 2.

[0066] S101, the polarimetric synthetic aperture radar (PSAR) image is sequentially processed by sub-strip separation, fine orbit correction, and radiation correction to obtain the backscattering coefficient. Then, the PSAR image is further processed by sub-pulse merging, sub-merging, multi-view processing, speckle noise filtering, and terrain correction to obtain the source temporal phase (T0) PSAR image data and the target temporal phase (T1) PSAR image data.

[0067] S102, band normalization processing is performed on the multispectral image to obtain the source temporal phase (T0) optical image data of the multispectral image.

[0068] S103, sample differentiation is performed on source temporal synthetic aperture radar image data, target temporal synthetic aperture radar image data and source temporal optical image data to obtain simple scene datasets and complex scene datasets.

[0069] In this embodiment, the selected data are the red, green, and blue bands of multispectral imagery and the VH and VV polarization bands of synthetic aperture radar imagery. Because the original images are too large, the low-resolution and high-resolution images are cropped into matching high-resolution-low-resolution data pairs according to latitude and longitude.

[0070] Optical images are generated by processing a dual-temporal dataset using a conditional generative adversarial network model. This conditional generative adversarial network model includes a temporal mutual attention head 1, a multi-scale generator 2, and a multi-scale discriminator 3.

[0071] The use of conditional generative adversarial network models to process bi-temporal datasets includes:

[0072] S2, using the temporal mutual attention head 1 to extract the change features of the dual temporal dataset and fuse them to obtain the fused features (Att.Map).

[0073] Specifically, the temporal mutual attention head 1 is used to extract and fuse the variation features of the dual-temporal dataset, resulting in the following fused features:

[0074] S200 uses the time mutual attention head 1 to extract the change features of the dual temporal dataset, and fuses the change features with the source temporal optical image features to obtain the pre-fused features.

[0075] S201 fuses the pre-fused features and the target temporal synthetic aperture radar image features to obtain the fused features.

[0076] The calculation method for this fusion feature is as follows:

[0077]

[0078] O = Norm(θX + MV)

[0079] S = O + Y

[0080] Where M represents variation features, O represents pre-fused features, S represents fused features, γ represents trainable parameters, X represents source temporal optical image features, Y represents target temporal synthetic aperture radar image features, Q represents query tensor, K represents encoding tensor, and V represents value tensor.

[0081] S3 extracts multi-scale features of the fusion features to obtain a simulated temporal optical image of the target.

[0082] The multiscale generator 2 includes a first generator 21 and a second generator 22.

[0083] Among them, the multi-scale features extracted from the fusion features to obtain the temporal optical image of the simulated target include:

[0084] S300 extracts large-scale features through the first generator 21.

[0085] S301, small-scale features are extracted through the second generator 22.

[0086] The fusion features input to the first generator 21 are twice the fusion features input to the second generator 22.

[0087] S4. Generate a target temporal optical image based on the discrimination results of the simulated target temporal optical image.

[0088] The generation of target temporal optical images, based on the discrimination results of simulated target temporal optical images, includes:

[0089] The discrimination results of the simulated target temporal optical image are obtained by the multi-scale discriminator. When the discrimination result of the simulated target temporal optical image is equal to the discrimination result of the target temporal optical image, the simulated target temporal optical image is output.

[0090] The optical image generation method provided in this invention utilizes a conditional generative adversarial network model to process a dual-temporal dataset formed from synthetic aperture radar image data and multispectral image data to generate optical images, which can significantly reduce the quality gap between optical images generated in simple and complex scenes.

[0091] Based on the above embodiments, the conditional generative adversarial network model further includes a weight balancing module, wherein the weight balancing module balances the weights of the source temporal synthetic aperture radar image features, the target temporal synthetic aperture radar image features, and the source temporal optical image features; the weights of the source temporal synthetic aperture radar image features and the target temporal synthetic aperture radar image features are twice the weight of the source temporal optical image features.

[0092] The calculation method for the weight balancing module is as follows:

[0093]

[0094] Where, m t σ represents the mean value of the target after linear transformation of synthetic aperture radar image data and multispectral image data. t This represents the standard deviation of synthetic aperture radar image data and multispectral image data after linear transformation.

[0095] As a preferred embodiment, in this embodiment, for the optical image band, m t and σ t For example, they can be set to 0 and 0.5 respectively. For synthetic aperture radar image bands, m t and σ t For example, they can be set to 0 and 1 respectively, so as to give the SAR feature prior a weight twice that of the optical feature, and avoid the model from being insensitive to the changes between the two time phases due to over-reliance on the optical feature.

[0096] Based on the above embodiments, the conditional generative adversarial network model also includes an optimizer. This optimizer optimizes the conditional generative adversarial network model using adversarial loss and perceptual loss. In each training iteration, the multi-scale discriminator 3 is optimized first, followed by the multi-scale generator 2, to ensure that the discriminative ability of the multi-scale discriminator 3 is stronger than the generative ability of the multi-scale generator 2.

[0097] The total loss function of the multi-scale generator 2 is:

[0098]

[0099] The total loss function of the multi-scale discriminator 3 is:

[0100]

[0101] Where G1 represents the first generator 21, G2 represents the second generator 22, D represents the multi-scale discriminator 3, and X t S represents the temporal optical image of the target, S represents the fusion feature, λ represents the parameter controlling the weights of the sensing loss, and F represents the target temporal optical image. (i) M represents layers 1 to i of a VGG19 network pre-trained on the ImageNet dataset (a large visualization database for research on visual object recognition software). i N is the size of the feature map of the i-th layer of the VGG19 (Oxford University Visual Geometry Group 19-layer convolutional network model) network, and N is the number of layers of the VGG19 network.

[0102] Figure 3 The diagram illustrates the structure of the conditional generative adversarial network model provided in an embodiment of the present invention.

[0103] like Figure 3 As shown, the conditional generative adversarial network model provided in this embodiment of the invention may include: a temporal mutual attention head 1, a multi-scale generator 2, and a multi-scale discriminator 3.

[0104] Among them, the temporal mutual attention head 1 is used to extract the change features of the dual temporal dataset and fuse them to obtain fused features.

[0105] Multiscale generator 2 is used to extract multiscale features of the fusion features to obtain a simulated target temporal optical image.

[0106] The multi-scale discriminator 3 is used to generate the discrimination result and generate an optimized generator based on the discrimination result.

[0107] Figure 4 The diagram illustrates the structure of the temporal mutual attention head in the conditional generative adversarial network model provided in an embodiment of the present invention.

[0108] like Figure 4 As shown, the temporal mutual attention head 1 of the optical image generation device provided in this embodiment of the invention may include three input terminals, a normalization layer (Norm), and an output terminal.

[0109] Among them, the source temporal synthetic aperture radar image data, the target temporal synthetic aperture radar image data, and the source temporal optical image data are input through the input end of the temporal mutual attention head 1. The query tensor (Q), the encoding tensor (K), and the value tensor (V) are extracted and fused into a change feature (Att.Map). The change feature is further fused into a pre-fusion feature, and the pre-fusion feature is then fused into a fusion feature, which is finally output by the output end of the temporal mutual attention head 1.

[0110] Figure 5The diagram illustrates the structure of a multi-scale generator in a conditional generative adversarial network model provided in an embodiment of the present invention.

[0111] like Figure 5 As shown, the multi-scale generator 2 of the optical image generation apparatus provided in this embodiment of the invention may include a first generator 21 and a second generator 22.

[0112] The first generator 21 and the second generator 22 are respectively composed of three parts: a convolution block, a residual block, and a transpose convolution block.

[0113] Figure 6 The schematic diagram illustrates the internal structure of the first or second generator in the conditional generative adversarial network model provided in the embodiments of the present invention.

[0114] like Figure 6 As shown, the convolutional block is composed of a series of convolutional layers, the residual block is composed of a series of residual convolutional layers, and the transpose convolutional block is composed of a series of transpose convolutional layers. The output of the first generator 21 and the residual block of the second generator 22 are spliced ​​together.

[0115] Figure 7 The diagram illustrates the structure of the multi-scale discriminator in the conditional generative adversarial network model provided in an embodiment of the present invention.

[0116] like Figure 7 As shown, the multi-scale discriminator 3 of the optical image generation device provided in this embodiment of the invention consists of three discriminators with the same structure, having two input terminals, five convolutional layers and one output terminal. Similar to the multi-scale generator 2, each discriminator has a different input scale. Here, real represents the real target temporal optical image, fake represents the generated simulated target temporal optical image, condition represents the generator input corresponding to the real target temporal optical image, Concat represents channel stitching, and Norm represents the normalization layer.

[0117] In this embodiment, the process of generating optical images using a conditional generative adversarial network (GAN) model to process a dual-temporal dataset (i.e., the GAN model training process) includes:

[0118] First, the source temporal synthetic aperture radar image data, target temporal synthetic aperture radar image data, and source temporal optical image data are sampled to obtain a simple scene dataset and a complex scene dataset. Then, the simple scene dataset and the complex scene dataset are each divided into 90% training data and 10% validation data. The simple scene dataset is trained first, and then the complex scene dataset is trained.

[0119] During the training phase, samples are input into the conditional generative adversarial network model in mini-batch mode. Each batch consists of two samples, and each sample contains three source temporal optical image data bands, three target temporal optical image data bands, two source temporal synthetic aperture radar image data bands, and two target temporal synthetic aperture radar image data bands.

[0120] For each training batch, multi-scale discriminator 3 is trained first, followed by multi-scale generator 2. The multi-scale generator is fixed during the training of multi-scale discriminator 3. Source temporal optical image data, source temporal synthetic aperture radar (SAR) image data, and target temporal SAR image data are input into temporal mutual attention head 1 and multi-scale generator 2 to generate simulated target temporal optical images. This simulated target temporal optical image, source temporal optical image data, source temporal SAR image data, and target temporal SAR image data are then input into multi-scale discriminator 3 to obtain a discrimination result. This result is then input into the loss function to obtain the loss L1 for training multi-scale discriminator 3 in this batch. Then, the target temporal optical image, source temporal optical image data, source temporal synthetic aperture radar image data, and target temporal synthetic aperture radar image data are input into the multi-scale discriminator 3 to obtain another discrimination result. This result is then input into the loss function to obtain the loss L2 for this batch of multi-scale discriminator 3 training, and finally the total loss L1+L2 of the multi-scale discriminator 3.

[0121] At this point, the multi-scale discriminator 3 is adjustable, while the multi-scale generator 2 is fixed. During this process, the gradients of each parameter in the multi-scale discriminator 3 are recorded, and then the gradient is backpropagated using the optimizer (Adam), thus automatically adjusting the parameters in the multi-scale discriminator 3.

[0122] Next, the multi-scale generator 2 is trained, while the multi-scale discriminator 3 is fixed during training. Source temporal optical image data, source temporal synthetic aperture radar (SAR) image data, and target temporal SAR image data are input into the temporal mutual attention head 1 and the multi-scale generator 2 to generate simulated target temporal optical images. This simulated target temporal optical image, source temporal optical image data, source temporal SAR image data, and target temporal SAR image data are then input into the multi-scale discriminator 3 to obtain a discrimination result. This result is used to calculate a loss. After training one batch, the next batch is trained, and the above steps are repeated until the conditional generative adversarial network model reaches its optimum.

[0123] After training, the optimized conditional generative adversarial network model was tested. During testing, multi-scale generator 2 was kept fixed, and test samples were input into multi-scale generator 2 in batches. Multi-scale generator 2 would then output batches of simulated target temporal optical images. Because multi-scale generator 2 had been optimized, the simulated target temporal optical images output during the test were very close to the target temporal optical images.

[0124] The detailed configuration for training is shown in the table below:

[0125]

[0126] Figure 8 The illustration shows the effect generated by the optical image generation method provided in the embodiments of the present invention.

[0127] like Figure 8 As shown, the optical image generation method provided in this embodiment of the invention generates a real dual-temporal optical image sample and a corresponding generated image. The source temporal phase T0 is April 2021, and the target temporal phase T1 is May 2022. The green box area is a region with more changes. It can be seen that some buildings and bare land were added during the time period from T0 to T1, which can be basically reflected in the generated optical image.

[0128] Figure 9 The illustration shows a comparison of the generation results of the optical image generation method provided in the embodiments of the present invention in a simple scene.

[0129] like Figure 9As shown in the figure, the optical image generation method provided by this embodiment generates images in a simple scene, which are compared with two other methods: PixPix (an image translation model based on conditional generative adversarial networks) and Pix2PixHD (a conditional generative adversarial network model for high-resolution image generation with semantic manipulation). The results show that the optical image generation method provided by this embodiment performs better in terms of detail generation and color restoration. In the figure, the local areas in rows (a), (c), and (e) are outlined and correspondingly displayed in rows (b), (d), and (f). Rows (b) and (f) show that the optical image generation method provided by this embodiment restores details such as field roads and plot boundaries very well, while the other two methods basically fail to restore these details. In addition, row (d) shows that the optical image generation method provided by this embodiment can better restore the color of the plots.

[0130] Figure 10 The illustration shows a comparison of the generation results of the optical image generation method provided in the embodiments of the present invention in complex scenes.

[0131] like Figure 10 As shown, the optical image generation method provided in this embodiment of the invention generates images in complex scenes that, compared with methods such as Pix2PixHD (semantically manipulated high-resolution image generation conditional generative adversarial network), PSP (style vector-based image translation model), Selection-GAN (cascaded semantically guided multi-channel attention generative adversarial network), CHAN (complementary heterogeneous image translation generative adversarial network), and VQGAN Transformer (latent space quantization full attention generative adversarial network), the results show that the optical image generation method provided in this embodiment of the invention has better restoration of details and changes, such as the vegetation degradation reflected in row (d), the newly built buildings reflected in row (b), and the newly built roads reflected in row (f).

[0132] The optical image generation method provided in this invention utilizes a conditional generative adversarial network model to process a dual-temporal dataset formed from synthetic aperture radar (SAR) image data and multispectral image data to generate optical images. This significantly reduces the quality gap between optical images generated in simple and complex scenes. Specifically, it leverages a temporal mutual attention head 1 and a multi-scale generator 2 to fully enhance feature extraction capabilities. Simultaneously, a weight balancing strategy is employed to fully utilize the optical image features of the source temporal phase and the SAR image features of the dual temporal phases. This ensures that the generated target temporal optical image possesses realistic optical image texture while maintaining the sensitivity of the optical image during the transition from the source temporal phase to the target temporal phase. Even in complex scenes, it maintains good generation results and can provide data support for various applications requiring continuous high-resolution optical remote sensing observation data.

[0133] Although the invention has been illustrated and described in detail in the accompanying drawings and the foregoing description, such illustrations and descriptions should be considered illustrative or exemplary rather than restrictive.

[0134] Those skilled in the art will understand that the features described in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0135] Although the invention has been shown and described with reference to specific exemplary embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined by the appended claims and their equivalents. Therefore, the scope of the invention should not be limited to the above embodiments, but should be determined not only by the appended claims but also by their equivalents.

Claims

1. A method for generating an optical image, characterized in that, include: Synthetic aperture radar image data and multispectral image data are processed to obtain a dual-temporal dataset; The dual-temporal dataset is processed using a conditional generative adversarial network model to generate optical images; The conditional generative adversarial network model includes a temporal mutual attention head, a multi-scale generator, and a multi-scale discriminator, wherein... The change features of the dual-temporal dataset are extracted using the aforementioned time mutual attention head and then fused to obtain fused features; The multi-scale features of the fused features are extracted using the multi-scale generator to obtain a temporal optical image of the simulated target; The multi-scale discriminator generates a target temporal optical image based on the discrimination results of the simulated target temporal optical image; The conditional generative adversarial network model also includes a weight balancing module, wherein... The weight balancing module is used to balance the weights of the source temporal synthetic aperture radar image features, the target temporal synthetic aperture radar image features, and the source temporal optical image features. The weights of the source temporal synthetic aperture radar image features and the target temporal synthetic aperture radar image features are twice the weights of the source temporal optical image features. The calculation method of the weight balancing module is as follows: in, This represents the mean value of the target after linear transformation of the aperture radar image data and the multispectral image data. The standard deviation represents the linear transformation of the aperture radar image data and the multispectral image data.

2. The method for generating optical images according to claim 1, characterized in that, The step of extracting and fusing the variation features of the dual-temporal dataset using the time mutual attention head to obtain the fused features includes: The temporal mutual attention head is used to extract the variation features of the dual temporal dataset, and the variation features are fused with the source temporal optical image features to obtain pre-fused features; The pre-fused features and the target temporal synthetic aperture radar image features are fused to obtain the fused features.

3. The method for generating optical images according to claim 2, characterized in that, The method for extracting and fusing the variation features of the dual-temporal dataset using the time mutual attention head to obtain the fused features is as follows: in, Represents the characteristics of change. Represents pre-fusion characteristics, Representing the characteristics of integration, Represents trainable parameters, Representative temporal optical image characteristics. Representative temporal synthetic aperture radar image features of the target. Represents a query tensor. Represents the encoded tensor. Representative value tensor.

4. The method for generating optical images according to claim 1, characterized in that, The multi-scale generator includes a first generator and a second generator; wherein... Large-scale features are extracted using the first generator; Small-scale features are extracted using the second generator; The fusion feature input to the first generator is twice the fusion feature input to the second generator.

5. The method for generating optical images according to claim 1, characterized in that, The conditional generative adversarial network model also includes an optimizer, wherein... The optimizer optimizes the conditional generative adversarial network model by utilizing adversarial loss and perceptual loss. The optimization order is as follows: first optimize the multi-scale discriminator, then optimize the multi-scale generator.

6. The method for generating optical images according to claim 1, characterized in that, The process of processing synthetic aperture radar image data and multispectral image data to obtain a dual-temporal dataset includes: Polarimetric synthetic aperture radar imagery and multispectral imagery acquired using Sentinel satellites; The polarimetric synthetic aperture radar (PSAR) image is subjected to sub-strip separation processing, fine orbit correction processing, and radiation correction processing to obtain the source temporal PSAR image data and the target temporal PSAR image data. The multispectral image is subjected to band normalization processing to obtain the source temporal optical image data of the multispectral image; The source temporal synthetic aperture radar image data, the target temporal synthetic aperture radar image data, and the source temporal optical image data are sampled to obtain a simple scene dataset and a complex scene dataset.

7. An apparatus for generating optical images, characterized in that, include: The first processing module is used to process synthetic aperture radar image data and multispectral image data to obtain a dual-temporal dataset. The second processing module is used to process the dual-temporal dataset using a conditional generative adversarial network model to generate optical images. The conditional generative adversarial network model includes a temporal mutual attention head, a multi-scale generator, and a multi-scale discriminator, wherein the second processing module includes: The first extraction module is used to extract the change features of the dual-temporal dataset using the time mutual attention head and fuse them to obtain fused features; The second extraction module is used to extract the multi-scale features of the fused features using the multi-scale generator to obtain a simulated target temporal optical image. The generation module is used to generate a target temporal optical image based on the discrimination result of the simulated target temporal optical image using the multi-scale discriminator; A weight balancing module is used to balance the weights of source temporal synthetic aperture radar (SAR) image features, target temporal SAR image features, and source temporal optical image features; the weights of the source temporal SAR image features and the target temporal SAR image features are twice the weight of the source temporal optical image features. The calculation method of the weight balancing module is as follows: in, This represents the mean value of the target after linear transformation of the aperture radar image data and the multispectral image data. The standard deviation represents the linear transformation of the aperture radar image data and the multispectral image data.

Citation Information

Patent Citations

  • Regional crop classification method and system based on multi-dimensional feature fusion

    CN112183209A

  • Pathological image classification method and system based on multi-scale domain adversarial network

    CN114299324A