Multi-source remote sensing image dodging and color dodging method based on style migration

Through the multi-source remote sensing image uniform light and color uniform method based on style migration, the problems of cloud layer interference and color differences in multi-source remote sensing image processing are solved, the seamless inlay of images and high-quality data foundation are realized, and the application value is enhanced.

CN120031708AActive Publication Date: 2025-05-23WUHAN UNIV

Patent Information

Application Number
CN202510120385.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-23
Estimated Expiration
2045-01-25

AI Technical Summary

Technical Problem

When dealing with multi-source remote sensing images, the existing technology faces problems such as cloud layer interference, significant color differences, insufficient automation level, difficulty in extracting splicing seams and unsatisfactory splicing seams, which affects the mosaic effect and application value of the image.

Method used

The multi-source remote sensing image uniform and uniform color method is adopted based on style transfer. Through deep learning technology, three technical modules including cloud removal, semantic extraction and style transfer, the color and brightness of the image are intelligently adjusted to achieve uniform and uniform color.

Benefits of technology

On the basis of retaining land features, the visual presentation of images is optimized, the application value of remote sensing images is enhanced, and the seamless mosaic and high-quality data foundation of multi-source remote sensing images is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031708A_ABST
    Figure CN120031708A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source remote sensing image dodging and color dodging method based on style migration, and belongs to the technical field of computer vision and remote sensing science. Comprising the steps of making a remote sensing satellite cloud image data set based on a standard optical model; constructing positive and negative cues through category semantic tags, introducing a degraded cloud image as a condition into a noise estimation network to estimate conditional noise distribution, and training a conditional diffusion model; constructing a remote sensing image semantic extraction model combining a convolutional network and Transform, performing semantic extraction on a multi-source remote sensing image, and outputting a ground object category label of each pixel point; according to content and style information of a multi-source remote sensing image, a style migration model based on Transform is constructed, and light uniformity and color uniformity of the remote sensing image are realized by referring to a reference image with appropriate radiation characteristics. The method can be applied to realizing tone coordination and moderate contrast between adjacent images, so that a color image approaches to a natural real color.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for uniform light and color of multi-source remote sensing images based on style migration, and belongs to the field of computer vision and remote sensing science and technology. Background Art

[0002] With the rapid advancement of remote sensing technology, remote sensing images play an increasingly important role in many key areas such as urban planning, land resource management, and environmental monitoring. In the context of the country's promotion of smart urban construction, the application of remote sensing technology not only promotes the rational use of land resources, but also provides strong support for the high-quality development of the economy. As the core tool for obtaining geospatial information, the color consistency and authenticity of remote sensing images are crucial for subsequent analysis and application. However, due to the influence of various factors such as lighting conditions, sensor characteristics, and atmospheric effects, the acquired remote sensing images often have color differences and unevenness. These problems seriously affect the mosaic effect and practical application value of the image.

[0003] In the field of remote sensing image processing, the research on uniform light and color methods has far-reaching significance. It not only significantly improves the availability and analysis accuracy of remote sensing images, but is also particularly important in processing large-scale, multi-source, and multi-temporal remote sensing images. By eliminating color differences caused by different sensors, imaging conditions, and environmental changes, the uniform light and color method ensures the consistency of remote sensing images in color and brightness, providing a solid high-quality data foundation for subsequent image analysis and application. At present, the research on uniform light and color of multi-source remote sensing images still faces many challenges, such as interference from cloud and fog layers, significant color differences, insufficient automation level, difficulty in extracting seams, and unsatisfactory seam elimination effects. Summary of the invention

[0004] The present invention provides a method for uniform lighting and coloring of multi-source remote sensing images based on style transfer, which solves the problems disclosed in the background technology. The method can be divided into three technical modules: cloud removal based on style transfer, semantic extraction, and style transfer. Through deep learning technology, the color and brightness of remote sensing images can be intelligently adjusted to achieve the purpose of uniform lighting and coloring. First, a remote sensing cloud image dataset is produced based on crawled OSM images and standard optical models, and a conditional diffusion model is trained to perform cloud removal preprocessing on multi-source remote sensing images. Then, semantic extraction is performed on the preprocessed remote sensing images to achieve semantic segmentation of the images. Finally, style transfer technology is used to ensure the consistency of the styles of objects between different images and achieve seamless mosaicking. This method optimizes the visual presentation of images and enhances the application value of remote sensing images while retaining the features of objects.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] A method for uniform lighting and color of multi-source remote sensing images based on style migration, comprising:

[0007] Based on the standard optical model, images with different degrees of cloud and fog are synthesized to produce a remote sensing satellite cloud and fog image dataset;

[0008] The prompt words are constructed through category semantic labels to generate text conditions; the degraded fog image in the remote sensing satellite fog image dataset is introduced as a condition into the noise estimation network to estimate the conditional noise distribution and generate fog control conditions. The latent diffusion model is trained based on the text conditions and fog control conditions; the result after applying the fog image control condition is added to the result output by the latent diffusion model, and the remote sensing image with fog removed is output by the VAE decoder;

[0009] Through the pre-built remote sensing image semantic extraction model, semantic extraction is performed on the remote sensing images after cloud and fog removal preprocessing, and the content and style information of multi-source remote sensing images is output;

[0010] A Transformer-based style migration model is constructed based on the content and style information of multi-source remote sensing images. The style migration model refers to a benchmark image with suitable radiation characteristics to achieve uniform light and color of the remote sensing images.

[0011] Furthermore, the standard optical model is: ;

[0012] in, is the coordinate value of the image pixel, For foggy images, is the dehazed image to be restored, is the transmittance, It is the global atmospheric light component. In cloudy weather conditions, the reflected energy is reduced, causing transmission attenuation. , which in turn reduces the image brightness; the surrounding lighting scatters to form air light , which increases the image brightness and reduces the saturation; the cloud image is formed by the light reflected by the object weakened by the cloud and the atmospheric light reflected by the cloud, It is the weakening and scattering of light caused by clouds and fog.

[0013] Furthermore, the training objective of the potential diffusion model is: ;

[0014] in, Represents the true noise value, which follows the standard normal distribution , represents the noise value predicted by the model; represents the time step in the diffusion process, is the representation vector in the latent space, Represents the potential variable in the diffusion process at a certain moment in the diffusion process status, represents the conditional embedding of degraded foggy images, represents text conditional embedding, Indicates expected value.

[0015] Furthermore, the method for constructing the remote sensing image semantic extraction model includes:

[0016] Construct a convolutional network-based encoder, using ResNet18 as the encoder to extract multi-scale semantic features. It consists of four stages of residual blocks, and each stage reduces the scale factor of the feature map by 2 times through downsampling; the feature map generated in each stage is fused with the corresponding feature map in the decoder through 1x1 convolution; the semantic features generated by the residual block and the features generated by the global-local conversion block in the decoder are aggregated through a weighted sum operation;

[0017] Building a Transformer-based decoder that utilizes three global-local attention Transformer blocks and a feature refinement head to build a lightweight Transformer-based decoder;

[0018] The global-local attention Transformer block constructs two parallel branches to extract global and local context respectively, including:

[0019] The local branch adopts two parallel convolutional layers with core sizes of 3 and 1 to extract local context; two batch normalization operations are appended before the final summation operation;

[0020] The global branch uses a window-based multi-head self-attention mechanism to capture the global context; the channel dimension of the input 2D feature map is expanded to three times using a standard 1x1 convolution; and a window split operation is then applied to split the 1D sequence into query Q, key K, and value V vectors.

[0021] Furthermore, the formula for the weighted summation operation is: ;

[0022] in, represents the fusion feature, represents the features generated by the residual block, represents the features generated by the global local Transformer block, Represents weight.

[0023] Furthermore, the method for constructing a remote sensing image semantic extraction model also includes constructing a cross-shaped window context interaction module, which captures the global context by fusing two feature maps generated by a horizontal average pooling layer and a vertical average pooling layer; the horizontal average pooling layer establishes a horizontal relationship between windows; for any point in window 1 , and the corresponding point in window 2 The dependency relationship is modeled as:

[0024] ;

[0025] ;

[0026] ;

[0027] ;

[0028] in, Represents the index, is the window size, represents the self-attention computation, which can model the dependencies between pairs of pixels in a local window.

[0029] Furthermore, the Transformer-based style transfer model includes:

[0030] Transformer encoder, which is used to feed the image sequence to be processed into the Transformer encoder. Each layer contains a multi-head self-attention module and a feed-forward network to convert the input sequence into a query Q, a key K, and a value V;

[0031] The multi-head self-attention module is used to process different heads in parallel through a multi-head self-attention mechanism, calculate attention, and effectively encode the input sequence;

[0032] The Transformer decoder is used to translate the encoded content sequence in a regressive manner according to the encoded style sequence. All sequence blocks are input and predicted at once. The feature sequence is generated by using the content sequence to generate the query Q, and the style sequence to generate the key K and value V for sequence translation. The output sequence of the Transformer decoder is further refined by a three-layer CNN decoder, including convolution, ReLU activation and upsampling.

[0033] Furthermore, the style transfer model refers to a reference image with suitable radiation characteristics, and the method for achieving uniform light and color of the remote sensing image includes:

[0034] After the input content sequence is embedded, it is fed into the Transformer encoder, and the input content sequence is converted into a representation of query Q, key K, and value V:

[0035] ;

[0036] in, Represents the input content sequence, , represents the sequence length, Indicates the number of attention heads;

[0037] The multi-head self-attention mechanism calculates attention by processing different heads in parallel, achieving efficient encoding of the input sequence:

[0038] ;

[0039] The Transformer decoder is based on the encoding style sequence , translating the encoded content sequence in a recursive manner ; The input of the Transformer decoder includes the encoded content sequence and style sequence ; Generate queries using content sequences , use the style sequence to generate the key Sum :

[0040] ;

[0041] Compute the output sequence of the Transformer decoder :

[0042] ;

[0043] The output sequence of the Transformer decoder With shape ,in and Represents the height and width of the output respectively, represents the number of channels; a three-layer CNN decoder is used to further refine the output of the Transformer decoder; each layer of the CNN decoder is scaled up, including 3x3 convolution, ReLU activation function and 2x upsampling, and finally, a resolution of The output image is made to match the input image in terms of spatial resolution. Each pixel of the output image contains information of three color channels, generating a remote sensing image with uniform light and color.

[0044] Accordingly, the present invention also provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by a computing device, the computing device performs any of the above methods.

[0045] Accordingly, the present invention also provides a computing device, comprising:

[0046] One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described above.

[0047] The beneficial effects achieved by the present invention are:

[0048] The present invention provides a method for uniform lighting and coloring of multi-source remote sensing images based on style transfer. By crawling images of some towns in the New York area at the OSM16 level and the category semantic labels in the images, a remote sensing satellite cloud image dataset is produced based on a standard optical model. The conditional diffusion model is trained with text conditions and degraded cloud image conditions to achieve cloud removal. On the basis of the preprocessing of cloud removal, a remote sensing image semantic extraction model combining a convolutional network and a Transformer is constructed to perform semantic extraction on the remote sensing image and achieve semantic segmentation of different categories. According to the content and style information of multi-source remote sensing images, a style transfer model based on Transformer is constructed to extract the content information and style information of the multi-source remote sensing images, realize the fusion of content and style information between the remote sensing images to be uniformly lit and colored and the reference images with suitable radiation characteristics, and realize the color coordination and moderate contrast between adjacent images, aiming to enhance the application value of remote sensing images in key fields such as land use monitoring, urban planning, and resource management. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the overall process of the multi-source remote sensing image uniform light and color uniformity method based on style transfer of the present invention;

[0050] Figure 2 It is a schematic diagram of the framework of the remote sensing image cloud removal model of the present invention;

[0051] Figure 3 It is a schematic diagram of the framework of the multi-source remote sensing image uniform light and color model based on style transfer of the present invention. DETAILED DESCRIPTION

[0052] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0053] like Figure 1 As shown, the present invention provides a method for uniform light and color of multi-source remote sensing images based on style migration, comprising the following steps:

[0054] Step 1: crawl OSM16-level images of some towns in the New York area and the category semantic labels in the images, and create a remote sensing satellite cloud image dataset based on the standard optical model;

[0055] Furthermore, crawling OSM16-level images of some towns in the New York area and the category semantic labels in the images, and making a remote sensing satellite cloud image dataset based on a standard optical model, includes:

[0056] Based on existing industry standards or national standards, such as "Map Features-OSM Wiki" and "Basic Geographic Information 1:50000 Terrain Feature Data Specification", as well as experience and statistics from other projects, a land feature classification system is constructed, including 54 land feature categories such as roads, rivers, grasslands, bare land, and buildings. Ensure that the dataset images produced contain a rich range of land feature categories, so that the trained model can be more widely applied to diverse terrains;

[0057] Based on the established feature classification system, we crawled OSM16 images of some towns in the New York area. Some categories were not reflected in the pictures, such as bridges, ships, water towers, etc. Therefore, we crawled images centered on these objects to supplement them;

[0058] Crawl the category semantic labels in each image for use in constructing prompt words for subsequent model training;

[0059] The standard optical model used, such as formula (1), is used to synthesize clouds and fog in remote sensing images. Based on the center point synthesis cloud and fog method, the image brightness, cloud and fog concentration, and fog size are randomly initialized, and clouds and fog at different locations, different ranges, and different degrees are randomly synthesized to ensure the diversity of samples.

[0060] (1);

[0061] in is the coordinate value of the image pixel, For foggy images, is the dehazed image to be restored, is the transmittance, It is the global atmospheric light component. In cloudy weather conditions, the reflected energy is reduced, causing transmission attenuation. , which results in a decrease in image brightness. Ambient lighting scatters to form air light , which increases the image brightness and reduces the saturation. The cloud image is formed by the light reflected by the object weakened by the cloud and the atmospheric light reflected by the cloud. It can be understood as the weakening and scattering of light by clouds and fog.

[0062] Step 2: construct positive and negative cue words through category semantic labels, and introduce the degraded fog image as a condition into the noise estimation network to estimate the conditional noise distribution and train the conditional diffusion model, such as Figure 2 As shown;

[0063] Furthermore, constructing positive and negative prompt words through category semantic labels, and introducing the degraded fog image as a condition into the noise estimation network to estimate the conditional noise distribution, and training the conditional diffusion model, includes:

[0064] Construct label prompt words according to the object category labels of OSM images, such as "river, forest, meadow, motorway, buildings, and..." and construct positive and negative image prompt words based on the label prompt words. The positive image prompt word is "A bright and clear remote sensing satellite image including -[class types]", where [class types] is the label prompt word, and the negative image prompt word is "A remote sensing satellite image obscured by clouds and haze." The label prompt words and image prompt words are used as text conditions and input into the network through the text encoder of clip;

[0065] Latent Diffusion Model (LDM) compresses the input image into a low-dimensional vector , and performs the diffusion process in the latent space, thereby reducing the cost of diffusion model training and inference. The generated samples Reconstructed into a high-resolution image in pixel space.

[0066] The stable diffusion model LDM is fine-tuned based on the text conditions and cloud image conditions of public and self-made cloud remote sensing datasets to avoid the time and computing cost of training from scratch, and the risk of "catastrophic forgetting", so as to utilize rich domain-specific knowledge without changing the pre-trained weights. The LDM model is copied into two identical parts, namely the "locked" copy and the "trainable" copy. The "locked" copy is frozen, that is, the weight remains unchanged, retaining the ability of the LDM model after training on large-scale text and image datasets. Using degraded cloud image conditions The “trainable” copy is fine-tuned by introducing the degraded fog image as an image condition into the noise estimation network to estimate the conditional noise distribution. The degraded fog image condition is passed through a learnable embedding layer and The training objective can be expressed as formula (2). Then the result after applying the cloud image control condition is added to the result of the original LDM model through the VAE decoder to obtain the final output clear remote sensing image.

[0067] (2);

[0068] in, Represents the true noise value, which follows the standard normal distribution , Represents the noise value predicted by the model. represents the time step in the diffusion process, is the representation vector in the latent space, Represents the potential variable in the diffusion process at a certain moment in the diffusion process status, represents the conditional embedding of degraded foggy images, represents text conditional embedding, Indicates expected value.

[0069] Step 3: Build a remote sensing image semantic extraction model that combines convolutional networks and Transformer to perform semantic extraction on multi-source remote sensing images and output the object category label for each pixel.

[0070] Furthermore, the construction combines the convolutional network and the Transformer remote sensing image semantic extraction model to perform semantic extraction on multi-source remote sensing images and output the ground object category label of each pixel point, including:

[0071] Construct an encoder based on a convolutional network, using ResNet18 as the encoder to extract multi-scale semantic features. It consists of four stages of residual blocks, and each stage reduces the scale factor of the feature map by 2 times through downsampling. The feature map generated in each stage is fused with the corresponding feature map in the decoder through 1x1 convolution. The semantic features generated by the residual block and the features generated by the global local transformation block (GLTB) in the decoder are aggregated through a weighted sum operation. The weighted sum operation selectively weights the two features according to their contribution to the segmentation accuracy, thereby learning a more generalized fusion feature. The formula for the weighted sum operation is as follows (3):

[0072] (3);

[0073] in, represents the fusion feature, represents the features generated by the residual block, represents the features generated by the global local Transformer block, represents weight;

[0074] Construct a Transformer-based decoder, that is, use three global-local Transformer blocks and a feature refinement head to build a lightweight Transformer-based decoder. Through this layered and lightweight design, the decoder can capture global and local information at multiple scales while maintaining high efficiency.

[0075] Specifically, a global-local attention Transformer block is proposed, which constructs two parallel branches to extract global and local context respectively, including:

[0076] The local branch uses two parallel convolutional layers with core sizes of 3 and 1 to extract local context. Two batch normalization operations are appended before the final summation operation;

[0077] The global branch uses a window-based multi-head self-attention mechanism to capture global context. The channel dimension of the input 2D feature map is expanded to three times using a standard 1x1 convolution. A window split operation is then applied to split the 1D sequence into query (Q), key (K), and value (V) vectors;

[0078] In order to capture cross-window relations with high computational efficiency, a cross-window context interaction module is proposed. This module captures the global context by fusing two feature maps produced by the horizontal average pooling layer and the vertical average pooling layer. The horizontal average pooling layer establishes the horizontal relations between windows. For any point in window 1 , which is the same as the The dependencies can be modeled as:

[0079] (4);

[0080] (5);

[0081] (6);

[0082] (7);

[0083] in, Represents an index, is the window size, represents the self-attention calculation, which can model the dependencies between pixel pairs in the local window, is the corresponding point in window 2.

[0084] Shallow features retain rich spatial details but lack semantic content. Deep features provide precise semantic information but have low spatial resolution. Directly summing these two features is fast but may reduce segmentation accuracy. Therefore, a feature refinement head is added to construct the decoder, including:

[0085] The shallow and deep features are weighted and summed to make full use of precise semantic information and spatial details;

[0086] The fused features are used as the input of the feature refinement head to strengthen the feature representation;

[0087] The channel path uses a global average pooling layer to generate a channel attention map, and adjusts the channel dimension through reduction and expansion operations;

[0088] The spatial path uses deep convolution to produce a spatial attention map to reflect the spatial resolution of the feature map;

[0089] The attention features generated by the two paths are further fused through the summation operation. A 1x1 convolution layer and upsampling operation are applied to generate the final segmentation map. At the same time, residual connections are introduced to prevent network performance degradation.

[0090] Step 4: Based on the content and style information of multi-source remote sensing images, a Transformer-based style transfer model is constructed, and reference images with appropriate radiation characteristics are used to achieve uniform light and color of remote sensing images.

[0091] Furthermore, the Transformer-based style migration model is constructed for the content and style information of multi-source remote sensing images, and reference images with suitable radiation characteristics are referred to to achieve uniform light and color of remote sensing images, including:

[0092] After the input content sequence is embedded, it is sent to the Transformer encoder. Each layer of the encoder contains a multi-head self-attention module and a feedforward network. The input sequence is converted into the representation of query (Q), key (K) and value (V). The calculation formula is shown in (8):

[0093] (8);

[0094] in, Represents the input content sequence, , represents the sequence length, Represents the number of attention heads.

[0095] The multi-head self-attention mechanism calculates attention by processing different heads in parallel, thereby achieving effective encoding of the input sequence. The calculation formula is shown in (7):

[0096] (9);

[0097] The decoder is used to sequence the encoding style , translating the encoded content sequence in a recursive manner . Unlike the autoregressive process in traditional natural language processing tasks, all sequence blocks are taken as input at one time to predict the output. Each Transformer decoder layer contains two multi-head self-attention modules and a feedforward network. The input of the Transformer decoder includes the encoded content sequence and style sequence As shown in formula (10), the query is generated using the content sequence , use the style sequence to generate the key Sum :

[0098] (10);

[0099] Then, the output sequence of the Transformer decoder can be calculated as follows :

[0100] (11);

[0101] Transformer output sequence With shape ,in and Represents the height and width of the output respectively, Represents the number of channels. To construct the final result, a three-layer CNN decoder is used to further refine the output of the Transformer decoder. For each layer of the CNN decoder, the scale is enlarged through a series of operations, including 3x3 convolution, ReLU activation function, and 2x upsampling. Finally, a resolution of This means that the output data not only matches the input image in spatial resolution, but also contains information from three color channels at each pixel, enabling the generation of high-quality images with natural colors.

[0102] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

[0103] A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by a computing device, the computing device executes a multi-source remote sensing image uniformity method based on style migration.

[0104] A computing device includes one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing a multi-source remote sensing image uniformity method based on style migration.

[0105] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0106] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0107] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0109] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the claims of the present invention to be approved.

Claims

1. A method for uniform lighting and color of multi-source remote sensing images based on style transfer, characterized in that: include: Based on the standard optical model, images with different degrees of cloud and fog are synthesized to produce a remote sensing satellite cloud and fog image dataset; The prompt words are constructed through category semantic labels to generate text conditions; the degraded fog image in the remote sensing satellite fog image dataset is introduced as a condition into the noise estimation network to estimate the conditional noise distribution and generate fog control conditions. The latent diffusion model is trained based on the text conditions and fog control conditions; the result after applying the fog image control condition is added to the result output by the latent diffusion model, and the remote sensing image with fog removed is output by the VAE decoder; Through the pre-built remote sensing image semantic extraction model, semantic extraction is performed on the remote sensing images after cloud and fog removal preprocessing, and the content and style information of multi-source remote sensing images is output; A Transformer-based style migration model is constructed based on the content and style information of multi-source remote sensing images. The style migration model refers to a benchmark image with suitable radiation characteristics to achieve uniform light and color of the remote sensing images.

2. The method for uniform lighting and color of multi-source remote sensing images based on style transfer according to claim 1, characterized in that: The standard optical model is: ; in, is the coordinate value of the image pixel, For foggy images, is the dehazed image to be restored, is the transmittance, It is the global atmospheric light component. In cloudy weather conditions, the reflected energy is reduced, causing transmission attenuation. , which in turn reduces the image brightness; the surrounding lighting scatters to form air light , which increases the image brightness and reduces the saturation; the cloud image is formed by the light reflected by the object weakened by the cloud and the atmospheric light reflected by the cloud, It is the weakening and scattering of light caused by clouds and fog.

3. The method for uniform lighting and color of multi-source remote sensing images based on style transfer according to claim 1, characterized in that: The training objective of the potential diffusion model is: ; in, Represents the true noise value, which follows the standard normal distribution , represents the noise value predicted by the model; represents the time step in the diffusion process, is the representation vector in the latent space, Represents the potential variable in the diffusion process at a certain moment in the diffusion process status, represents the conditional embedding of degraded foggy images, represents text conditional embedding, Indicates expected value.

4. The method for uniform lighting and color of multi-source remote sensing images based on style transfer according to claim 1, characterized in that: The construction method of remote sensing image semantic extraction model includes: Construct a convolutional network-based encoder, using ResNet18 as the encoder to extract multi-scale semantic features. It consists of four stages of residual blocks, and each stage reduces the scale factor of the feature map by 2 times through downsampling; the feature map generated in each stage is fused with the corresponding feature map in the decoder through 1x1 convolution; the semantic features generated by the residual block and the features generated by the global-local conversion block in the decoder are aggregated through a weighted sum operation; Building a Transformer-based decoder that utilizes three global-local attention Transformer blocks and a feature refinement head to build a lightweight Transformer-based decoder; The global-local attention Transformer block constructs two parallel branches to extract global and local context respectively, including: The local branch adopts two parallel convolutional layers with core sizes of 3 and 1 to extract local context; two batch normalization operations are appended before the final summation operation; The global branch uses a window-based multi-head self-attention mechanism to capture the global context; the channel dimension of the input 2D feature map is expanded to three times using a standard 1x1 convolution; and a window split operation is then applied to split the 1D sequence into query Q, key K, and value V vectors.

5. The method for uniform lighting and color of multi-source remote sensing images based on style transfer according to claim 4, characterized in that: The formula for the weighted summation operation is: ; in, represents the fusion feature, represents the features generated by the residual block, represents the features generated by the global local Transformer block, Represents weight.

6. The method for uniform lighting and color of multi-source remote sensing images based on style transfer according to claim 4, characterized in that: The method for constructing a remote sensing image semantic extraction model also includes constructing a cross-shaped window context interaction module, which captures the global context by fusing two feature maps generated by a horizontal average pooling layer and a vertical average pooling layer; the horizontal average pooling layer establishes a horizontal relationship between windows; for any point in window 1 , and the corresponding point in window 2 The dependency relationship is modeled as: ; ; ; ; in, Represents an index, is the window size, represents the self-attention computation, which can model the dependencies between pairs of pixels in a local window.

7. The method for uniform lighting and color of multi-source remote sensing images based on style transfer according to claim 1, characterized in that: Transformer-based style transfer models include: Transformer encoder, which is used to feed the image sequence to be processed into the Transformer encoder. Each layer contains a multi-head self-attention module and a feed-forward network to convert the input sequence into a query Q, a key K, and a value V; The multi-head self-attention module is used to process different heads in parallel through a multi-head self-attention mechanism, calculate attention, and effectively encode the input sequence; The Transformer decoder is used to translate the encoded content sequence in a regressive manner according to the encoded style sequence. All sequence blocks are input and predicted at once. The feature sequence is generated by using the content sequence to generate the query Q, and the style sequence to generate the key K and value V for sequence translation. The output sequence of the Transformer decoder is further refined by a three-layer CNN decoder, including convolution, ReLU activation and upsampling.

8. The method for uniform lighting and color of multi-source remote sensing images based on style transfer according to claim 7, characterized in that: The style migration model refers to a reference image with suitable radiation characteristics, and the method for achieving uniform light and color of remote sensing images includes: After the input content sequence is embedded, it is fed into the Transformer encoder, and the input content sequence is converted into a representation of query Q, key K, and value V: ; in, Represents the input content sequence, , represents the sequence length, Indicates the number of attention heads; The multi-head self-attention mechanism calculates attention by processing different heads in parallel, achieving efficient encoding of the input sequence: ; The Transformer decoder is based on the encoding style sequence , translating the encoded content sequence in a recursive manner ; The input of the Transformer decoder includes the encoded content sequence and style sequence ; Generate queries using content sequences , use the style sequence to generate the key Sum : ; Compute the output sequence of the Transformer decoder : ; The output sequence of the Transformer decoder With shape ,in and Represents the height and width of the output respectively, represents the number of channels; a three-layer CNN decoder is used to further refine the output of the Transformer decoder; each layer of the CNN decoder is scaled up, including 3x3 convolution, ReLU activation function and 2x upsampling, and finally, a resolution of The output image is made to match the input image in terms of spatial resolution. Each pixel of the output image contains information of three color channels, generating a remote sensing image with uniform light and color.

9. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions which, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1 to 8.

10. A computing device, characterized in that include: One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods according to claims 1 to 8.

Citation Information

Patent Citations

  • Image defogging method and system based on style migration network

    CN109934791A

  • Wide remote sensing image ship target rapid detection method based on intra-domain transfer learning

    CN114627372A

  • Convolutional neural network-based remote sensing image dodging and color dodging method and device

    CN116703744A

  • Noctilucent remote sensing image cloud and fog removing method and device

    CN118247176A

  • Spatial simulation method for assessment of direct economic losses of typhoon flood based on remote sensing

    US20240242289A1

Cited By

  • Multi-source remote sensing image radiation normalization method and system based on frequency domain decomposition

    CN120451592A

  • Multi-source remote sensing image radiometric normalization method and system based on frequency domain decomposition

    CN120451592B