Double-layer multi-scale diffusion color homogenizing method and device based on fusion channel contrast attention
By adopting a double-layer multi-scale diffusion uniform color method based on fusion channel comparison attention in remote sensing image processing, the problem of insufficient recovery ability of traditional methods when the color difference is too large is solved, high-precision color restoration and texture reconstruction are achieved, and the visual effect and texture fidelity of the image are significantly improved.
Patent Information
- Application Number
- CN202510182360.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
The traditional remote sensing image uniform color method lacks recovery ability when the color difference is too large and lacks capture of multi-scale features, resulting in poor visual effects and texture fidelity of the image.
The double-layer multi-scale diffusion uniform color method based on fusion channel contrast attention is adopted, and the multi-scale features of the image are captured and optimized through the double-layer multi-scale U-shaped network and the supervised attention module to achieve high-precision color restoration and texture reconstruction.
It effectively overcomes the lack of recovery ability of traditional methods in scenarios with excessive color difference, and achieves high-precision color restoration and structural consistency of high-resolution remote sensing images, which significantly improves the visual effect and texture fidelity of the image.
Smart Images

Figure CN120125482A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image color equalization, and particularly to a double-layer multi-scale diffusion equalization method and device based on fused channel contrast attention. Background Art
[0002] With the rapid development of remote sensing technology and the popularization of high-resolution remote sensing imaging equipment, remote sensing images play an increasingly important role in the fields of land resource management, environmental monitoring, urban planning, etc. However, due to the limitations of the imaging environment, sensors, and the changes in lighting conditions, remote sensing images are often affected by various degradation factors such as color difference and noise during the acquisition process, resulting in distortion of colors and textures in the images. This distortion not only affects the visual effect of the images but also reduces the accuracy of subsequent analysis (such as ground object classification and target detection). Therefore, color equalization and texture reconstruction of remote sensing images have become important research topics in remote sensing image processing.
[0003] Traditional remote sensing image color equalization methods mainly rely on techniques such as histogram matching, color correction, and multispectral fusion. Although these methods have achieved certain effects in specific scenarios, they generally have the following limitations:
[0004] Insufficient recovery ability when the color difference is too large: For images in strong light, shadow areas, or under different lighting conditions, these methods often cannot accurately restore the true colors. Lack of capture of multi-scale features: Traditional methods usually only process the global or local information of the images and cannot effectively fuse the detailed features of the images at different scales. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the above defects in the prior art and propose a double-layer multi-scale diffusion equalization method and device based on fused channel contrast attention. The double-layer multi-scale diffusion equalization architecture based on fused channel contrast attention can well achieve color restoration and texture reconstruction of degraded images with excessive color difference.
[0006] The present invention adopts the following technical solutions:
[0007] A double-layer multi-scale diffusion equalization method based on fused channel contrast attention, comprising:
[0008] Obtain a remote sensing image to be equalized, and perform cropping to obtain image patches;
[0009] Input the image patches into a pre-trained double-layer multi-scale diffusion equalization model to obtain equalized image patches;
[0010] Calculate the noise estimation loss and singular value loss for the equalized image patches, and optimize the model parameters;
[0011] Stitch the image patches after uniform colorization to obtain the complete remote sensing image after uniform color restoration.
[0012] The double-layer multi-scale diffusion uniform colorization model includes a first processing module, a noise level embedding module, a first U-shaped network, a supervised attention module, a second processing module, a second U-shaped network, and an adaptive feature aggregation module;
[0013] The first processing module is used to sequentially perform feature dimensionality reduction and enhanced feature expression on the image patches to be uniformly colored and then input them into the first U-shaped network; the noise level embedding module is used to train learnable noise and embed it into the first U-shaped network; the first U-shaped network trains the input image to obtain a feature map and inputs it into the supervised attention module; the supervised attention module is used to generate an attention map to suppress features with less information in the current stage to obtain the first-stage output result map;
[0014] The second processing module is used to sequentially perform feature learning and enhanced feature expression on the image patches to be uniformly colored, splice them with the first-stage output result map, and then input them into the second U-shaped network to obtain the second-stage output result map;
[0015] The adaptive feature aggregation module aggregates the features of the second-stage output result map and the first-stage output result map to obtain the image patches after uniform colorization.
[0016] The first processing module includes a first convolution module and a first feature enhancement module. The convolution kernel of the operation unit of the first convolution module is a 3*3 convolution, the number of convolution kernels is 64, and the stride is 1; the first feature enhancement module includes two convolutions, a PReLU activation function, and a channel attention layer; the channel attention layer includes a global average pooling, which compresses the spatial dimension of the feature map to 1*1, only retains the channel dimension information, then uses two convolutions with a convolution kernel of 1*1 and a padding size of 1 for dimensionality reduction and dimensionality increase, and finally outputs after using the Sigmoid activation function.
[0017] The second processing module includes a second convolution module and a second feature enhancement module. The convolution kernel of the convolution operation unit of the second convolution module is a 3*3 convolution, the number of convolution kernels is 64, and the stride is 1; the second feature enhancement module includes two convolutions, a PReLU activation function, and a channel attention layer; the channel attention layer includes a global average pooling, which compresses the spatial dimension of the feature map to 1*1, only retains the channel dimension information, then uses two convolutions with a convolution kernel of 1*1 and a padding size of 1 for dimensionality reduction and dimensionality increase, and finally outputs after using the Sigmoid activation function.
[0018] The adaptive feature aggregation module includes a global average pooling, a dimensionality reduction convolution, a Softmax activation function, and global feature aggregation. The global average pooling is used to compress the image into a 1*1 feature map. The convolution operation unit of a dimensionality reduction convolution uses a 1*1 convolution kernel with 64 kernels, and its function is to generate attention weights for each feature map. Global feature aggregation is used to perform weighted summation on two input features in the channel dimension to obtain the final fused feature.
[0019] The noise level embedding module uses a multi-layer perceptron to generate a noise with 3 channels and a size of 256*256.
[0020] The supervised attention module includes three convolutional layers and a Sigmoid activation function. Among them, the convolution operation unit of the first convolutional layer uses a 1*1 convolution kernel with 64 kernels and a stride of 1; the convolution operation unit of the second convolutional layer uses a 1*1 convolution kernel with 6 kernels and a stride of 1; the convolution operation unit of the third convolutional layer uses a 1*1 convolution kernel with 64 kernels and a stride of 1.
[0021] The first U-shaped network includes five groups of encoding structures connected in sequence and five groups of decoding structures connected in sequence through an upsampling structure. Feature maps are extracted through the five groups of encoding structures, and the extracted feature map information is upsampled through the five groups of structures.
[0022] The second U-shaped network includes five groups of encoding structures connected in sequence and five groups of decoding structures connected in sequence through an upsampling structure. Feature maps are extracted through the five groups of encoding structures, and the extracted feature map information is upsampled through the five groups of structures.
[0023] A double-layer multi-scale diffusion color homogenization device based on fused channel contrast attention includes:
[0024] A cutting module, which obtains a remote sensing image to be color homogenized and performs cropping to obtain image cuts.
[0025] A color homogenization module, which inputs the image cuts into a pre-trained double-layer multi-scale diffusion color homogenization model to obtain color homogenized image cuts; calculates the noise estimation loss and the singular value loss for the color homogenized image cuts, and optimizes the model parameters.
[0026] A splicing module, which splices the color homogenized image cuts to obtain a complete remote sensing image after color homogenization restoration.
[0027] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects:
[0028] The method and device of the present invention combine a fusion channel contrast attention mechanism with a double-layer multi-scale U-shaped network, achieving the capture and optimization of multi-scale features of images, overcoming the limitation of insufficient restoration ability of traditional methods in scenarios with excessive color differences. It ensures high-precision color restoration and structural consistency when processing complex degraded remote sensing images, and can well perform color homogenization processing on high-resolution remote sensing images with large color differences. The processed images have significant improvements in visual effects and texture fidelity.
[0029] In the present invention, deep learning-based diffusion models and attention mechanisms have demonstrated powerful potential in tasks such as image enhancement, super-resolution, and image restoration. Among them, diffusion models can achieve excellent performance in complex image degradation problems by simulating the gradual diffusion and denoising process of noise. And the attention mechanism, especially the channel attention mechanism, can adaptively capture the dependence relationships between channels and improve the model's understanding ability of different channel features. Brief Description of the Drawings
[0030] Figure 1 is a schematic flowchart of the method of the present invention.
[0031] Figure 2 is a network structure diagram of the double-layer multi-scale diffusion color homogenization architecture of the present invention.
[0032] Figure 3 is a training flowchart of the model of the present invention.
[0033] Figure 4 is a comparison diagram before and after color homogenization processing.
[0034] The following further details the present invention in conjunction with the drawings and specific embodiments. Detailed Embodiments
[0035] The following further describes the present invention through specific embodiments.
[0036] See Figure 1 , a double-layer multi-scale diffusion color homogenization method based on fusion channel contrast attention, mainly used to solve the problems of color difference and texture distortion caused by imaging environment, sensor limitations, and lighting changes during the acquisition of remote sensing images. The constructed double-layer multi-scale diffusion color homogenization model is used to efficiently restore the color and reconstruct the texture of remote sensing images. The processed images have significant improvements in visual effects and texture fidelity. The method of the present invention specifically includes the following steps:
[0037] S1 Obtain the remotely sensed image to be color homogenization processed, and perform cropping to obtain image chunks. In this step, the remotely sensed image to be color homogenization processed is cut into image chunks of a preset size and a preset number. For example, the preset size is 256*256. In other embodiments, the size and overlapping area of the image chunks can be set and adjusted according to requirements, and no specific limitations are made here.
[0038] Figure 4 In the figure, a is an example of the remotely sensed image to be color homogenization processed, which is a screenshot of the image with color difference obtained by the GF-1 satellite and to be color homogenization processed.
[0039] S2 Input the image chunks into a pre-trained double-layer multi-scale diffusion color homogenization model to obtain the color homogenized image chunks.
[0040] The core innovation of the double-layer multi-scale diffusion color homogenization model in the embodiments of the present invention is to combine the fusion channel contrast attention mechanism with the double-layer multi-scale U-shaped network, realizing the capture and optimization of multi-scale features of the image, and overcoming the limitation of the insufficient restoration ability of traditional methods in scenarios with excessive color difference.
[0041] Among them, referring to Figure 2 , the double-layer multi-scale diffusion color homogenization model includes a first processing module, a noise level embedding module, a first U-shaped network, a supervised attention module, a second processing module, a second U-shaped network, and an adaptive feature aggregation module, etc.
[0042] The first processing module is used to sequentially perform feature dimension reduction and enhanced feature expression on the image chunks to be color homogenized and then input them into the first U-shaped network; the noise level embedding module uses a multi-layer perception mechanism (MLP) to train learnable noise and embeds it into the first U-shaped network for training; the first U-shaped network trains the input image to obtain a feature map; the supervised attention module is used to generate an attention map to suppress the features with less information in the current stage, and only allows useful features to propagate to the next stage. The supervised attention module processes the feature map output by the first U-shaped network to obtain the output result map of the first stage; the second processing module is used to sequentially perform feature learning and enhanced feature expression on the image chunks to be color homogenized, splice them with the output result map of the first stage, and then input them into the second U-shaped network to obtain the output result map of the second stage. The adaptive feature aggregation module aggregates the features of the output result map of the second stage and the output result map of the first stage to obtain the color homogenized image chunks.
[0043] Specifically, the first processing module includes a first convolutional module and a first feature enhancement module. The first convolutional module is used to perform feature dimensionality reduction on the image patches to be color-equalized and output feature maps to the first feature enhancement module. The convolutional kernel of the operation unit of the first convolutional module is a 3*3 convolution, the number of convolutional kernels is 64, and the stride is 1. The first feature enhancement module is a module for enhancing feature expression, focusing on capturing channel-level attention information of the input features. The first feature enhancement module includes two convolutions, a PReLU activation function, and a channel attention layer; the channel attention layer includes a global average pooling that compresses the spatial dimension of the feature map to 1*1, only retaining the channel dimension information, and then uses two convolutions with a convolutional kernel of 1*1 and a padding size of 1 for dimensionality reduction and dimensionality increase. Finally, after using the Sigmoid activation function, the feature map is output to the first U-shaped network. The noise-level embedding module uses a multi-layer perceptron to generate a noise with 3 channels and a size of 256*256.
[0044] The first U-shaped network includes five groups of encoding structures connected in sequence and five groups of decoding structures connected in sequence through an upsampling structure. Feature maps are extracted through the five groups of encoding structures, and the extracted feature map information is upsampled through the five groups of structures to realize the training and generation of images.
[0045] In this embodiment, for the five groups of encoding structures, specifically, they are as follows: The first group of encoding structures includes a convolutional layer. Preferably, the convolutional kernel of the convolutional operation unit of this convolutional layer is a 3*3 convolution, the number of convolutional kernels is 64, the stride is 1, and the padding size is 1. The first group of encoding structures outputs 64 feature maps with a size of 128*128. The second group of encoding structures to the fifth group of encoding structures each include a convolutional layer connected in sequence. Preferably, the convolutional kernel of the convolutional operation unit of the second group of encoding structures is a 3*3 convolution, the number of convolutional kernels is 128, the stride is 1, and the padding size is 1. The second group of structures outputs 128 feature maps with a size of 64*64. The convolutional kernel of the convolutional operation unit of the third group of encoding structures is a 3*3 convolution, the number of convolutional kernels is 256, the stride is 1, and the padding size is 1. The third group of structures outputs 256 feature maps with a size of 32*32. The convolutional kernel of the convolutional operation unit of the fourth group of encoding structures is a 3*3 convolution, the number of convolutional kernels is 512, the stride is 1, and the padding size is 1. The fourth group of structures outputs 512 feature maps with a size of 16*16. The convolutional kernel of the convolutional operation unit of the fifth group of encoding structures is a 3*3 convolution, the number of convolutional kernels is 1024, the stride is 1, and the padding size is 1. The fifth group of structures outputs 1024 feature maps with a size of 8*8.
[0046] In this embodiment, for the five groups of decoding structures, specifically, they are as follows:
[0047] The first group of decoding structures includes an upsampling module connected to the fifth group of encoding structures. Preferably, the upsampling module of the first group of decoding structures uses a transposed convolution with a 3*3 convolution kernel, 512 convolution kernels, a stride of 1, and a padding size of 1. The first group of structures outputs 512 feature maps with a size of 16*16.
[0048] The second to fifth groups of decoding structures each include a transposed convolution layer connected in sequence. Preferably, the convolution kernels of the transposed convolution layer and the convolution operation unit in the second group of decoding structures are 3*3 convolutions, with 256 convolution kernels, a stride of 1, and a padding size of 1. The second group of decoding structures outputs 256 feature maps with a size of 32*32. The convolution kernels of the transposed convolution layer and the convolution operation unit in the third group of decoding structures are 3*3 convolutions, with 128 convolution kernels, a stride of 1, and a padding size of 1. The third group of decoding structures outputs 128 feature maps with a size of 64*64. The convolution kernels of the transposed convolution layer and the convolution operation unit in the fourth group of decoding structures are 3*3 convolutions, with 64 convolution kernels, a stride of 1, and a padding size of 1. The fourth group of decoding structures outputs 64 feature maps with a size of 128*128. The convolution kernels of the transposed convolution layer and the convolution operation unit in the fifth group of decoding structures are 3*3 convolutions, with 3 convolution kernels, a stride of 1, and a padding size of 1. The fifth group of decoding structures outputs 3 feature maps with a size of 256*256. Such a method enables each layer structure of the decoder to obtain the feature map information features output by the previous layer structure of the decoder, effectively restoring details such as the color and texture of the image.
[0049] The supervised attention module includes three convolution layers and a Sigmoid activation function; among them, the convolution operation unit of the first convolution layer has a 1*1 convolution kernel, 64 convolution kernels, and a stride of 1; the convolution operation unit of the second convolution layer has a 1*1 convolution kernel, 6 convolution kernels, and a stride of 1; the convolution operation unit of the third convolution layer has a 1*1 convolution kernel, 64 convolution kernels, and a stride of 1. After the feature maps output by the first U-shaped network are processed by the supervised attention module, 64 feature maps with a size of 256*256 are output, denoted as the output result diagram of the first stage, and are used for splicing with the input features of the second stage.
[0050] The second processing module consists of a second convolutional module and a second feature enhancement module. The second convolutional module is used for feature learning of the image patches to be color-equalized. The convolutional kernel of the convolutional operation unit in the second convolutional module is a 3*3 convolution, the number of convolutional kernels is 64, and the stride is 1. The feature map output by the second convolutional module is fed into the second feature enhancement module. The second feature enhancement module includes two convolutions, a PReLU activation function, and a channel attention layer; the channel attention layer includes a global average pooling, which compresses the spatial dimension of the feature map to 1*1, only retaining the channel dimension information, and then uses two convolutions with a kernel size of 1*1 and a padding size of 1 for dimensionality reduction and dimensionality increase, and finally outputs after using the Sigmoid activation function. Then, the feature map processed by the second feature enhancement module and the output result map of the first stage are concatenated and then input into the second U-shaped network.
[0051] In this embodiment, the structure of the second U-shaped network is the same as that of the first U-shaped network. That is, the second U-shaped network includes five groups of encoding structures connected in sequence and five groups of decoding structures connected in sequence through an upsampling structure. Feature maps are extracted through the five groups of encoding structures, and the extracted feature map information is upsampled through the five groups of structures.
[0052] The second U-shaped network finally outputs 3 feature maps with a size of 256*256 as the output result map of the second stage. The output result map of the second stage and the output result map of the first stage are concatenated and then fed into the adaptive feature aggregation module together.
[0053] The adaptive feature aggregation module includes a global average pooling, a dimensionality reduction convolution, a Softmax activation function, and global feature aggregation; the global average pooling is used to compress the image into a 1*1 feature map; the convolutional kernel of the convolutional operation unit of a dimensionality reduction convolution is a 1*1 convolution, and the number of convolutional kernels is 64, and its function is to generate attention weights for each feature map; global feature aggregation is used to perform weighted summation of the two input features in the channel dimension to obtain the final fused feature, and an RGB three-channel color remote sensing image is output.
[0054] The training process of the double-layer multi-scale diffusion color equalization model in this embodiment is as Figure 4 shown, specifically including steps A1 to A6.
[0055] A1. Select training images and their corresponding original complete images. The training images are made into a dataset and input into the generator. Specifically, there are 2 sub-images in the training images, namely the training image with color difference and the color difference training image after histogram equalization, and the image size is 256*256.
[0056] A2. The encoder in the generator learns the regions with excessive color difference in the training images and extracts the global and local information features of the training images.
[0057] A3. In the generator, by simulating the gradual diffusion and denoising process of noise, the dependence relationship between the channels of the feature map is captured.
[0058] A4. The decoder in the generator effectively fuses the detailed features of the image to generate a preliminary color uniformity result image.
[0059] A5. Receive the preliminary color uniformity result image output by the generator, combine it with the original target complete image, calculate the noise estimation loss and the singular value loss, and return to the generator.
[0060] In this embodiment, step A5 further includes: passing the 256*256-sized semantic information result map output by the decoder through an adaptive feature aggregation module to generate an RGB three-channel 256*256-sized picture.
[0061] S3. Calculate the noise estimation loss and the singular value loss for the color-uniformed image patches to optimize the model parameters. This step specifically includes the following:
[0062] S31. Calculate the noise estimation loss between the color-uniformed result map obtained in step S2 and the label. The main loss is calculated based on the L2 norm, that is, the error between the output of the model and the actual noise.
[0063] S32. Calculate the singular value loss between the color-uniformed result map obtained in step S2 and the label. It maintains the structural consistency of the image by aligning the ranks of similar patches.
[0064] Among them, the noise estimation loss and the singular value loss are used as the joint loss to jointly optimize the model parameters.
[0065] S4. Stitch the color-uniformed image patches to obtain the complete remotely sensed image after color uniformity restoration.
[0066] Figure 4 b in [reference] is the image processed by the method of the present invention. Figure 4 c in [reference] is the true target image corresponding to the image to be processed.
[0067] Based on this, the present invention also proposes a double-layer multi-scale diffusion color uniformity device based on fused channel contrast attention for implementing the above-mentioned double-layer multi-scale diffusion color uniformity method based on fused channel contrast attention, including:
[0068] A patch module that acquires the remotely sensed image to be color-uniformed and performs cropping to obtain image patches. This patch module is used to execute step S1 of the above method.
[0069] The color homogenization module inputs the image chunks into a pre-trained double-layer multi-scale diffusion color homogenization model to obtain the color-homogenized image chunks, calculates the noise estimation loss and the singular value loss for the color-homogenized image chunks, and optimizes the model parameters. This color homogenization module is used to execute steps S2 - S3 of the above method.
[0070] The splicing module splices the color-homogenized image chunks to obtain the complete remote sensing image after color homogenization restoration. This splicing module is used to execute step S4 of the above method.
[0071] It can be understood that the color homogenization device of the present invention can be an electronic device with computing performance such as a portable notebook computer, a desktop computer, a server, a smart phone, or a tablet computer, etc.
[0072] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0073] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0074] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure.
[0075] The above are only the specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made using this concept to the present invention shall fall within the scope of infringement of the protection of the present invention.
Claims
1. A dual-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention, characterized in that: include: Acquire the remote sensing image to be processed for uniform color, and crop it to obtain image blocks; Inputting the image blocks into a pre-trained double-layer multi-scale diffusion color-homogenizing model to obtain color-homogenized image blocks; Calculate the noise estimation loss and singular value loss for the image blocks after color uniformity, and optimize the model parameters; The image blocks after uniform color are stitched together to obtain a complete remote sensing image after uniform color restoration.
2. A double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 1, characterized in that: The double-layer multi-scale diffusion color model includes a first processing module, a noise level embedding module, a first U-shaped network, a supervised attention module, a second processing module, a second U-shaped network and an adaptive feature aggregation module; The first processing module is used to perform feature dimension reduction and feature expression enhancement on the image blocks to be uniformly colored and then input them into the first U-shaped network; the noise level embedding module is used to train learnable noise and embed it into the first U-shaped network; the first U-shaped network trains the input image to obtain a feature map and inputs it into the supervised attention module; the supervised attention module is used to generate an attention map to suppress features with less information in the current stage to obtain a first stage output result map; The second processing module is used to perform feature learning and enhanced feature expression on the image blocks to be color-uniformed in sequence, and splice the blocks with the output result graph of the first stage, and then input the blocks into the second U-shaped network to obtain the output result graph of the second stage; The adaptive feature aggregation module performs feature aggregation on the second stage output result image and the first stage output result image to obtain the uniformly colored image blocks.
3. A double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 2, characterized in that: The first processing module includes a first convolution module and a first feature enhancement module. The convolution kernel of the operating unit of the first convolution module is a 3*3 convolution, the number of convolution kernels is 64, and the step size is 1; the first feature enhancement module contains two convolutions, a PRelu activation function and a channel attention layer; the channel attention layer includes a global average pooling, which compresses the spatial dimension of the feature map to 1*1, retaining only the channel dimension information, and then uses two convolution kernels of 1*1 and a padding size of 1 to reduce and increase the dimension, and finally uses the Sigmoid activation function and outputs.
4. A double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 2, characterized in that: The second processing module includes a second convolution module and a second feature enhancement module. The convolution kernel of the convolution operation unit of the second convolution module is a 3*3 convolution, the number of convolution kernels is 64, and the step size is 1; the second feature enhancement module contains two convolutions, a PRelu activation function and a channel attention layer; the channel attention layer includes a global average pooling, which compresses the spatial dimension of the feature map to 1*1, retaining only the channel dimension information, and then uses two convolution kernels of 1*1 and a padding size of 1 to perform dimensionality reduction and dimensionality increase, and finally uses the Sigmoid activation function and outputs.
5. The double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 1, characterized in that: The adaptive feature aggregation module includes a global average pooling, a dimensionality reduction convolution, a Softmax activation function and global feature aggregation; the image is compressed into a 1*1 feature map using the global average pooling; the convolution kernel of the convolution operation unit using a dimensionality reduction convolution is a 1*1 convolution, and the number of convolution kernels is 64, which is used to generate attention weights for each feature map; Global feature aggregation is used to perform weighted summation of the two input features in the channel dimension to obtain the final fused feature.
6. A double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 2, characterized in that: The noise level embedding module uses a multi-layer perceptron to generate a noise with 3 channels and a size of 256*256.
7. A double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 2, characterized in that: The supervised attention module includes three convolutional layers and a Sigmoid activation function; wherein, the convolution kernel of the convolution operation unit of the first convolution layer is 1*1 convolution, the number of convolution kernels is 64, and the step length is 1; the convolution kernel of the convolution operation unit of the second convolution layer is 1*1 convolution, the number of convolution kernels is 6, and the step length is 1; the convolution kernel of the convolution operation unit of the third convolution layer is 1*1 convolution, the number of convolution kernels is 64, and the step length is 1.
8. A double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 2, characterized in that: The first U-type network includes five groups of encoding structures connected in sequence and five groups of decoding structures connected in sequence through upsampling structures. Feature graphs are extracted through the five groups of encoding structures, and upsampling operations are performed on the extracted feature graph information through the five groups of structures.
9. A double-layer multi-scale diffusion color uniformity method based on fusion channel contrast attention as claimed in claim 2, characterized in that: The second U-shaped network includes five groups of encoding structures connected in sequence and five groups of decoding structures connected in sequence through upsampling structures. Feature graphs are extracted through the five groups of encoding structures, and upsampling operations are performed on the extracted feature graph information through the five groups of structures.
10. A dual-layer multi-scale diffusion color-leveling device based on fusion channel contrast attention, characterized in that: include: A cutting module obtains the remote sensing image to be processed for uniform color and cuts it to obtain image blocks; A color-leveling module inputs the image blocks into a pre-trained double-layer multi-scale diffusion color-leveling model to obtain color-leveled image blocks; calculates noise estimation loss and singular value loss for the color-leveled image blocks, and optimizes model parameters; The stitching module stitches the image blocks after uniform color restoration to obtain a complete remote sensing image after uniform color restoration.