Panchromatic sharpening method based on double-granularity semantic guide diffusion
By constructing a full-color sharpening method based on two-particle semantic-guided diffusion, using the scene and land object information of remote sensing images, dynamically selecting experts to adjust the convolutional layer characteristics, solving the scene dependence problem of full-color sharpening technology in different scenarios, and achieving a more accurate full-color sharpening effect.
Patent Information
- Application Number
- CN202510307100.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-16
- Publication Date
- 2025-07-18
AI Technical Summary
The existing full-color sharpening technology shows scene dependence in different remote sensing scenarios, limiting its actual application scope and making it difficult to achieve generalization.
Using a two-particle semantic guided diffusion method, the scene information and landform information of the remote sensing image are used to build a denoising network through the semantic guided expert mixing module and the Transformer module, dynamically select experts to adjust the convolutional layer channel characteristics, and realize joint training across multiple remote sensing data sets.
It effectively alleviates the problem of domain differences caused by different scenarios and land objects, achieves a more accurate full-color sharpening effect, and improves the generalization ability of the model in different scenarios.
Smart Images

Figure CN120339115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a panchromatic sharpening method based on dual-granularity semantic-guided diffusion. Background Art
[0002] The multi-spectral and panchromatic image fusion technology, also known as panchromatic sharpening, its core idea is to use signal processing and machine learning methods to combine the low-spatial-resolution multi-spectral image (Multispectral, MS) and the high-spatial-resolution panchromatic image (Panchromatic, PAN) obtained from the same area but different sensors, to generate a high-spatial-resolution multi-spectral image (High-Resolution Multispectral, HRMS) that not only contains rich spatial details but also retains the original spectral characteristics. It can not only accurately capture the geometric features such as the size and shape of the target ground object, but also finely reflect its internal physical properties, greatly improving the application value of remote sensing images and greatly enhancing the capabilities of applications such as change detection, target detection, and land classification.
[0003] Since remote sensing satellite images are acquired from a high-altitude vertical perspective, covering a wide range of types of scenes and ground objects, including cities, vegetation, water bodies, mountains, etc., the panchromatic sharpening technology based on deep learning often shows "scene dependence", that is, it shows a preference for specific types of scenes, and may encounter challenges when applied to new scenes, restricting the actual application scope of the panchromatic sharpening technology.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The present invention provides a panchromatic sharpening method based on dual-granularity semantic-guided diffusion, which effectively utilizes the scene information and ground object information of remote sensing images to provide semantic guidance for general panchromatic sharpening, so as to improve the generalization ability of the model in different scenes.
[0006] Other features and advantages of the present invention will become apparent through the following detailed description, or be learned in part through the practice of the present invention.
[0007] According to the first aspect of the present invention, there is provided a panchromatic sharpening method based on dual-granularity semantic-guided diffusion, the method comprising:
[0008] Obtain the original multi-spectral image and the panchromatic image that are mutually registered, and after preprocessing the original multi-spectral image and the panchromatic image, divide them into a training set and a test set;
[0009] The multi-spectral image, panchromatic image, and generated noisy residual image in the training set are stitched together as the input image, and the input image is input into a pre-trained denoising network, which includes a first convolutional layer, multiple Transformer modules, and a second convolutional layer; among them, the first convolutional layer is used to extract features, and the multiple Transformer modules are used to encode and decode the features, and the second convolutional layer outputs a predicted residual image without noise.
[0010] The test set is input into the trained denoising network to obtain a predicted residual image, and the predicted residual image is added to the multi-spectral image in the test set to obtain a predicted fusion image.
[0011] In some exemplary embodiments, the preprocessing includes:
[0012] Normalize the original multi-spectral image and panchromatic image;
[0013] Process the original multi-spectral image and panchromatic image according to the Wald protocol;
[0014] Use bicubic interpolation algorithm to upsample the original multi-spectral image by four times.
[0015] In some exemplary embodiments, the noisy residual image is obtained by subtracting the reference image and the multi-spectral image, denoted as X0 = Y - M. During the forward noise addition process of the diffusion model, Gaussian noise is gradually added to X0 to obtain a series of noisy residual images X1, X2, …, X T ; where the reference image is the original multi-spectral image.
[0016] In some exemplary embodiments, the Transformer module is composed of a 3D multi-head transposed attention module and a semantic-guided mixture-of-experts module cascaded in sequence.
[0017] In some exemplary embodiments, the semantic-guided mixture-of-experts module includes Geochat, a CLIP text encoder, an MLP, and a 3D convolutional layer;
[0018] The Geochat is used to obtain scene classification and ground description from the multi-spectral image and panchromatic image respectively; the scene classification and ground description are input into the CLIP text encoder to output corresponding text encodings respectively; the text encodings are input into the MLP layer and a SoftMax operation is performed to obtain the potential mixture weights e s of the scene category and e g ;
[0019] The mixture weights e s 、e gInput into the Linear layer and the ReLU activation function layer to generate gating scores for selecting and adjusting the channel weights of each layer of the 3D convolutional layer.
[0020] In some exemplary embodiments, the 3D multi-head transposed attention module performs self-attention operations in the feature channel dimension and introduces time step information; starting from a layer-normalized tensor X ∈ R N×C×H×W At the beginning, time steps t ~ 1, …, T are encoded as one-dimensional vectors and then passed through two Linear transformation layers, and element-wise multiplication and element-wise addition are performed with the tensor X in sequence.
[0021] In some exemplary embodiments, the training process of the denoising network includes:
[0022] Starting from the clean residual map X0 = Y - M, randomly sample a time step t and noise Generate the noisy image X t , and optimize the parameters θ of the denoising network f θ (X t , P, M, e s , e g ) using the following reconstruction loss:
[0023]
[0024] where k represents the number of datasets;
[0025] In addition, add sparse regularization to the gating output to promote the diversity of potential mixing weights and encourage the sparsity of expert selection, and define the final loss function as:
[0026]
[0027] where μ is used to control the sparsity of the gating scores, and G l (e s ), G l (e g ) are the gating scores.
[0028] According to the second aspect of the present invention, there is provided a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the panchromatic sharpening method based on dual-granularity semantic-guided diffusion described in the first aspect above.
[0029] According to the third aspect of the present invention, there is provided a computer program product having a computer program stored thereon, and when the computer program is executed by a processor, it implements the panchromatic sharpening method based on dual-granularity semantic-guided diffusion described in the first aspect above.
[0030] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:
[0031] a processor; and
[0032] a memory for storing executable instructions of the processor;
[0033] wherein, the processor is configured to implement the pan-sharpening method based on dual-granularity semantic-guided diffusion described in the first aspect above when executing the executable instructions.
[0034] The pan-sharpening method based on dual-granularity semantic-guided diffusion provided by the embodiments of the present invention utilizes a vision-language multi-modal model (GeoChat) optimized for earth sciences to extract dual-granularity semantic information of scene descriptions and ground object features. Through a semantic-guided expert mixture module, the most suitable expert for the current input is dynamically selected to adjust the channel features of the convolutional layer. This method constructs an efficient denoising network, supports joint training across multiple remote sensing datasets, obtains a general knowledge representation, and effectively distinguishes different scenes and ground object features, thereby alleviating the domain difference problem and achieving accurate pan-sharpening.
[0035] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0037] Figure 1 Schematic diagram of the semantic-guided expert mixture module proposed by the present invention;
[0038] Figure 2 Schematic diagram of the detailed architecture of the denoising network proposed by the present invention;
[0039] Figure 3 Overall inference process framework diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0041] In addition, the drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0042] In view of the drawbacks and deficiencies of the prior art, a pan-sharpening method based on dual-granularity semantic-guided diffusion is provided in this example embodiment. With the help of GeoChat, a vision-language multi-modal model (VLM) designed specifically for earth sciences and optimized by a remote sensing multi-modal instruction dataset, it is used to extract dual-granularity semantic information covering scene descriptions and ground object features. This information is input as a condition into the diffusion model to provide high-level guidance. Specifically, the present invention introduces a semantic-guided mixture-of-experts module and designs a dynamic activation mechanism to select the expert most suitable for the current input, thereby adjusting the channel features of the input to the convolutional layer. This helps to construct an efficient denoising network architecture that supports joint training across multiple remote sensing datasets to obtain a general knowledge representation, while being able to effectively identify and distinguish various scenes and ground object features in different satellite datasets. By fully utilizing the semantic priors provided by multiple remote sensing datasets and the vision-language multi-modal model, the present invention effectively alleviates the domain difference problem caused by different scenes and ground objects and achieves a more accurate pan-sharpening effect.
[0043] The pan-sharpening method based on dual-granularity semantic-guided diffusion may specifically include the following steps:
[0044] Step 1: Dataset preparation;
[0045] Intercept image patches from paired and registered large-scale remote sensing multi-spectral MS images and panchromatic images PAN in the order from left to right and from top to bottom, and divide these image patches into a training set, a validation set, and a test set;
[0046] First, normalize the training set, validation set, and test set; then process the image patches in the training set, validation set, and test set according to the Wald protocol, and then use the processed image patches as the input of the model; among them, the Wald protocol processing means first filtering the original MS and PAN images with a Gaussian smoothing kernel of size 5×5, and then downsampling to 1 / 4 of the original spatial resolution.
[0047] Use the original MS image patch as the reference image.
[0048] Exemplarily, the data comes from satellite sensors of QuickBird (QB), WorldView-4 (WV-4), and WorldView-2 (WV-2). For QB data, the spatial resolution of PAN is 0.6 meters, the spatial resolution of MS is 2.4 meters, and MS contains 4 spectral bands: blue, green, red, and near-infrared bands. For WV-4 data, the spatial resolution of PAN is 0.3 meters, the spatial resolution of MS is 1.2 meters, and MS contains 4 spectral bands: blue, green, red, and near-infrared bands. For WV-2 data, the spatial resolution of PAN is 0.5 meters, the spatial resolution of MS is 2 meters, and MS contains 8 spectral bands: cyan, yellow, infrared, blue, green, red, near-infrared 1 band, and near-infrared 2 band. The spatial resolution ratio between the MS and PAN images in the three datasets is 4.
[0049] Specifically, the size of the image patches in the training set and validation set is 256×256 (PAN) / 64×64×4 (MS), the size of the image patches in the test set is 1024×1024 (PAN) / 256×256×4 (MS), and the ratio of the training, validation, and test data volumes is 8:1:1. Since there is no reference image, process the MS image patches and PAN image patches in the training set, validation set, and test set according to the Wald protocol, and then use the processed images as the input of the network. The processed PAN image patch is denoted as P, the processed MS image patch is upsampled four times using the bicubic interpolation algorithm and denoted as M, and the original MS image is used as the reference image and denoted as Y.
[0050] The above three satellite datasets form a joint dataset for model training.
[0051] Step 2: Construction of the semantic-guided mixture of experts module;
[0052] As Figure 1As shown in the figure, the semantic-guided mixture-of-experts module includes a pre-trained Geochat, a pre-trained CLIP text encoder, an MLP, and two 3D convolutional layers; the MLP is composed of a Linear linear transformation layer, a LeakyReLU activation function, and a Linear linear transformation layer in sequence; this module regards the input channels of the convolutional layer as experts, and dynamically activates the corresponding experts as the input information of the convolutional layer through the text encoding generated by the semantic prior, and finally outputs the processed features.
[0053] Step 2-1: Generation of text prior information;
[0054] As Figure 1 shown, the semantic-guided mixture-of-experts module uses a pre-trained multi-modal large model GeoChat with frozen parameters to extract semantic priors from remote sensing images. This module includes two branches, a scene branch and a ground branch, which receive a panchromatic image P and a multi-spectral image M as input images respectively; among them, the upsampled multi-spectral image M refers to the MS image patches being upsampled to the spatial resolution of the PAN image patches by bicubic interpolation.
[0055] In the scene branch, the multi-spectral image M is input into GeoChat and the prompt 'Classify the image in one of the following classes. Classes: River, Agricultural, Buildings, Freeway, Forest, Harbor, Mountains. Answer in one word or a short phrase.' is given. The text response output by the GeoChat model is used as a coarse-grained scene description; in the ground branch, the panchromatic image PAN is input into GeoChat and the prompt '[Grounding] Describe the image in detail.' is given. The text response output by the GeoChat model is used as a fine-grained ground description; the scene description and the ground description are input into the pre-trained CLIP text encoder with frozen parameters and the text encoding is output. The text encoding is input into the MLP layer and a SoftMax operation is performed to obtain the latent mixture weights in the predefined set of latent experts, denoted as e s for the scene category and e g for the ground description, defined as follows:
[0056] e s = SoftMax(MLP(CLIP(L scene ))) (1)
[0057] e g = SoftMax(MLP(CLIP(L grounding))) (2)
[0058] Among them, L scene and L grounding respectively represent the scenario category of GeoChat and the ground description response; the potential mixing weights e s and e g are input into the Linear linear transformation layer and the ReLU activation function layer to select and adjust the channel weights of each layer; the gating score of the l-th layer of the denoising network is:
[0059]
[0060] Among them, and are the Linear linear transformation layers, and the ReLU activation function is used to map the mixing weights to sparse, layer-specific gating scores, denoted as G l (e s ) and G l (e g ).
[0061] Step 2-2: Construction of the sparse routing mechanism;
[0062] As Figure 1 shown, the sparse routing mechanism regards each input feature channel of the 3D convolutional layer as an activatable expert, and the activation of the experts at a specific layer is determined by the above-mentioned sparse gating scores; the gating scores G l (e s ) and G l (e g ) obtained from the scenario category and the ground description are integrated into the routing process to dynamically adjust the output of the semantic guidance hybrid expert module; let x represent the input features of the semantic guidance hybrid expert module, and the routing strategy of the l-th layer can be expressed as:
[0063]
[0064] Among them, C in and C out are the numbers of input and output channels respectively. "*" represents the 3D convolution operation, K is the convolution kernel, F s and F g respectively represent the feature output guided by the scenario and the feature output guided by the ground information; by applying the gating scores G l (e s ) and G l (e g ) to adjust the input channels, the expertise is balanced in the expert selection, and shared expert isolation is adopted to enhance the specialization ability of individual experts and reduce the knowledge redundancy between the activated experts; the so-called shared expert isolation is to always set the gating score Gl (e s ) and G l (e g ) The top 10% of the elements are 1.0.
[0065] Step 3: Construction of the 3D multi-head attention module;
[0066] As Figure 2 shown, the 3D multi-head attention module performs self-attention operations in the feature channel dimension and introduces the time step information; starting from a layer-normalized tensor X ∈ R N×C×H×W , the time steps t ~ 1,…, T are encoded as one-dimensional vectors and then passed through two Linear linear transformation layers, and element-wise multiplication and element-wise addition are performed with the tensor X in sequence; the time step encoding algorithm comes from the sine-cosine encoding algorithm of Transformer; it is achieved by applying 1×1×1 convolution to aggregate pixel-level cross-channel context and then using 3×3×3 depth convolution to encode channel-level spatial context, thus obtaining and where, represents 1×1×1 pointwise convolution, while represents 3×3×3 depth convolution. Next, the query and key projections are reshaped so that their dot product operation produces a transposed attention map A of shape (R N×N ). Generally speaking, the 3D multi-head attention process is defined as follows:
[0067]
[0068] where, X and are the input and output feature maps respectively; and matrices are obtained after reshaping the tensor of the original size R N×C×H×W ; α is a learnable scaling parameter used to control the size of the dot product between and before applying the Softmax function.
[0069] Step 4: Construction of the denoising model;
[0070] Step 4-1: Construction of the input of the denoising model;
[0071] As Figure 3 shown, the denoising model expands the panchromatic image P, the multispectral image M, and the noisy residual image X t to the shape R 1 ×C×H×W and then concatenates them as the image input I t ∈ R 3×C×H×W; The scene category and ground description are used as text inputs; among them, the residual map is obtained by subtracting the reference image Y from the multispectral image M, expressed as X0 = Y - M. During the forward noise addition process of the diffusion model, Gaussian noise is gradually added to X0 to obtain a series of noisy residual images X1, X2, …, X T , where X t is the noisy residual image at time step t, and the forward noise addition process is defined as follows:
[0072]
[0073] where q represents the probability distribution of the forward noise addition process; the noise scale α t adopts predefined fixed parameters, with a value range of (0, 1), and α is initialized using the cosine noise strategy t ; the cosine noise strategy is:
[0074]
[0075] where s = 0.008; the noisy images obtained through formulas (9) and (10) follow a Gaussian distribution. If the total number of noise addition steps T is large enough, X T will follow a standard Gaussian distribution I represents a matrix of all 1s, represents a Gaussian distribution; the text input gating scores are obtained through the semantic-guided mixture-of-experts module in step 2 to obtain the gating scores G l (e s ) and G l (e g );
[0076] Step 4-2: Network construction of the denoising model;
[0077] As Figure 2 shown, the denoising model first uses a 3×3×3 convolutional layer to extract features F0 ∈ R N×C×H×W ; F0 passes through an encoder-decoder architecture containing multiple Transformer modules and outputs a noise-free predicted residual image through another 3×3×3 convolutional layer The prediction result will be used for the reverse inference of the diffusion model; specifically, the Transformer block is sequentially cascaded by a 3D multi-head transposed attention module in step 3 and a semantic-guided mixture-of-experts module in step 2;
[0078] Step 5: Model training;
[0079] The semantic-guided routing diffusion model prepares the training set by combining K satellite datasets Each dataset consists of composed of, C krepresents the number of images in the k-th dataset; train the denoising model f θ (X t , P, M, e s , e g ), starting from the clean residual map X0 = Y - M, randomly sample a time step t and noise Generate the noisy image X based on Equation (9) t , where I is the identity matrix, and optimize the denoising model parameters θ using the following reconstruction loss:
[0080]
[0081] In one embodiment, the number of satellite datasets K = 3.
[0082] In addition, add sparse regularization to the gated output to promote the diversity of potential mixing weights and encourage the sparsity of expert selection. Define the final loss function as:
[0083]
[0084] where μ is used to control the sparsity of the gated scores; preferably, μ = 10 -6 .
[0085] Step 6: Model inference;
[0086] As Figure 3 shown, after the denoising network training converges, sample an image from the standard Gaussian distribution and denote it as Based on the images P and M to be fused, obtain the potential mixing weights e s , e g through Step 2, and perform T-step iteration using the following formula. When t = 0, the iteration terminates, and finally obtain the predicted residual image Predicted residual image Finally, add it to M to obtain the predicted fused image, denoted as
[0087]
[0088] where,
[0089] Example 1:
[0090] (1) Dataset preparation:
[0091] Use panchromatic images and multispectral images with a width-to-height ratio of 4:1 and registered with each other, and perform the following processing:
[0092] ① Read the image, divide the original image into two parts, which are used as training data image and test data image respectively. The division principle is that the widths of the two parts are the same and the height ratio is 9:1. Do this processing for both PAN and MS;
[0093] ② For the training data part, intercept the corresponding image patches at the corresponding positions of the matching PAN and MS training images from left to right and top to bottom. Among them, the size of the PAN image patch is 256×256, and the size of the MS image patch is 64×64×4 (4 is the number of channels, when the number of MS channels is 8, it can be changed to 8 accordingly). The test data part is constructed in a similar way, where the size of the PAN image patch is 1024×1024, and the size of the MS image patch is 256×256×4.
[0094] ③ Randomly divide 1 / 9 of the training data part as the validation set data.
[0095] So far, the training set, validation set, and test set data are obtained. For QB, the training set contains 6943 pairs of images, the validation set contains 743 pairs of images, and the test set contains 156 pairs of images; for WV-4, the training set contains 7166 pairs of images, the validation set contains 772 pairs of images, and the test set contains 271 pairs of images; for WV-2, the training set contains 9641 pairs of images, the validation set contains 945 pairs of images, and the test set contains 136 pairs of images.
[0096] ④ When processing according to the Wald protocol, the PAN image and the MS image are Gaussian blurred with a Gaussian kernel of 5×5 and a standard deviation of 2, and then downsampled by 4 times to form a new image as the training set. The same operation is performed on the validation set and the test set.
[0097] So far, the dataset preparation steps are completed.
[0098] (2) Network model construction
[0099] The network structure diagram is shown in Figure 3 , and the important parameters for constructing the network include:
[0100] ① The 3D convolutional layer or 3D transposed convolutional layer used in the entire network: the window size of all convolutional layers is 3×3×3; spatial downsampling and spatial upsampling are implemented using the PixelShuffle and PixelUnShuffle algorithms for upsampling or downsampling by a factor of two respectively;
[0101] Figure 3 The output channel numbers of the convolutional layers are 32, 64, 128, 256, 128, 64, 32, 32, 1 respectively; the skip connection occurs in the upsampling module, connecting the input information with the output of the downsampling module in the channel dimension, and then halving the channel number through a 1×1×1 convolutional layer;
[0102] ② The time step encoding algorithm uses the sine-cosine encoding in the literature "Attention is all you need" to encode the time step t into a one-dimensional vector of length 32. This vector passes through a Linear layer, and the vector length is consistent with the input feature channels of the 3D transposed attention module for channel-level operations;
[0103] (3) Network training
[0104] ① Input images: Panchromatic image P with a size of 64×64 (height×width) and upsampled multispectral image M with a size of 64×64×4 (height×width×number of channels). Here, the upsampled multispectral M is obtained by bilinearly interpolating the multispectral image with a size of 16×16×4 by 4 times. The input images to the network are normalized, which means dividing the pixel values of the input images by 2047.0.
[0105] ② Other relevant settings: Use the Adam optimizer to update the parameters. The number of training epochs is set to 200, the batch size for batch training is set to 32, and the initial learning rate is set to 0.0003. The network performance is tested using the validation set at the end of every 50 epochs, and the network parameters with the best performance are saved. The joint dataset sequentially samples batch size data pairs from the training data of different datasets to train the model;
[0106] ③ Stopping condition for training: The loss function of the network reaches the convergence state.
[0107] (4) Network testing
[0108] ① Input images: Panchromatic image P with a size of 256×256 (height×width) and upsampled multispectral image M with a size of 256×256×4 (height×width×number of channels). Here, the upsampled multispectral M is obtained by bilinearly interpolating the multispectral image with a size of 64×64×4 by 4 times.
[0109] ② Load the network parameters of the denoising network with the best performance saved during the training stage. Select a pair of panchromatic image P and multispectral image M from the test set, obtain the scene category and ground object information description based on Geochat, get the dual-grained gating score after encoding, sample an image from Gaussian noise, and perform T-step repeated iterations based on the output of the denoising network to obtain the fusion result. In one embodiment, T = 1000.
[0110] It should be noted that, on the other hand, the present application also provides a storage medium, which may be included in an electronic device; or may exist separately without being assembled into the electronic device. The above storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the methods described in the following embodiments.
[0111] In one embodiment, the present application provides a computer program product, including a computer program, which when executed by a processor implements the steps in the above method embodiments.
[0112] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously in, for example, multiple modules.
[0113] Those skilled in the art will readily think of other embodiments of the present invention after considering the specification and practicing the invention herein. The present application aims to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.
[0114] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A panchromatic sharpening method based on dual-granularity semantic-guided diffusion, characterized in that, The method includes: Obtain the original multi-spectral image and the panchromatic image that are mutually registered, and after preprocessing the original multi-spectral image and the panchromatic image, divide them into a training set and a test set; Stitch the multi-spectral image, the panchromatic image, and the generated noisy residual image in the training set as the input image, and input the input image into a pre-trained denoising network, where the denoising network includes a first convolutional layer, multiple Transformer modules, and a second convolutional layer; among them, the first convolutional layer is used to extract features, the multiple Transformer modules are used to encode and decode the features, and the second convolutional layer outputs a predicted residual image without noise; Input the test set into the trained denoising network to obtain a predicted residual image, and add the predicted residual image to the multi-spectral image in the test set to obtain a predicted fusion image.
2. The pan-sharpening method based on dual-granularity semantic-guided diffusion according to claim 1, wherein The preprocessing includes: Perform normalization processing on the original multi-spectral image and the panchromatic image; Process the original multi-spectral image and the panchromatic image according to the Wald protocol; Use the bicubic interpolation algorithm to upsample the original multi-spectral image by four times.
3. The pan-sharpening method based on dual-granularity semantic-guided diffusion according to claim 1, wherein The noise-added residual image is the subtraction of the reference image and the multispectral image, expressed as X0 = Y - M. During the forward noise-adding process of the diffusion model, Gaussian noise is gradually added to X0 to obtain a series of noise-added residual images X1, X2, …, X T ; where the reference image is the original multispectral image.
4. The pan-sharpening method based on dual-granularity semantic-guided diffusion according to claim 1, characterized in that The Transformer module is sequentially cascaded by a 3D multi-head transposed attention module and a semantic-guided mixture-of-experts module.
5. The pan-sharpening method based on dual-granularity semantic-guided diffusion according to claim 4, characterized in that, The semantic-guided mixture-of-experts module includes Geochat, a CLIP text encoder, an MLP, and a 3D convolutional layer; The Geochat is used to obtain scene classification and ground description from a multispectral image and a panchromatic image respectively; the scene classification and ground description are input into a CLIP text encoder to respectively output corresponding text encodings; the text encodings are input into an MLP layer and a SoftMax operation is performed to obtain the potential mixed weights e of the scene categories s and the potential mixed weights e of the ground description g ; The mixed weight e s and e g are input into the Linear layer and the ReLU activation function layer to generate gating scores for selecting and adjusting the channel weights of each layer of the 3D convolutional layer.
6. The pan-sharpening method based on dual-granularity semantic-guided diffusion according to claim 4, wherein The 3D multi-head transposed attention module performs self-attention operations in the feature channel dimension and introduces time step information; starting from a layer-normalized tensor X ∈ R N×C×H×W At the beginning, the time steps t~1,…,T are encoded as one-dimensional vectors and then passed through two Linear layers, and the element-wise multiplication and element-wise addition are performed with the tensor X in sequence.
7. The pan-sharpening method based on dual-granularity semantic-guided diffusion according to claim 4, wherein The training process of the denoising network includes: Starting from the clean residual map X0 = Y - M, randomly sample a time step t and noise Generate the noise image X t , and optimize the denoising network f using the following reconstruction loss θ (X t , P, M, e s , e g ) with respect to the parameters θ: Where k represents the number of data sets; In addition, add sparse regularization to the gated output to promote the diversity of potential mixing weights and encourage the sparsity of expert selection, and define the final loss function as: Among them, μ is used to control the sparsity of the gating score, and G l (e s ) and G l (e g ) are the gating scores.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the panchromatic sharpening method based on dual-granularity semantic-guided diffusion according to any one of claims 1 to 7.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the panchromatic sharpening method based on dual-granularity semantic-guided diffusion according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Including: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the panchromatic sharpening method based on dual-granularity semantic-guided diffusion according to any one of claims 1 to 7 by executing the executable instructions.