W-transformer-based multispectral image generation method and system, and medium

By using the W-Transformer network to encode features and refine spatial information of multispectral and panchromatic images, the performance degradation of deep learning methods under noise pollution is solved, and multispectral image generation with high spatial resolution and spectral fidelity is achieved, which is applicable to fields such as remote sensing and geological exploration.

CN120997060BActive Publication Date: 2026-02-03HUANTIAN SMART TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511516766.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-02-03
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing deep learning-based panchromatic sharpening methods ignore noise and blur that may occur during image transmission when processing multispectral and panchromatic images, resulting in performance degradation when processing degraded original images.

Method used

A W-Transformer-based multispectral image generation method is adopted. By encoding features of contaminated multispectral and panchromatic images, the W-Transformer network is used to refine the spatial information of multispectral image features based on panchromatic image features at the same scale. Finally, the multispectral image features refined by spatial information at different scales are fused to generate a multispectral image.

Benefits of technology

It effectively improves the performance degradation of deep learning methods when faced with noise pollution, and can simultaneously process large-scale structures and minute details in images, generating multispectral images with improved spatial resolution and excellent spectral fidelity, which are suitable for fields such as remote sensing, geological exploration, environmental monitoring and military reconnaissance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997060B_ABST
    Figure CN120997060B_ABST
Patent Text Reader

Abstract

The application discloses a multispectral image generation method and system based on W-Transformer, and a medium; relates to the technical field of remote sensing image processing; is improved on the basis of the prior art, features of a polluted multispectral image and a panchromatic image are encoded based on a W-Transformer network to obtain multispectral image features and panchromatic image features of multiple different scales; under a preset condition, the multispectral image features are spatially refined based on the panchromatic image features; and the multispectral image features are fused after spatial information refinement under different scales to generate a multispectral image, effectively improving the performance degradation of the existing deep learning method when facing original panchromatic images and multispectral images polluted by noise during the acquisition process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, specifically to a method, system, and medium for generating multispectral images based on W-Transformer. Background Technology

[0002] Multispectral images are of great significance for important applications such as geological exploration, environmental protection, and urban planning. However, although the spectral resolution of multispectral images is constantly improving, their spatial resolution is often limited. With the development of remote sensing technology in recent years, multispectral and high spatial resolution panchromatic images can be acquired from the same imaging platform. Therefore, fusing multispectral (MS) and high spatial resolution panchromatic (PAN) images, also known as panchromatic sharpening, has become one of the main methods for obtaining high spatial resolution multispectral images.

[0003] Current sharpening methods are mainly divided into two categories: traditional fusion methods and deep learning methods. The former includes component substitution, multi-resolution analysis, and variational optimization. Component substitution is the simplest method to solve the sharpening problem. It transforms the multispectral data to another representation domain, replaces the spatial components in the representation domain with panchromatic data, and obtains the sharpened result through inverse transformation. Principal component analysis, band-dependent spatial detail methods, and Gram-Schmidt adaptive methods all belong to component substitution methods. The high fidelity of component substitution methods in spatial information is a significant advantage, but they often suffer from severe spectral distortion. Multi-resolution analysis decomposes the multispectral and panchromatic data into multiple high-frequency and low-frequency coefficients at different spatial scales, and fuses these coefficients together according to different fusion rules to generate a denoised and enhanced image. Methods based on multi-resolution analysis include the Laplacian pyramid method, additive wavelet method, and curve transform method. By injecting high-frequency detail information into the fusion result, multi-resolution analysis-based methods can maintain good spectral information, but may produce some spatial distortion. Variational optimization-based methods are usually based on models of the denoised and enhanced result and the input image. These methods rely on constraints based on prior assumptions to derive energy functionals, transforming denoising and enhancement into an optimization problem. Currently, the most popular variational optimization-based methods include sparse representation, compressed sensing, and gradient prior methods.

[0004] In recent years, deep learning-based panchromatic sharpening methods have become increasingly popular due to their convenience and excellent performance. Some scholars have innovatively proposed using a three-layer convolutional neural network to solve the panchromatic sharpening problem. However, due to the simple network structure, this method cannot fully extract image features. Alternatively, training the proposed panchromatic sharpening network in the high-frequency domain can preserve spatial information. At the same time, residual connections are used to achieve spectral fidelity. Combining the advantages of traditional machine learning, a deep convolutional neural network based on detail injection can be established to preserve spatial information as much as possible, and the convolutional neural network can be used to adjust the extracted image details. Since existing deep learning-based panchromatic sharpening methods usually acquire multispectral and panchromatic images with good imaging quality and ignore noise and blurring that may occur during image transmission, existing methods will experience performance degradation when processing degraded original images. Summary of the Invention

[0005] The technical problem this invention aims to solve is that existing deep learning-based panchromatic sharpening methods typically acquire high-quality multispectral and panchromatic images, ignoring noise and blurring that may occur during image transmission. This leads to performance degradation when processing degraded original images. This invention aims to provide a W-Transformer-based multispectral image generation method, system, and medium. Building upon existing technologies, it improves upon existing methods by using a W-Transformer network to encode features from contaminated multispectral and panchromatic images, obtaining multispectral and panchromatic image features at various scales. At the same scale, it refines the spatial information of multispectral image features based on panchromatic image features, fusing the refined multispectral image features from different scales to generate a multispectral image. This effectively improves the performance degradation of existing deep learning methods when dealing with noise contamination during the acquisition of original panchromatic and multispectral images.

[0006] This invention is achieved through the following technical solution:

[0007] This solution provides a multispectral image generation method based on W-Transformer, including:

[0008] Acquire contaminated multispectral and panchromatic images;

[0009] Input the multispectral and panchromatic images into the W-Transformer network to generate a multispectral image:

[0010] Feature encoding is performed on multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales.

[0011] Under preset conditions, spatial information refinement of multispectral image features is performed based on panchromatic image features; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale.

[0012] Multispectral image features refined by fusing spatial information from different scales;

[0013] Multispectral images are generated based on the features of the fused multispectral images.

[0014] A further optimized solution is that the pollution includes: fuzzy pollution, noise pollution, or fuzzy noise pollution;

[0015] The multispectral image has a first resolution, and the panchromatic image has a second resolution;

[0016] The first resolution is lower than the second resolution.

[0017] A further optimization scheme is that the W-Transformer network includes a first branch and a second branch, which are used to extract features from multispectral images and panchromatic images, respectively;

[0018] Both the first and second branches include four feature extraction stages. The first, second, and third feature extraction stages each include a Transformer block and a patch merging block, and the fourth feature extraction stage includes a Transformer block. The Transformer block divides the input features into fixed-size and non-overlapping patches based on a window multi-head self-attention mechanism, and calculates the self-attention within each window.

[0019] The first feature extraction stage of the first branch has a first spatial resolution, and the first feature extraction stage of the second branch has a second spatial resolution; the first spatial resolution is greater than the second spatial resolution.

[0020] A further optimized solution is that the calculation process for the self-attention includes:

[0021] Calculate the self-attention within each window using the following formula;

[0022] ;

[0023] Where Q, K, and V represent the query matrix, key matrix, and value matrix, d represents the dimension of the query matrix and key matrix, and B represents the offset, B∈M 2 ×M 2 M 2 This indicates the number of tiles in the window; SoftMax() represents the normalization function.

[0024] A further optimized solution is that, under preset conditions, the spatial information refinement of multispectral image features based on panchromatic image features includes the following method:

[0025] Obtain the multispectral image features extracted in the i-th feature extraction stage. and panchromatic image features ;

[0026] For the multispectral image features and panchromatic image features The first multispectral image features are obtained after flattening and shaping. and features of the first panchromatic image Among them, the first multispectral image features The height is greater than the features of the first panchromatic image. Height;

[0027] Features of the first multispectral image and features of the first panchromatic image Global adaptive pooling is used to obtain global features of the multispectral image. and global features of panchromatic images ;

[0028] Global attention is calculated based on Transformer blocks, and global features of the first multispectral image are obtained. and global features of the first panchromatic image Among them, the global features of the first multispectral image have the same size as the features of the multispectral image, and the global features of the first panchromatic image have the same size as the features of the panchromatic image.

[0029] The multispectral image features are refined and fused based on the global features of the first multispectral image and the global features of the first panchromatic image.

[0030] A further optimized solution involves refining and fusing the multispectral image features based on the global features of the first multispectral image and the global features of the first panchromatic image; including the following method:

[0031] The following formula is used to refine and fuse multispectral image features;

[0032] ;

[0033] in, This represents the multispectral image features after spatial refinement; Conv() represents convolution processing; Cat(a, b) represents superimposing a and b; Upsample() represents upsampling; reshape() represents deformation.

[0034] A further optimized solution is that the multispectral image features refined by fusing spatial information at different scales include the following methods:

[0035] After convolution processing of the multispectral image features refined by spatial information, activation processing is performed to obtain the first target feature; after Transformer processing of the multispectral image features refined by spatial information, the second target feature is obtained.

[0036] The first target feature and the second target feature are superimposed and then output through a convolutional layer.

[0037] A further optimized solution is that the method for generating a multispectral image based on the fused multispectral image features includes:

[0038] The features of the fused multispectral image are first upsampled, and then convolved through a 3×3 convolutional layer to generate a multispectral image.

[0039] This solution also provides a W-Transformer-based multispectral image generation system for implementing the aforementioned W-Transformer-based multispectral image generation method; the system includes:

[0040] The acquisition module is used to acquire contaminated multispectral and panchromatic images;

[0041] W-Transformer network for generating multispectral images;

[0042] The W-Transformer network includes:

[0043] The feature encoding module is used to encode features of multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales.

[0044] The feature refinement module is used to refine the spatial information of multispectral image features based on panchromatic image features under preset conditions; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale.

[0045] The fusion module is used to fuse multispectral image features refined from spatial information at different scales.

[0046] The generation module is used to generate multispectral images based on the features of the fused multispectral images.

[0047] This solution also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, can implement the W-Transformer-based multispectral image generation method described above.

[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0049] 1. The present invention provides a multispectral image generation method, system, and medium based on W-Transformer; it improves upon existing technologies by using a W-Transformer network to encode features of contaminated multispectral and panchromatic images, obtaining multispectral and panchromatic image features at various scales; when the panchromatic and multispectral image features belong to the same scale, the multispectral image features are spatially refined based on the panchromatic image features; and multispectral image features refined with spatial information at different scales are fused to generate a multispectral image, effectively improving the performance degradation of existing deep learning methods when faced with noise contamination during the acquisition of original panchromatic and multispectral images.

[0050] 2. The multispectral image generation method, system, and medium based on W-Transformer provided by this invention: By refining features at multiple scales, the model can simultaneously process large-scale structures (such as mountains and urban blocks) and minute details (such as rooftops and vehicles) in images. This mechanism enables the final generated multispectral image to have good recovery capabilities for spatial information at different levels, is insensitive to defects and noise in the input image, and produces more stable and reliable output results.

[0051] 3. The multispectral image generation method, system, and medium based on W-Transformer provided by this invention successfully achieve both extremely high spatial resolution enhancement and excellent spectral fidelity through its unique multi-scale feature encoding and attention-based spatial information refinement process. Its powerful global modeling capabilities and excellent generalization performance make it highly valuable and promising for applications in remote sensing, geological exploration, environmental monitoring, and military reconnaissance, representing an important direction in the development of multispectral image fusion technology.

[0052] 4. The present invention provides a multispectral image generation method, system, and medium based on W-Transformer; the self-attention mechanism in the W-Transformer network can calculate the relationship weight between any two pixels in the image. When performing "spatial information refinement", the model can make more informed decisions based on global context information, and the generated spatial details are more coherent and more consistent with the real spatial distribution of ground objects, effectively avoiding local artifacts or structural distortions. At the same time, since W-Transformer is a derivative architecture of U-Net network, it has the denoising advantages of U-Net network, which can enhance the network's resistance to noise in the original data. Attached Figure Description

[0053] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:

[0054] Figure 1 This is a schematic diagram of the process for generating multispectral images based on W-Transformer;

[0055] Figure 2 This is a schematic diagram of the W-Transformer network principle. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0057] Currently, deep learning-based panchromatic sharpening methods can effectively achieve super-resolution of low-resolution multispectral images. However, existing deep learning-based panchromatic sharpening methods mainly consider ideal panchromatic / multispectral images, which are easily degraded by different types of noise during the acquisition process of multispectral and panchromatic images. In view of this, this solution provides the following embodiments to solve the above-mentioned technical problems:

[0058] Example 1: This example provides a multispectral image generation method based on W-Transformer, such as... Figure 1 and Figure 2 As shown, it includes:

[0059] Step 1: Acquire contaminated multispectral and panchromatic images; contamination includes: blur contamination, noise contamination, or blur-noise contamination.

[0060] The multispectral image represents the first resolution, and the panchromatic image represents the second resolution; the first resolution is lower than the second resolution.

[0061] The contaminated multispectral and panchromatic images in this scheme can be understood as low-resolution multispectral images and high-resolution panchromatic images that have been contaminated by noise during the acquisition process.

[0062] The following steps involve inputting multispectral and panchromatic images into the W-Transformer network to generate a multispectral image:

[0063] Step 2 involves feature encoding of the multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales. The W-Transformer network includes a first branch and a second branch, which are used to extract features from the multispectral and panchromatic images, respectively.

[0064] Both the first and second branches include four feature extraction stages. The first, second, and third feature extraction stages each include a Transformer block and a patch merging block, while the fourth feature extraction stage includes a Transformer block. The Transformer block divides the input features into fixed-size, non-overlapping patches based on a window multi-head self-attention mechanism and calculates self-attention within each window. The self-attention calculation process includes:

[0065] Calculate the self-attention within each window using the following formula:

[0066] ;

[0067] Where Q, K, and V represent the query matrix, key matrix, and value matrix, d represents the dimension of the query matrix and key matrix, and B represents the offset, B∈M 2 ×M 2 M 2 This indicates the number of tiles in the window; SoftMax() represents the normalization function.

[0068] The first feature extraction stage of the first branch has a first spatial resolution, and the first feature extraction stage of the second branch has a second spatial resolution; the first spatial resolution is greater than the second spatial resolution.

[0069] Multispectral and panchromatic images are first segmented into fixed-size images through convolutional layers. For multispectral images, the spatial resolution of the first feature extraction stage is H×W, and the number of bands is C. For panchromatic images, the spatial resolution of the first feature extraction stage is (H / 4)×(W / 4), and the number of bands is 1.

[0070] Step 3: Under preset conditions, refine the spatial information of the multispectral image features based on the panchromatic image features; the preset conditions include that the panchromatic image features and the multispectral image features belong to the same scale; this step specifically includes the following methods:

[0071] S31, Obtain the multispectral image features extracted in the i-th feature extraction stage. and panchromatic image features ;

[0072] S32, for the multispectral image features and panchromatic image features The first multispectral image features are obtained after flattening and shaping. and features of the first panchromatic image Among them, the first multispectral image features The size is larger than the features of the first panchromatic image. Size; First multispectral image features The size is HW / 16 i ×2 i-1 L times; H represents the height of the multispectral image, W represents the width of the multispectral image, and L represents the number of encoding / decoding operations; features of the first panchromatic image. The size is HW / 4 i ×2 i-1 L;

[0073] S33, in order to enable the panchromatic image features to fully refine and guide the spatial information of the spectral image features, while reducing the computational load of the attention mechanism, the first multispectral image features... and features of the first panchromatic image Global adaptive pooling is used to obtain global features of the multispectral image. and global features of panchromatic images ;

[0074] ;

[0075] ;

[0076] Wherein, GApool() represents global adaptive pooling;

[0077] S34, calculate global attention based on Transformer blocks and obtain the global features of the first multispectral image. and global features of the first panchromatic image Among them, the global features of the first multispectral image have the same size as the features of the multispectral image, and the global features of the first panchromatic image have the same size as the features of the panchromatic image.

[0078] ;

[0079] ;

[0080] S35, refine and fuse the multispectral image features based on the global features of the first multispectral image and the global features of the first panchromatic image; this step specifically includes the following method:

[0081] The following formula is used to refine and fuse multispectral image features:

[0082] ;

[0083] in, This represents the multispectral image features after spatial refinement; Conv() represents convolution processing; Cat(a, b) represents superimposing a and b; Upsample() represents upsampling; reshape() represents deformation.

[0084] Step four involves fusing multispectral image features refined from spatial information at different scales; this step specifically includes the following methods:

[0085] After convolution processing of the multispectral image features refined by spatial information, activation processing is performed to obtain the first target feature; after Transformer processing of the multispectral image features refined by spatial information, the second target feature is obtained.

[0086] The first target feature and the second target feature are superimposed and then passed through a convolutional layer for output.

[0087] The main structure of the decoder in the W-Transformer network is as follows: two branches process the multispectral image features refined with spatial information at different scales. One branch includes convolutional layers and activation layers, while the other branch includes a Transformer module. The processed features are then superimposed and output through a convolutional layer. This process can be represented by the following formula:

[0088] ;

[0089] ;

[0090] ;

[0091] Step 5: Generate a multispectral image based on the fused multispectral image features. This step specifically includes the following methods:

[0092] The features of the fused multispectral image are first upsampled, and then processed by a 3×3 convolutional layer to generate the multispectral image. The fusion module mainly consists of an upsampling layer and a convolutional layer. The final fused image is obtained after passing through the upsampling layer and the 3×3 convolutional layer.

[0093] ;

[0094] Here, fused represents the final fused image result; Conv3*3() represents a convolution operation with a 3×3 kernel.

[0095] Example 2

[0096] This embodiment provides a W-Transformer-based multispectral image generation system for implementing the W-Transformer-based multispectral image generation method of Embodiment 1; the system includes:

[0097] The acquisition module is used to acquire contaminated multispectral and panchromatic images;

[0098] W-Transformer network for generating multispectral images;

[0099] The W-Transformer network includes:

[0100] The feature encoding module is used to encode features of multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales.

[0101] The feature refinement module is used to refine the spatial information of multispectral image features based on panchromatic image features under preset conditions; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale.

[0102] The fusion module is used to fuse multispectral image features refined from spatial information at different scales.

[0103] The generation module is used to generate multispectral images based on the features of the fused multispectral images.

[0104] Example 3

[0105] This embodiment provides a computer-readable medium storing a computer program, which, when executed by a processor, can implement the multispectral image generation method based on W-Transformer as described in Embodiment 1; specifically, it performs the following steps:

[0106] Step 1: Acquire the contaminated multispectral and panchromatic images;

[0107] Step two involves feature encoding of the multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales; for example... Figure 2 As shown, the encoder of the W-Transformer network mainly consists of two branches, each with four feature extraction stages (cross-modal feature extraction) to extract features from the degraded panchromatic image and multispectral image respectively. The first three feature extraction stages all contain Transformer blocks and patch merging blocks. The extracted panchromatic image feature sizes are as follows: , , Where H represents the height of the panchromatic image, W represents the width of the panchromatic image, and C represents the number of channels; the extracted multispectral image feature dimensions are as follows: , , Where h represents the height of the multispectral image and w represents the width of the multispectral image; the last feature extraction stage contains only Transformer blocks; the Transformer blocks use localization and offset windows, specifically, the Transformer blocks divide the input features into fixed-size and non-overlapping patches and compute self-attention within each window.

[0108] Step 3: Under preset conditions, refine the spatial information of the multispectral image features based on the panchromatic image features; the preset conditions include that the panchromatic image features and the multispectral image features belong to the same scale;

[0109] Step four: Fusing multispectral image features refined from spatial information at different scales; the W-Transformer network decoder includes two branches, which process panchromatic and multispectral features at different scales respectively. One branch includes convolutional layers and activation layers, and the other branch includes Transformer blocks. The processed features are superimposed and output through a convolutional layer. The image sizes during the fusion process are as follows: , , and .

[0110] Step 5: Generate a multispectral image based on the features of the fused multispectral image.

[0111] Table 1 shows the down-resolution metrics for multispectral images generated based on the W-Transformer network. These metrics include: Global Relative Minimum Error (ERGAS), Band Average Exponent (Q4), SCC correlation coefficient, and Spectral Angle Error (SAM). Table 2 shows the full-resolution evaluation metrics for multispectral images generated based on the W-Transformer network, including the degree of spectral distortion (D) at full resolution. λ Spatial distortion level D S And the comprehensive evaluation index QNR;

[0112] Table 1. Results of the resolution reduction experiment

[0113]

[0114] Table 2 Full-resolution experimental results indicators

[0115]

[0116] Tables 1 and 2 show the performance of six deep learning-based multispectral panchromatic sharpening methods in Experiments 1 and 2, respectively. Among them, the multispectral image generated by this scheme achieves the lowest ERGAS and SAM and the highest Q4 and SCC when down-resolution is used; and the lowest D at full resolution. λ D S And the highest QNR. Overall, the multispectral images obtained by this method have good performance and a certain degree of stability in both resolution evaluation and full-resolution evaluation.

[0117] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multispectral image generation method based on W-Transformer, characterized in that, include: Acquire contaminated multispectral and panchromatic images; Input the multispectral and panchromatic images into the W-Transformer network to generate a multispectral image: Feature encoding is performed on multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales. Under preset conditions, spatial information refinement of multispectral image features is performed based on panchromatic image features; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale. Multispectral image features refined by fusing spatial information from different scales; Generate multispectral images based on the features of the fused multispectral images; The W-Transformer network includes a first branch and a second branch, which are used for feature extraction from multispectral images and panchromatic images, respectively; Both the first and second branches include four feature extraction stages. The first, second, and third feature extraction stages each include a Transformer block and a patch merging block, and the fourth feature extraction stage includes a Transformer block. The Transformer block divides the input features into fixed-size and non-overlapping patches based on a window multi-head self-attention mechanism, and calculates the self-attention within each window. The first feature extraction stage of the first branch has a first spatial resolution, and the first feature extraction stage of the second branch has a second spatial resolution; the first spatial resolution is greater than the second spatial resolution. The method for spatial information refinement of multispectral image features based on panchromatic image features under preset conditions includes: Obtain the multispectral image features extracted in the i-th feature extraction stage. and panchromatic image features ; For the multispectral image features and panchromatic image features The first multispectral image features are obtained after flattening and shaping. and features of the first panchromatic image Among them, the first multispectral image features The height is greater than the features of the first panchromatic image. Height; Features of the first multispectral image and features of the first panchromatic image Global adaptive pooling is used to obtain global features of the multispectral image. and global features of panchromatic images ; Global attention is calculated based on Transformer blocks, and global features of the first multispectral image are obtained. and global features of the first panchromatic image Among them, the global features of the first multispectral image have the same size as the features of the multispectral image, and the global features of the first panchromatic image have the same size as the features of the panchromatic image. The multispectral image features are refined and fused based on the global features of the first multispectral image and the global features of the first panchromatic image.

2. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The pollution includes: fuzzy pollution, noise pollution, or fuzzy noise pollution; The multispectral image has a first resolution, and the panchromatic image has a second resolution; The first resolution is lower than the second resolution.

3. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The calculation process for self-attention includes: Calculate the self-attention within each window using the following formula; ; Where Q, K, and V represent the query, key, and value matrix, d represents the dimension of the query and key, and B represents the offset, B∈M 2 ×M 2 M 2 This indicates the number of tiles in the window; SoftMax() represents the normalization function.

4. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The method involves refining and fusing the multispectral image features based on the global features of the first multispectral image and the global features of the first panchromatic image; including: The following formula is used to refine and fuse multispectral image features; ; in, This represents the multispectral image features after spatial refinement; Conv() represents convolution processing; Cat(a, b) represents superimposing a and b; Upsample() represents upsampling; reshape() represents deformation.

5. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The multispectral image features refined by fusing spatial information at different scales include the following methods: After convolution processing of the multispectral image features refined by spatial information, activation processing is performed to obtain the first target features; The second target feature is obtained by processing the multispectral image features refined with spatial information using a Transformer. The first target feature and the second target feature are superimposed and then output through a convolutional layer.

6. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The method for generating a multispectral image based on the fused multispectral image features includes: The features of the fused multispectral image are first upsampled, and then convolved through a 3×3 convolutional layer to generate a multispectral image.

7. A multispectral image generation system based on W-Transformer, characterized in that, The system is used to implement the W-Transformer-based multispectral image generation method according to any one of claims 1-6; the system comprises: The acquisition module is used to acquire contaminated multispectral and panchromatic images; W-Transformer network for generating multispectral images based on multispectral and panchromatic images; The W-Transformer network includes: The feature encoding module is used to encode features of multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales. The feature refinement module is used to refine the spatial information of multispectral image features based on panchromatic image features under preset conditions; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale. The fusion module is used to fuse multispectral image features refined from spatial information at different scales. The generation module is used to generate multispectral images based on the features of the fused multispectral images.

8. A computer-readable medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, can implement the W-Transformer-based multispectral image generation method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Panchromatic sharpening method and system based on cross-resolution adversarial learning and Mama network

    CN120689241A