Multi-spectral image generation method and system based on W-Transform and medium

By using the W-Transformer network to encode features and refine spatial information of multispectral and panchromatic images, the performance degradation problem of deep learning methods under noise pollution is solved, and multispectral image generation with high spatial resolution and spectral fidelity is achieved.

CN120997060AActive Publication Date: 2025-11-21HUANTIAN SMART TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511516766.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing deep learning-based panchromatic sharpening methods ignore noise and blur that may occur during image transmission when processing multispectral and panchromatic images, resulting in performance degradation when processing degraded original images.

Method used

A W-Transformer-based multispectral image generation method is adopted. By encoding features of multispectral and panchromatic images, the W-Transformer network is used to refine the spatial information of multispectral image features based on panchromatic image features at the same scale, and the spatial information refinement features at different scales are fused to generate multispectral images.

Benefits of technology

It effectively improves the performance degradation of deep learning methods when faced with noise pollution, and can simultaneously process large-scale structures and minute details in images, generating multispectral images with high spatial resolution and good spectral fidelity, and has good recovery ability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997060A_ABST
    Figure CN120997060A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral image generation method and system based on a W-Transform, and a medium. Relates to the technical field of remote sensing image processing. Improvement is carried out on the basis of the prior art, feature coding is carried out on a polluted multispectral image and a polluted panchromatic image based on a W-Transform network, and multispectral image features and panchromatic image features of various different scales are obtained; performing spatial information refinement on the multispectral image features based on the panchromatic image features under a preset condition; and the multi-spectral image is generated by fusing the multi-spectral image features after spatial information refinement under different scales, so that the condition that the performance is degraded due to noise pollution in the acquisition process of the original panchromatic image and the multi-spectral image in the existing deep learning method is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image processing, and particularly relates to a multispectral image generation method and system based on W-Transformer and a medium. BACKGROUND

[0002] Multispectral images are of great importance for important applications such as geological exploration, environmental protection, and urban planning. However, although the spectral resolution of multispectral images is constantly improving, their spatial resolution is often limited. With the development of remote sensing technology in recent years, multispectral and high-spatial-resolution panchromatic images can be obtained from the same imaging platform. Therefore, fusing multispectral resolution multispectral images (Multispectral, MS) and high-spatial-resolution panchromatic images (Panchromatic, PAN) together, also known as panchromatic sharpening, has become one of the main ways to obtain high-spatial-resolution multispectral images.

[0003] Current sharpening methods can be mainly divided into two categories: traditional fusion methods and deep learning methods. The former includes component substitution, multi-resolution analysis, and variational optimization. Component substitution is the simplest method to solve the sharpening problem. It converts multispectral to another representation domain, substitutes the spatial components in the representation domain with panchromatic, and obtains the sharpening result through inverse transformation. Principal component analysis, band-dependent spatial detail-based methods, and Gram-Schmidt adaptive methods all belong to component substitution methods. The high fidelity of spatial information of component substitution methods is a significant advantage, but they often suffer from severe spectral distortion. Multi-resolution analysis decomposes multispectral and panchromatic into multiple high-frequency and low-frequency coefficients at different spatial scales and fuses these coefficients together according to different fusion rules, thus generating a denoised and enhanced image. Methods based on multi-resolution analysis include Laplacian pyramid methods, additive wavelet methods, and curvelet transform methods. By injecting high-frequency detail information into the fusion result, methods based on multi-resolution analysis can maintain good spectral information but may produce some spatial distortion. Methods based on variational optimization are usually based on models of denoised and enhanced results and input images. They rely on constraints based on prior assumptions to propose an energy functional, converting denoised and enhanced into an optimization problem. The most popular methods based on variational optimization include sparse representation, compressed sensing, and gradient prior methods.

[0004] In recent years, the panchromatic sharpening method based on deep learning is more and more popular due to its convenience and excellent performance, and some scholars have innovatively proposed a three-layer convolutional neural network to solve the panchromatic sharpening problem. However, due to the simple network structure, this method cannot completely extract image features, or the proposed panchromatic sharpening network is trained in the high-frequency domain to retain spatial information, and residual connection is used to achieve spectral fidelity, and the advantages of traditional machine learning are combined to establish a deep convolutional neural network based on detail injection to protect spatial information as much as possible and adjust the extracted image details using a convolutional neural network. Since the existing panchromatic sharpening method based on deep learning usually acquires multispectral images and panchromatic images with good imaging quality, it ignores the noise and blur that may occur during image transmission, resulting in performance degradation when processing degraded original images. SUMMARY

[0005] The technical problem to be solved by the present application is that the existing panchromatic sharpening method based on deep learning usually acquires multispectral images and panchromatic images with good imaging quality, ignores the noise and blur that may occur during image transmission, and causes performance degradation when processing degraded original images. The present application aims to provide a multispectral image generation method and system based on W-Transformer and a medium, which improves the existing technology, encodes the polluted multispectral image and panchromatic image based on the W-Transformer network, and obtains multispectral image features and panchromatic image features of different scales. Under the same scale, the multispectral image features are refined based on the panchromatic image features, and the multispectral image features after spatial information refinement under different scales are fused to generate a multispectral image, effectively improving the performance degradation of the existing deep learning method when facing original panchromatic images and multispectral images polluted by noise during acquisition.

[0006] The present application is realized by the following technical solutions:

[0007] The present application provides a multispectral image generation method based on W-Transformer, which includes:

[0008] Obtain the polluted multispectral image and panchromatic image;

[0009] Input the multispectral image and panchromatic image into the W-Transformer network to generate a multispectral image:

[0010] Encode the multispectral image and panchromatic image to obtain multispectral image features and panchromatic image features of different scales;

[0011] Refine spatial information of the multispectral image features based on the panchromatic image features under preset conditions; the preset conditions include that the panchromatic image features and the multispectral image features belong to the same scale;

[0012] Fuse the multispectral image features refined under different scales of spatial information;

[0013] Generate the multispectral image based on the fused multispectral image features.

[0014] Further optimization scheme is that the pollution includes: blur pollution, noise pollution or blur noise pollution;

[0015] The multispectral image is of a first resolution, and the panchromatic image is of a second resolution;

[0016] The first resolution is lower than the second resolution.

[0017] Further optimization scheme is that the W-Transformer network includes a first branch and a second branch, respectively used for feature extraction of the multispectral image and the panchromatic image;

[0018] The first branch and the second branch both include four feature extraction stages, wherein the first feature extraction stage, the second feature extraction stage and the third feature extraction stage all include a Transformer block and a patch merging block, and the fourth feature extraction stage includes a Transformer block; the Transformer block divides input features into fixed window size and non-overlapping patches based on a window multi-head self-attention mechanism, and calculates self-attention within each window;

[0019] The first feature extraction stage of the first branch is of a first spatial resolution, and the first feature extraction stage of the second branch is of a second spatial resolution; the first spatial resolution is greater than the second spatial resolution.

[0020] Further optimization scheme is that the calculation process of the self-attention includes:

[0021] The self-attention within each window is calculated according to the following formula:

[0022] ;

[0023] Wherein, Q, K and V represent query matrix, key matrix and value matrix, d represents the dimension of the query matrix and the key matrix, B represents the offset, B∈M 2 ×M 2 ;M 2 represents the number of tiles in the window; SoftMax() represents a normalization function.

[0024] Further optimization scheme is that the spatial information refinement of the multispectral image features based on the panchromatic image features under the preset condition comprises the following steps:

[0025] The multispectral image features extracted in the i-th feature extraction stage and the panchromatic image features ;

[0026] After flattening and shaping the multispectral image features and the panchromatic image features , first multispectral image features and first panchromatic image features are obtained; wherein the height of the first multispectral image features is greater than the height of the first panchromatic image features ;

[0027] Global adaptive pooling is performed on the first multispectral image features and the first panchromatic image features to obtain multispectral image global features and panchromatic image global features ;

[0028] The global attention is calculated based on the Transformer block, and first multispectral image global features and first panchromatic image global features are obtained; wherein the size of the first multispectral image global features is the same as that of the multispectral image features, and the size of the first panchromatic image global features is the same as that of the panchromatic image features;

[0029] The multispectral image features are refined and fused based on the first multispectral image global features and the first panchromatic image global features.

[0030] Further optimization scheme is that the multispectral image features are refined and fused based on the first multispectral image global features and the first panchromatic image global features; comprising the following steps:

[0031] The multispectral image features are refined and fused based on the following formula:

[0032] ;

[0033] wherein, represents the multispectral image features after spatial information refinement; Conv() represents convolution processing; Cat(a, b) represents stacking a and b; Upsample() represents up-sampling; reshape() represents reshaping.

[0034] Further optimization scheme is that the multi-spectral image features after spatial information refinement under different scales are fused, including the method:

[0035] After the multi-spectral image features after spatial information refinement are convolved, the first target feature is obtained after activation processing; the second target feature is obtained after the multi-spectral image features after spatial information refinement are processed by the Transformer.

[0036] The first target feature and the second target feature are superimposed and output after a convolution layer.

[0037] Further optimization scheme is that the multi-spectral image is generated based on the fused multi-spectral image features, including the method:

[0038] The fused multi-spectral image features are first up-sampled, and then convolved by a convolution layer with a convolution kernel of 3x3 to generate the multi-spectral image.

[0039] The present scheme also provides a multi-spectral image generation system based on W-Transformer, which is used to realize the above-mentioned multi-spectral image generation method based on W-Transformer; the system comprises:

[0040] The acquisition module is used to acquire the multi-spectral image and the panchromatic image that have been contaminated;

[0041] The W-Transformer network is used to generate the multi-spectral image;

[0042] The W-Transformer network comprises:

[0043] The feature encoding module is used to encode the multi-spectral image and the panchromatic image to obtain multi-spectral image features and panchromatic image features of different scales;

[0044] The feature refinement module is used to refine the spatial information of the multi-spectral image features based on the panchromatic image features under a preset condition; the preset condition includes that the panchromatic image features and the multi-spectral image features belong to the same scale;

[0045] The fusion module is used to fuse the multi-spectral image features after spatial information refinement under different scales;

[0046] The generation module is used to generate the multi-spectral image based on the fused multi-spectral image features.

[0047] The present scheme also provides a computer readable medium having a computer program stored thereon, wherein the computer program is executed by a processor to realize the above-mentioned multi-spectral image generation method based on W-Transformer.

[0048] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0049] 1. The W-Transformer-based multispectral image generation method, system and medium provided by the present application improve the existing deep learning method by encoding the polluted multispectral image and panchromatic image based on the W-Transformer network to obtain multispectral image features and panchromatic image features of different scales, and refining the spatial information of the multispectral image features based on the panchromatic image features when the panchromatic image features and the multispectral image features belong to the same scale.

[0050] 2. The W-Transformer-based multispectral image generation method, system and medium provided by the present application can process large-scale structures (such as mountains and city blocks) and small details (such as roofs and vehicles) in the image at the same time by refining the features at multiple scales. This mechanism enables the generated multispectral image to have good recovery ability for different levels of spatial information and be insensitive to defects and noise in the input image, making the output result more stable and reliable.

[0051] 3. The W-Transformer-based multispectral image generation method, system and medium provided by the present application successfully realize high spatial resolution enhancement and excellent spectral fidelity at the same time through its unique multi-scale feature encoding and attention mechanism-based spatial information refinement process. Its strong global modeling capability and excellent generalization performance make it have high application value and market prospect in the fields of remote sensing, geological exploration, environmental monitoring and military reconnaissance, representing an important direction of multispectral image fusion technology development.

[0052] 4. The W-Transformer-based multispectral image generation method, system and medium provided by the present application. The self-attention mechanism in the W-Transformer network can calculate the relationship weight between any two pixel points in the image. When performing "spatial information refinement", the model can make more intelligent decisions based on global context information, generate more coherent spatial details that conform to the real spatial distribution of ground objects, effectively avoid local artifacts or structural distortion, and enhance the network's resistance to noise in the original data due to the denoising advantage of U-Net network. BRIEF DESCRIPTION OF DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor. In the drawings:

[0054] Figure 1 A flowchart of a method for generating a multispectral image based on a W-Transformer;

[0055] Figure 2 A schematic diagram of the principle of a W-Transformer network. DETAILED DESCRIPTION

[0056] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the present application in detail with reference to the embodiments and drawings, the exemplary embodiments of the present application and their descriptions are only used to explain the present application, and should not be regarded as limiting the present application.

[0057] At present, the panchromatic sharpening method based on deep learning can effectively realize the super-resolution of low-resolution multispectral images. However, the research object of the existing panchromatic sharpening method based on deep learning mainly considers ideal panchromatic / multispectral images, and in practice, multispectral images and panchromatic images are easily degraded by different types of noise in the acquisition process; In view of this, the present scheme provides the following embodiments to solve the above technical problems:

[0058] Embodiment 1: The present embodiment provides a method for generating a multispectral image based on a W-Transformer, as shown in Figure 1 and Figure 2 , comprising:

[0059] Step 1, obtaining a contaminated multispectral image and a panchromatic image; the contamination includes blur contamination, noise contamination or blur noise contamination;

[0060] The multispectral image is of a first resolution, and the panchromatic image is of a second resolution; the first resolution is lower than the second resolution;

[0061] The contaminated multispectral image and the panchromatic image in the present scheme can be understood as a low-resolution multispectral image and a high-resolution panchromatic image contaminated by noise in the acquisition process.

[0062] The multispectral image and the panchromatic image will be input into the W-Transformer network to generate a multispectral image:

[0063] Step two, feature coding is performed on the multispectral image and the panchromatic image to obtain multispectral image features and panchromatic image features of multiple different scales; the W-Transformer network comprises a first branch and a second branch, which are respectively used for feature extraction of the multispectral image and the panchromatic image;

[0064] Both the first branch and the second branch comprise four feature extraction stages, wherein the first feature extraction stage, the second feature extraction stage and the third feature extraction stage all comprise a Transformer block and a patch merging block, and the fourth feature extraction stage comprises a Transformer block; the Transformer block divides input features into fixed window sizes and non-overlapping patches based on a window multi-head self-attention mechanism, and calculates self-attention in each window; the calculation process of the self-attention comprises:

[0065] The self-attention in each window is calculated according to the following formula:

[0066] ;

[0067] Wherein, Q, K and V represent a query matrix, a key matrix and a value matrix, d represents the dimension of the query matrix and the key matrix, B represents an offset, B ∈ M 2 ×M 2 ; M 2 represents the number of patches in the window; SoftMax() represents a normalization function.

[0068] The first feature extraction stage of the first branch is of a first spatial resolution, and the first feature extraction stage of the second branch is of a second spatial resolution; the first spatial resolution is greater than the second spatial resolution.

[0069] The multispectral image and the panchromatic image are first cut into fixed-size images by a convolution layer; for the multispectral image, the spatial resolution of the first feature extraction stage is HxW, and the number of bands is C; for the panchromatic image, the spatial resolution of the first feature extraction stage is (H / 4)x(W / 4), and the number of bands is 1.

[0070] Step three, spatial information refinement is performed on the multispectral image features based on the panchromatic image features under a preset condition; the preset condition comprises that the panchromatic image features and the multispectral image features belong to the same scale; this step specifically comprises the following method:

[0071] S31, obtaining multispectral image features and panchromatic image features extracted by the i-th feature extraction stage;

[0072] S32, performing spatial information refinement on the multispectral image features and the panchromatic image features The first multi-spectral image feature is obtained after flattening and shaping and the first panchromatic image feature ; wherein the size of the first multi-spectral image feature is greater than the size of the first panchromatic image feature ; the size of the first multi-spectral image feature is H W / 16 i ×2 i-1 L; H represents the height of the multi-spectral image, W represents the width of the multi-spectral image, and L represents the number of encoding / decoding times; the size of the first panchromatic image feature is H W / 4 i ×2 i-1 L;

[0073] S33, in order to enable the panchromatic image feature to fully refine and guide the spatial information of the spectral image feature, while reducing the amount of calculation of the attention mechanism, global adaptive pooling is performed on the first multi-spectral image feature and the first panchromatic image feature to obtain multi-spectral image global features and panchromatic image global features ;

[0074] ;

[0075] ;

[0076] wherein GApool() represents global adaptive pooling;

[0077] S34, the global attention is calculated based on the Transformer block to obtain first multi-spectral image global features and first panchromatic image global features ; wherein the size of the first multi-spectral image global features is the same as that of the multi-spectral image features, and the size of the first panchromatic image global features is the same as that of the panchromatic image features;

[0078] ;

[0079] ;

[0080] S35, the multi-spectral image features are refined and fused based on the first multi-spectral image global features and the first panchromatic image global features; the step specifically comprises the following method:

[0081] The multi-spectral image features are refined and fused based on the following formula:

[0082] ;

[0083] wherein, denotes the spatial information refined multispectral image features; Conv() denotes convolution processing; Cat(a, b) denotes stacking a and b; Upsample() denotes up-sampling; reshape() denotes reshaping.

[0084] Step four, fusing the spatial information refined multispectral image features at different scales; this step specifically includes the method:

[0085] After the spatial information refined multispectral image features are processed by convolution, the first target features are obtained after activation processing; after the spatial information refined multispectral image features are processed by the Transformer, the second target features are obtained;

[0086] The first target features and the second target features are stacked and output after a convolution layer;

[0087] The main structure of the decoder of the W-Transformer network is: two branches process the spatial information refined multispectral image features at different scales, one branch includes a convolution layer and an activation layer, and the other branch includes a Transformer module, and the processed features are stacked and output after a convolution layer, which can be represented by the following formula:

[0088] ;

[0089] ;

[0090] ;

[0091] Step five, generating a multispectral image based on the fused multispectral image features, which specifically includes the method:

[0092] The fused multispectral image features are first up-sampled and then processed by a convolution layer with a convolution kernel of 3x3 to generate a multispectral image. The fusion module mainly includes an up-sampling layer and a convolution layer, and the final fusion image result is obtained after the up-sampling layer and the convolution layer with a convolution kernel of 3x3:

[0093] ;

[0094] wherein fused denotes the final fusion image result; Conv3*3() denotes convolution operation with a convolution kernel of 3x3.

[0095] Example 2

[0096] This embodiment provides a W-Transformer-based multispectral image generation system for implementing the W-Transformer-based multispectral image generation method of Embodiment 1; the system includes:

[0097] The acquisition module is used to acquire contaminated multispectral and panchromatic images;

[0098] W-Transformer network for generating multispectral images;

[0099] The W-Transformer network includes:

[0100] The feature encoding module is used to encode features of multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales.

[0101] The feature refinement module is used to refine the spatial information of multispectral image features based on panchromatic image features under preset conditions; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale.

[0102] The fusion module is used to fuse multispectral image features refined from spatial information at different scales.

[0103] The generation module is used to generate multispectral images based on the features of the fused multispectral images.

[0104] Example 3

[0105] This embodiment provides a computer-readable medium storing a computer program, which, when executed by a processor, can implement the multispectral image generation method based on W-Transformer as described in Embodiment 1; specifically, it performs the following steps:

[0106] Step 1: Acquire the contaminated multispectral and panchromatic images;

[0107] Step two involves feature encoding of the multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales; for example... Figure 2 As shown, the encoder of the W-Transformer network mainly consists of two branches, each with four feature extraction stages (cross-modal feature extraction) to extract features from the degraded panchromatic image and multispectral image respectively. The first three feature extraction stages all contain Transformer blocks and patch merging blocks. The extracted panchromatic image feature sizes are as follows: , , Where H represents the height of the panchromatic image, W represents the width of the panchromatic image, and C represents the number of channels; the extracted multispectral image feature dimensions are as follows: , , ; wherein h represents the height of the multispectral image and w represents the width of the multispectral image; the last feature extraction stage only contains a Transformer block; the Transformer block uses a positional and offset window, specifically, the Transformer block divides the input features into fixed window size and non-overlapping patches and calculates the self-attention within each window.

[0108] Step three, under a preset condition, spatial information of the multispectral image features is refined based on the panchromatic image features; the preset condition includes that the panchromatic image features and the multispectral image features belong to the same scale;

[0109] Step four, multispectral image features after spatial information refinement under different scales are fused; the decoder of the W-Transformer network includes two branches, the two branches process panchromatic features and multispectral features of different scales respectively, one branch includes a convolution layer and an activation layer, and the other branch includes a Transformer block, the processed features are superimposed and output after a convolution layer, and the image size in the fusion process is: , , and .

[0110] Step five, based on the fused multispectral image features, a multispectral image is generated.

[0111] Table 1 shows a reduced resolution index table for generating a multispectral image based on the W-Transformer network, and the resolution index includes: global relative minimum error ERGAS, band average index Q4, SCC correlation coefficient and spectral angle error SAM; Table 2 shows a full resolution evaluation index table for generating a multispectral image based on the W-Transformer network, and the index includes full resolution spectral distortion D λ , spatial distortion D S and comprehensive evaluation index QNR;

[0112] Table 1 reduced resolution experimental result index

[0113]

[0114] Table 2 full resolution experimental result index

[0115]

[0116] Table 1 and Table 2 show the performance of the six deep learning based multispectral panchromatic sharpening methods in Experiment 1 and Experiment 2 respectively; among them, the multispectral images generated by the present scheme obtain the lowest ERGAS, SAM and the highest Q4, SCC in the down-resolution; obtain the lowest D λ , D S and the highest QNR in the full-resolution. From the overall results, the multispectral images obtained by the present scheme method have good performance and certain stability whether in the resolution evaluation or the full-resolution evaluation.

[0117] The above specific embodiments further illustrate the purpose, technical scheme and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A multispectral image generation method based on W-Transformer, characterized in that, include: Acquire contaminated multispectral and panchromatic images; Input the multispectral and panchromatic images into the W-Transformer network to generate a multispectral image: Feature encoding is performed on multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales. Under preset conditions, spatial information refinement of multispectral image features is performed based on panchromatic image features; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale. Multispectral image features refined by fusing spatial information from different scales; Multispectral images are generated based on the features of the fused multispectral images.

2. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The pollution includes: fuzzy pollution, noise pollution, or fuzzy noise pollution; The multispectral image has a first resolution, and the panchromatic image has a second resolution; The first resolution is lower than the second resolution.

3. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The W-Transformer network includes a first branch and a second branch, which are used for feature extraction from multispectral images and panchromatic images, respectively; Both the first and second branches include four feature extraction stages. The first, second, and third feature extraction stages each include a Transformer block and a patch merging block, and the fourth feature extraction stage includes a Transformer block. The Transformer block divides the input features into fixed-size and non-overlapping patches based on a window multi-head self-attention mechanism, and calculates the self-attention within each window. The first feature extraction stage of the first branch has a first spatial resolution, and the first feature extraction stage of the second branch has a second spatial resolution; the first spatial resolution is greater than the second spatial resolution.

4. The multispectral image generation method based on W-Transformer according to claim 3, characterized in that, The calculation process for self-attention includes: Calculate the self-attention within each window using the following formula; ; Where Q, K, and V represent the query, key, and value matrix, d represents the dimension of the query and key, and B represents the offset, B∈M 2 ×M 2 M 2 This indicates the number of tiles in the window; SoftMax() represents the normalization function.

5. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The method for spatial information refinement of multispectral image features based on panchromatic image features under preset conditions includes: Obtain the multispectral image features extracted in the i-th feature extraction stage. and panchromatic image features ; For the multispectral image features and panchromatic image features The first multispectral image features are obtained after flattening and shaping. and features of the first panchromatic image Among them, the first multispectral image features The height is greater than the features of the first panchromatic image. Height; Features of the first multispectral image and features of the first panchromatic image Global adaptive pooling is used to obtain global features of the multispectral image. and global features of panchromatic images ; Global attention is calculated based on Transformer blocks, and global features of the first multispectral image are obtained. and global features of the first panchromatic image Among them, the global features of the first multispectral image have the same size as the features of the multispectral image, and the global features of the first panchromatic image have the same size as the features of the panchromatic image. The multispectral image features are refined and fused based on the global features of the first multispectral image and the global features of the first panchromatic image.

6. The multispectral image generation method based on W-Transformer according to claim 5, characterized in that, The method involves refining and fusing the multispectral image features based on the global features of the first multispectral image and the global features of the first panchromatic image; including: The following formula is used to refine and fuse multispectral image features; ; in, This represents the multispectral image features after spatial refinement; Conv() represents convolution processing; Cat(a, b) represents superimposing a and b; Upsample() represents upsampling; reshape() represents deformation.

7. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The multispectral image features refined by fusing spatial information at different scales include the following methods: After convolution processing of the multispectral image features refined by spatial information, activation processing is performed to obtain the first target features; The second target feature is obtained by processing the multispectral image features refined with spatial information using a Transformer. The first target feature and the second target feature are superimposed and then output through a convolutional layer.

8. The multispectral image generation method based on W-Transformer according to claim 1, characterized in that, The method for generating a multispectral image based on the fused multispectral image features includes: The features of the fused multispectral image are first upsampled, and then convolved through a 3×3 convolutional layer to generate a multispectral image.

9. A multispectral image generation system based on W-Transformer, characterized in that, The system is used to implement the W-Transformer-based multispectral image generation method according to any one of claims 1-8; the system comprises: The acquisition module is used to acquire contaminated multispectral and panchromatic images; W-Transformer network for generating multispectral images based on multispectral and panchromatic images; The W-Transformer network includes: The feature encoding module is used to encode features of multispectral and panchromatic images to obtain multispectral and panchromatic image features at various scales. The feature refinement module is used to refine the spatial information of multispectral image features based on panchromatic image features under preset conditions; the preset conditions include that the panchromatic image features and multispectral image features belong to the same scale. The fusion module is used to fuse multispectral image features refined from spatial information at different scales. The generation module is used to generate multispectral images based on the features of the fused multispectral images.

10. A computer-readable medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, can implement the W-Transformer-based multispectral image generation method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Image classification method and image classification system based on sparse feature fusion

    CN117994579A

  • Hyperspectral image super-resolution reconstruction method and system

    CN118247145A

  • Panchromatic sharpening method based on multi-resolution panchromatic feature guidance

    CN120013808A

  • Multi-spectral image and panchromatic image fusion method and system based on full-spectrum space

    CN120070195A

  • Panchromatic sharpening method and system based on cross-resolution adversarial learning and Mama network

    CN120689241A