Remote sensing image fusion method and system, terminal and medium
By adopting a bidirectional information flow network and an enhanced spatial attention module in remote sensing image fusion, the problems of insufficient feature extraction and unsatisfactory fusion effect in the prior art are solved, and the image fusion effect with high spatial resolution and rich spectral information is achieved.
Patent Information
- Application Number
- CN202510173056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-20
AI Technical Summary
The existing remote sensing image fusion method ignores the data source differences between multispectral images and full-color image features, only considers local features, lacks the synergistic effect between long-range and short-range features of different scales, resulting in insufficient feature extraction and unsatisfactory fusion effect.
The bidirectional information flow network structure is adopted to extract multi-resolution spatial features through full-color branches, and these features are gradually injected into the multi-spectral image using the enhanced spatial attention module to achieve the fusion of the differentiation of multi-source data and the synergistic effect between local features and long-range information.
By enhancing the spatial attention mechanism of the spatial attention module, the fusion image greatly retains and enhances spatial details, makes full use of context information, and generates high spatial resolution multispectral images, which significantly improves the fusion effect.
Smart Images

Figure CN120182104A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image fusion, and particularly relates to a remote sensing image fusion method, system, terminal and medium. Background Art
[0002] Multispectral images have rich spectral information and can provide surface features in different bands, such as vegetation, soil types, etc. However, due to their relatively low spatial resolution, they may not be able to capture fine surface features. Panchromatic images, on the other hand, often have high spatial resolution but only one band and almost no ability to provide spectral colors. Therefore, by fusing the spectral information of multispectral images with the high-resolution information of panchromatic images, the deficiencies of each other can be made up, and an image with both rich spectral information and high-resolution spatial information can be obtained.
[0003] Traditional remote sensing image fusion techniques include component replacement methods and multiresolution analysis methods. These two methods simply linearly combine panchromatic images as components and incorporate them into multispectral images. Although the fusion speed is relatively fast, the fusion effect is poor.
[0004] In recent years, the feature extraction ability of deep learning methods has been greatly enhanced, and the fusion effect of deep learning-based fusion methods is often stronger than that of traditional methods. However, many current deep learning-based fusion methods only use stacked convolutional layers for feature extraction and fusion, ignoring the data source differences between the features of multispectral images and panchromatic images, and only considering local features, lacking the synergistic effect between long-range and short-range features at different scales, making it difficult for the fusion result to utilize context information, resulting in insufficient feature extraction and unsatisfactory fusion effect. Summary of the Invention
[0005] To solve the above problems, the present invention provides a remote sensing image fusion method, system, terminal and medium, extracts the multiresolution spatial features of panchromatic images, and gradually injects each spatial feature into the multispectral image through an enhanced spatial attention module, realizes a fusion strategy for the differences of multi-source data and the synergistic effect between local features and long-range information, obtains a high-spatial-resolution fusion image, and effectively improves the fusion effect.
[0006] In the first aspect, the technical solution of the present invention provides a remote sensing image fusion method, which is implemented by a bidirectional information flow network structure including a panchromatic branch and a multispectral branch, and includes the following steps S1 - step S4, where the panchromatic branch implements step S1, and the multispectral branch implements steps S3 - S4; S1, perform multiresolution spatial information extraction on the panchromatic image to obtain the first-resolution spatial feature and the second-resolution spatial feature of the panchromatic image, and the resolution of the second-resolution spatial feature is higher than that of the first-resolution spatial feature; S2. Upsample the original multispectral image for the first time to obtain the first multispectral image, and use the enhanced spatial attention model to fuse the first multispectral image with the second-resolution spatial features to obtain the second multispectral image; S3. Upsample the second multispectral image for the second time to obtain the third multispectral image, and use the enhanced spatial attention model to fuse the third multispectral image with the first-resolution spatial features to obtain the fourth multispectral image; S4. Upsample the original multispectral image for the third time to obtain the fifth multispectral image, and perform element-wise addition fusion on the fourth multispectral image and the fifth multispectral image to obtain the final remote sensing image fusion result.
[0007] In an optional embodiment, extracting multi-resolution spatial information from the panchromatic image to obtain the first-resolution spatial features and the second-resolution spatial features of the panchromatic image specifically includes: Input the panchromatic image into the first convolutional block for processing to obtain the first-resolution spatial features; Perform max pooling operation on the first-resolution spatial features and then input them into the second convolutional block for processing to obtain the second-resolution spatial features; Among them, the first convolutional block and the second convolutional block have the same structure, both including 4 convolutional layers, 1 Relu activation function and residuals. The first 3 convolutional layers are connected in sequence, and the output result is processed by the Relu activation function and then input into the last convolutional layer through residual connection.
[0008] In an optional embodiment, the convolutional layers of the first convolutional block and the second convolutional block are 3×3 convolutional layers.
[0009] In an optional embodiment, using the enhanced spatial attention model to fuse the multispectral image with the spatial features is expressed as:
[0010] In the formula, represents the Swin-Transformer block, represents a group of convolutional layers, represents the enhanced spatial attention mechanism, represents the result of the convolutional operation of the shallow features required for the synthesized fused image;
[0011] In the formula, represents the feature map of the panchromatic image, that is, the spatial features, represents the upsampled multispectral image; The fusion process specifically includes: The spatial features are directly superimposed on the spectral features of the upsampled multispectral image to obtain fused shallow features; The fused shallow features are subjected to a convolution operation to obtain , and the enhanced spatial attention mechanism is used to process to obtain intermediate features; The intermediate features are input into the Swin-Transformer block for further integration to fuse the spectral and spatial information of the features; The integrated features are processed through a group of convolutional layers to obtain the fusion result.
[0012] In an optional embodiment, the enhanced spatial attention mechanism is used to process to obtain intermediate features, specifically including: Use a 1×1 convolutional layer to perform convolution processing to reduce the embedding dimension of the features to obtain , denoted as,
[0013] where, represents a 1×1 convolutional layer; Use a 3×3 convolutional layer with a stride of 2 to perform convolution processing, and perform a pooling operation on the convolution processing result; Use a group of 3×3 convolutional layers to perform convolution processing on the pooling operation result, with an activation function Relu after each convolution, and perform upsampling on the convolution processing result to obtain , denoted as
[0014] where, represents upsampling, represents the pooling operation, represents a group of 3×3 convolutional layers, with an activation function Relu after each convolution, represents a 3×3 convolutional layer with a stride of 2; Add and element-wise, use a 1×1 convolutional layer to perform convolution processing on the addition result, and use the normalization function to process the convolution processing result and then multiply it element-wise with to obtain the intermediate feature , denoted as,
[0015] where, represents a 1×1 convolutional layer, and × represents element-wise multiplication calculation.
[0016] In an optional embodiment, the intermediate features are input into the Swin-Transformer block for further integration to fuse the spectral and spatial information of the features, expressed as
[0017] where represents the input of the Swin-Transformer block, that is, the intermediate features , , , , represents the feature maps at different scales; LN represents the normalization layer, W-MSA represents the window-based multi-head self-attention layer, SW-MSA represents the sliding window multi-head self-attention layer, and MLP represents the multi-layer perceptron layer.
[0018] In an optional embodiment, both the first upsampling and the second upsampling are 2× upsampling, and the third upsampling is 4× upsampling.
[0019] In a second aspect, the technical solution of the present invention provides a remote sensing image fusion system, which is implemented based on a bidirectional information flow network structure including a panchromatic branch and a multispectral branch. The spatial feature extraction unit is implemented based on the panchromatic branch, and the first fusion unit, the second fusion unit, and the third fusion unit are implemented based on the multispectral branch, including: A spatial feature extraction unit, configured to extract multi-resolution spatial information from the panchromatic image to obtain the first-resolution spatial feature and the second-resolution spatial feature of the panchromatic image, and the resolution of the second-resolution spatial feature is higher than that of the first-resolution spatial feature; The first fusion unit: performs first upsampling on the original multispectral image to obtain the first multispectral image, and uses the enhanced spatial attention model to fuse the first multispectral image with the second-resolution spatial feature to obtain the second multispectral image; The second fusion unit: performs second upsampling on the second multispectral image to obtain the third multispectral image, and uses the enhanced spatial attention model to fuse the third multispectral image with the first-resolution spatial feature to obtain the fourth multispectral image; The third fusion unit: performs third upsampling on the original multispectral image to obtain the fifth multispectral image, and performs element-wise addition fusion on the fourth multispectral image and the fifth multispectral image to obtain the final remote sensing image fusion result.
[0020] In a third aspect, the technical solution of the present invention provides a terminal, including: A memory, configured to store a remote sensing image fusion program; A processor for implementing the steps of the remote sensing image fusion method as described in any one of the above when executing the remote sensing image fusion program.
[0021] In a fourth aspect, the technical solution of the present invention provides a computer-readable storage medium, on which a remote sensing image fusion program is stored. When the remote sensing image fusion program is executed by a processor, the steps of the remote sensing image fusion method as described in any one of the above are implemented.
[0022] The remote sensing image fusion method, system, terminal and storage medium provided by the present invention have the following beneficial effects compared with the prior art: First, the multi-resolution spatial features of the panchromatic image are extracted, and then each spatial feature is gradually injected into the multispectral image through the enhanced spatial attention module. Thanks to the enhanced spatial attention mechanism of the enhanced spatial attention module, the fused image greatly retains and enhances the spatial details. Through the Swin-Transformer and the convolutional neural network, the model takes into account the local features and long-range information modeling, so that the fusion result makes full use of the context information. Through the structure of the bidirectional flow information network, the spatial features of the panchromatic image are gradually injected into the multispectral image, and the features of the two are gradually fused. More source image features can be retained than directly fusing the image features of the two to generate a high-spatial-resolution multispectral image, realizing a fusion strategy for the differences of multi-source data and the synergistic effect between local features and long-range information. There are obvious improvements compared with traditional methods in terms of indicators such as spatial correlation coefficient, spectral mapping angle, peak signal-to-noise ratio, and ERGAS, effectively improving the fusion effect. Description of the Drawings In order to more clearly illustrate the technical solution of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic flowchart of a remote sensing image fusion method provided by an embodiment of the present invention.
[0024] Figure 2 It is a schematic diagram of the bidirectional information flow network structure.
[0025] Figure 3 It is a schematic diagram of the Conv Blocks structure of the PAN branch.
[0026] Figure 4 It is a schematic diagram of the ESAFormer structure.
[0027] Figure 5 It is a schematic diagram of the enhanced spatial attention mechanism structure.
[0028] Figure 6 It is a schematic structural diagram of a Swin-Transformer Block (STB).
[0029] Figure 7 It is a schematic block diagram of the structure of a remote sensing image fusion system provided by an embodiment of the present invention.
[0030] Figure 8 It is a schematic structural diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners
[0031] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the specific embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0033] The following explains the key terms that appear in the present invention.
[0034] ESA: Enhanced Spatial Attention, enhanced spatial attention.
[0035] Figure 1 It is a schematic flowchart of a remote sensing image fusion method provided by an embodiment of the present invention. Among them, Figure 1 The execution subject can be a remote sensing image fusion system. The remote sensing image fusion method provided by the embodiment of the present invention is executed by a computer device. Correspondingly, the remote sensing image fusion system runs in the computer device.
[0036] The remote sensing image fusion method of this embodiment is implemented by a two-way information flow network structure. The two-way information flow network structure includes an MS (multispectral) branch and a PAN (panchromatic) branch. The overall process is to extract the multi-resolution spatial information of the PAN image and gradually inject it into the MS image. By gradually integrating the multi-spectral and spatial information of the MS and PAN images, a fused image is generated. As Figure 1As shown, the method includes the following steps S1 - S4, where the panchromatic branch implements step S1, and the multispectral branch implements steps S3 - S4. It should be noted that according to different requirements, the order of the steps in this flowchart can be changed.
[0037] S1, perform multi - resolution spatial information extraction on the panchromatic image to obtain the first - resolution spatial feature and the second - resolution spatial feature of the panchromatic image. The resolution of the second - resolution spatial feature is higher than that of the first - resolution spatial feature.
[0038] The panchromatic branch performs multi - resolution spatial information extraction on the panchromatic image to extract multi - resolution spatial features. In this embodiment, two different - resolution spatial features are obtained, the first - resolution spatial feature and the second - resolution spatial feature, where the resolution of the second - resolution spatial feature is higher than that of the first - resolution spatial feature, that is, the second - resolution spatial feature contains more detailed spatial information.
[0039] S2, perform the first up - sampling on the original multispectral image to obtain the first multispectral image, and use the enhanced spatial attention model to fuse the first multispectral image with the second - resolution spatial feature to obtain the second multispectral image.
[0040] The multispectral branch processes the multispectral image, including up - sampling and image fusion. This step realizes the fusion of the initial up - sampling of the multispectral image and the spatial feature. First, perform the first up - sampling operation on the original multispectral image to improve its resolution and obtain the first multispectral image. Then, use the enhanced spatial attention model to fuse the first multispectral image obtained by up - sampling with the second - resolution spatial feature to generate the second multispectral image. The enhanced spatial attention model can focus on the important spatial information in the image, thereby promoting the fusion of the spatial features of the multispectral image and the panchromatic image to generate the second multispectral image. This process can enhance the spatial resolution and detail performance of the multispectral image.
[0041] S3, perform the second up - sampling on the second multispectral image to obtain the third multispectral image, and use the enhanced spatial attention model to fuse the third multispectral image with the first - resolution spatial feature to obtain the fourth multispectral image.
[0042] This step realizes the further up - sampling and spatial feature fusion of the multispectral image. The second multispectral image is second - up - sampled to obtain the high - resolution third multispectral image, and then the enhanced spatial attention model is used to fuse the third multispectral image with the first - resolution spatial feature to generate the fourth multispectral image. This step further improves the spatial resolution of the multispectral image and fuses a wider range of spatial information, so that the final image has higher spatial details while maintaining the spectral information.
[0043] S4. Upsample the original multi-spectral image for the third time to obtain the fifth multi-spectral image, and perform element-wise addition fusion on the fourth multi-spectral image and the fifth multi-spectral image to obtain the final remote sensing image fusion result.
[0044] This step realizes the final fusion and output. Perform the third upsampling on the initial multi-spectral image to obtain the fifth multi-spectral image with the same resolution as the fourth multi-spectral image, and then perform element-wise addition fusion on the fourth multi-spectral image and the fifth multi-spectral image. This step combines the advantages of the two multi-spectral images, retaining both the spectral information of the initial multi-spectral image and integrating the high-resolution spatial information of the panchromatic image. Finally, through this fusion step, the fusion result of the remote sensing image is obtained, making the result have higher spatial resolution and richer spectral information.
[0045] The remote sensing image fusion method of this embodiment first extracts the multi-resolution spatial features of the panchromatic image, and then gradually injects each spatial feature into the multi-spectral image through the enhanced spatial attention module. Thanks to the enhanced spatial attention mechanism of the enhanced spatial attention module, the fused image greatly retains and enhances the spatial details. Through the Swin-Transformer and convolutional neural network, the model takes into account local features and long-range information modeling, making the fusion result fully utilize the context information. Through the structure of the bidirectional flow information network, the spatial features of the panchromatic image are gradually injected into the multi-spectral image, gradually fusing the features of the two, and more source image features can be retained than directly fusing the image features of the two to generate a multi-spectral image with high spatial resolution, realizing the fusion strategy of the differences of multi-source data and the synergy between local features and long-range information. There are obvious improvements compared with traditional methods in indicators such as spatial correlation coefficient, spectral mapping angle, peak signal-to-noise ratio, and ERGAS, effectively improving the fusion effect.
[0046] To further understand the present invention, the following provides specific embodiments to further elaborate on the present invention in detail.
[0047] Figure 2 It is a schematic diagram of the bidirectional information flow network structure, including a PAN branch and an MS branch. The bidirectional information flow network structure attaches importance to the differences in data sources, and gradually injects the PAN image information features into the MS image in layers, including gradually extracting the PAN image information features, extracting feature maps of two resolutions, and integrating them into the ESAFormer module for feature fusion of the two images, so as to better generate fusion features and utilize the information advantages of the two to generate a fused image.
[0048] The PAN branch adopts two Conv Blocks composed of four convolutional layers and a residual connection to generate spatial features of different resolutions. Figure 3Schematic diagram of the Conv Blocks structure for the PAN branch, including 4 convolutional layers, 1 Relu activation function, and a residual. The first 3 convolutional layers are connected in sequence, and the output result is processed by the Relu activation function and then input into the last convolutional layer through residual connection.
[0049] In some alternative embodiments, during step S1, the PAN branch extracts multi-resolution spatial information from the panchromatic image to obtain the first-resolution spatial feature and the second-resolution spatial feature of the panchromatic image, which specifically includes the following steps.
[0050] S1.1, input the panchromatic image into the first convolutional block for processing to obtain the first-resolution spatial feature.
[0051] S1.2, perform a max-pooling operation on the first-resolution spatial feature and then input it into the second convolutional block for processing to obtain the second-resolution spatial feature.
[0052] It should be noted that the first convolutional block and the second convolutional block are the Conv Blocks of the PAN branch.
[0053] Through the sequential processing of the first convolutional block and the second convolutional block, combined with the max-pooling operation, this method can efficiently extract the spatial features of the panchromatic image at different resolutions. This multi-resolution feature extraction method can capture the detailed information and global structure in the image to understand the details and structures at different scales of the image. From the large-scale image contour to the small-scale texture details can be effectively captured, enriching the image information and providing a spatial feature basis for subsequent image fusion. Among them, the first-resolution spatial feature mainly retains the details and local features of the image, and the second-resolution spatial feature further emphasizes the global structure and key information of the image through the max-pooling operation. When extracting the second-resolution spatial feature, a max-pooling operation is first performed on the first-resolution spatial feature to reduce the data volume, lower the computational cost of subsequent processing by the second convolutional block, improve the computational efficiency, and avoid the overfitting problem caused by excessive data volume.
[0054] The MS branch includes two steps, namely, upsampling and ESAFormer (Enhanced Spatial Attention Model) fusion, corresponding to different-resolution features of the PAN image respectively. In each step, ESAFormer will fuse the spectral and spatial features to generate a fusion result with the corresponding resolution. Finally, the multi-resolution fusion result and the MS image are upsampled by 4 times to obtain the HRMS (High Spatial Resolution Multispectral) image.
[0055] Before ESAFormer fusion, the input image is upsampled by 2 times. Figure 4It is a schematic diagram of the ESAFormer structure. The ESAFormer integrates the spatial information of the PAN image and the spectral information of the MS image and generates fused features. First, the feature maps of the upsampled MS and PAN images are directly superimposed to obtain fused shallow features. Second, the ESA is used to increase the receptive field of the model. Then, the intermediate features are input into the Swin-Transformer Block to further integrate the spectral and spatial information of the fused features. Finally, the initial fused result is generated through two convolutional layers while avoiding excessive spatial artifacts.
[0056] Specifically, using the enhanced spatial attention model to fuse the multispectral image with the spatial features includes the following steps.
[0057] Step 1: Directly superimpose the spatial features and the spectral features of the upsampled multispectral image to obtain fused shallow features.
[0058] Step 2: Perform a convolutional operation on the fused shallow features to obtain and use the enhanced spatial attention mechanism to process to obtain intermediate features.
[0059] Step 3: Input the intermediate features into the Swin-Transformer block to further integrate the spectral and spatial information of the fused features.
[0060] Step 4: Process the integrated features through a group of convolutional layers to obtain the fused result.
[0061] It should be noted that in the above step S2, using the enhanced spatial attention model to fuse the first multispectral image with the second-resolution spatial features to obtain the second multispectral image, and in step S3, using the enhanced spatial attention model to fuse the third multispectral image with the first-resolution spatial features to obtain the fourth multispectral image are both achieved by performing the above steps 1 - step 4.
[0062] Based on the above steps, using the enhanced spatial attention model to fuse the multispectral image with the spatial features is expressed as:
[0063] In the formula, represents the Swin-Transformer block, represents a group of convolutional layers, represents the enhanced spatial attention mechanism (ESA), represents the result of the convolutional operation of the shallow features required for the synthesized fused image.
[0064]
[0065] In the formula, represents the feature map of the panchromatic image, that is, the spatial feature, represents the upsampled multispectral image.
[0066] Figure 5 is the schematic diagram of the enhanced spatial attention mechanism structure. In the figure, represents element-wise multiplication, represents element-wise addition, represents the normalization function. Using the enhanced spatial attention mechanism to process to obtain intermediate features, which specifically include the following steps.
[0067] Step 2.1, use a 1×1 convolutional layer to perform convolutional processing on to reduce the embedding dimension of the features to obtain , expressed as,
[0068] where, represents a 1×1 convolutional layer.
[0069] Step 2.2, use a 3×3 convolutional layer with a stride of 2 to perform convolutional processing on and perform a pooling operation on the convolutional processing result.
[0070] Step 2.3, use a group of 3×3 convolutional layers to perform convolutional processing on the pooling operation result. There is an activation function Relu after each convolutional layer, and perform upsampling on the convolutional processing result to obtain , expressed as:
[0071] where, represents upsampling, represents the pooling operation, represents a group of 3×3 convolutional layers, with an activation function Relu after each convolutional layer, represents a 3×3 convolutional layer with a stride of 2.
[0072] In this step, both the pooling layer and the convolutional layer with a stride of 2 will reduce the spatial dimension, and then the features are restored by the upsampling layer.
[0073] Step 2.4, add and element-wise, use a 1×1 convolutional layer to perform convolutional processing on the addition result, and use the normalization function to process the convolutional processing result and then with Perform element-wise multiplication to obtain intermediate features , which is expressed as:
[0074] Among them, represents a 1×1 convolutional layer used to restore the embedding dimension of the features, is the normalization function, and × represents element-wise multiplication calculation. Due to the operation of the enhanced spatial attention mechanism, these features are more concentrated in the regions of interest in spatial details at the beginning of ESAFormer. After this series of operations, the key regions of the image are aggregated together, greatly retaining and enhancing the spatial details, which is beneficial to the subsequent image fusion work.
[0075] Figure 6 Figure Figure 6 is the structural schematic diagram of the Swin-Transformer Block (STB), including a normalization layer (LN), a window-based multi-head self-attention layer (W-MSA), a sliding window multi-head self-attention layer (SW-MSA), and a multi-layer perceptron layer (MLP).
[0076] Input the intermediate features into the Swin-Transformer block for further integration to fuse the spectral and spatial information of the features, which is expressed as,
[0077] Among them, represents the input of the Swin-Transformer block, that is, the intermediate features , , , , represents the feature maps at different scales. In the process of Swin-Transformer, the input image first passes through an embedding layer to convert the pixel values into an embedding representation that can be processed by the Transformer model. Secondly, the image is divided into a series of windows, and there is some overlap between these windows. The self-attention mechanism within each window is used to capture local features, while the cross-attention mechanism between windows is used to capture long-distance context information. In cross-attention, a window shifting strategy is adopted to further improve information transmission. The LN and MLP components are used to normalize and map the features to better present the features. Finally, residual connections are used to retain and transmit information. In the final stage of ESAFormer, the spectral and spatial information in the fused features will be integrated through two-layer convolutional processing.
[0078] In some alternative embodiments, the remote sensing image fusion method is deployed on the cloud to achieve flexible adjustment of computing resources for specific business requirements and monitoring of the algorithm running status.
[0079] In the above, embodiments of a remote sensing image fusion method have been described in detail. Based on the remote sensing image fusion method described in the above embodiments, an embodiment of the present invention further provides a remote sensing image fusion system corresponding to this method.
[0080] Figure 7 FIG. is a schematic block diagram of the structure of a remote sensing image fusion system provided by an embodiment of the present invention. In this embodiment, the remote sensing image fusion system 700 can be divided into multiple functional units according to the functions it performs, such as Figure 7 shown. The unit referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory.
[0081] The system 700 is implemented based on a bidirectional information flow network structure including a panchromatic branch and a multispectral branch. The spatial feature extraction unit 710 is implemented based on the panchromatic branch, and the first fusion unit 720, the second fusion unit 730, and the third fusion unit 740 are implemented based on the multispectral branch.
[0082] The spatial feature extraction unit 710 is used to extract multi-resolution spatial information from the panchromatic image to obtain the first-resolution spatial feature and the second-resolution spatial feature of the panchromatic image. The resolution of the second-resolution spatial feature is higher than that of the first-resolution spatial feature.
[0083] The first fusion unit 720: performs a first upsampling on the original multispectral image to obtain a first multispectral image, and uses an enhanced spatial attention model to fuse the first multispectral image with the second-resolution spatial feature to obtain a second multispectral image.
[0084] The second fusion unit 730: performs a second upsampling on the second multispectral image to obtain a third multispectral image, and uses an enhanced spatial attention model to fuse the third multispectral image with the first-resolution spatial feature to obtain a fourth multispectral image.
[0085] The third fusion unit 740: performs a third upsampling on the original multispectral image to obtain a fifth multispectral image, and performs an element-wise addition fusion on the fourth multispectral image and the fifth multispectral image to obtain the final remote sensing image fusion result.
[0086] The remote sensing image fusion system of this embodiment is used to implement the aforementioned remote sensing image fusion method. Therefore, the specific implementation manners in this system can be seen in the embodiment part of the remote sensing image fusion method in the foregoing text. Therefore, its specific implementation manners can be referred to the descriptions of the corresponding parts of each embodiment, and will not be elaborated here.
[0087] In addition, since the remote sensing image fusion system of this embodiment is used to implement the aforementioned remote sensing image fusion method, its functions correspond to those of the above method, and will not be elaborated here.
[0088] Figure 8 FIG. 4 is a schematic structural diagram of a terminal 800 provided by an embodiment of the present invention, including: a processor 810, a memory 820, and a communication unit 830. When the processor 810 is used to implement the remote sensing image fusion program stored in the memory 820, the following steps are implemented: S1, performing multi-resolution spatial information extraction on the panchromatic image to obtain a first-resolution spatial feature and a second-resolution spatial feature of the panchromatic image, and the resolution of the second-resolution spatial feature is higher than that of the first-resolution spatial feature; S2, performing first upsampling on the original multi-spectral image to obtain a first multi-spectral image, and using an enhanced spatial attention model to fuse the first multi-spectral image with the second-resolution spatial feature to obtain a second multi-spectral image; S3, performing second upsampling on the second multi-spectral image to obtain a third multi-spectral image, and using an enhanced spatial attention model to fuse the third multi-spectral image with the first-resolution spatial feature to obtain a fourth multi-spectral image; S4, performing third upsampling on the original multi-spectral image to obtain a fifth multi-spectral image, and performing element-wise addition fusion on the fourth multi-spectral image and the fifth multi-spectral image to obtain a final remote sensing image fusion result.
[0089] The present invention also provides a computer storage medium, and the storage medium herein may be a magnetic disk, an optical disk, a read-only memory (abbreviation: ROM), a random access memory (abbreviation: RAM), etc.
[0090] The computer storage medium stores a remote sensing image fusion program, and when the remote sensing image fusion program is executed by a processor, the following steps are implemented: S1, performing multi-resolution spatial information extraction on the panchromatic image to obtain a first-resolution spatial feature and a second-resolution spatial feature of the panchromatic image, and the resolution of the second-resolution spatial feature is higher than that of the first-resolution spatial feature; S2, performing first upsampling on the original multi-spectral image to obtain a first multi-spectral image, and using an enhanced spatial attention model to fuse the first multi-spectral image with the second-resolution spatial feature to obtain a second multi-spectral image; S3, performing second upsampling on the second multi-spectral image to obtain a third multi-spectral image, and using an enhanced spatial attention model to fuse the third multi-spectral image with the first-resolution spatial feature to obtain a fourth multi-spectral image; S4. Thirdly, up-sample the original multi-spectral image to obtain a fifth multi-spectral image, and perform element-wise addition fusion on the fourth multi-spectral image and the fifth multi-spectral image to obtain the final remote sensing image fusion result.
[0091] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image fusion method, characterized in that: The method is implemented by a bidirectional information flow network structure including a panchromatic branch and a multispectral branch, and includes the following steps S1 to S4, wherein the panchromatic branch implements step S1, and the multispectral branch implements steps S3 to S4; S1, performing multi-resolution spatial information extraction on a full-color image to obtain a first resolution spatial feature and a second resolution spatial feature of the full-color image, wherein the resolution of the second resolution spatial feature is higher than that of the first resolution spatial feature; S2, performing a first upsampling on the original multispectral image to obtain a first multispectral image, and using an enhanced spatial attention model to fuse the first multispectral image with the second resolution spatial features to obtain a second multispectral image; S3, performing a second upsampling on the second multispectral image to obtain a third multispectral image, and using an enhanced spatial attention model to fuse the third multispectral image with the first resolution spatial feature to obtain a fourth multispectral image; S4, performing a third upsampling on the original multispectral image to obtain a fifth multispectral image, and performing element-level addition fusion on the fourth multispectral image and the fifth multispectral image to obtain a final remote sensing image fusion result.
2. The remote sensing image fusion method according to claim 1, characterized in that: The panchromatic image is subjected to multi-resolution spatial information extraction to obtain the first resolution spatial feature and the second resolution spatial feature of the panchromatic image, specifically comprising: Inputting the full-color image into the first convolution block for processing to obtain a first-resolution spatial feature; Performing a maximum pooling operation on the first resolution spatial features and then inputting the features into the second convolution block for processing to obtain the second resolution spatial features; Among them, the structures of the first convolution block and the second convolution block are the same, both including 4 convolution layers, 1 Relu activation function and residual. The first 3 convolution layers are connected in sequence, and the output results are processed by the Relu activation function and then input into the last convolution layer through the residual connection.
3. The remote sensing image fusion method according to claim 2, characterized in that: The convolutional layers of the first convolutional block and the second convolutional block are 3×3 convolutional layers.
4. The remote sensing image fusion method according to claim 1 or 2, characterized in that: The enhanced spatial attention model is used to fuse multispectral images with spatial features as follows: In the formula, represents the Swin-Transformer block, represents a set of convolutional layers. represents the enhanced spatial attention mechanism, It represents the result of shallow feature convolution operation required for the synthesized fused image; In the formula, It represents the feature map of the full-color image, that is, the spatial feature. represents an upsampled multispectral image; The fusion process specifically includes: The spatial features are directly superimposed with the spectral features of the upsampled multispectral image to obtain the fused shallow features; The fused shallow features are convolved to obtain , and use the enhanced spatial attention mechanism to Processing to obtain intermediate features; The intermediate features are input into the Swin-Transformer block for further integration to fuse the spectral and spatial information of the features; The integrated features are processed by a set of convolutional layers to obtain the fusion results.
5. The remote sensing image fusion method according to claim 4, characterized in that: Using enhanced spatial attention mechanism Processing is performed to obtain intermediate features, including: Using a 1×1 convolutional layer pair Convolution is performed to reduce the embedding dimension of the feature , expressed as, in, represents a 1×1 convolutional layer; Use a 3×3 convolutional layer with a stride of 2 Perform convolution processing and perform pooling operation on the convolution processing result; A set of 3×3 convolutional layers are used to convolve the pooling operation results. Each convolution layer is followed by an activation function Relu, and the convolution result is upsampled to obtain , expressed as, in, represents upsampling, represents the pooling operation. It represents a set of 3×3 convolutional layers, each of which is followed by an activation function Relu. represents a 3×3 convolutional layer with a stride of 2; Will and Perform element-by-element addition, use a 1×1 convolutional layer to convolve the addition result, and use The normalization function processes the convolution result and Perform element-by-element multiplication to obtain intermediate features , expressed as, in, It represents a 1×1 convolutional layer, and × represents element-by-element multiplication.
6. The remote sensing image fusion method according to claim 5, characterized in that: The intermediate features are input into the Swin-Transformer block for further integration to fuse the spectral and spatial information of the features, expressed as, in, Represents the input of the Swin-Transformer block, i.e., the intermediate features , , , , It represents feature maps at different scales; LN represents the normalization layer, W-MSA represents the window-based multi-head self-attention layer, SW-MSA represents the sliding window multi-head self-attention layer, and MLP represents the multi-layer perception layer.
7. The remote sensing image fusion method according to claim 1 or 2, characterized in that: The first upsampling and the second upsampling are both 2-fold upsampling, and the third upsampling is 4-fold upsampling.
8. A remote sensing image fusion system, characterized in that: The method is implemented based on a bidirectional information flow network structure including a panchromatic branch and a multispectral branch, the spatial feature extraction unit is implemented based on the panchromatic branch, and the first fusion unit, the second fusion unit, and the third fusion unit are implemented based on the multispectral branch, including: A spatial feature extraction unit, configured to perform multi-resolution spatial information extraction on the full-color image to obtain a first resolution spatial feature and a second resolution spatial feature of the full-color image, wherein the resolution of the second resolution spatial feature is higher than that of the first resolution spatial feature; A first fusion unit: performing a first upsampling on the original multispectral image to obtain a first multispectral image, and using an enhanced spatial attention model to fuse the first multispectral image with a second resolution spatial feature to obtain a second multispectral image; A second fusion unit: performing a second upsampling on the second multispectral image to obtain a third multispectral image, and using an enhanced spatial attention model to fuse the third multispectral image with the first resolution spatial feature to obtain a fourth multispectral image; The third fusion unit performs a third upsampling on the original multispectral image to obtain a fifth multispectral image, and performs element-level addition fusion on the fourth multispectral image and the fifth multispectral image to obtain a final remote sensing image fusion result.
9. A terminal, characterized in that: include: A memory, used for storing a remote sensing image fusion program; A processor is used to implement the steps of the remote sensing image fusion method as described in any one of claims 1 to 7 when executing the remote sensing image fusion program.
10. A computer-readable storage medium, characterized in that: The readable storage medium stores a remote sensing image fusion program, and when the remote sensing image fusion program is executed by the processor, the steps of the remote sensing image fusion method according to any one of claims 1 to 7 are implemented.