A panchromatic sharpening method for remote sensing images based on spatial-spectral information enhancement
By combining dynamic deformable convolution and triple attention mechanism with residual network blocks, the problem of adapting to scale changes and texture features in panchromatic sharpening of remote sensing images is solved, achieving efficient spectral and spatial information fusion, and improving the reconstruction quality of remote sensing images and the generalization performance of the model.
Patent Information
- Application Number
- CN202511205974.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing neural network-based panchromatic image sharpening methods lack adaptive mechanisms when dealing with image scale variations and complex texture features, resulting in decreased generalization performance of the model in cross-modal datasets. Furthermore, traditional methods ignore the gradient information of panchromatic images and cannot dynamically adjust static feature fusion strategies, leading to blurred edges and spectral distortion in the reconstructed images.
A feature extraction module is constructed using dynamic deformable convolution, normalization, and ReLU activation function. It combines a triple attention mechanism and residual dense blocks to perform feature fusion through a coupling mechanism. The model is optimized using perceptual loss function and pixel loss function to enhance the extraction and fusion of spectral and spatial information.
It improves the model's adaptability to features at different scales, enhances the ability to capture image details and edge features, improves image edge clarity and spectral fidelity, reduces the risk of overfitting, and improves the model's generalization ability.
Smart Images

Figure CN120689243B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and image processing technology, and in particular to a method for panchromatic sharpening of remote sensing images based on spatial-spectral information enhancement. Background Technology
[0002] Panchromatic sharpening of remote sensing images is an important technique in the field of remote sensing image processing. Its purpose is to fuse multispectral and panchromatic images obtained from remote sensing satellites to generate high spatial resolution multispectral images with rich spectral information. It is widely used in remote sensing of the Earth, and can be used to monitor land erosion, geological disasters, and national defense security in my country. Although significant progress has been made in recent years, many problems still urgently need to be solved due to various issues in the design methods, such as spectral distortion and spatial deformation.
[0003] Deep learning methods based on neural networks have demonstrated superior performance in panchromatic sharpening of remote sensing images. The proposed neural network-based approach achieves dual fidelity in panchromatic sharpening of remote sensing images, preserving both spectral and spatial information. First, neural networks are used to extract spectral features from the multispectral image and spatial features from the panchromatic image, respectively. Then, based on the obtained features and a loss function, a high-resolution multispectral image is learned.
[0004] While neural network-based deep learning methods have made significant progress in panchromatic sharpening of remote sensing images, they still have the following limitations: these methods often ignore the impact of image scale variations and complex texture features. In practical applications, the diversity of target sizes and complex texture structures in scenes (such as densely packed similar objects, irregular surface textures, etc.) require networks to dynamically adjust feature extraction and feature fusion capabilities. Without such an adaptive mechanism, the network will struggle to accurately capture feature changes at different scales, leading to a significant decrease in the model's generalization performance in cross-modal datasets.
[0005] Specifically, these limitations manifest in three main aspects: First, traditional bilinear / bicubic interpolation methods ignore the gradient information of panchromatic images, resulting in blurred edges in the reconstructed images, and fixed-scale convolutional kernels are ill-suited to the needs of multi-scale panchromatic sharpening; second, static feature fusion strategies cannot dynamically adjust feature weights based on texture complexity; and finally, existing dual-branch architectures employ simple addition or concatenation operations, failing to dynamically balance the contributions of spectral and spatial features, leading to spectral distortion in homogeneous areas such as farmland. These factors collectively constrain the robustness of the model in real-world complex scenarios. Summary of the Invention
[0006] In view of the above, the main objective of this invention is to propose a panchromatic sharpening method for remote sensing images based on spatial-spectral information enhancement, so as to solve the above-mentioned technical problems.
[0007] This invention proposes a panchromatic sharpening method for remote sensing images based on spatial-spectral information enhancement, the method comprising the following steps:
[0008] Step 1: Construct a feature extraction module based on dynamic deformable convolution, normalization, and ReLU activation function. Introduce a coupling mechanism into the triple attention mechanism to obtain a fine-tuned triple attention mechanism. Construct a feature fusion module based on residual dense blocks, gating mechanism, and Sobel operator. The feature extraction module, the fine-tuned triple attention mechanism, the residual network block, and the feature fusion module constitute a remote sensing image enhancement fusion model.
[0009] Step 2: Acquire remote sensing images, perform upsampling on the low-resolution multispectral image in the remote sensing image to obtain an upsampled multispectral image, extract features from the upsampled multispectral image and the panchromatic image in the remote sensing image respectively, and stitch them together to obtain a feature map after stitching the low-resolution multispectral image and the panchromatic image.
[0010] Step 3: Use the feature extraction module to extract features from the upsampled multispectral image and panchromatic image respectively, so as to obtain the spectral features of the low-resolution multispectral image and the spatial features of the panchromatic image respectively;
[0011] Step 4: The feature maps obtained by stitching the low-resolution multispectral image and the panchromatic image are processed by a fine-tuned triple attention mechanism and residual network blocks to obtain refined spectral features and refined spatial features respectively.
[0012] Step 5: Use the feature fusion module to process the refined spectral features and refined spatial features to obtain a high-resolution multispectral image;
[0013] Step 6: Construct the perceptual loss function and pixel loss function based on the high-resolution multispectral image, and optimize the remote sensing image enhancement and fusion model using the perceptual loss function and pixel loss function to obtain the optimized remote sensing image enhancement and fusion model. Input the remote sensing image into the optimized remote sensing image enhancement and fusion model for processing to obtain the final high-resolution multispectral image.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0015] 1. The dynamic deformable convolution constructed in this invention can dynamically adjust the shape and weight of the convolution kernel according to the input image, thereby improving the network's adaptability to features of different scales during feature extraction and fusion. This adaptability enables the network to capture image details and edge features more effectively, thereby improving its generalization ability.
[0016] 2. This invention trains the model by constructing pixel loss function and perceptual loss function, enabling the model to better capture the structure and content of the image at the feature level while finely adjusting the pixel details of the image. This method enhances the model's adaptability to different types of remote sensing images and reduces the risk of overfitting.
[0017] 3. The remote sensing image enhancement and fusion model proposed in this invention enhances the spatial and spectral information of the remote sensing image before image fusion, and then fuses the enhanced information, thereby improving the generalization ability of the model. At the same time, by using the gradient of the panchromatic image to enhance the high-frequency details of the fusion features, it overcomes the edge blurring problem in traditional image reconstruction and significantly improves the edge clarity of the image.
[0018] 4. By introducing a coupling mechanism into the triple attention mechanism, this invention can accurately distinguish subtle differences between different channels, especially in areas that are spatially similar but have weak spectral differences. At the same time, it effectively reduces the excessive attention to specific features by a single dimension, suppresses background interference and artifacts, and effectively maintains the spectrum of the source image, thereby improving the accuracy and generalization ability of feature representation.
[0019] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by means of embodiments of the invention. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the steps of a remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement proposed in this invention.
[0021] Figure 2 This is a diagram illustrating the architecture of a remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement proposed in this invention. Detailed Implementation
[0022] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0023] These and other aspects of the embodiments of the present invention will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention; however, it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0024] See also Figure 1 This embodiment provides a method for panchromatic sharpening of remote sensing images based on spatial-spectral information enhancement, the method comprising the following steps:
[0025] Step 1: Construct a feature extraction module based on dynamic deformable convolution, normalization, and ReLU activation function. Introduce a coupling mechanism into the triple attention mechanism to obtain a fine-tuned triple attention mechanism. Construct a feature fusion module based on residual dense blocks, gating mechanism, and Sobel operator. The feature extraction module, the fine-tuned triple attention mechanism, the residual network block, and the feature fusion module constitute a remote sensing image enhancement fusion model.
[0026] It should be noted that, as Figure 2 As shown, the remote sensing image enhancement and fusion model proposed in this invention extracts and fuses features through spectral enhancement and spatial enhancement branches, respectively. This branching structure enables the model to effectively utilize spatial and spectral information in low-resolution multispectral and panchromatic images, thereby enhancing the model's expressive power and adaptability. In the spectral enhancement branch, a fine-tuned triple attention mechanism is employed, allowing the model to adaptively focus on important regions and channels of features, improving the effectiveness and accuracy of feature selection and making it suitable for different scenarios. The residual network block used in the spatial enhancement branch effectively fuses features, reduces information loss, and promotes deep feature learning.
[0027] Furthermore, since the kernel settings of dynamic deformable convolution are more flexible and can adapt to different pixels and image features, this paper uses dynamic deformable convolution to extract image features in the spectral and spatial enhancement branches, making the model more robust when processing images of different scales.
[0028] Step 2: Acquire remote sensing images. Upsample the low-resolution multispectral image in the remote sensing image to obtain an upsampled multispectral image. Extract features from the upsampled multispectral image and the panchromatic image in the remote sensing image respectively, and then stitch them together to obtain a feature map after stitching the low-resolution multispectral image and the panchromatic image.
[0029] See also Figure 2In step 2, a remote sensing image is acquired, and the low-resolution multispectral image in the remote sensing image is upsampled to obtain an upsampled multispectral image. Features are extracted from the upsampled multispectral image and the panchromatic image in the remote sensing image, and then stitched together to obtain a feature map after stitching the low-resolution multispectral image and the panchromatic image. Specifically, the steps include the following sub-steps:
[0030] To obtain an upsampled multispectral image, the low-resolution multispectral image from the remote sensing image is sequentially processed by deconvolution and ReLU activation function. The following relationship exists in the corresponding process:
[0031] ;
[0032] in, This represents an upsampled multispectral image. This indicates that the device has undergone ReLU activation function processing. This indicates that a deconvolution operation has been performed. Represents low-resolution multispectral images;
[0033] The upsampled multispectral image is sequentially processed through convolution and ReLU activation functions to obtain the initial spectral features of the low-resolution multispectral image. The following relationship exists in the corresponding process:
[0034] ;
[0035] in, This represents the initial spectral features of a low-resolution multispectral image. This indicates that a convolution operation has been performed;
[0036] The panchromatic image from the remote sensing image is sequentially processed through convolution and ReLU activation functions to obtain the initial spatial features of the panchromatic image. The following relationship exists in the corresponding process:
[0037] ;
[0038] in, Represents the initial spatial features of a panchromatic image. Represents a panchromatic image;
[0039] The initial spectral features of the low-resolution multispectral image are concatenated with the initial spatial features of the panchromatic image to obtain a feature map after concatenation of the low-resolution multispectral image and the panchromatic image. The following relationship exists in the correspondence process:
[0040] ;
[0041] in, This represents the feature map after stitching together a low-resolution multispectral image and a panchromatic image. This indicates that a splicing operation has been performed.
[0042] It should be noted that, in Figure 2 middle, It indicates addition.
[0043] Step 3: Use the feature extraction module to extract features from the upsampled multispectral image and panchromatic image respectively, so as to obtain the spectral features of the low-resolution multispectral image and the spatial features of the panchromatic image respectively.
[0044] In step 3, the feature extraction module is used to extract features from the upsampled multispectral image and the panchromatic image respectively, so as to obtain the spectral features of the low-resolution multispectral image and the spatial features of the panchromatic image, which specifically includes the following sub-steps:
[0045] The upsampled multispectral image is sequentially processed through dynamic deformable convolution, normalization, and ReLU activation to obtain the spectral features of the low-resolution multispectral image. The following relationship exists in the corresponding process:
[0046] ;
[0047] in, Representing the spectral features of low-resolution multispectral images, This indicates that the normalization operation has been performed. This indicates that the operation has undergone a dynamically deformable convolution.
[0048] The panchromatic image is processed sequentially through dynamic deformable convolution, normalization, and ReLU activation to obtain its spatial features. The following relationship exists in the corresponding process:
[0049] ;
[0050] in, Represents the spatial features of a panchromatic image.
[0051] Furthermore, the dynamically deformable convolution operation includes the following sub-steps:
[0052] Dynamic offsets are generated based on the spatial structure of the input feature map. The sampling positions are adaptively adjusted according to the dynamic offsets using a learnable offset field to obtain the offset convolution positions. Continuous sampling is achieved using bilinear interpolation to obtain the weights of the offset positions. The following relationship exists in the corresponding process:
[0053] ;
[0054] in, This represents the pixel value at the offset position. Indicates the position of the convolution after the offset. Indicates the center position of the convolution kernel. In the standard convolution kernel, the first... Each sampling location Indicates the index of the sampling point. In the standard convolution kernel, the first... The dynamic offset of each position. Indicates the interpolation weights. Indicates the first position closest to the offset position neighboring pixels, This indicates that the input feature map is at the location Pixel values;
[0055] The weights of the convolution kernel are calculated using the softmax activation function, and the following relationship exists in the process:
[0056] ;
[0057] in, Indicates the The weights of each convolutional kernel, Indicates the index of the convolution kernel. This represents the weight parameters generated by the dynamic sensing mechanism;
[0058] The output of the convolution kernel is weighted and fused with its weights to obtain the output feature map. The following relationship exists in this process:
[0059] ;
[0060] in, Indicates the output feature map at position Pixel value at that location, This indicates the number of sampling points in the convolution kernel. Indicates the The convolutional kernel at the _th ... The weight values of each sampling point.
[0061] It should be noted that this invention uses a 5*5 convolution kernel when performing deconvolution, convolution, and dynamically deformable convolution operations.
[0062] Step 4: The feature maps obtained by stitching together the low-resolution multispectral image and the panchromatic image are processed by a fine-tuned triple attention mechanism and residual network blocks to obtain refined spectral features and refined spatial features, respectively.
[0063] In step 4, the feature maps obtained by stitching the low-resolution multispectral image and the panchromatic image are processed by a fine-tuned triple attention mechanism and residual network blocks to obtain refined spectral features and refined spatial features, respectively. This includes the following sub-steps:
[0064] The feature map obtained by stitching the low-resolution multispectral image and the panchromatic image is processed through a fine-tuned triple attention mechanism, then added to the spectral features of the low-resolution multispectral image, and processed again through the fine-tuned triple attention mechanism to obtain refined spectral features. The following relationship exists in the corresponding process:
[0065] ;
[0066] in, Indicates refined spectral characteristics, This indicates a finely tuned triple attention mechanism.
[0067] The feature map obtained by stitching the low-resolution multispectral image and the panchromatic image is processed by a residual network block, added to the spatial features of the panchromatic image, and then processed again by a residual network block to obtain refined spatial features. The following relationship exists in the corresponding process:
[0068] ;
[0069] in, Indicates refined spatial characteristics, This indicates that the data has been processed by residual network blocks.
[0070] Furthermore, the fine-tuned triple attention mechanism specifically includes the following sub-steps:
[0071] The original feature maps are rearranged in terms of channel-height and channel-width dimensions to obtain height rearranged feature maps and width rearranged feature maps, respectively.
[0072] Max pooling and average pooling operations are performed on the original feature maps to obtain a first max pooling feature map and a first average pooling feature map, respectively. These two feature maps are then concatenated and subsequently processed by convolution and the sigmoid activation function to obtain a first attention map. The following relationship exists in this process:
[0073] ;
[0074] in, This represents the first max-pooling feature map. This indicates that the max pooling operation has been performed. Represents the original feature map, and ; Indicates batch size, Indicates the number of channels. Indicates altitude, Indicates width, This represents the first average pooling feature map. This indicates that the average pooling operation has been performed. This represents the first attention map. This indicates that the process has been performed using the Sigmoid activation function;
[0075] Max pooling and average pooling operations are performed on the height rearranged feature maps to obtain second max pooling feature maps and second average pooling feature maps, respectively. These two feature maps are then concatenated and subsequently processed through convolution and the sigmoid activation function to obtain the second attention map. The following relationship exists in this process:
[0076] ;
[0077] in, This represents the second max-pooling feature map. This represents the second average pooling feature map; This represents a highly rearranged feature map, and ; This represents the second attention map;
[0078] Max pooling and average pooling operations are performed on the width rearranged feature map to obtain a third max pooling feature map and a third average pooling feature map, respectively. The third max pooling feature map and the third average pooling feature map are then concatenated and processed by convolution and the sigmoid activation function to obtain a third attention map. The following relationship exists in the corresponding process:
[0079] ;
[0080] in, This represents the third max-pooling feature map. This represents the third average pooling feature map; This represents the width rearranged feature map, and ; This represents the third attention map;
[0081] The first attention map, second attention map, and third attention map are processed separately using a coupling mechanism to obtain enhanced first attention map, enhanced second attention map, and enhanced third attention map, respectively. The following relationship exists in the corresponding process:
[0082] ;
[0083] in, This represents the enhanced first attention map. This represents the enhanced second attention map. This represents an enhanced third attention map;
[0084] The original feature map is weighted element-wise using the enhanced first attention map, and then subjected to an inverse dimension permutation operation to obtain an enhanced original spectral feature map. The height-rearranged feature map is weighted element-wise using the enhanced second attention map, and then subjected to an inverse dimension permutation operation to obtain an enhanced height spectral feature map. The width-rearranged feature map is weighted element-wise using the enhanced third attention map, and then subjected to an inverse dimension permutation operation to obtain an enhanced width spectral feature map. The following relationship exists in this process:
[0085] ;
[0086] in, This represents the enhanced original spectral features. This indicates an enhanced hyperspectral feature map. This indicates an enhanced width spectral feature map. This indicates that the dimensional inverse permutation operation has been performed. This represents element-wise multiplication;
[0087] The enhanced original spectral feature map, the enhanced height spectral feature map, and the enhanced width spectral feature map are weighted and fused to obtain the final output feature map. The following relationship exists in the corresponding process:
[0088] ;
[0089] in, This represents the final output feature map. Indicates the first An enhanced spectral feature map.
[0090] Step 5: Use the feature fusion module to process the refined spectral features and refined spatial features to obtain a high-resolution multispectral image.
[0091] In step 5, the refined spectral features and refined spatial features are processed using the feature fusion module to obtain a high-resolution multispectral image. This process includes the following sub-steps:
[0092] The refined spectral features are input into the residual dense block for feature extraction to obtain the residual features. The following relationship exists in the corresponding process:
[0093] ;
[0094] in, Representing residual characteristics, This indicates that the data has undergone residual dense block processing;
[0095] The residual features are concatenated with the refined spatial features to obtain the gated features. The following relationship exists in the corresponding process:
[0096] ;
[0097] in, Indicates gating characteristics;
[0098] The gated features are sequentially processed through convolution and the Sigmoid function to obtain the gated weights. The following relationship exists in the process:
[0099] ;
[0100] in, Indicates the gating weight;
[0101] The residual features and refined spatial features are weighted and fused based on gated weights to obtain fused features. The following relationship exists in the corresponding process:
[0102] ;
[0103] in, Indicates fusion characteristics;
[0104] The fused features are then subjected to a convolution operation to obtain convolutional features. The following relationship exists in the corresponding process:
[0105] ;
[0106] in, Represents convolutional features;
[0107] The Sobel operator is used to process the panchromatic image to obtain the gradient map. The following relationship exists in the corresponding process:
[0108] ;
[0109] in, Represents the gradient plot. This indicates that the process has been performed using the Sobel operator;
[0110] The convolutional features are multiplied element-wise with the gradient map to obtain high-frequency detail features. The following relationship exists in the corresponding process:
[0111] ;
[0112] in, Indicates high-frequency detail features;
[0113] High-frequency detail features are added to the fused features and then convolutional is performed to obtain a high-resolution multispectral image. The following relationship exists in the corresponding process:
[0114] ;
[0115] in, This represents a high-resolution multispectral image.
[0116] Step 6: Construct the perceptual loss function and pixel loss function based on the high-resolution multispectral image, and optimize the remote sensing image enhancement and fusion model using the perceptual loss function and pixel loss function to obtain the optimized remote sensing image enhancement and fusion model. Input the remote sensing image into the optimized remote sensing image enhancement and fusion model for processing to obtain the final high-resolution multispectral image.
[0117] In step 6, a perceptual loss function and a pixel loss function are constructed based on the high-resolution multispectral image. These functions are then used to optimize the remote sensing image enhancement and fusion model, resulting in an optimized model. The remote sensing image is then input into this optimized model for processing, yielding the final high-resolution multispectral image. The expression for the perceptual loss function is as follows:
[0118] ;
[0119] in, Indicates perceived loss. Indicates the dimension of the feature layer. Represents the target image. This indicates that it has been processed by the VGG16 network. Indicates taking the 2-norm;
[0120] The expression for the pixel loss function is:
[0121] ;
[0122] in, Indicates pixel loss, This represents the total number of pixels in the image;
[0123] Furthermore, the expression for the total loss function is:
[0124] ;
[0125] in, Indicates the total loss. This represents a hyperparameter used to control the weights of perceptual loss and pixel loss.
[0126] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0127] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0128] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0129] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for panchromatic sharpening of remote sensing images based on spatial-spectral information enhancement, characterized in that, The method includes the following steps: Step 1: Construct a feature extraction module based on dynamic deformable convolution, normalization and ReLU activation function, introduce a coupling mechanism into the triple attention mechanism to obtain a fine-tuned triple attention mechanism, and construct a feature fusion module based on residual dense blocks, gating mechanism and Sobel operator. The feature extraction module, the fine-tuned triple attention mechanism, the residual network block, and the feature fusion module constitute the remote sensing image enhancement and fusion model; Step 2: Acquire remote sensing images, perform upsampling on the low-resolution multispectral image in the remote sensing image to obtain an upsampled multispectral image, extract features from the upsampled multispectral image and the panchromatic image in the remote sensing image respectively, and stitch them together to obtain a feature map after stitching the low-resolution multispectral image and the panchromatic image. Step 3: Use the feature extraction module to extract features from the upsampled multispectral image and panchromatic image respectively, so as to obtain the spectral features of the low-resolution multispectral image and the spatial features of the panchromatic image respectively; Step 4: The feature maps obtained by stitching the low-resolution multispectral image and the panchromatic image are processed by a fine-tuned triple attention mechanism and residual network blocks to obtain refined spectral features and refined spatial features respectively. Step 5: Use the feature fusion module to process the refined spectral features and refined spatial features to obtain a high-resolution multispectral image; Step 6: Based on the high-resolution multispectral image, construct the perceptual loss function and the pixel loss function respectively. Optimize the remote sensing image enhancement and fusion model using the perceptual loss function and the pixel loss function to obtain the optimized remote sensing image enhancement and fusion model. Input the remote sensing image into the optimized remote sensing image enhancement and fusion model for processing to obtain the final high-resolution multispectral image. In step 5, the refined spectral features and refined spatial features are processed using a feature fusion module to obtain a high-resolution multispectral image. This process includes the following sub-steps: The refined spectral features are input into the residual dense block for feature extraction to obtain the residual features; The residual features are concatenated with the refined spatial features to obtain the gated features; The gated features are processed sequentially through convolution and the Sigmoid function to obtain the gated weights. The residual features and refined spatial features are weighted and fused based on gating weights to obtain fused features; The fused features are then subjected to a convolution operation to obtain convolutional features; The Sobel operator is used to process the panchromatic image to obtain the gradient map; The convolutional features are multiplied element-wise with the gradient map to obtain high-frequency detail features; High-frequency detail features are added to the fused features and then convolutional is performed to obtain a high-resolution multispectral image.
2. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 1, characterized in that, In step 2, a remote sensing image is acquired, and the low-resolution multispectral image in the remote sensing image is upsampled to obtain an upsampled multispectral image. Features are extracted from the upsampled multispectral image and the panchromatic image in the remote sensing image, and then stitched together to obtain a feature map after stitching the low-resolution multispectral image and the panchromatic image. Specifically, this includes the following sub-steps: The remote sensing image is acquired, and the low-resolution multispectral image in the remote sensing image is processed by deconvolution and ReLU activation function in sequence to obtain the upsampled multispectral image. The upsampled multispectral image is processed sequentially through convolution and ReLU activation functions to obtain the initial spectral features of the low-resolution multispectral image. The panchromatic image in the remote sensing image is processed sequentially through convolution and ReLU activation function to obtain the initial spatial features of the panchromatic image; The initial spectral features of the low-resolution multispectral image are stitched together with the initial spatial features of the panchromatic image to obtain a feature map after stitching the low-resolution multispectral image and the panchromatic image.
3. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 2, characterized in that, In the process of acquiring remote sensing images and sequentially processing the low-resolution multispectral images from the remote sensing images through deconvolution and ReLU activation functions to obtain upsampled multispectral images, the following relationship exists: ; in, This represents an upsampled multispectral image. This indicates that the device has undergone ReLU activation function processing. This indicates that a deconvolution operation has been performed. Represents low-resolution multispectral images; In the process of sequentially processing the upsampled multispectral image through convolution and ReLU activation functions to obtain the initial spectral features of the low-resolution multispectral image, the following relationship exists: ; in, This represents the initial spectral features of a low-resolution multispectral image. This indicates that a convolution operation has been performed; In the process of sequentially processing the panchromatic image from the remote sensing image through convolution and ReLU activation to obtain the initial spatial features of the panchromatic image, the following relationship exists: ; in, Represents the initial spatial features of a panchromatic image. Represents a panchromatic image; In the step of concatenating the initial spectral features of a low-resolution multispectral image with the initial spatial features of a panchromatic image to obtain a feature map after concatenation of the low-resolution multispectral image and the panchromatic image, the following relationship exists: ; in, This represents the feature map after stitching together a low-resolution multispectral image and a panchromatic image. This indicates that a splicing operation has been performed.
4. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 3, characterized in that, In step 3, the feature extraction module is used to extract features from the upsampled multispectral image and the panchromatic image respectively, so as to obtain the spectral features of the low-resolution multispectral image and the spatial features of the panchromatic image, which specifically includes the following sub-steps: The upsampled multispectral image is sequentially processed through dynamic deformable convolution, normalization, and ReLU activation to obtain the spectral features of the low-resolution multispectral image. The following relationship exists in the corresponding process: ; in, Representing the spectral features of low-resolution multispectral images, This indicates that the normalization operation has been performed. This indicates that the operation has undergone a dynamically deformable convolution. The panchromatic image is processed sequentially through dynamic deformable convolution, normalization, and ReLU activation to obtain its spatial features. The following relationship exists in the corresponding process: ; in, Represents the spatial features of a panchromatic image.
5. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 4, characterized in that, In step 4, the feature maps obtained by stitching the low-resolution multispectral image and the panchromatic image are processed by a fine-tuned triple attention mechanism and residual network blocks to obtain refined spectral features and refined spatial features, respectively. Specifically, this includes the following sub-steps: The feature map obtained by stitching the low-resolution multispectral image and the panchromatic image is processed through a fine-tuned triple attention mechanism, then added to the spectral features of the low-resolution multispectral image, and processed again through the fine-tuned triple attention mechanism to obtain refined spectral features. The following relationship exists in the corresponding process: ; in, Indicates refined spectral characteristics, This indicates a finely tuned triple attention mechanism. The feature map obtained by stitching the low-resolution multispectral image and the panchromatic image is processed by a residual network block, added to the spatial features of the panchromatic image, and then processed again by a residual network block to obtain refined spatial features. The following relationship exists in the corresponding process: ; in, Indicates refined spatial characteristics, This indicates that the data has been processed by residual network blocks.
6. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 5, characterized in that, The fine-tuned triple attention mechanism specifically includes the following sub-steps: The original feature maps are rearranged in terms of channel-height and channel-width dimensions to obtain height rearranged feature maps and width rearranged feature maps, respectively. Max pooling and average pooling operations are performed on the original feature maps to obtain a first max pooling feature map and a first average pooling feature map, respectively. The first max pooling feature map and the first average pooling feature map are concatenated and then processed by convolution and sigmoid activation function to obtain a first attention map. Max pooling and average pooling operations are performed on the height rearranged feature map to obtain a second max pooling feature map and a second average pooling feature map, respectively. The second max pooling feature map and the second average pooling feature map are concatenated and then processed by convolution and sigmoid activation function to obtain a second attention map. Max pooling and average pooling operations are performed on the width rearranged feature map to obtain the third max pooling feature map and the third average pooling feature map, respectively. The third max pooling feature map and the third average pooling feature map are concatenated and then processed by convolution and sigmoid activation function to obtain the third attention map. The first attention map, the second attention map, and the third attention map are processed separately using a coupling mechanism to obtain the enhanced first attention map, the enhanced second attention map, and the enhanced third attention map, respectively. The original feature map is weighted element-wise using the enhanced first attention map and subjected to inverse dimension permutation to obtain an enhanced original spectral feature map; the height rearranged feature map is weighted element-wise using the enhanced second attention map and subjected to inverse dimension permutation to obtain an enhanced height spectral feature map; the width rearranged feature map is weighted element-wise using the enhanced third attention map and subjected to inverse dimension permutation to obtain an enhanced width spectral feature map. The enhanced original spectral feature map, the enhanced height spectral feature map, and the enhanced width spectral feature map are weighted and fused to obtain the final output feature map.
7. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 6, characterized in that, In the process of performing max pooling and average pooling operations on the original feature maps to obtain a first max pooling feature map and a first average pooling feature map, respectively, concatenating the first max pooling feature map and the first average pooling feature map, and then processing them through convolution and the Sigmoid activation function to obtain the first attention map, the following relationship exists: ; in, This represents the first max-pooling feature map. This indicates that the max pooling operation has been performed. Represents the original feature map. This represents the first average pooling feature map. This indicates that the average pooling operation has been performed. This represents the first attention map. This indicates that the process has been performed using the Sigmoid activation function; In the process of performing max pooling and average pooling operations on the height rearranged feature map to obtain a second max pooling feature map and a second average pooling feature map, respectively, concatenating the second max pooling feature map and the second average pooling feature map, and then processing them through convolution and the Sigmoid activation function to obtain the second attention map, the following relationship exists: ; in, This represents the second max-pooling feature map. This represents the second average pooling feature map; This represents a highly rearranged feature map. Represents the second attention map; In the process of performing max pooling and average pooling operations on the width rearranged feature map to obtain the third max pooling feature map and the third average pooling feature map respectively, concatenating the third max pooling feature map and the third average pooling feature map, and then processing them through convolution and the Sigmoid activation function to obtain the third attention map, the following relationship exists: ; in, This represents the third max-pooling feature map. This represents the third average pooling feature map; This represents the width rearrangement feature map. This represents the third attention map; In the steps of processing the first attention map, the second attention map, and the third attention map separately using the coupling mechanism to obtain the enhanced first attention map, the enhanced second attention map, and the enhanced third attention map, the following relationship exists: ; in, This represents the enhanced first attention map. This represents the enhanced second attention map. This represents an enhanced third attention map; In the steps of using the enhanced first attention map to perform element-wise weighting and inverse dimension permutation on the original feature map to obtain the enhanced original spectral feature map; using the enhanced second attention map to perform element-wise weighting and inverse dimension permutation on the height rearranged feature map to obtain the enhanced height spectral feature map; and using the enhanced third attention map to perform element-wise weighting and inverse dimension permutation on the width rearranged feature map to obtain the enhanced width spectral feature map, the following relationship exists: ; in, This represents the enhanced original spectral features. This indicates an enhanced hyperspectral feature map. This indicates an enhanced width spectral feature map. This indicates that the dimensional inverse permutation operation has been performed. This represents element-wise multiplication; In the step of weighted averaging and fusing the enhanced original spectral feature map, the enhanced height spectral feature map, and the enhanced width spectral feature map to obtain the final output feature map, the following relationship exists: ; in, This represents the final output feature map. Indicates the first An enhanced spectral feature map.
8. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 1, characterized in that, In the step of inputting refined spectral features into residual dense blocks for feature extraction to obtain residual features, the following relationship exists: ; in, Representing residual characteristics, This indicates that the data has undergone residual dense block processing; In the step of concatenating residual features with refined spatial features to obtain gated features, the following relationship exists: ; in, Indicates gating characteristics; In the process of sequentially processing the gated features through convolution and the Sigmoid function to obtain the gated weights, the following relationship exists: ; in, Indicates the gating weight; In the step of weighted fusion of residual features and refined spatial features based on gated weights to obtain fused features, the following relationship exists: ; in, Indicates fusion characteristics; In the step of performing a convolution operation on the fused features to obtain convolutional features, the following relationship exists: ; in, Represents convolutional features; In the step of processing a panchromatic image using the Sobel operator to obtain a gradient map, the following relationship exists: ; in, Represents the gradient plot. This indicates that the process has been performed using the Sobel operator; In the step of element-wise multiplication of convolutional features with gradient maps to obtain high-frequency detail features, the following relationship exists: ; in, Indicates high-frequency detail features; In the step of adding high-frequency detail features and fused features and performing convolution to obtain a high-resolution multispectral image, the following relationship exists: ; in, This represents a high-resolution multispectral image.
9. The remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement according to claim 8, characterized in that, In step 6, a perceptual loss function and a pixel loss function are constructed based on the high-resolution multispectral image. These functions are then used to optimize the remote sensing image enhancement and fusion model, resulting in an optimized model. The remote sensing image is then input into the optimized model for processing to obtain the final high-resolution multispectral image. The expression for the perceptual loss function is as follows: ; in, Indicates perceived loss. Indicates the dimension of the feature layer. Represents the target image. This indicates that it has been processed by the VGG16 network. Indicates taking the 2-norm; The expression for the pixel loss function is as follows: ; in, Indicates pixel loss, This represents the total number of pixels in the image.
Citation Information
Patent Citations
Remote sensing image fusion method and system based on multi-scale dynamic convolutional neural network
CN111080567A
Remote sensing panchromatic sharpening method and system based on cross spectrum-space fusion network
CN117274093A