A Hyperspectral Panchromatic Sharpening Method and System Based on U-Shaped Convolutional Neural Network
Through the design of the U-shaped convolutional neural network, the problems of spatial information loss and high computational complexity in the full-color sharpening of hyperspectral images are solved, efficient spatial resolution improvement and feature expression are achieved, and computing resource consumption and running time are reduced.
Patent Information
- Application Number
- CN202311100155.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-08-29
AI Technical Summary
When the existing hyperspectral image full-color sharpening methods improve spatial resolution, there are problems such as spatial information loss, spectral distortion, high computational complexity and long inference running time, and fail to make full use of multi-scale information and correlation.
Using a hyperspectral full-color sharpening method based on U-shaped convolutional neural network, the UNet backbone network, channel cross-connection and space-spectral attention network are designed, and the loss function is combined to achieve the fusion of low-spatial resolution hyperspectral images and full-color images.
While retaining more spatial and spectral information, the network computing complexity and inference running time are reduced, spatial resolution is improved, feature expression ability and training stability are enhanced.
Smart Images

Figure CN117237212B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a hyperspectral panchromatic sharpening method and system based on a U-shaped convolutional neural network. Background Art
[0002] Hyperspectral images (HSIs) have very high spectral resolution and can well reflect the characteristics of different materials, and have important application values in many fields such as fine classification of ground objects, medical diagnosis, anomaly detection, cultural relic identification and protection. However, due to the hardware limitations of hyperspectral sensors, HSIs often have low spatial resolution, which limits their application in scenarios with high spatial resolution requirements. Therefore, the hyperspectral image super-resolution technology aiming at improving spatial resolution has become a research hotspot in recent years. At present, some studies have tried to fuse HSIs with PANs to improve the spatial resolution of HSIs. Such methods are called hyperspectral panchromatic sharpening. Although some achievements have been made, there are still some deficiencies, such as: the inability to fully utilize multi-scale spatial and spectral information, there is a certain degree of spatial information loss and spectral distortion, the computational complexity of the network model is too high, and the inference running time of the model is too long, etc.
[0003] Existing panchromatic sharpening methods can generally be divided into traditional panchromatic sharpening methods and deep learning-based methods. Traditional panchromatic sharpening methods are further divided into component substitution methods, multi-scale analysis methods, Bayesian estimation methods, and matrix decomposition methods. The component substitution method can better retain spatial information, but the spectral information will be distorted to a certain extent. The multi-scale analysis method will lose some spatial information and cause ringing phenomena. Both the Bayesian estimation method and the matrix decomposition method belong to methods based on model optimization and solution. Compared with the component substitution method and the multi-scale analysis method, such methods can better retain the spatial and spectral information of the image. However, due to the very large amount of calculation, this method also requires more computing resources.
[0004] In recent years, some studies have gradually introduced various deep learning methods into the field of hyperspectral pan-sharpening. For the patent "Hyperspectral Image Pansharpening Method Based on Spectral Dimension Controlled Convolutional Neural Network" (Application No.: 201811377992.4) applied by South China University of Technology, spectral features are first extracted from the interpolated and upsampled LR-HSI through two convolutional layers, and then combined with the PAN. Then, fusion and reconstruction are carried out through multiple convolutional layers. This method successfully extracts different features of each input image, but does not fully consider the correlation between the LR-HSI and the PAN during the feature fusion process. Yuxuan Zheng et al. from Xidian University proposed a hyperspectral pan-sharpening method based on Deep Hyperspectral Prior (DHP) and Spatial-Spectral Dual Attention Residual Network (DARN) in the paper "Hyperspectral Pansharpening Using DeepPrior and Dual Attention Residual Network" (IEEE Transactions on Geoscience and Remote Sensing, 2020, 58(11): 8059-8076.). This method only uses spectral constraints during the DHP process and does not consider spatial constraints. In the fusion network, only feature maps of a single scale are used for fusion, without considering multi-scale feature information. Bandara et al. from Johns Hopkins University in the United States proposed a new spatial constraint for the upsampling process of Deep Image Prior (DIP) and a network HyperKite for residual reconstruction in the paper "Hyperspectral Pansharpening Based on Improved Deep Image Prior and Residual Reconstruction" (IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 1-16.). However, the DIP network results in extremely long model inference running times, and the network architecture design of the HyperKite network, which upsamples layer by layer and then downsamples layer by layer, also causes a huge computational burden.
[0005] To overcome these deficiencies, this chapter proposes a hyperspectral pan-sharpening method and system based on a U-shaped convolutional neural network, which fuses a hyperspectral image (HSI) with a lower spatial resolution and a panchromatic image (PAN) with a higher spatial resolution to improve the spatial resolution of the hyperspectral image with a lower spatial resolution. Summary of the Invention
[0006] The purpose of this application is to provide a hyperspectral panchromatic sharpening method and system based on a U-shaped convolutional neural network, aiming to solve the above problems.
[0007] To achieve the above objective, this application provides the following technical solutions:
[0008] This application provides a hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network, including:
[0009] Obtain a hyperspectral image with low spatial resolution;
[0010] Use the hyperspectral image as input to generate a training set and a test set;
[0011] Design a hyperspectral panchromatic sharpening neural network model;
[0012] Train the hyperspectral panchromatic sharpening neural network model;
[0013] Input the test images of the test set into the trained hyperspectral panchromatic sharpening neural network model to obtain a hyperspectral image with high spatial resolution.
[0014] Further, in the step of using the hyperspectral image as input to generate a training set and a test set, it specifically includes the following steps:
[0015] Select the upper left corner A×B pixel area of the hyperspectral image and cut it into n non-overlapping H×W pixel area image sub-blocks without overlapping to form the reference HR-HSI image of the data set;
[0016] Use a Gaussian filter with a kernel size of 8×8 and a standard deviation of σ to blur the HR-HSI image, perform 4-fold downsampling to obtain an LR-HSI image with a spatial resolution of h×w, and obtain the PAN image by averaging the panchromatic band of the HR-HSI image in the spectral dimension;
[0017] Randomly select 75% of the image pairs as the training set, and the remaining 25% as the test set;
[0018] The standard deviation σ of the Gaussian filter is calculated by the following formula:
[0019]
[0020] where β is the downsampling scale factor, that is, the linear spatial resolution ratio of the reference HR-HSI image to the generated LR-HSI.
[0021] Further, in the step of designing the hyperspectral panchromatic sharpening neural network model, it specifically includes the following steps:
[0022] Design a UNet backbone network;
[0023] Design a channel cross-connection method;
[0024] Design a spatial-spectral attention network;
[0025] Design a loss function.
[0026] Furthermore, the UNet backbone network consists of an encoder part, a decoder part, and a bottleneck layer, including four scales. The encoders and decoders of the first three scales are connected using a spatial-spectral attention network, and the fourth scale consists of a bottleneck layer;
[0027] The encoder part consists of 3 encoders, each encoder including 1 convolutional module and 1 downsampling module; the decoder part consists of 3 decoders, each decoder including 1 upsampling module and 1 convolutional module, and the bottleneck layer includes 1 convolutional module;
[0028] The formula of the convolutional module is:
[0029] CB out = f CB (CB in ) = δ(BN(Conv 3×3 (CB in )))
[0030] where CB in and CB out are the input and output of the convolutional module respectively, f CB is the function representation of the convolutional module, δ is the LeakyReLU activation function layer, BN is the batch normalization layer, and Conv 3×3 is a 3×3 convolutional layer;
[0031] Design a 1×1 convolutional layer for residual map reconstruction after the decoder part, and the formula is expressed as:
[0032] X res = Conv 1×1 (D3)
[0033] where Conv 1×1 is a 1×1 convolutional layer, X res ∈R H×W×C is the hyperspectral residual image obtained after reconstruction, and D3 is the output feature map of the third decoder.
[0034] Furthermore, the channel cross-connection method includes two channel cross-connection methods: Input CCC and Feature CCC;
[0035] The input of the Input CCC is two different source images, Up-HSI and PAN. Up-HSI is split into m parts along the channel dimension, and the number of spectral bands is:
[0036]
[0037] When m ≥ 2, C1 = … = C m-1 ≥ C m ; when m = 1, C1 = C, and at this time, Up-HSI is not split;
[0038] The formula representation of the splitting process is:
[0039] U1, U2, …, U m-1 , U m = Split(U)
[0040] Among them, is the tensor corresponding to Up-HSI, Split is the operation of tensor splitting, are the sub-tensors obtained by splitting U;
[0041] After splitting, the sub-tensor U i with C i channel numbers is inserted with the PAN tensor P with C p = 1 channel number, and they are sequentially connected in the channel dimension to obtain the output tensor O. Then the number of channels of the output tensor O is:
[0042]
[0043] The process of sequentially connecting in the channel dimension is represented by the formula:
[0044] O = Concat(U1, P, U2, P, …, U m , P)
[0045] Among them, is the tensor corresponding to PAN, Concat is the operation of channel connection, is the tensor output by channel cross-connection;
[0046] The operation process of cross-connecting Up-HSI and PAN in the channel dimension is represented by the formula:
[0047] O = InputCCC(U, P)
[0048] Among them, InputCCC is the operation process of Input CCC;
[0049] The input of the Feature CCC is two feature maps Feature1 and Feature2 at different levels. Feature1 and Feature2 are respectively evenly divided into n = 2 k (0 ≤ k ≤ 5) equal parts, that is:
[0050]
[0051]
[0052] where C qi is the number of spectral bands of the feature map Feature1, and C rj is the number of spectral bands of the feature map Feature1; when n = 1, that is, k = 0, the feature maps Feature1 and Feature2 do not perform the splitting operation;
[0053] The formula for the splitting process is expressed as:
[0054] Q1, Q2,..., Q n = Split(Q)
[0055] R1, R2,..., R n = Split(R)
[0056] where, is the tensor corresponding to Feature1, is the tensor corresponding to Feature2, Split is the operation of tensor splitting, are the sub-tensors obtained by splitting Q, are the sub-tensors obtained by splitting R;
[0057] After the sub-tensor Q qi with C i channels is split, the sub-tensor R rj with C j channels is inserted, and the output tensor O' is obtained by connecting them in sequence along the channel dimension. Then the number of channels of the output tensor O' is:
[0058]
[0059] The process of connecting them in sequence along the channel dimension is expressed by the formula:
[0060] O' = Concat(Q1, R1, Q2, R2,..., Q n , R n )
[0061] where Concat is the operation of channel connection, is the tensor output by channel cross-connection;
[0062] The operation process of cross - connecting the feature map Feature1 and the feature map Feature2 in the channel dimension is expressed by the formula:
[0063] O′ = FeaCCC(Q, R)
[0064] Where FeaCCC represents the operation process of Feature CCC.
[0065] Furthermore, the spatial - spectral attention network is located between the encoder and the decoder of the UNet backbone network and is composed of N sequentially stacked Res - SSA modules based on the residual spatial - spectral attention mechanism, which is expressed by the formula:
[0066]
[0067] Where F0 is the input feature map of SSA - Net, F N is the output feature map of SSA - Net, is the function representation of the k - th Res - SSA module.
[0068] Furthermore, the loss function adopts to represent the mean absolute error between the reconstructed image and the reference hyperspectral image. The calculation formula is:
[0069]
[0070] Where, and Y d are the d - th reconstructed HR - HSI and the reference HSI respectively, D is the number of images, Θ is the parameter of the hyperspectral pan - sharpening neural network model, Φ is the hyperspectral pan - sharpening neural network model, and ‖·‖1 is the l1 norm.
[0071] Furthermore, in the step of training the hyperspectral pan - sharpening neural network model, it specifically includes the following steps:
[0072] Input a training set {[X1, P1, Y1], …, [X D , P D , Y D} with D image pairs, and process it through the hyperspectral pan - sharpening neural network model Φ to obtain an output image set
[0073] Optimize the parameter Θ of the hyperspectral pan - sharpening neural network model through an optimization algorithm. The training process of the hyperspectral pan - sharpening neural network model is expressed by the formula:
[0074]
[0075] Among them, is the optimized parameter, and Loss(a, b) refers to the loss function adopted by the network, which evaluates the gap between a and b.
[0076] Furthermore, in the step of inputting the test images of the test set into the trained hyperspectral panchromatic sharpening neural network model to obtain a hyperspectral image with high spatial resolution, the following steps are specifically included:
[0077] Input the test image pair [X t , P t , and process it through the trained hyperspectral panchromatic sharpening neural network model to output the fused image The test process is expressed by the formula:
[0078]
[0079] Among them, X t and P t are the input LR-HSI and PAN respectively, is the hyperspectral image HR-HSI with high spatial resolution output by the hyperspectral panchromatic sharpening neural network model.
[0080] This application also provides a hyperspectral panchromatic sharpening system based on a U-shaped convolutional neural network, including:
[0081] Acquisition module: Acquire a hyperspectral image with low spatial resolution;
[0082] Input module: Take the hyperspectral image as input to generate a training set and a test set;
[0083] Design module: Design a hyperspectral panchromatic sharpening neural network model;
[0084] Training module: Train the hyperspectral panchromatic sharpening neural network model;
[0085] Output module: Input the test images of the test set into the trained hyperspectral panchromatic sharpening neural network model to obtain a hyperspectral image with high spatial resolution.
[0086] This application provides a hyperspectral panchromatic sharpening method and system based on a U-shaped convolutional neural network, having the following beneficial effects:
[0087] This application constructs a U-shaped convolutional neural network CCC-SSA-UNet for hyperspectral panchromatic sharpening, which can reduce the size of intermediate feature maps, lower the computational complexity of the network, reduce the consumption of computing resources, and shorten the inference running time of the network model; receives low-resolution hyperspectral images and panchromatic images as initial inputs, uses bilinear interpolation to upsample the LR-HSI to obtain the image Up-HSI, and then cross-connects Up-HSI and PAN in the channel direction and inputs them into an encoder-decoder network with a spatial-spectral residual attention mechanism to learn the residual image, which can further enhance the expression ability of spatial and spectral features; among them, the spectral attention module can filter out the spectral information in the feature tensor that is not important for the fusion result and enable the network to adaptively select important spectral information; the spatial attention module can adaptively enable the network to pay more attention to the features in the regions closely related to the enhancement of hyperspectral image spatial details; by combining channel attention and spatial attention and embedding them into the basic residual module, it is possible to improve the spatial-spectral feature expression ability of the network while improving the stability of network training and accelerating the convergence speed of the network model. Finally, the residual image is added to Up-HSI pixel by pixel to obtain the final fusion result; by fusing the hyperspectral image with a lower spatial resolution and the panchromatic image with a higher spatial resolution, the spatial resolution of the hyperspectral image with a lower spatial resolution is improved, and it has a lower computational complexity and a shorter inference running time while retaining more spatial and spectral information. Description of the Drawings
[0088] Figure 1 It is a schematic flowchart of a hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to Embodiment 1 of this application;
[0089] Figure 2 It is a schematic structural diagram of the hyperspectral panchromatic sharpening neural network model CCC-SSA-UNet according to Embodiment 1 of this application;
[0090] Figure 3 It is a schematic diagram of the input source image channel cross-connection method Input CCC according to Embodiment 1 of this application;
[0091] Figure 4 It is a schematic diagram of the feature map channel cross-connection method Feature CCC according to Embodiment 1 of this application;
[0092] Figure 5 It is a schematic structural diagram of the spatial-spectral attention network SSA-Net according to Embodiment 1 of this application;
[0093] Figure 6 It is a schematic structural diagram of the Res-SSA module based on the residual spatial-spectral attention mechanism according to Embodiment 1 of this application;
[0094] Figure 7 Schematic diagram of the training stage of the hyperspectral panchromatic sharpening neural network model according to Embodiment 1 of the present application;
[0095] Figure 8 Schematic diagram of the testing stage of the hyperspectral panchromatic sharpening neural network model according to Embodiment 1 of the present application;
[0096] Figure 9 Schematic diagram of the structure of a hyperspectral panchromatic sharpening system based on a U-shaped convolutional neural network according to Embodiment 2 of the present application. Specific embodiments
[0097] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0098] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0099] Embodiment 1
[0100] Please refer to Figure 1 , which is a schematic flow chart of a hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to Embodiment 1 of the present application; the specific steps include:
[0101] S1: Obtain a hyperspectral image with low spatial resolution.
[0102] In this embodiment, the hyperspectral image with low spatial resolution is an existing publicly available hyperspectral remote sensing image collected on the network, or a hyperspectral image collected by a hyperspectral sensor is used as test data.
[0103] S2: Use the hyperspectral image as input to generate a training set and a test set.
[0104] In this embodiment, for the obtained hyperspectral image with low spatial resolution, an A×B pixel area at the upper left corner of the hyperspectral image is selected and cut into n non-overlapping image sub-blocks of H×W pixels to form a reference HR-HSI image of the data set;
[0105] Generate a PAN image and an LR-HSI image corresponding to each HR-HSI image using the Wald protocol: First, blur the HR-HSI image with a Gaussian filter with a kernel size of 8×8 and a standard deviation of σ, then perform 4-fold downsampling to obtain an LR-HSI image with a spatial resolution of h×w, and obtain the PAN image by averaging the panchromatic band of the HR-HSI image in the spectral dimension; randomly select 75% of the image pairs as the training set, and the remaining 25% as the test set.
[0106] The standard deviation σ of the Gaussian filter used to generate the LR-HSI is calculated by the following formula:
[0107]
[0108] where β is the downsampling scale factor, that is, the ratio of the linear spatial resolution of the reference HR-HSI image to the generated LR-HSI.
[0109] S3: Design a hyperspectral pan-sharpening neural network model.
[0110] In this embodiment, in the step of designing the hyperspectral pan-sharpening neural network model, it specifically includes the following steps S31 to S34, and the implementation methods of each step are described in detail below.
[0111] Please refer to Figure 2 , which is a schematic diagram of the structure of the hyperspectral pan-sharpening neural network model CCC-SSA-UNet of Embodiment 1 of this application; build a U-shaped convolutional neural network CCC-SSA-UNet for hyperspectral pan-sharpening, receive a low-spatial-resolution hyperspectral image (LR-HSI) and a panchromatic image (PAN) as initial inputs, and use a high-spatial-resolution hyperspectral image (HR-HSI) as the final output. The designed hyperspectral pan-sharpening neural network model CCC-SSA-UNet is mainly composed of four parts: UNet based on the encoder-decoder architecture, channel cross-connection methods Input CCC and Feature CCC, SSA-Net based on the residual spatial-spectral attention module, and the loss function.
[0112] S31: Design the UNet backbone network.
[0113] The UNet backbone network consists of an encoder part, a decoder part, and a bottleneck layer, including four scales. The encoders and decoders of the first three scales are connected by a spatial-spectral attention network, and the fourth scale consists of a bottleneck layer. The encoder part is composed of 3 encoders, each encoder includes 1 convolutional module and 1 downsampling module. The decoder part is composed of 3 decoders, each decoder includes 1 upsampling module and 1 convolutional module. The bottleneck layer includes 1 convolutional module. The skip connection between the encoder and decoder at each scale is composed of SSA-Net. There is also a bottleneck layer between the last downsampling module and the first upsampling module, which is composed of 1 convolutional module.
[0114] The convolutional module is composed of a 3×3 convolutional layer with a stride of 1, a batch normalization layer, and a LeakyReLU activation function layer connected in sequence. It is used to extract feature representations in the encoder and to reconstruct features in the decoder. The formula for the convolutional module is:
[0115] CB out =f CB (CB in )=δ(BN(Conv 3×3 (CB in )))
[0116] Where CB in and CB out are the input and output of the convolutional module respectively, f CB is the function representation of the convolutional module, δ is the LeakyReLU activation function layer, BN is the batch normalization layer, and Conv 3×3 is the 3×3 convolutional layer.
[0117] In the encoder, the output image of InputCCC first passes through the first convolutional module to obtain the feature map of the first level then passes through the downsampling layer to obtain a feature map with a spatial scale half of the original E d1 then passes through the second convolutional module to obtain the feature map of the second level then passes through the downsampling layer to obtain the feature map E d2 then passes through the second convolutional module to obtain the feature map of the third level then passes through the downsampling layer to obtain the feature map E d3 then passes through the fourth convolutional module to obtain the feature map The above process is expressed by the formula:
[0118] E1 = f CB1 (O)
[0119] E2 = f CB2 (E d1 ) = f CB2 (↓(E1))
[0120] E3 = f CB3 (E d2 ) = f CB3 (↓(E2))
[0121] B = f CB4 (E d3 ) = f CB4 (↓(E3))
[0122] where f CBi , i = 1, 2, 3, 4 is the function representation of the i-th convolutional module, and ↓ represents a 2×2 max-pooling downsampling layer.
[0123] The feature maps E1, E2, and E3 of the first three levels are respectively passed through the SSA-Net of three levels for enhanced expression of spatial and spectral features to obtain the feature maps and Expressed by the formula as:
[0124] Q1 = f SSA-Net1 (E1)
[0125] Q2 = f SSA-Net2 (E2)
[0126] Q3 = f SSA-Net3 (E3)
[0127] where f SSA-Net1 , f SSA-Net2 and f SSA-Net3 are respectively the function representations of the SSA-Net of three levels.
[0128] In the decoder, the feature map B first passes through an upsampling layer to obtain the feature map Then use the Feature CCC method to connect R3 with the output feature map Q3 of the SSA-Net of the third level in the channel dimension to obtain O′3 then passes through the fifth convolutional module to obtain the feature map Then passes through an upsampling layer to obtain the feature map R2 is connected with the output feature map Q2 of the SSA-Net of the second level in the channel dimension to obtain O′2 then passes through the sixth convolutional module to obtain the feature map Then passes through an upsampling layer to obtain the feature map Finally, R1 is concatenated with the output feature map Q1 of the first-level SSA-Net in the channel dimension to obtain O′1 then passes through the seventh convolutional module to obtain the feature map The above process is expressed by the formula:
[0129] O′3 = FeaCCC(Q3, R3) = FeaCCC(Q3, ↑(B))
[0130] D1 = f CB5 (O′3)
[0131] O′2 = FeaCCC(Q2, R2) = FeaCCC(Q2, ↑(D1))
[0132] D2 = f CB6 (O′2)
[0133] O′1 = FeaCCC(Q1, R1) = FeaCCC(Q1, ↑(D2))
[0134] D3 = f CB7 (O′1)
[0135] where f CBi , i = 5, 6, 7 represents the function of the i-th convolutional module, ↑ represents the bilinear interpolation upsampling operation with a scale of 2, and FeaCCC() represents the operation process of Feature CCC.
[0136] Design a 1×1 convolutional layer for residual map reconstruction after the decoder, which is expressed by the formula:
[0137] X res = Conv 1×1 (D3)
[0138] where Conv 1×1 is a 1×1 convolutional layer, is the hyperspectral residual image obtained after reconstruction, and D3 is the output feature map of the third decoder.
[0139] S32: Design a channel cross-connection method.
[0140] To improve the fusion ability of different information sources, this application proposes a novel channel cross-connection method, simply referred to as CCC. According to different input information sources, the channel cross-connection method includes two types of channel cross-connection methods: Input CCC and Feature CCC; Input CCC refers to the channel cross-connection method between two input images, HSI and PAN, aiming to enhance the fusion ability of images from different input sources; Feature CCC refers to the channel cross-connection method between two feature maps, aiming to enhance the fusion ability between feature maps at different levels.
[0141] Please refer to Figure 3 , which is a schematic diagram of the Input CCC of the input source image channel cross-connection method in Embodiment 1 of this application; the input of Input CCC is two different source images, Up-HSI and PAN. The spatial resolution of Up-HSI is H×W pixels, and the number of spectral bands is C; the spatial resolution of PAN is H×W pixels, and the number of spectral bands is C p .
[0142] First, Up-HSI is divided into m parts along the channel dimension, then the number of spectral bands is:
[0143]
[0144] When m≥2, C1 = … = C m-1 ≥C m ; when m = 1, C1 = C, and at this time, Up-HSI is not divided;
[0145] The formula for the division process is:
[0146] U1, U2…, U m-1 , U m = Split(U)
[0147] where is the tensor corresponding to Up-HSI, Split is the operation of tensor division, are the sub-tensors obtained by dividing U;
[0148] After dividing to obtain sub-tensors U i with C i channels, a PAN tensor P with C p = 1 channel is inserted, and the output tensor O is obtained by connecting them in sequence along the channel dimension. Then the number of channels of the output tensor O is:
[0149]
[0150] The process of connecting in sequence along the channel dimension is expressed by the formula:
[0151] O = Concat(U1, P, U2, P, …, U m , P)
[0152] wherein, is the tensor corresponding to PAN, and Concat is the operation of channel concatenation, is the tensor output by channel cross - connection;
[0153] In summary, the operation process of cross - connection of Up - HSI and PAN in the channel dimension is expressed by the formula:
[0154] O = InputCCC(U, P)
[0155] where InputCCC is the operation process of Input CCC.
[0156] Please refer to Figure 4 , which is a schematic diagram of the Feature CCC method for cross - connection of feature map channels in Embodiment 1 of this application; the inputs of Feature CCC are two feature maps Feature1 and Feature2 at different levels. Feature1 and Feature2 are respectively evenly divided into n = 2 k (0 ≤ k ≤ 5) equal parts, that is:
[0157]
[0158]
[0159] where C qi is the number of spectral bands of the feature map Featurel, and C rj is the number of spectral bands of the feature map Feature1; when n = 1, that is, k = 0, the feature maps Feature1 and Feature2 do not perform the splitting operation;
[0160] The formula for the splitting process is expressed as:
[0161] Q1, Q2, …, Q n = Split(Q)
[0162] R1, R2, …, R n = Split(R)
[0163] wherein, is the tensor corresponding to Feature1, is the tensor corresponding to Feature2, Split is the operation of tensor splitting, are the sub - tensors obtained by splitting Q, are the sub - tensors obtained by splitting R;
[0164] The sub-tensor Q with qi the number of channels C i is obtained by splitting and then a sub-tensor R with rj the number of channels C j is inserted. The output tensor O' is obtained by successively connecting them in the channel dimension. Then the number of channels of the output tensor O' is:
[0165]
[0166] The process of successively connecting in the channel dimension is expressed by the formula:
[0167] O′ = Concat(Q1, R1, Q2, R2, …, Q n , R n )
[0168] where Concat is the operation of channel connection, is the tensor output by channel cross-connection;
[0169] In summary, the operation process of cross-connecting the feature maps Feature1 and Feature2 in the channel dimension is expressed by the formula:
[0170] O′ = FeaCCC(Q, R)
[0171] where FeaCCC represents the operation process of Feature CCC.
[0172] S33: Design the spatial-spectral attention network.
[0173] Please refer to Figure 5 , which is the structural schematic diagram of the spatial-spectral attention network SSA-Net in Embodiment 1 of this application; the spatial-spectral attention network is located between the encoder and the decoder of the UNet backbone network and is composed of N sequentially stacked Res-SSA modules based on the residual spatial-spectral attention mechanism, which is expressed by the formula:
[0174]
[0175] where F0 is the input feature map of SSA-Net, corresponding to the encoder output feature maps E1, E2, and E3 of the first three levels of CCC-SSA-UNet; F N is the output feature map of SSA-Net, corresponding to the decoder input feature maps Q1, Q2, and Q3 of the first three levels of CCC-SSA-UNet; is the function representation of the k-th Res-SSA module.
[0176] Please refer toFigure 6 , which is a schematic structural diagram of the Res-SSA module based on the residual space-spectral attention mechanism in Embodiment 1 of the present application;
[0177] The Res-SSA module embeds channel attention and spatial attention in parallel into the basic residual module, thereby improving the stability of network training and accelerating the convergence speed while enhancing the spatial-spectral feature representation. For the Nth Res-SSA module, its input is the feature map F N-1 , F N-1 First, it passes through a 3×3 convolutional layer to change the number of channels to 64, and then through a ReLU activation function layer and a 3×3 convolutional layer to further extract the feature map F U as the input of the attention module. F U is divided into two paths. One path passes through the spectral attention module to obtain the spectral mask M CA and then multiplies it pixel by pixel with the feature map F U to get F CA . The other path passes through the spatial attention module to obtain the spatial mask M SA and then multiplies it pixel by pixel with the feature map F U to get F SA . Finally, F CA is added pixel by pixel with F SA and then added with the input F N-1 to obtain the output F N of the Res-SSA module. It can be described by the formula:
[0178] F U = Conv 3×3 (δ′(Conv 3×3 (F N-1 )))
[0179]
[0180]
[0181] F N = F CA + F SA + F N-1
[0182] where δ′() is the ReLU activation function layer, Conv 3×3 () is the 3×3 convolutional layer, M CA and M SA are the spectral mask and the spatial mask respectively, is the pixel-by-pixel multiplication operation.
[0183] Specifically, the backbone of the spectral attention module consists of a global average pooling layer along the spatial dimension, a 1×1 convolutional layer that reduces the number of channels from 64 to 64 / r, a ReLU activation function layer, a 1×1 convolutional layer that increases the number of channels from 64 / r to 64, and a sigmoid activation function layer in sequence. r is called the channel shrinkage rate and is used to reduce the computational complexity of the network model. The backbone of the spatial attention module consists of a globally average pooling layer and a global max pooling layer in parallel, a 1×1 convolutional layer, and a sigmoid activation function layer in sequence. It can be described by the formula:
[0184] M CA = σ(Conv 1×1 (δ′(Conv 1×1 (GAP(F U )))))
[0185] M SA = σ(Conv 1×1 (Concat(GAP(F U ), GMP(F U ))))
[0186] Wherein, σ is the sigmoid activation function, Conv 1×1 () is the 1×1 convolutional layer, GAP() is the global average pooling operation, GMP() is the global max pooling operation, and Concat() is the operation of channel concatenation.
[0187] S34: Design the loss function.
[0188] The present invention uses loss to evaluate the fusion effect of the network because loss can effectively penalize small errors and make the training have better convergence. The loss is the mean absolute error (MAE) between all the reconstructed images and the reference hyperspectral image in a training batch, and the calculation formula is:
[0189]
[0190] Wherein, and Y d are the d-th reconstructed HR-HSI and the reference HSI respectively, D is the number of images, Θ is the parameter of the hyperspectral pan-sharpening neural network model, Φ is the hyperspectral pan-sharpening neural network model, and ‖·‖1 is the l1 norm.
[0191] S4: Train the hyperspectral pan-sharpening neural network model.
[0192] Please refer to Figure 7, which is a schematic diagram of the training stage of the hyperspectral panchromatic sharpening neural network model in Embodiment 1 of the present application.
[0193] In this embodiment, let denote the LR-HSI, whose spatial resolution is h×w pixels and whose number of spectral bands is C; let denote the PAN, whose spatial resolution is H×W pixels and whose number of spectral bands is 1; let denote the hyperspectral image (HR-HSI) obtained after fusion and reconstruction, whose spatial resolution is H×W pixels and whose number of spectral bands is C; let denote the reference true high-spatial-resolution hyperspectral image (Ref-HR-HSI), whose spatial resolution is H×W pixels and whose number of spectral bands is C. And H>h, W>w, C>>1 holds.
[0194] The training process of the hyperspectral panchromatic sharpening neural network is as follows: Input a training set {[X1, P1, Y1], …, [X D , P D , Y D} with D image pairs, and process it through the hyperspectral panchromatic sharpening neural network model Φ to obtain an output image set Optimize the parameters Θ of the hyperspectral panchromatic sharpening neural network model through an optimization algorithm.
[0195] The training process of the hyperspectral panchromatic sharpening neural network model is expressed by the formula:
[0196]
[0197] where are the optimized parameters, and Loss(a, b) refers to the loss function adopted by the network to evaluate the gap between a and b.
[0198] S5: Input the test images of the test set into the trained hyperspectral panchromatic sharpening neural network model to obtain a high-spatial-resolution hyperspectral image.
[0199] Please refer to Figure 8 , which is a schematic diagram of the test stage of the hyperspectral panchromatic sharpening neural network model in Embodiment 1 of the present application.
[0200] In this embodiment, in the test stage, input the test image pair [X t , P t , and process it through the trained hyperspectral panchromatic sharpening neural network model to output a fused image The test process is expressed by the formula:
[0201]
[0202] Where X t and P t are the input LR-HSI and PAN respectively, and HR-HSI is the high-spatial-resolution hyperspectral image output by the hyperspectral pan-sharpening neural network model.
[0203] In summary, in Embodiment 1 of the present application, a U-shaped convolutional neural network CCC-SSA-UNet for hyperspectral pan-sharpening is built, which receives a low-resolution hyperspectral image and a panchromatic image as initial inputs. The LR-HSI is upsampled using bilinear interpolation to obtain the image Up-HSI. Then, after cross-connecting Up-HSI and PAN in the channel direction, the result is input into an encoder-decoder network with a spatial-spectral residual attention mechanism to learn the residual image. Finally, the residual image is added to Up-HSI pixel by pixel to obtain the final fusion result. By fusing a hyperspectral image with a low spatial resolution and a panchromatic image with a high spatial resolution, the spatial resolution of the hyperspectral image with a low spatial resolution is improved, while more spatial and spectral information is retained, and it has a low computational complexity and a short inference running time.
[0204] Embodiment 2
[0205] Please refer to Figure 9 , which is a schematic structural diagram of a hyperspectral pan-sharpening system based on a U-shaped convolutional neural network according to Embodiment 2 of the present application. The specific content includes:
[0206] Acquisition module: Acquire a hyperspectral image with a low spatial resolution;
[0207] Input module: Use the hyperspectral image as an input to generate a training set and a test set;
[0208] Design module: Design a hyperspectral pan-sharpening neural network model;
[0209] Training module: Train the hyperspectral pan-sharpening neural network model;
[0210] Output module: Input the test image of the test set into the trained hyperspectral pan-sharpening neural network model to obtain a hyperspectral map with a high spatial resolution.
[0211] In this embodiment, the design module integrates the U-Net based on the encoder-decoder architecture with the Spatial-Spectral Attention Network (SSA-Net) in the network model design, further enhancing the ability to extract spatial and spectral features. By designing a novel channel cross-connection method, Input CCC and Feature CCC, the fusion ability of different input source images and the fusion ability between feature maps of different levels can be effectively enhanced without significantly increasing the computational complexity. In the training module, the training images are input and processed by the neural network model to obtain the output image, i.e., the fusion result. The trainable parameters in the neural network are continuously optimized and adjusted through the optimization algorithm, so that the loss function between the output image and the reference ground truth image gradually decreases until it converges to a certain value. In the output module, the test images are input and processed by the neural network model with the trained parameters to obtain the finally output fused image.
[0212] In summary, in this embodiment 2, the hyperspectral image is obtained through the acquisition module. After cross-connecting Up-HSI and PAN in the channel direction, it is input into the encoder-decoder network with a spatial-spectral residual attention mechanism to learn the residual image. Finally, the residual image is added to Up-HSI pixel by pixel to obtain the final fusion result. By combining channel attention and spatial attention and embedding them into the basic residual module, the spatial-spectral feature expression ability of the network can be improved while enhancing the stability of network training and accelerating the convergence speed of the network model.
[0213] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, apparatus, article or method including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, apparatus, article or method including that element.
[0214] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
[0215] Although the embodiments of the present application have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.
[0216] Of course, the present invention may have many other embodiments. Based on this embodiment, all other embodiments obtained by those of ordinary skill in the art without any creative work fall within the scope of protection of the present invention.
Claims
1. A hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network, characterized in that Including: Obtain a hyperspectral image with low spatial resolution; Use the hyperspectral image as input to generate a training set and a test set; Design a hyperspectral pan-sharpening neural network model; Train the hyperspectral pan-sharpening neural network model; Input the test images of the test set into the trained hyperspectral pan-sharpening neural network model to obtain a hyperspectral image with high spatial resolution; In the step of designing the hyperspectral pan-sharpening neural network model, it specifically includes the following steps: Design a UNet backbone network; Design a channel cross-connection method; Design a spatial-spectral attention network; Design a loss function; The channel cross-connection method includes two channel cross-connection methods: Input CCC and Feature CCC; The input of the Input CCC is two different source images Up-HSI and PAN. If Up-HSI is divided into m parts along the channel dimension, the number of spectral bands is: When m ≥ 2, C1 = … = C m-1 ≥ C m ; when m = 1, C1 = C, and at this time, Up-HSI is not segmented; The formula for the segmentation process is expressed as: U1, U2, …, U m-1 , U m = Split(U) Among them, is the tensor corresponding to Up-HSI, and Split is the operation of tensor splitting, are the sub-tensors obtained by splitting U; Split to obtain a sub-tensor U with C i number of channels i Then insert a PAN tensor P with C p = 1 number of channels, and connect them in sequence in the channel dimension to obtain an output tensor O. Then the number of channels of the output tensor O is: The process of sequentially connecting in the channel dimension is expressed by the formula: O = Concat(U1, P, U2, P, …, U m , P) Among them, is the tensor corresponding to PAN, and Concat is the operation of channel concatenation. is the tensor output by channel cross-connection; The operation process of cross-connecting Up-HSI and PAN in the channel dimension is expressed by the formula: O = InputCCC(U, P) where InputCCC is the operation process of Input CCC; The input of the Feature CCC is two feature maps Feature1 and Feature2 at different levels. Feature1 and Feature2 are respectively evenly divided into n = 2 k , where 0 ≤ k ≤ 5 equal parts, that is: Among them, C q is the number of spectral bands of the feature map Feature1, and C r is the number of spectral bands of the feature map Feature2; when n = 1, that is, k = 0, no splitting operation is performed on the feature map Feature1 and the feature map Feature2; The formula for the segmentation process is expressed as: Q1, Q2, …, Q n = Split(Q) R1, R2, …, R n = Split(R) Among them, is the tensor corresponding to Feature1, is the tensor corresponding to Feature2, and Split is the operation of tensor splitting. are the sub-tensors obtained by splitting Q, are the sub-tensors obtained by splitting R, where H and W are the pixels of the spatial resolution, and C is the number of spectral bands; The sub-tensor Q with C qi number of channels i After that, insert the sub-tensor R with C rj number of channels j , and connect them in sequence along the channel dimension to obtain the output tensor O'. Then the number of channels of the output tensor O' is: The process of sequentially connecting in the channel dimension is expressed by the formula: O′ = Concat(Q1, R1, Q2, R2, …, Q n , R n ) Among them, Concat is an operation for channel connection, which is the tensor output by channel cross-connection; The operation process of cross-connecting the feature map Feature1 and the feature map Feature2 in the channel dimension is expressed by the formula: O′ = FeaCCC(Q, R) where FeaCCC represents the operation process of Feature CCC.
2. The hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to claim 1, wherein In the step of using the hyperspectral image as input to generate a training set and a test set, it specifically includes the following steps: Select the upper left A×B pixel area of the hyperspectral image and cut it into n non-overlapping H×W pixel area image sub-blocks to form the reference HR-HSI image of the data set; Blur the HR-HSI image with a Gaussian filter with a kernel size of 8×8 and a standard deviation of σ, and perform 4-fold downsampling to obtain an LR-HSI image with a spatial resolution of h×w. Obtain the PAN image by averaging the panchromatic band of the HR-HSI image in the spectral dimension; Randomly select 75% of the image pairs as the training set, and the remaining 25% as the test set; The standard deviation σ of the Gaussian filter is calculated by the following formula: where β is the downsampling scale factor, that is, the linear spatial resolution ratio of the reference HR-HSI image to the generated LR-HSI.
3. A hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to claim 1, characterized in that The UNet backbone network consists of an encoder part, a decoder part, and a bottleneck layer, including four scales. The encoders and decoders of the first three scales are connected by a spatial-spectral attention network, and the fourth scale consists of a bottleneck layer; The encoder part consists of 3 encoders, each encoder includes 1 convolutional module and 1 downsampling module; the decoder part consists of 3 decoders, each decoder includes 1 upsampling module and 1 convolutional module, and the bottleneck layer includes 1 convolutional module; The formula for the convolutional module is: CB out = f CB (CB in ) = δ(BN(Conv 3×3 (CB in ))) Among them, CB in and CB out are the input and output of the convolutional module respectively, f CB is the functional representation of the convolutional module, δ is the LeakyReLU activation function layer, BN is the batch normalization layer, and Conv 3×3 is a 3×3 convolutional layer; Design a 1×1 convolutional layer for residual map reconstruction after the decoder part, which is expressed by the formula: X res = Conv 1×1 (D3) Among them, Conv 1×1 is a 1×1 convolutional layer, is the hyperspectral residual image obtained after reconstruction, H and W are the pixels of the spatial resolution, C is the number of spectral bands; D3 is the output feature map of the third decoder.
4. A hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to claim 1, characterized in that The spatial-spectral attention network is located between the encoder and the decoder of the UNet backbone network and is composed of N Res-SSA modules based on the residual spatial-spectral attention mechanism stacked in sequence, which is expressed by the formula: Among them, F0 is the input feature map of SSA-Net, and F N is the output feature map of SSA-Net, is the function representation of the k-th Res-SSA module.
5. A hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to claim 1, characterized in that, The loss function uses to represent the mean absolute error between the reconstructed image and the reference hyperspectral image, and the calculation formula is: wherein, and Y d are the d-th reconstructed HR-HSI and the reference HSI respectively, D is the number of images, Θ is the parameter of the hyperspectral pan-sharpening neural network model, Φ is the hyperspectral pan-sharpening neural network model, and ‖·‖1 is the l1 norm.
6. The hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to claim 1, wherein In the step of training the hyperspectral pan-sharpening neural network model, it specifically includes the following steps: Input a training set of D image pairs {[X1, P1, Y1], …, [X D , P D , Y D}, and process it through the hyperspectral panchromatic sharpening neural network model Φ to obtain an output image set Optimize the parameters Θ of the hyperspectral pan-sharpening neural network model through an optimization algorithm. The training process of the hyperspectral pan-sharpening neural network model is expressed by the formula: Among them, is the optimized parameter, and Loss(a, b) refers to the loss function adopted by the network, which evaluates the gap between a and b.
7. A hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to claim 1, characterized in that, In the step of inputting the test image of the test set into the trained hyperspectral pan-sharpening neural network model to obtain a high-spatial-resolution hyperspectral image, it specifically includes the following steps: Input the test image pair [X t , P t , and process it through the trained hyperspectral-panchromatic sharpening neural network model to output the fused image The test process is expressed by the formula as follows: Where X t and P t are the input LR-HSI and PAN respectively, is the high-spatial-resolution hyperspectral image HR-HSI output by the hyperspectral panchromatic sharpening neural network model.
8. A system for a hyperspectral panchromatic sharpening method based on a U-shaped convolutional neural network according to any one of claims 1-7, characterized in that Include: Acquisition module: Acquire a low-spatial-resolution hyperspectral image; Input module: Take the hyperspectral image as input and generate a training set and a test set; Design module: Design a hyperspectral pan-sharpening neural network model; Training module: Train the hyperspectral pan-sharpening neural network model; Output module: Input the test image of the test set into the trained hyperspectral pan-sharpening neural network model to obtain a high-spatial-resolution hyperspectral image.
Citation Information
Patent Citations
A Hyperspectral Image Panchromatic Sharpening Method Based on Spectral Dimension Controlled Convolutional Neural Network
CN109859110B
Hyperspectral image panchromatic sharpening method, device and medium
CN114638761A
Systems and Methods for Multi-Spectral Image Super-Resolution
US20190287216A1