Remote Sensing Image Fusion Method Based on Compressed Panchromatic Light Images
By using a deep learning network to remove the compressed block effect and using a network model of multi-scale hollow residual block group to perform remote sensing image fusion, the problems of spectral distortion, spatial blur and high computing costs in the prior art are solved, and high-quality and low-cost remote sensing image fusion effect is achieved.
Patent Information
- Application Number
- CN202110466705.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-04-28
AI Technical Summary
The existing remote sensing image fusion methods have problems of spectral distortion, spatial blur and high computational costs, making it difficult to achieve high-quality and low-cost fusion effects.
The remote sensing image fusion is carried out by lossy compressed full-color light images, the compressed block effect is removed through the deep learning network, and the network model of multi-scale hollow residual block groups is used for fusion, to construct a remote sensing image fusion method based on compressed full-color light images.
While ensuring the fusion quality, the amount of fusion data is significantly reduced, low-cost remote sensing image fusion is achieved, and the compression block effect is removed, and the image clarity and detail recovery effect is improved.
Smart Images

Figure CN115249222B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image fusion technology, and particularly to a remote sensing image fusion method based on compressed panchromatic light images, belonging to the field of image processing. Background Art
[0002] Image fusion is a research hotspot in the field of information fusion. Its main goal is to use different fusion technologies to realize the recombination of multi-source image information. Essentially, it is an image enhancement technology. By merging image data from different sensors or different parameter settings, and using their correlation and complementarity to reconstruct a new image with rich information and strong robustness. Image fusion technology is widely used in fields such as remote sensing mapping, military target detection, and medical image analysis. Remote sensing image fusion technology fuses the image information of panchromatic light images and multi-spectral images, improves the spatial resolution of multi-spectral images while retaining spectral information, and provides rich image information for subsequent analysis processes.
[0003] Remote sensing image fusion methods mainly include three categories: fusion methods based on component substitution, fusion methods based on multi-resolution analysis, and fusion methods based on learning. The image fusion algorithm based on component substitution is simple and effective, and performs component substitution through image domain transformation; the fusion method based on multi-resolution analysis extracts details from the panchromatic light image and injects them into the upsampled multi-spectral image; the fusion method based on learning trains a convolutional neural network model to achieve remote sensing image fusion. The fusion method based on component substitution usually has serious spectral distortion phenomena, the fusion method based on multi-resolution analysis causes spatial blurring in the fused result image due to spatial information redundancy in the decomposition process, and the method based on learning requires a large amount of training data. It is very challenging to achieve high-quality and low-cost remote sensing image fusion. Summary of the Invention
[0004] The purpose of the present invention is to use lossily compressed panchromatic light images for remote sensing image fusion, so as to greatly reduce the amount of fused data and achieve low-cost remote sensing image fusion while ensuring the fusion quality. The lossy compression format of the present invention is JPEG compression. The input compressed panchromatic light image eliminates the compression block effect through a deep learning network that removes the compression effect, and then realizes remote sensing image fusion through a deep learning-based fusion network, thereby constructing a remote sensing image fusion method based on compressed panchromatic light images.
[0005] The remote sensing image fusion method based on compressed panchromatic light images proposed by the present invention mainly includes the following operating steps:
[0006] (1) In the discrete cosine transform (DCT) domain, a network model for removing compression block artifacts in joint dual-domain learning is constructed with a local wide activation residual block group as the main building unit, and in the pixel domain, a residual module with high-pass filtering skip connections is used as the main building unit.
[0007] (2) Using the network in step (1), a deblocking network model with a compression quality factor of 60 is trained.
[0008] (3) For the input compressed panchromatic image, image restoration is performed through the trained network, and a panchromatic image with compression block artifacts removed is output.
[0009] (4) A network model for remote sensing image fusion is constructed with a multi-scale dilated residual block group as the main building unit.
[0010] (5) Using the training image dataset, the network constructed in step (4) is trained.
[0011] (6) Using the panchromatic image and multi-spectral image after removing compression block artifacts as network inputs, the final fusion result is output. Description of the Drawings
[0012] Figure 1 is the principle block diagram of the remote sensing image fusion method based on compressed panchromatic images of the present invention.
[0013] Figure 2 is the network structure diagram of the deblocking network for joint dual-domain learning of the present invention.
[0014] Figure 3 is the network structure diagram of the local wide activation residual block group of the present invention.
[0015] Figure 4 is the network structure diagram of the remote sensing image fusion network based on multi-scale dilated convolutional residual blocks of the present invention.
[0016] Figure 5 is a comparison chart of the fusion results of the present invention and five methods for the test image "Stadium" (quality factor is 60, downsampling factor is 4, Gaussian blur kernel size is 5×5, standard deviation is 1.5): Among them, (a) is the reference image, (b) is the panchromatic image, (c) is the multi-spectral image bicubic interpolation magnified four times, and (d)(e)(f)(g)(h)(i) are the reconstruction results of method 1, method 2, method 3, method 4, method 5 and the present invention respectively.
[0017] Figure 6This is a comparison chart of the fusion results of the present invention and five methods for the test image "Factory" (quality factor is 60, downsampling factor is 4, Gaussian blur kernel size is 5×5, standard deviation is 1.5): Among them, (a) is the reference image, (b) is the panchromatic light image, (c) is the multi-spectral image bicubic interpolation magnified four times, and (d)(e)(f)(g)(h)(i) are the reconstruction results of Method 1, Method 2, Method 3, Method 4, Method 5 and the present invention respectively. Detailed implementation manner
[0018] The present invention will be further described below with reference to the accompanying drawings:
[0019] Figure 1 Among them, the remote sensing image fusion method based on the compressed panchromatic light image can be specifically divided into the following steps:
[0020] (1) In the discrete cosine transform (DCT) domain, use the local wide activation residual block group as the main construction unit, and in the pixel domain, use the high-pass filter skip connection residual module as the main construction unit to build a network model for removing the compression image block effect by joint dual-domain learning;
[0021] (2) Use the network in step (1) to train the deblocking effect network model with a compression quality factor of 60;
[0022] (3) For the input compressed panchromatic light image, perform image restoration through the trained network, and output the panchromatic light image with the compression block effect removed;
[0023] (4) Use the multi-scale dilated residual block group as the main construction unit to build a network model for remote sensing image fusion;
[0024] (5) Use the training image data set to train the network constructed in step (4);
[0025] (6) Use the panchromatic light image and the multi-spectral image with the compression block effect removed as the network input, and output the final fusion result.
[0026] Specifically, in the step (1), the network model for removing the compression image block effect by joint dual-domain learning mainly includes: the local wide activation residual block group DCT domain decompression network and the high-pass filter skip connection residual block group pixel domain decompression network.
[0027] The local wide activation residual block group DCT domain decompression network mainly includes an image stretching layer, a DCT convolutional layer, a local wide activation residual block group, an IDCT convolutional layer and an image recombination layer; the high-pass filter skip connection residual block group pixel domain decompression network mainly includes a conventional residual block group and a high-pass filter skip connection. Figure 2It shows that after deblocking the compressed panchromatic light image in the DCT domain and the pixel domain respectively and then performing stacking and cascading, it is input into a 1×1 convolutional layer without bias for adaptive processing of dual-domain features to obtain the final panchromatic light image after deblocking.
[0028] In the local wide activation residual block group DCT domain de-compression network, both the image stretching and recombination layers are composed of filters with 64 channels and a size of 8×8 to achieve the extraction and recombination of image blocks. The image block pixels are stretched into a 64-dimensional image tensor through the stretching layer. Both the DCT convolutional layer and the IDCT convolutional layer are composed of filters with 64 channels and a size of 1×1. The DCT layer is used to perform a convolutional operation on the output tensor of the image stretching layer. After the weights of the DCT layer and the IDCT layer are initialized, they are fixed and no longer updated with the gradient during network training, so as to implement the DCT transformation and IDCT transformation of the image on the convolutional layer. In the DCT domain sub-network, except for the image stretching and recombination layer, the DCT layer and the IDCT layer, the sizes of the remaining convolutional kernels are all set to 3×3. Figure 3 It shows the local wide activation residual block group, which is mainly composed of four layers of wide activation convolutional layers, four layers of excitation layers and skip connections. The role of the skip connection is to accelerate the convergence speed of the network while ensuring the flow of information. The four layers of wide activation convolutional layers double the number of feature channels from 64 to 256 for the first and third layer convolutions respectively, and restore the number of feature channels from 256 to 64 for the second and fourth layer convolutions, so as to improve the learning performance of the network without increasing the computational complexity. The excitation layer uses the ReLU activation function.
[0029] In the high-pass filtering skip connection residual block group pixel domain de-compression network, after the image is encoded by a 64-channel filter with a size of 3×3, it passes through two conventional residual blocks. Each residual block is composed of two 64-channel filters with a size of 3×3 and a skip connection. Finally, it is decoded by a 64-channel filter with a size of 3×3. The overall input and output ends are connected by a long skip connection through a high-pass filtering information transfer. The high-pass filtering information is obtained by subtracting the low-pass information from the original image information, and the low-pass information is generated by mean filtering. The excitation layer also uses the ReLU activation function.
[0030] The loss function adopted by the network model for removing the compression block effect of joint dual-domain learning is expressed as:
[0031] L = L mse + ηL quantization
[0032] L mse is the mean square error loss function and can be expressed as:
[0033] L mse = ||Y - X||2
[0034] Where Y is the original uncompressed reference image, and X is the predicted image patch after removing the compression effect.
[0035] L quantization is the quantization coefficient loss function, η is the weight value set to 0.5, and L quantization can be expressed as:
[0036]
[0037] Where, and represent the DCT coefficients of the reference original image patch and the image patch after removing the compression effect, respectively. The subscripts k and l represent the row and column where the pixel is located in the image patch, and the range is [1, 8]. M is the quantization matrix, and round represents rounding. This loss function fully considers the accuracy of DCT coefficient learning in the network, which is beneficial to training a model with better performance.
[0038] In the step (4), the network model for remote sensing image fusion built mainly consists of a multi-scale dilated convolutional residual block group.
[0039] As Figure 4 shows, the multi-scale dilated convolutional residual block group consists of three dilated convolutional residual blocks. Each block combines three 3×3 dilated convolutional layers with dilation factors d = 1, d = 4, and d = 8 respectively to form a dilated convolutional group. Then, a 1×1 convolutional layer is used to fuse the features extracted by the three dilated convolutional layers. Finally, the three dilated convolutional groups are cascaded in a residual manner to form a dilated convolutional residual block.
[0040] The loss function adopted by the network model for remote sensing image fusion based on the multi-scale dilated convolutional residual block is expressed as:
[0041] L mse = ||Y fusion - X ms ||2
[0042] Where Y fusion is the fused remote sensing image, and X ms is the reference multi-spectral image.
[0043] To verify the effectiveness of the present invention, a large number of comparative experiments were carried out on the commonly used remote sensing satellites Pléiades (including 7480 training images and 10 test images) and QuickBird (including 3424 training images and 10 test images). In the experiment, the present invention was compared with 5 typical remote sensing image fusion methods, which include traditional component substitution fusion methods and convolutional neural network-based fusion algorithms. These 5 compression image deblocking effect methods for comparison are:
[0044] Method 1: The method proposed by Aiazzi et al., reference "B. Aiazzi, S. Baronti, and M. Selva, “Improving component substitution pansharpening through multivariate regression of MS+Pan data,” IEEE Trans. Geoscience and Remote Sensing. 45(10), 3230-3239(2007).”
[0045] Method 2: The method proposed by Song et al., reference "Y. Song, W. Wu, and Z. Liu, “An adaptive pansharpening method by using weighted least squares filter,” IEEE Geoscience and Remote Sensing Letters. 13(1), 18-22(2015).”
[0046] Method 3: The method proposed by Yang et al., reference "J. Yang, X. Fu, Y. Hu, Y. Huang, X. Ding, and J. Paisley, “PanNet: A deep network architecture for pan-sharpening,” in Proc. IEEE Int. Conf. Comput. Vision. 1753-1761(2017).”
[0047] Method 4: The method proposed by Scarpa et al., reference "G. Scarpa, S. Vitale, and D. Cozzolino, “Target-adaptive CNN-based pansharpening,” IEEE Trans. Geoscience and Remote Sensing. 56(9), 5443-5457(2018).”
[0048] Method 5: The method proposed by Vitale et al., reference "S. Vitale, and G. Scarpa, “A detail-preserving cross-scale learning strategy for CNN-based pan sharpening,” Remote Sensing. 12(3), 348-367(2020).”
[0049] The content of the comparative experiment is as follows:
[0050] Experiment 1: The multispectral images and panchromatic images simulated from 10 test images of the remote sensing satellite Pléiades were fused and reconstructed using Method 1, Method 2, Method 3, Method 4, Method 5, and the method of the present invention, respectively. In this experiment, the blur kernel was taken as a Gaussian blur kernel with a size of 5×5 and a standard deviation of 1.5. Table 1 shows the CC (Correlation Coefficient), ERGAS (Relative Dimensionless Global Error in Synthesis), UIQI (Universal Image Quality Indexes), SAM (Spectral Angle Mapper), RASE (Relative Average Spectral Error), and RMSE (Root Mean Squared Error) parameters of the reconstruction results of each method. Among them, the best values of the parameters CC and UIQI are 1, and the best values of other indicators are 0. In addition, for visual comparison, the results of the "Stadium" image are given. The original image, panchromatic image, bicubic interpolation quadruple magnification of the multispectral image, and the reconstruction results of each method of the "Stadium" are respectively as Figure 5 (a), Figure 5 (b), Figure 5 (c), Figure 5 (d), Figure 5 (e), Figure 5 (f), Figure 5 (g), Figure 5 (h), and Figure 5 (i) shown.
[0051] Table 1
[0052] Evaluation Index Method 1 Method 2 Method 3 Method 4 Method 5 The present invention CC 0.9518 0.9600 0.9866 0.9875 0.9908 0.9931 ERGAS 3.6379 3.3131 1.8512 1.7762 1.6397 1.4228 UIQI 0.9478 0.9578 0.9864 0.9874 0.9906 0.9919 SAM 3.4039 3.4055 3.1430 3.0940 2.6320 2.4491 RASE 14.7207 13.4510 7.9138 7.6334 6.6027 6.1092 RMSE 103.2254 94.3220 55.4938 53.5277 46.2997 42.8394
[0053] Experiment 2: The multispectral images and panchromatic images simulated from 10 test images of the remote sensing satellite QuickBird were fused and reconstructed using Method 1, Method 2, Method 3, Method 4, Method 5 and the method of the present invention, respectively. In this experiment, the blur kernel was taken as a Gaussian blur kernel with a size of 5×5 and a standard deviation of 1.5. Table 1 shows the CC (Correlation Coefficient), ERGAS (Relative Dimensionless Global Error in Synthesis), UIQI (Universal Image Quality Indexes), SAM (Spectral Angle Mapper), RASE (Relative Average Spectral Error) and RMSE (Root Mean Squared Error) parameters of the reconstruction results of each method. Among them, the best values of the parameters CC and UIQI are 1, and the best values of other indicators are 0. In addition, for visual comparison, the results of the "Factory" image are given. The original image, panchromatic image, bicubic interpolation quadruple magnification of the multispectral image and the reconstruction results of each method of the "Factory" are shown respectively as Figure 6 (a), Figure 6 (b), Figure 6 (c), Figure 6 (d), Figure 6 (e), Figure 6 (f), Figure 6 (g), Figure 6 (h) and Figure 6 (i).
[0054] Table 2
[0055] Evaluation Index Method 1 Method 2 Method 3 Method 4 Method 5 The present invention CC 0.9338 0.9310 0.9594 0.9640 0.9774 0.9818 ERGAS 2.0644 1.4480 1.1728 1.2910 1.0740 0.7731 UIQI 0.9295 0.9288 0.9593 0.9623 0.9774 0.9817 SAM 2.7336 1.6593 1.5774 1.3866 1.1287 0.9417 RASE 7.7396 5.4480 4.9090 4.3340 3.6018 2.8648 RMSE 19.1517 13.4811 12.1473 10.7246 8.9128 7.0890
[0056] From Figure 5 and Figure 6It can be seen from the shown experimental results that there are obvious blurs in the results of the traditional fusion methods 1 and 2, and the visual effect of the images is poor; for the fusion methods 3, 4, and 5 based on convolutional neural networks, all three comparison methods can achieve a certain resolution improvement. The details of method 5 are restored relatively fully, but some compression artifacts are retained during image restoration, and obvious compression block effects can be seen in the magnified image details. There is no obvious compression effect in the results of the present invention, and the image is relatively clear, the details are restored to a certain extent, the edges are better preserved, and the visual effect is better. In addition, from the various parameters given in Table 1 and Table 2, compared with other conventional fusion methods, the fusion image of the present invention has been processed to remove compression block effects, so the highest values are obtained in all indicators and the improvement is obvious. Therefore, by comprehensively comparing the subjective visual effects and objective parameters of the fusion results of each method, it can be seen that the fusion effect of the method of the present invention is better. In summary, the present invention is an effective remote sensing image fusion method for compressed panchromatic light images.
Claims
1. A remote sensing image fusion method based on compressed panchromatic light images, characterized in that It includes the following steps: Step 1: Construct a network model for removing the compression image block effect with joint dual-domain learning, taking the local wide activation residual block group as the construction unit in the discrete cosine transform (DCT) domain and the high-pass filtering skip connection residual module as the construction unit in the pixel domain; Step 2: Use the network in Step 1 to train the image deblocking model with a compression quality factor of 60; Step 3: For the input compressed panchromatic image, perform image restoration through the trained network for removing the compression block effect, and output the panchromatic image with the compression block effect removed; Step 4: Construct a network model for remote sensing image fusion with the multi-scale dilated residual block group as the construction unit; Step 5: Use the training image dataset to train the network constructed in Step 4; Step Six: Use the panchromatic image and the multispectral image after removing the compression block effect as the network input, and output the final fusion result; it is characterized in that the high-pass filtering skip connection residual block group in the joint dual-domain learning compression image block effect removal network described in Step One includes, in the pixel-domain decompression network, after the image is encoded by a 64-channel filter with a size of 3×3, it passes through two conventional residual blocks. The residual block is composed of two 64-channel filters with a size of 3×3 and a skip connection. Finally, it is decoded by a 64-channel filter with a size of 3×3. The overall input-output end is connected by a high-pass filtering information transfer for long skip connection. The high-pass filtering information is obtained by subtracting the low-pass information from the original image information. The low-pass information is generated by mean filtering. The skip connection of the high-pass filtering information provides the transfer of high-frequency detail information in the pixel domain, accelerating the convergence speed of the network, so as to reduce the training difficulty while better learning features; it is characterized in that the quantization coefficient loss function L in the joint dual-domain learning compression image block effect removal network described in Step One quantization : wherein, and represent the DCT coefficients of the reference original image block and the image block with compression artifacts removed, respectively; the subscripts k and l represent the row and column where the pixel is located in the image block, and the range is [1, 8]; M is the quantization matrix, and M k.l is the quantization value of the k-th row and l-th column in the image block, and round represents rounding.
Citation Information
Patent Citations
JPEG compressed image decompression effect method combining DCT domain and pixel domain learning
CN112188217A