Panchromatic image and multispectral image joint compression method based on deep learning
Through a deep learning-based method, the remote sensing image is jointly compressed, which solves the problem of independent processing of MS and PAN images in the prior art, and realizes efficient and flexible image compression and maintains image quality.
Patent Information
- Application Number
- CN202510009165.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
The existing remote sensing image compression technology has the problem of independently processing MS and PAN images, and fails to fully utilize the complementary information between the two, resulting in low compression efficiency and difficult to maintain the reconstruction quality of the image, especially at high compression rates.
The combined compression method of full-color images and multi-spectral images based on deep learning is adopted, and efficient joint compression of MS and PAN images is achieved through preprocessing, joint feature extraction, multi-level feature fusion, joint probability modeling and entropy coding.
It realizes efficient joint compression of MS and PAN images, significantly improves compression efficiency, maintains image reconstruction quality, and has flexible compression rate control and robustness.
Smart Images

Figure CN120075475A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image compression, and particularly to a method for jointly compressing panchromatic images and multispectral images based on deep learning. Background Art
[0002] With the rapid development of remote sensing technology, new-generation remote sensing satellites can simultaneously acquire high-resolution panchromatic images and multispectral images. Panchromatic images have high spatial resolution but only contain single-band information, while multispectral images, although having lower spatial resolution, contain rich spectral information. These two complementary image data provide important support for applications such as ground object recognition and change detection. However, the massive remote sensing data also brings huge pressure to storage and transmission. Currently, the mainstream remote sensing image compression technologies have the following problems and deficiencies:
[0003] First, most of the existing compression methods independently compress MS images and PAN images respectively. For example, traditional compression algorithms such as TEPG2000 and BPG are used to encode the two types of images separately. This independent processing method ignores the inherent correlation between MS images and PAN images and cannot make full use of the complementary information of the two types of images, resulting in low compression efficiency. Although there are also studies attempting to perform image fusion before compression, the fusion process will introduce additional computational overhead and information loss.
[0004] Second, traditional compression methods based on transform coding, such as wavelet transform and DTC transform, often have problems such as image detail loss, edge blurring, and texture distortion when processing high-resolution remote sensing images. These methods usually use fixed transform basis functions and are difficult to adapt to extract and preserve the key features of images. At the same time, simple quantization and entropy coding strategies also limit the further improvement of compression performance. Although deep learning methods have made significant progress in the field of image compression, most of the existing end-to-end compression networks are designed for natural images and do not fully consider the characteristics of remote sensing images.
[0005] Third, remote sensing images face the trade-off problem between compression ratio and reconstruction quality. At high compression ratios, existing methods are difficult to maintain the geometric structure and spectral characteristics of images, which will affect subsequent analysis and processing tasks. Some improved methods improve the reconstruction quality by introducing perceptual quality loss or structural similarity constraints, but often reduce the compression efficiency or increase the computational complexity. In addition, different scenarios have different requirements for compression ratio and reconstruction quality, and existing methods lack a flexible adjustment mechanism.
[0006] Fourth, in practical applications, the acquisition process of remote sensing images is often affected by factors such as atmospheric conditions and urban-rural angles, resulting in quality problems such as noise and blurring in the images. Most of the existing compression methods do not consider these degradation factors and directly compress the noisy images, which will cause the noise to be encoded together, not only reducing the compression efficiency but also affecting the quality of the reconstructed images. Although image denoising can be performed first and then compression, this sequential processing method increases the computational overhead and may cause secondary loss of information.
[0007] Fifth, when the downlink transmission of remote sensing satellite data usually uses wireless communication methods, the change of channel conditions will affect the reliability of data transmission. Traditional compression methods lack a fault-tolerant mechanism for transmission errors. Once a bit error occurs, it may cause the image to be unable to be decoded correctly. Although the reliability can be improved by adding error correction coding, this will reduce the transmission efficiency of the effective data. Therefore, the issue of transmission reliability needs to be considered at the compression coding stage.
[0008] Sixth, with the popularization of remote sensing applications, users have put forward higher requirements for the access efficiency of compressed images. Traditional compression methods usually require decoding the entire image to view a local area, which is less efficient when dealing with large-scale remote sensing images. Although some methods support block-based encoding and decoding, the block effect and the reduction of compression efficiency are still problems that need to be solved.
[0009] In summary, there is an urgent need to develop a joint compression method that can make full use of the complementary information of MS and PAN images, ensure the reconstruction quality, and has strong robustness and practicability. Summary of the Invention
[0010] To solve the existing problems, the present invention provides a joint compression method for panchromatic images and multispectral images based on deep learning. The specific scheme is as follows:
[0011] A joint compression method for panchromatic images and multispectral images based on deep learning, comprising the following steps:
[0012] S1, preprocess the input panchromatic image PAN and multispectral image MS, and upsample the multispectral image to the same spatial resolution as the panchromatic image;
[0013] S2, construct a joint feature extraction network to extract the complementary features of the panchromatic image and the multispectral image simultaneously;
[0014] S3, remove the spatial-spectral redundancy through a multi-level feature fusion mechanism;
[0015] S4, perform joint probability modeling and entropy coding on the fused features to achieve efficient compression;
[0016] S5. Decoding and reconstruction: A dual-branch decoder architecture is adopted to reconstruct the panchromatic image and the multispectral image respectively.
[0017] Preferably, in step S1, the preprocessing of the multispectral image uses mean shift operation to remove the statistical deviation of each channel of the image, specifically subtracting the pre-calculated means of the RGB three channels, and then performing feature extraction; the preprocessing of the panchromatic image is to expand it into a three-channel image to match the feature dimension of the multispectral image.
[0018] Preferably, the way to perform channel expansion on the PAN image is to copy the single-channel data three times to form a three-channel image with the same spatial resolution.
[0019] Preferably, when performing feature extraction on the multispectral image in step S2, first use a 5×5 convolutional layer with reflection padding for initial feature extraction, and then perform deep feature learning through multiple residual blocks. Each residual block contains two convolutional layers and a skip connection; when performing feature extraction on the panchromatic image, use three consecutive 3×3 convolutional layers for feature extraction, and use downsampling operations with a stride of 2 between each layer to gradually reduce the spatial size of the feature map.
[0020] Preferably, the multi-level feature fusion mechanism in step S3 is a three-level feature fusion mechanism; the first-level fusion focuses on the integration of low-level features, aligns the feature channels through a 1×1 convolutional layer, and then uses a 3×3 convolutional layer for feature fusion; the second-level fusion is performed on the middle-level features, adding an attention mechanism to highlight important features; the third-level fusion processes high-level semantic features, adopting an improved coordinate attention mechanism that can simultaneously focus on information in both the spatial and channel dimensions.
[0021] Preferably, the multi-level feature fusion mechanism in step S3 uses the high-resolution information of the panchromatic image to guide the feature extraction of the multispectral image in the spatial dimension; in the spectral dimension, uses the rich spectral information of the multispectral image to enhance the features of the panchromatic image; and realizes the dynamic fusion of features through adaptive weight learning to remove information redundancy.
[0022] Preferably, removing the spatial-spectral redundancy in step S3 includes: using the local correlation of the image to remove spatial redundancy; using the correlation between multispectral bands to remove spectral redundancy; and using the complementary information between the multispectral and panchromatic images to remove cross-modal redundancy.
[0023] Preferably, step S4 specifically includes the following steps:
[0024] S41. Considering the statistical correlation between the two image features, for the features after fusion in step S3, capture the spatial-spectral joint probability distribution through a context model;
[0025] S42. Adopt a threshold processing mechanism for quantization processing;
[0026] S43. Implement precise entropy coding based on the Gaussian mixture model.
[0027] Preferably, achieve efficient compression through the following optimization objectives: minimize the reconstruction error, minimize the rate loss, and maximize the degree of feature reuse.
[0028] Preferably, in step S5, for the MS image decoder, a multi-level upsampling structure is adopted. Each level includes a transposed convolutional layer and an inverse GDN layer. After each upsampling level, a residual learning module is added, which contains multiple residual blocks for restoring the detailed information of the image. For the PAN image decoder, while adopting the upsampling module, a spatial attention mechanism and a high-frequency detail enhancement module are added. The spatial attention mechanism guides feature reconstruction by generating an attention map. The high-frequency detail enhancement module contains two convolutional layers with skip connections for restoring high-frequency texture information.
[0029] A system based on the above method, including: an image input module, an image preprocessing module, a joint feature extraction module, a joint probability coding module, and a decoding and reconstruction module;
[0030] The image input module includes a multi-spectral image input module and a panchromatic image input module. The multi-spectral input module is used to input an RGB three-channel multi-spectral image with a size of H×W×3, perform format standardization processing on the image, and ensure that the image data type is float32 and the pixel value range is [0,1]. The panchromatic image input module is used to input a single-channel panchromatic image with a size of 4H×4W×1, perform format standardization processing on the image, and ensure that the image data type is float32 and the pixel value range is [0,1].
[0031] The image preprocessing module processes the multi-spectral image and the panchromatic image respectively. The preprocessing of the multi-spectral image includes: performing mean shift - subtracting the preset RGB channel mean, upsampling to the same resolution as the PAN image, and edge padding processing. The preprocessing of the panchromatic image includes: channel expansion - copying the single channel to three channels, normalization processing, and edge padding processing.
[0032] The joint feature extraction module is used to jointly extract and fuse the features of MS and PAN images, including a multi-spectral image feature extraction module, a panchromatic image feature extraction module, and a feature fusion module. The multi-spectral image feature extraction module contains multiple residual blocks and attention modules, and is used to extract multi-scale features using a deep convolutional network and generate multi-level feature maps; the panchromatic image feature extraction module uses a dedicated feature extraction network to maintain spatial detail information and generate corresponding feature maps; the feature fusion module adopts a three-level progressive fusion mechanism to perform adaptive feature alignment and fusion and remove redundant information;
[0033] The joint probability encoding module is used to achieve efficient compression encoding of features, including a context model and an entropy encoding module; the context model is used to extract context features using 3D masked convolution, establish statistical dependencies between features, and generate probability distribution parameters; the entropy encoding module, based on the probability estimation of the mixture Gaussian model, uses arithmetic coding to achieve lossless compression and generate the final bitstream;
[0034] The decoding and reconstruction module is used to reconstruct the panchromatic image and the multi-spectral image respectively using a dual-branch decoder architecture.
[0035] The beneficial effects of the present invention are as follows:
[0036] 1. The present invention sets up an effective feature fusion mechanism to realize the joint expression of multi-spectral images and panchromatic images, enabling the two modal images to complement and enhance each other.
[0037] 2. The present invention constructs a joint probability model to improve the compression encoding efficiency, accurately models the joint probability distribution of the compressed features of MS and PAN images, and designs an efficient context model to capture the correlation between features, reducing the computational complexity of probability estimation while maintaining the compression performance.
[0038] 3. The present invention maintains the characteristics of each image under the joint compression framework, avoids spectral information aliasing and distortion during the feature fusion process, maintains the high-frequency details and edge structures of the PAN image, and constructs a suitable reconstruction mechanism to keep the decoded image with its original characteristics.
[0039] 4. The present invention realizes flexible control of the compression ratio during the joint compression process, realizes independent adjustment of the compression ratios of MS and PAN images, maintains the effectiveness of feature fusion at different compression ratios, and balances the bit allocation of the two images in the joint compression.
[0040] In summary, the present invention realizes the efficient joint compression of MS and PAN images by constructing innovative modules such as a multi-level feature fusion mechanism, an improved joint probability model, and a dual-branch decoder. Experimental results show that compared with independent compression methods, the present invention significantly improves the compression efficiency while maintaining the image quality, and has important practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 is a flowchart of the method of the present invention;
[0043] Figure 2 is a structural diagram of the model network of the present invention;
[0044] Figure 3 is a block diagram of the overall principle of system compression coding;
[0045] Figure 4 is a structural diagram of multi-level feature fusion;
[0046] Figure 5 is a structural diagram of the joint probability coding module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0048] As Figure 1 and Figure 2 show, the present invention realizes the joint feature extraction, coding compression, and decoding reconstruction of multi-spectral (MS) images and panchromatic (PAN) images by designing an end-to-end deep learning framework, and can be applied to fields such as remote sensing satellite data transmission, remote sensing image storage management, and geographic information systems.
[0049] The joint compression method for multi - resolution remote sensing images proposed by the present invention mainly includes four core stages: the pre - processing stage, the feature fusion stage, the entropy coding stage, and the decoding and reconstruction stage. These four stages form a complete end - to - end compression framework, which can effectively achieve the joint compression and reconstruction of MS and PAN images.
[0050] As Figure 1 , a joint compression method for panchromatic images and multi - spectral images based on deep learning includes the following steps:
[0051] S1, pre - process the input panchromatic image PAN and multi - spectral image MS, and upsample the multi - spectral image to the same spatial resolution as the panchromatic image.
[0052] Specifically, the present invention first pre - processes the input MS image and PAN image. For the MS image, a mean - shift operation is used to remove the statistical bias of each channel of the image, specifically by subtracting the pre - calculated mean values of the RGB three channels. This step helps to standardize the image data distribution and facilitate subsequent feature extraction. For the PAN image, since it is a single - channel image, it needs to be extended to three channels to match the feature dimension of the MS image. The extension method is to copy the single - channel data three times to form a three - channel image with the same spatial resolution.
[0053] In addition, due to a 4 - fold difference in the spatial resolution between the MS image and the PAN image, the MS image needs to be upsampled. The present invention uses the bilinear interpolation method for upsampling to maintain the smoothness of the image. After such pre - processing, the two images are consistent in both the number of channels and the spatial resolution, laying a foundation for subsequent feature fusion.
[0054] As Figure 4 shown, S2, construct a joint feature extraction network to extract complementary features of the panchromatic image and the multi - spectral image simultaneously.
[0055] Feature fusion is one of the core innovations of the present invention, which adopts a unique multi - level fusion architecture. First, feature extraction is performed on the MS image and the PAN image respectively. For the MS image, an initial feature extraction is performed using a 5×5 convolutional layer with reflection padding, which can effectively avoid edge effects. The extracted features are then subjected to deep feature learning through multiple residual blocks. Each residual block contains two convolutional layers and a skip connection, and this structure can effectively prevent the problem of gradient disappearance.
[0056] For the PAN image, considering its high spatial resolution, three consecutive 3×3 convolutional layers are used for feature extraction, and a downsampling operation with a stride of 2 is used between each layer to gradually reduce the spatial size of the feature map. This progressive feature extraction method can effectively capture the multi - scale spatial information in the PAN image.
[0057] S3. Remove the spatial-spectral redundancy through a multi-level feature fusion mechanism.
[0058] After the respective feature extractions are completed, the present invention designs a three-level feature fusion mechanism. The first-level fusion mainly focuses on the integration of low-level features. The feature channels are aligned through a 1×1 convolutional layer, and then a 3×3 convolutional layer is used for feature fusion. The second-level fusion is carried out on the middle-level features, and an attention mechanism is added to highlight important features. The third-level fusion processes high-level semantic features and adopts an improved coordinate attention mechanism, which can simultaneously focus on the information in the spatial and channel dimensions.
[0059] The multi-level feature fusion mechanism uses the high-resolution information of the panchromatic image to guide the multi-spectral image feature extraction in the spatial dimension; in the spectral dimension, it uses the rich spectral information of the multi-spectral image to enhance the panchromatic image features; and realizes the dynamic fusion of features through adaptive weight learning to remove information redundancy.
[0060] Removing the spatial-spectral redundancy includes: using the local correlation of the image to remove spatial redundancy; using the correlation between multi-spectral bands to remove spectral redundancy; and using the complementary information between the multi-spectral and panchromatic images to remove cross-modal redundancy.
[0061] S4. Perform joint probability modeling and entropy coding on the fused features to achieve efficient compression. Specifically, it includes the following steps: S41. Considering the statistical correlation between the two image features, for the features fused in step S3, capture the spatial-spectral joint probability distribution through a context model; S42. Adopt a threshold processing mechanism for quantization processing; S43. Implement accurate entropy coding based on a mixture Gaussian model.
[0062] Specifically, the goal of the entropy coding stage is to achieve efficient compression of the features. The present invention first estimates the probability of the fused features and adopts a context model based on 3D masked convolution. This model uses 3D masked convolution of 11×11×11 for initial feature extraction, which can effectively capture the probability dependence relationship of the feature map in the spatial and channel dimensions. Subsequently, the features are mapped to the probability distribution space through multi-layer 1×1×1 point convolution.
[0063] The present invention innovatively adopts a mixture Gaussian distribution model to describe the probability distribution of the features. Specifically, 9 output channels are used to represent the mean, variance, and mixture weights of 3 Gaussian distributions respectively. This mixture distribution model has stronger expressive power than a single Gaussian distribution and can more accurately model complex feature distributions.
[0064] In terms of quantization, during the training phase, adding uniform noise is adopted to simulate the quantization effect. This soft quantization method enables end-to-end training of the entire model. During the testing phase, rounding is directly used for quantization. To improve the compression efficiency, the present invention also introduces a threshold processing mechanism before quantization, which can suppress unimportant eigenvalue.
[0065] S5. Decoding and reconstruction: A dual-branch decoder architecture is adopted to reconstruct the panchromatic image and the multispectral image respectively.
[0066] During the decoding and reconstruction phase, the present invention adopts a dual-branch decoder architecture to reconstruct the MS image and the PAN image respectively. For the MS image decoder, a multi-level upsampling structure is adopted, and each level contains a transposed convolutional layer and an inverse GDN (Generalized Divisive Normalization) layer. The inverse GDN can effectively remove the statistical correlation of features and improve the reconstruction quality. After each upsampling level, a residual learning module containing multiple residual blocks is also added to restore the detailed information of the image.
[0067] The PAN image decoder adopts an improved structure. In addition to the basic upsampling module, a spatial attention mechanism and a high-frequency detail enhancement module are particularly added. The spatial attention mechanism guides feature reconstruction by generating an attention map, which can better restore the structural information of the image. The high-frequency detail enhancement module contains two convolutional layers with skip connections, which are specifically used to restore the high-frequency texture information.
[0068] Efficient compression is achieved through the following optimization objectives: minimizing the reconstruction error, minimizing the rate loss, and maximizing the degree of feature reuse.
[0069] In addition, the present invention adopts a phased training strategy. In the first phase, the reconstruction quality is the main objective, and a weighted combination of the MSE loss and the MS-SSIM loss is used. The MSE loss ensures pixel-level reconstruction accuracy, while the MS-SSIM loss focuses on maintaining the structural information. In the second phase, rate-distortion trade-off is introduced, and a bit rate term is added to the loss function, and the compression rate and the reconstruction quality are balanced through an adjustable weight coefficient.
[0070] The specific training parameter settings are as follows: The batch size is set to 16, the Adam optimizer is used, the initial learning rate is 1e-4, and a phased decay strategy is adopted. In terms of network parameters, the main feature channel number M is set to 192, and the bottleneck layer channel number N2 is set to 128. The selection of these parameters has been verified through a large number of experiments and can achieve a good balance between the model capacity and the computational efficiency.
[0071] A system based on the above method, comprising: an image input module, an image preprocessing module, a joint feature extraction module, a joint probability encoding module, and a decoding and reconstruction module.
[0072] The image input module includes a multispectral image input module and a panchromatic image input module; the multispectral input module is used to input an RGB three-channel multispectral image with a size of H×W×3, perform format standardization processing on the image, and ensure that the image data type is float32 and the pixel value range is [0,1]; the panchromatic image input module is used to input a single-channel panchromatic image with a size of 4H×4W×1, perform format standardization processing on the image, and ensure that the image data type is float32 and the pixel value range is [0,1].
[0073] The image preprocessing module processes the multispectral image and the panchromatic image respectively; the preprocessing of the multispectral image includes: performing mean shift - subtracting the preset RGB channel means, upsampling to the same resolution as the PAN image, and edge padding processing; the preprocessing of the panchromatic image includes: channel expansion - copying the single channel to three channels, normalization processing, and edge padding processing.
[0074] The joint feature extraction module is used to realize the joint extraction and fusion of MS and PAN image features, including a multispectral image feature extraction module, a panchromatic image feature extraction module, and a feature fusion module. The multispectral image feature extraction module contains multiple residual blocks and attention modules, and is used to extract multi-scale features and generate multi-level feature maps using a deep convolutional network; the panchromatic image feature extraction module uses a dedicated feature extraction network to retain spatial detail information and generate corresponding feature maps; the feature fusion module adopts a three-stage progressive fusion mechanism to perform adaptive feature alignment and fusion and remove redundant information.
[0075] As Figure 5 , the joint probability encoding module is used to realize the efficient compression encoding of features, including a context model and an entropy encoding module; the context model is used to extract context features using 3D masked convolution, establish statistical dependencies between features, and generate probability distribution parameters; the entropy encoding module, based on the probability estimation of the mixture Gaussian model, adopts arithmetic coding to achieve lossless compression and generate the final bitstream.
[0076] The decoding and reconstruction module is used to adopt a two-branch decoder architecture to reconstruct the panchromatic image and the multispectral image respectively.
[0077] As Figure 3For the overall system compression encoding far from the block diagram, the MS image (1-1) and the PAN image (1-2) are respectively input into their respective preprocessing modules (1-3, 1-4). The preprocessed features are simultaneously input into the joint feature extraction module (1-5). The extracted joint features are input into the probability encoding module (1-6) for compression. The whole process forms an end-to-end joint compression framework.
[0078] Example: This example uses a standard remote sensing image dataset for testing. The input data includes a multispectral MS image (256×256×3) and a panchromatic PAN image (1024×1024×1). The system implementation uses the PyTorch framework and is trained and tested on a workstation configured with an NVIDIA RTX 4090 GPU.
[0079] 1. Basic configuration
[0080] The network uses the following parameter configuration: the number of backbone feature channels M = 192, the number of bottleneck layer channels N2 = 128, the number of residual blocks is 6, and the number of attention modules is 4. The training uses the Adam optimizer, the batch size is 16, the initial learning rate is 1e-4, and the training is carried out for 1000 rounds.
[0081] 2. Experimental results
[0082] By adjusting the rate-distortion weight lambda, the following performance is obtained at different bitrates:
[0083] Low bitrate (lambda = 36, bpp = 0.096)
[0084] MS: PSNR = 40.96 dB, MS-SSIM = 21.95
[0085] PAN: PSNR = 34.83 dB, MS-SSIM = 12.46
[0086] Medium bitrate (lambda = 256, bpp = 0.2838)
[0087] MS: PSNR = 45.55 dB, MS-SSIM = 26.43
[0088] PAN: PSNR = 39.85 dB, MS-SSIM = 18.00
[0089] High bitrate (lambda = 1024, bpp = 0.955)
[0090] MS: PSNR = 44.62 dB, MS-SSIM = 27.39
[0091] PAN: PSNR = 39.50 dB, MS-SSIM = 19.33
[0092] Experiments show that when lambda = 256, the system reaches the optimal performance balance point. At this time, at a low bit rate of 0.2838 bpp, both the MS image and the PAN image achieve high reconstruction quality.
[0093] The actual scenario applications are as follows:
[0094] Based on the optimal configuration (lambda = 256) of the above embodiments, verification is carried out in two actual application scenarios:
[0095] 1. Real-time data transmission of remote sensing satellites
[0096] Test conditions: Transmission bandwidth 10 Mbps, latency requirement < 100 ms Measured results:
[0097] Compression ratio: 0.2838 bpp (approximate compression ratio of 35:1)
[0098] Encoding latency: 45 ms
[0099] Decoding latency: 38 ms
[0100] Image quality: MS-PSNR = 45.55 dB, PAN-PSNR = 39.85 dB
[0101] 2. Storage of large-scale remote sensing data
[0102] Test conditions: 100 TB of original data, requiring support for fast retrieval Measured results:
[0103] Reduction in storage space: Occupies 2.8 TB after compression
[0104] Retrieval speed: Average response time < 50 ms
[0105] Reconstruction quality: The same as in the real-time transmission scenario
[0106] The above embodiments verify the feasibility and superiority of the present invention in different application scenarios. While ensuring a high compression ratio, the system can maintain good reconstruction quality and meet the actual application requirements.
[0107] The present invention sets up an effective feature fusion mechanism to achieve the joint expression of multi-spectral images and panchromatic images, enabling the two modal images to complement and enhance each other.
[0108] The present invention constructs a joint probability model to improve the compression coding efficiency, accurately models the joint probability distribution of the compression features of MS and PAN images, and designs an efficient context model to capture the correlation between features, while maintaining the compression performance and reducing the computational complexity of probability estimation.
[0109] Meanwhile, the present invention maintains the characteristics of each image under the joint compression framework, avoids the aliasing and distortion of spectral information during the feature fusion process, preserves the high-frequency details and edge structures of the PAN image, constructs a suitable reconstruction mechanism, and enables the decoded image to maintain its original characteristics.
[0110] In addition, the present invention realizes the flexible control of the compression ratio during the joint compression process, achieves the independent adjustment of the compression ratios of the MS and PAN images, maintains the effectiveness of feature fusion at different compression ratios, and balances the bit allocation of the two images in the joint compression.
[0111] In summary, the present invention realizes the efficient joint compression of the MS and PAN images through innovative modules such as constructing a multi-level feature fusion mechanism, an improved joint probability model, and a dual-branch decoder. The experimental results show that, compared with the independent compression method, the present invention significantly improves the compression efficiency while maintaining the image quality, and has important practical application value.
[0112] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0113] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0114] Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A joint compression method for panchromatic images and multispectral images based on deep learning, characterized in that, it includes the following steps: S1. Preprocess the input panchromatic image and multispectral image, and upsample the multispectral image to the same spatial resolution as the panchromatic image; S2. Construct a joint feature extraction network to extract complementary features of the panchromatic image and multispectral image simultaneously; S3. Remove spatial-spectral redundancy through a multi-level feature fusion mechanism; S4. Perform joint probability modeling and entropy coding on the fused features to achieve efficient compression; S5. Decode and reconstruct, and adopt a dual-branch decoder architecture to reconstruct the panchromatic image and multispectral image respectively.
2. The method according to claim 1, characterized in that: In step S1, the preprocessing of the multispectral image uses a mean shift operation to remove the statistical bias of each channel of the image, specifically subtracting the pre-calculated RGB three-channel mean value, and then performing feature extraction; the preprocessing of the panchromatic image is to expand it into a three-channel image to match the feature dimension of the multispectral image.
3. The method according to claim 1, characterized in that: When performing feature extraction on the multispectral image in step S2, first use a 5×5 convolutional layer with reflection padding for initial feature extraction, and then perform deep feature learning through multiple residual blocks. Each residual block contains two convolutional layers and a skip connection; when performing feature extraction on the panchromatic image, use three consecutive 3×3 convolutional layers for feature extraction, and use a downsampling operation with a stride of 2 between each layer to gradually reduce the spatial size of the feature map.
4. The method according to claim 1, characterized in that: The multi-level feature fusion mechanism in step S3 is a three-level feature fusion mechanism; the first-level fusion focuses on the integration of low-level features, aligns the feature channels through a 1×1 convolutional layer, and then uses a 3×3 convolutional layer for feature fusion; The second-level fusion is performed on the middle-level features, and an attention mechanism is added to highlight important features; The third-level fusion processes high-level semantic features and adopts an improved coordinate attention mechanism, which can simultaneously focus on information in the spatial and channel dimensions.
5. The method according to claim 1, characterized in that: The multi-level feature fusion mechanism in step S3 uses the high-resolution information of the panchromatic image to guide the feature extraction of the multispectral image in the spatial dimension; in the spectral dimension, it uses the rich spectral information of the multispectral image to enhance the features of the panchromatic image; and realizes the dynamic fusion of features through adaptive weight learning to remove information redundancy.
6. The method according to claim 6, characterized in that, Removing spatial-spectral redundancy in step S3 includes: using the local correlation of the image to remove spatial redundancy; using the correlation between multispectral bands to remove spectral redundancy; using the complementary information between the multispectral and panchromatic images to remove cross-modal redundancy.
7. The method according to claim 1, characterized in that, Step S4 specifically includes the following steps: S41. Considering the statistical correlation between the two image features, for the features fused in step S3, capture the spatial-spectral joint probability distribution through a context model; S42. Adopt a threshold processing mechanism for quantization processing; S43. Implement precise entropy coding based on the Gaussian mixture model.
8. The method according to claim 1, wherein, efficient compression is achieved through the following optimization objectives: minimizing the reconstruction error, minimizing the rate loss, and maximizing the degree of feature reuse.
9. The method according to claim 1, wherein, in step S5, for the MS image decoder, a multi-level upsampling structure is adopted. Each level includes a transposed convolutional layer and an inverse GDN layer. And after each upsampling level, a residual learning module is added, which contains multiple residual blocks for restoring the detailed information of the image; for the PAN image decoder, while adopting an upsampling module, a spatial attention mechanism and a high-frequency detail enhancement module are added. The spatial attention mechanism generates an attention map to guide feature reconstruction, and the high-frequency detail enhancement module contains two convolutional layers with skip connections for restoring high-frequency texture information.
10. A system based on the method according to claims 1-9, wherein, it includes: an image input module, an image preprocessing module, a joint feature extraction module, a joint probability coding module, and a decoding and reconstruction module; the image input module includes a multi-spectral image input module and a panchromatic image input module; the multi-spectral input module is used to input an RGB three-channel multi-spectral image with a size of H×W×3, perform format standardization processing on the image, and ensure that the image data type is float32 and the pixel value range is [0,1]; the panchromatic image input module is used to input a single-channel panchromatic image with a size of 4H×4W×1, perform format standardization processing on the image, and ensure that the image data type is float32 and the pixel value range is [0,1]; the image preprocessing module processes the multi-spectral image and the panchromatic image respectively; the preprocessing of the multi-spectral image includes: performing mean shift - subtracting the preset RGB channel means, upsampling to the same resolution as the PAN image, and edge padding processing; the preprocessing of the panchromatic image includes: channel expansion - copying the single channel to three channels, normalization processing, and edge padding processing; the joint feature extraction module is used to realize the joint extraction and fusion of MS and PAN image features, including a multi-spectral image feature extraction module, a panchromatic image feature extraction module, and a feature fusion module. The multi-spectral image feature extraction module contains multiple residual blocks and an attention module for extracting multi-scale features using a deep convolutional network and generating multi-level feature maps; the panchromatic image feature extraction module uses a dedicated feature extraction network to retain spatial detail information and generate corresponding feature maps; the feature fusion module adopts a three-level progressive fusion mechanism for adaptive feature alignment and fusion and removing redundant information; The joint probability encoding module is used to achieve efficient compression encoding of features, including a context model and an entropy encoding module; the context model is used to extract context features using 3D masked convolution, establish statistical dependencies between features, and generate probability distribution parameters; the entropy encoding module, based on the probability estimation of the mixture Gaussian model, uses arithmetic coding to achieve lossless compression and generate the final bitstream; The decoding and reconstruction module is used to adopt a dual-branch decoder architecture to reconstruct the panchromatic image and the multispectral image respectively.
Citation Information
Cited By
Double-flow space-spectrum fusion method and device, electronic equipment and storage medium
CN121482545A
Hyperspectral compression imaging method and system, terminal and storage medium
CN121661515A