Remote sensing image compression method of dynamic feature enhancement network based on multi-dimensional collaborative side information guidance
By introducing a dynamic feature enhancement network guided by multi-dimensional collaborative edge information in remote sensing image compression, the problem of structural feature loss under high compression ratio is solved, efficient remote sensing image compression and reconstruction is achieved, and compression performance and image quality are significantly improved.
Patent Information
- Application Number
- CN202510034960.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing remote sensing image compression methods are prone to losing structural features under high compression ratios, resulting in artifacts, block effects and blurring, and thus losing important information about the image.
A dynamic feature enhancement network (DMENet) based on multi-dimensional collaborative edge information guidance is proposed. Through the multi-dimensional feature extraction module, slice dynamic pyramid module and potential representation space enhancement module guided by edge information, it realizes efficient compression and reconstruction of remote sensing images, and retains high-quality structural features.
While achieving high-fidelity remote sensing image compression, DMENet can effectively retain structural features, significantly improve the compression performance of the model, and perform superiorly in evaluation indicators such as PSNR and MS-SSIM.
Smart Images

Figure CN119941880A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a remote sensing image compression method. Background Art
[0002] Remote sensing images can reflect many features of land objects, such as terrain, temperature, and crop types. Therefore, remote sensing images have been widely used in many fields such as environmental monitoring, geological science, and military reconnaissance. However, with the development of sensor technology, the resolution of remote sensing images has been continuously improved. In addition, billions of remote sensing images are captured and transmitted every day. For the above reasons, a high-fidelity remote sensing image compression method with higher compression efficiency is urgently needed.
[0003] At present, traditional image compression methods have achieved some results. The classic JPEG and JPEG2000 are mainly composed of three parts: image transformation, quantization and entropy coding. First, the image is transformed and dequantized; then, important information is retained through quantization; finally, entropy coding is used to compress the decorrelation coefficient. In addition, BPG and WebP with superior performance have also been born in the field of image compression. Many scholars have made targeted research and improvements on the properties of remote sensing images with high information entropy, rich texture, and rich features of different scales. For example, Báscones et al. proposed a method to combine principal component analysis and JPEG2000 to compress hyperspectral image data, so as to achieve the effect of dimensionality reduction and retention of main spectral information. Li et al. used MDSI as a quality evaluation indicator to improve the BPG compression algorithm, and provided more accurate remote sensing image quality control through a two-step compression strategy, achieving consistency in compression efficiency and image quality. Traditional remote sensing image compression methods can be divided into predictive coding, transform coding and vector quantization. For example, 3D-MBLP uses prediction technology to first eliminate image spatial redundancy, then predict the current frequency band content, and finally efficiently encode the prediction error through an entropy decoder. As a transform compression method for three-dimensional images, 3D-SPIHT achieves efficient image compression by applying 3D wavelet transform in spatial and spectral domains. Qian developed an efficient fast vector quantization compression algorithm for multispectral images. Its core strategy is to directly map the input vector to the best matching codeword index in the codebook, thereby significantly improving the efficiency of data transmission and storage. However, traditional remote sensing image compression methods have the following limitations. Under high compression ratios, there are obvious artifacts and block effects in the reconstructed image. Therefore, for remote sensing images, it is difficult to obtain high-fidelity images at high compression ratios using traditional compression methods.
[0004] In order to seek breakthroughs, researchers focus on deep learning technology, which has been popular in recent years. Classic deep learning-based image compression frameworks mainly include autoencoders (AE) and variational autoencoders (VAE). SSCNet proposed by Riccardo's team uses deep convolutional autoencoder technology to efficiently compress space science and satellite image big data, while showing excellent performance in compression ratio and signal reconstruction. Alves' team designed a simplified variational autoencoder, especially for the computational resource constraints in satellite image compression. By reducing the network size and optimizing the entropy model, the encoder effectively reduces the computational complexity while ensuring compression efficiency. However, compared with AE, the VAE framework has a continuous mapping space, so it shows stronger image reconstruction capabilities, especially good at generating images with smooth transitions between pixels. In recent years, baseline networks based on VAE have performed well in image compression, surpassing traditional methods and achieving efficient and high-quality compression. VAE-based image compression networks usually include encoders, entropy encoders, and decoders. They first compress images through neural networks, then quantize pixel data, and finally use traditional coding techniques to generate efficient bit streams. In addition, in order to improve modeling accuracy and make full use of prior information, some compression models will introduce entropy models (Laplace entropy model, mixed Gaussian model, hierarchical entropy model, etc.) into the framework. Based on the above theory, some researchers have developed some deep learning-based compression networks for remote sensing images and achieved good rate-distortion performance.
[0005] Common deep learning techniques for remote sensing image compression mainly include three categories: image compression methods based on convolutional neural network (CNN), image compression methods based on Transformer, and image compression methods based on generative adversarial network (GAN). Among the CNN-based methods, Li et al. proposed a remote sensing image compression method (OF-RSIC) based on deep learning, which achieved high object fidelity compression at low bit rate by distinguishing object and background areas, optimizing global and local information encoding and bit rate allocation. In addition, Shao et al. proposed a compression network LR-CompNet for remote sensing images, which effectively captures extensive contextual information by fusing long-range convolution and improved non-local attention mechanism, achieving efficient compression while maintaining a lightweight design with low computational load. Among the Transformer-based image compression methods, Chuan et al. constructed a hyper-prior network framework that integrates Transformer and CNN for remote sensing image compression, taking into account local and non-local redundancy reduction, strengthening generalization ability through three-stage training, and significantly improving compression efficiency and quality. In addition, Li et al. proposed OF-RSIC, which uses deep neural networks to distinguish objects and backgrounds in remote sensing images and reduces bit rates by smoothing the background. On the other hand, it combines Transformer and patch local attention modules to optimize compression and balance bit allocation through regional differentiation loss. In the GAN-based image compression method, Han et al. proposed an edge-guided adversarial network that aims to simultaneously retain sharp texture information. In addition, Kan et al. proposed a remote sensing satellite image compression method based on conditional generative adversarial networks, which improves the quality and detail of reconstructed images by introducing Gaussian Laplace loss and perceptual metrics. Although the above methods have achieved good compression effects, they will have artifacts and blur in the reconstructed images under high compression ratios. The essence of this phenomenon is the loss of structural features, which also leads to the suboptimal rate-distortion performance of these methods.
[0006] Remote sensing images contain a wealth of structural features. Structural features mainly include edge features, texture features, structural information in various directions, etc. In the case of high compression ratios, the loss of structural features will lead to artifacts, block effects, and blurring, thereby losing important information of the image. Therefore, efficiently aligning the structural features between the original image and the reconstructed image has become a serious challenge that needs to be solved in the field of remote sensing image compression.
[0007] At present, there are some common problems in lossy remote sensing image compression methods, including blocking and blurring effects, which are particularly obvious at high compression ratios. Although some models have been developed, which apply local smoothness prior knowledge to probabilistic models to solve the above problems, it will lead to obvious loss of structural features. Summary of the invention
[0008] The purpose of the present invention is to solve the problem that in the current lossy remote sensing image compression method, remote sensing images are prone to lose structural features under high compression ratios, resulting in artifacts, block effects and blurring, thereby losing important image information, and propose a remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information.
[0009] The specific process of the remote sensing image compression method based on the dynamic feature enhancement network guided by multi-dimensional collaborative side information is as follows:
[0010] Build the DMENet network model;
[0011] The DMENet network model represents a dynamic feature enhancement network model guided by multi-dimensional collaborative side information;
[0012] The DMENet network model includes a multi-dimensional feature extraction module MDEI guided by side information, a compression module, a coding and decoding module, and a reconstruction module;
[0013] The side information guided multi-dimensional feature extraction module MDEI comprises a horizontal attention module HA, a vertical attention module VA and an edge feature extraction module EEM;
[0014] The compression module includes compression block 1, compression block 2, slice dynamic pyramid module SDPM, compression block 3, and compression block 4 in sequence;
[0015] The reconstruction module includes reconstruction block 4, reconstruction block 3, slice dynamic pyramid module SDPM, reconstruction block 2, and reconstruction block 1 in sequence;
[0016] The slicing dynamic pyramid module SDPM includes a slicing strategy SS, an irregular shape feature extraction block SEB, and a pyramid feature enhancement block PEB in sequence;
[0017] The coding and decoding module includes a probability model, a Q quantizer, an AE arithmetic coding and an AD arithmetic decoding;
[0018] The probability model includes a super encoder, a super decoder, a latent representation space enhancement module LSM, a Q quantizer, an AE arithmetic encoding, an AD arithmetic decoding, and a Factorized Entropy;
[0019] The working process of the DMENet network model is as follows:
[0020] The original remote sensing image data is input into the compression module, and the compression module outputs features;
[0021] The compression module outputs features that are input into the encoding and decoding module, and the encoding and decoding module outputs features;
[0022] The encoding and decoding module outputs features that are input into the reconstruction module, and the reconstruction module outputs the reconstructed image;
[0023] The original remote sensing image data and the reconstructed image data are respectively input into the side information guided multidimensional feature extraction module MDEI, which outputs the structural features of the original part and the structural features of the reconstructed part respectively, and calculates the structural difference Loss between the structural features of the original part and the structural features of the reconstructed part MDEI .
[0024] The beneficial effects of the present invention are:
[0025] The present invention discloses a dynamic feature enhancement network guided by multi-dimensional collaborative edge information for remote sensing image compression (DMENet), which can retain higher quality structural features while achieving high-fidelity remote sensing image compression. Firstly, a multi-dimensional feature extraction module guided by edge information (MDEI) is designed to extract horizontal structural features, vertical structural features and edge features in the image. These features are structurally aligned through loss to achieve high-quality structural feature recovery. Secondly, a slice dynamic pyramid feature enhancement module (SDPM) is designed to achieve dynamic extraction of irregular shape features and multi-scale features. Thirdly, a latent representation space enhancement module (LSM) is designed to solve the problem of deep-level feature loss caused by low information capacity of probabilistic models. Finally, the entire network performs high-quality remote sensing image compression under the guidance of a proposed rate-distortion optimization strategy (a constraint that pays more attention to structural features). Experimental results show that DMENet can compress remote sensing images more effectively compared with some advanced compression models.
[0026] The present invention achieves high-fidelity remote sensing image compression while retaining higher quality structural features. Remote sensing images contain rich structural features, which play a vital role in the reconstruction of high-fidelity images. In the present invention, structural features are divided into three categories: horizontal structural features, vertical structural features and edge features. For these three features, the present invention constructs horizontal attention (HA), vertical attention (VA) and edge feature extraction module (EEM), respectively. A three-branch multi-dimensional feature extraction module guided by edge information (MDEI) is constructed through HA, VA and EEM, and a multi-dimensional synergistic loss guided by edge information (Loss) is constructed based on this. MDSE ) is used to align the structural features between the original image and the reconstruction. Secondly, remote sensing images often contain irregular shape features (features after multiple objects overlap, often showing strange features) and multi-scale features. For this, the present invention respectively constructs a slicing strategy (SS) for improving computational efficiency and reducing memory usage, a strange feature extraction block (SEB) for enhancing the ability to extract strange features, and a pyramid feature enhancement block (PEB) for enhancing the ability to capture multi-scale features. Based on the above work, a slice dynamic pyramid module (SDPM) is constructed to realize the dynamic extraction of irregular shape features and multi-scale features. Finally, the potential space representation capability of conventional probabilistic models is insufficient. The feature maps in the probabilistic model are all deep features after multiple feature extractions, which contain a large number of spatial and channel features, while the convolution blocks in the conventional probabilistic model cannot accommodate a large number of features, which leads to the loss of features. To this end, the present invention proposes a latent representation space enhancement module (LSM) to enhance the deep feature representation capability of the probability model. MDSE, a high-performance DMENet was constructed.
[0027] The present invention has been fully experimented on three remote sensing image datasets: San Francisco, NWPU-RESISC45, and UC-Merced. The experimental results show that compared with some comparative methods, the proposed DMENet performs better in evaluation indicators such as Peak signal-to-noise ratio (PSNR) and multiscale structural similarity index metric (MS-SSIM). In addition, the reconstructed images are also used for remote sensing image scene classification to test the impact of compression methods on downstream tasks. The experimental results show that the remote sensing images reconstructed using the DMENet method have the best classification effect. The main contributions of the present invention are as follows:
[0028] 1) A multi-dimensional feature extraction module guided by edge information (MDEI) is proposed, and a multi-dimensional synergistic loss guided by edge information (Loss MDSE ). It realizes efficient extraction of multi-dimensional structural features by aligning horizontal structural features, vertical structural features and edge features.
[0029] 2) A slice dynamic pyramid module (SDPM) is designed, which efficiently realizes the dynamic extraction of irregular shape features and multi-scale features through an efficient slicing strategy, a dynamic feature capture mechanism and a multi-scale feature enhancement block.
[0030] 3) A latent representation space enhancement module (LSM) is constructed, which captures multi-level spatial features and channel features through spatial attention and channel attention respectively, effectively enhancing the deep feature representation capability of the probabilistic model.
[0031] 4) The present invention effectively embeds MDEI, SDPM, LSM and a rate-distortion optimization strategy for structural feature alignment to construct a dynamic feature enhancement network guided by multi-dimensional collaborative edgeinformation for remote sensing image compression (DMENet). Through a large number of experiments on the San Francisco, NWPU-RESISC45 and UC-Merced datasets, the superior performance of DMENet in multiple evaluation indicators is demonstrated.
[0032] In this paper, a DMENet is proposed to achieve high-fidelity remote sensing image compression while retaining higher quality structural features. First, an MDEI is designed to extract multidimensional structural features in the image. These features are structurally aligned through loss to achieve the purpose of restoring high-quality structural features. Secondly, an SDPM is designed to achieve dynamic extraction of weird features and multi-scale features. Thirdly, an LSM is designed to solve the problem of deep feature loss caused by low information capacity of probabilistic models. Finally, the entire network performs high-quality remote sensing image compression under the guidance of a rate-distortion optimization strategy that pays more attention to multidimensional structural features. Compared with other methods, the proposed DMENet achieves the best rate-distortion performance. Classification is used to evaluate the impact of reconstructed images obtained by different compression methods on applications, which proves that the proposed DMENet method can provide the best classification performance, which shows that the proposed method can more effectively retain important information in remote sensing images. In future work, we will explore variable-rate remote sensing image compression methods to greatly reduce the training cost of the model. In addition, we will further perform more detailed hierarchical processing on the compression and reconstruction process of remote sensing images to reduce the information gap between latent representation features and specific tasks, thereby further improving the rate-distortion performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 To propose the overall structure diagram of DMENet, Universal distortionD represents conventional distortion;
[0034] Figure 2It is a schematic diagram of the structure of MDEI (Origin), Input represents the input feature map, Output represents the output feature map, weight coefficients a, b, c, d, e, f, g represent 0.5, 0.5, 0.5, 0.5, 0.5, 0.1, 0.5 respectively; Avgpool represents global average pooling, Stdpool represents global standard deviation pooling; C, H, W represent the number of channels, width and height of the data block respectively; the convolution kernel shape of Barconvolution is set to 1×3; Permute represents dimensional transformation, Vector represents vector, Expand represents dimensional expansion, GaussianConv1 represents the first Gaussian convolution, Downsample represents interval sampling downsampling, Upsample represents interpolation upsampling, and the insertion value is 0; GaussianConv2 represents the second Gaussian convolution;
[0035] Figure 3 The structural diagram of EEM, Input Tensor represents the input feature map, Output Tensor represents the output feature map, M and N are 1 and 256 respectively; the convolution kernel initialized by Gaussian on the right is the convolution kernel of Gaussian convolution 1, the convolution kernel of Gaussian convolution 2 is the value of the convolution kernel of Gaussian convolution 1 multiplied by 4, and the convolution kernel size of Gaussian convolution is 5×5; the letters in each square represent the pixel value;
[0036] Figure 4 This is the structural diagram of SDPM. Input represents the input feature map, Output represents the output feature map, and in Conv2D3×3c-m cm 1, 3×3 represents the shape of the convolution kernel, the first cm represents the number of input channels, the second cm represents the number of output channels, and 1 represents the number of channel groups. The parameter settings of other convolution kernels are similar. c represents the number of input channels, 0.25c represents the number of output channels, and 1 represents the number of groups; m represents the number of channels;
[0037] Figure 5 It is a schematic diagram of the structure of LSM;
[0038] Figure 6 This is a picture of San Francisco, (a) houses, (b) coastline, (c) roads, (d) basketball courts, (e) tennis courts, (f) ports, (g) parking lots, (h) forests, (i) farmland, and (j) lakes;
[0039] Figure 7 For NWPU-RESISC45, (a) airport, (b) basketball court, (c) beach, (d) bridge, (e) desert, (f) church, (g) cloud, (h) forest, (i) port, (j) island;
[0040] Figure 8 For the UC-Merced image, (a) farmland, (b) airplane, (c) baseball field, (d) beach, (e) building, (f) forest, (g) highway, (h) golf course, (i) port, (j) overpass;
[0041] Fig. 9 The rate-distortion performance curves of different methods on San Francisco, (a) PSNR curve, (b) MS-SSIM curve, Rate (bpp) represents the bit rate, PSNR (dB) represents the peak signal-to-noise ratio, and MS-SSIM (dB) represents the multi-scale structural similarity index metric;
[0042] Fig.10 Rate-distortion performance curves of different methods on NWPU-RESISC45: (a) PSNR curve, (b) MS-SSIM curve;
[0043] Fig.11 Rate-distortion performance curves of different methods on UC-Merced: (a) PSNR curve, (b) MS-SSIM curve;
[0044] Fig.12 Visual comparison of the reconstructed images obtained by different methods on the SanFrancisco dataset, (a) original image, (b) Minnen et al. (bpp: 0.251; PSNR: 29.66; MS-SSIM: 7.55), (c) Balle et al. (hyperprior) (bpp: 0.251; PSNR: 29.74; MS-SSIM: 7.77), (d) Balle et al. (factorized-relu) (bpp: 0.250; PSNR: 29.06 MS-SSIM: 7.27), (e) Tong2023 (bpp: 0.249; PSNR: 30.25; MS-SSIM: 8.08), (f) Shi2024 (bpp: 0.251; PSNR: 30.43; MS_SSIM: 8.27), (g) JPEG2000 (b pp:0.265; PSNR:22.79; MS-SSIM:1.25), (h)Webp (bpp:0.38; PSNR:24.81; MS-SSIM:2.78), (i)BP G (bpp: 0.244; PSNR: 24.36; MS-SSIM: 1.94), (j) DMENet (bpp: 0.250; PSNR: 30.39; MS-SSIM: 8.24);
[0045] Fig.13Visual comparison of the reconstructed images obtained by different methods on the dataset UC-Merced, (a) original image, (b) Minnen et al. (bpp: 0.280; PSNR: 30.79; MS-SSIM: 6.42), (c) Balle et al. (hyperprior) (bpp: 0.278; PSNR: 30.90; MS-SSIM: 6.81), (d) Ballee et al. (factorized-relu) (bpp: 0.281; PSNR: 29.27; MS-SSIM: 5.75), (e) Tong2023 (bpp: 0.281; PSNR: 29.74; MS-SSIM: 5.07), (f) Shi2024 (bpp: 0.278; PSNR: 30 .43; MS-SSIM: 6.83), (g) JPEG2000 (bpp: 0.268; PSNR: 26.47; MS-SSIM: 1.21), (h) Webp (bpp: 0.260; PSNR: 26.88; MS- SSIM:1.88), (i) BPG (bpp: 0.271; PSNR: 27.19; MS-SSIM: 2.90), (j) DMENet (bpp: 0.281; PSNR: 31.55; MS-SSIM: 7.49);
[0046] Fig.14 The OA value diagram of the reconstructed images obtained by different compression methods in remote sensing scene classification. OverallAccuracy represents the overall accuracy. DETAILED DESCRIPTION
[0047] Specific implementation method 1: The specific process of the remote sensing image compression method based on the dynamic feature enhancement network guided by multi-dimensional collaborative side information in this implementation method is as follows:
[0048] Build the DMENet network model;
[0049] The DMENet network model represents a dynamic feature enhancement network model guided by multi-dimensional collaborative side information;
[0050] The DMENet network model includes a multi-dimensional feature extraction module MDEI guided by side information, a compression module, a coding and decoding module, and a reconstruction module;
[0051] The side information guided multi-dimensional feature extraction module MDEI comprises a horizontal attention module HA, a vertical attention module VA and an edge feature extraction module EEM;
[0052] The compression module includes compression block 1, compression block 2, slice dynamic pyramid module SDPM, compression block 3, and compression block 4 in sequence;
[0053] The reconstruction module includes reconstruction block 4, reconstruction block 3, slice dynamic pyramid module SDPM, reconstruction block 2, and reconstruction block 1 in sequence;
[0054] The slicing dynamic pyramid module SDPM includes a slicing strategy SS, an irregular shape feature extraction block SEB, and a pyramid feature enhancement block PEB in sequence;
[0055] The coding and decoding module includes a probability model, a Q quantizer, an AE arithmetic coding and an AD arithmetic decoding;
[0056] The probability model includes a hyper encoder, a hyper decoder, a latent representation space enhancement module LSM, a Q quantizer, an AE arithmetic encoding, an AD arithmetic decoding, and a Factorized Entropy;
[0057] The working process of the DMENet network model is as follows:
[0058] The original remote sensing image data is input into the compression module, and the compression module outputs features;
[0059] The compression module outputs features that are input into the encoding and decoding module, and the encoding and decoding module outputs features;
[0060] The encoding and decoding module outputs features that are input into the reconstruction module, and the reconstruction module outputs the reconstructed image;
[0061] The original remote sensing image data and the reconstructed image data are respectively input into the side information guided multidimensional feature extraction module MDEI, which outputs the structural features of the original part and the structural features of the reconstructed part respectively, and calculates the structural difference Loss between the structural features of the original part and the structural features of the reconstructed part MDEI .
[0062] The DMENet proposed in the present invention retains high-quality structural features from the perspective of aligning the multi-dimensional structural features between the original image and the reconstructed image, thereby comprehensively improving the compression performance of the model. It mainly achieves high-quality remote sensing image compression through MDEI for extracting structural features, SDPM for dynamically extracting irregular shape features and multi-scale features, LSM for enhancing the deep feature representation capabilities of the probabilistic model, and a rate-distortion optimization strategy focused on structural feature alignment. Among them, MDEI involves three submodules, including HA that can extract horizontal structural features, VA that can extract vertical structural features, and EEM that can extract edge features. Among them, SDPM involves three submodules, including SS for channel slicing, SEB that can extract irregular shape features, and PEB that can extract multi-scale features. Among them, LSM involves two parts of potential space representation capability enhancement, namely spatial feature enhancement and channel feature enhancement. In addition, the present invention constructs a Loss for calculating structural differences through MDEI. MDSE , and proposed a rate-distortion optimization strategy focusing on structural differences.
[0063] The overall framework of the proposed DMENet
[0064] The proposed DMENet retains high-quality structural features from the perspective of aligning the multi-dimensional structural features between the original image and the reconstructed image, thereby comprehensively improving the compression performance of the model. It mainly achieves high-quality remote sensing image compression through MDEI for extracting structural features, SDPM for dynamically extracting irregular shape features and multi-scale features, LSM for enhancing the deep feature representation capabilities of the probabilistic model, and a rate-distortion optimization strategy focused on structural feature alignment. Among them, MDEI involves three sub-modules, including HA that can extract lateral structural features, VA that can extract longitudinal structural features, and EEM that can extract edge features. Among them, SDPM involves three sub-modules, including SS for channel slicing, SEB that can extract irregular shape features, and PEB that can extract multi-scale features. Among them, LSM involves two parts of potential space representation capability enhancement, namely spatial feature enhancement and channel feature enhancement. In addition, the present invention constructs a Loss for calculating structural differences through MDEI. MDSE , and proposed a rate-distortion optimization strategy focusing on structural differences.
[0065] The overall structure of DMENet is shown in the figure Figure 1As shown. The present invention designs Compression block (C block1-4) for compression and Reconstruction block (Rblock1-4) for image reconstruction by reasonably selecting the convolution kernel size and reallocating the number of channels, thereby achieving excellent rate-distortion performance under low complexity. The specific parameters of C block1-4 and Rblock1-4 are shown in Table 1. Probabilitymodel is the probability model part, which mainly includes a hyper-prior network (hyper encoder, hyper decoder, LSM), Q quantizer, AE arithmetic coding and AD arithmetic decoding. The specific parameters of Hyper encoder and Hyper decoder are shown in Table 2. Between AE and AD is the minimum form (bit stream) of data in this model. In the present invention, the hyper-prior network is used to learn the probability model (i.e., entropy model) on which entropy coding depends, and is also used to generate parameters of the entropy model (i.e., mean parameter μ i and the scale parameter σ i 2 ), the entropy model is modeled as a conditional Gaussian. MDEI (where a = 0.5, b = 0.5, c = 0.5, d = 0.5, e = 0.5, f = 0.1, g = 0.5) is used to calculate the loss of the structural difference between the compressed part and the reconstructed part. MDSE , Loss MDSE Represents the difference between the structural features of the compressed part and the structural features of the reconstructed part. The smaller the loss value, the smaller the structural difference between the two sides, that is, the higher the quality of the extracted and reconstructed structural features. The loss here uses the Mean squared error (MSE), which can be expressed as Formula 1. In the Rate-Distortion Optimization part, R represents the entropy rate, λ represents the penalty coefficient used to control different bit rates, and D represents the distortion (obtained by MSE calculation).
[0066] Table 1 Specific parameters of Cblock1-4 and Rblock1-4
[0067]
[0068] Table 2 Specific parameters of Hyper encoder and Hyper decoder
[0069]
[0070] In Tables 1 and 2, N represents the number of channels, ↓ represents downsampling, ↑ represents upsampling, and RELU represents the linear rectification function. GDN represents the generalized split normalization function, and IGDN represents its inverse operation. They are nonlinear activation functions and are more suitable for normalization of image data than other normalization functions.
[0071] The overall working principle of DMENet is as follows. Image compression part: First, the remote sensing image data block passes through Cblock1-2 to obtain a shallow feature map. Then, the dynamic extraction of irregular shape features and multi-scale features is enhanced by SDPM, and the data is passed through C block3-4 to obtain a deep feature map. Then, the initially compressed deep feature map is further removed from the statistical redundancy through quantization, arithmetic coding, and the Probability model enhanced by LSM to obtain the minimum bit stream after data processing of this model. Image decoding part: The model combines the obtained bit stream with the mean μ in the Probability model i and scale σ i 2 Parameters, high-quality reconstruction of the image is performed through Rblock4, Rblock3, SDPM, Rblock2 and Rblock1. Finally, the reconstruction and the original image are input into MDEI for structural feature alignment. MDSE Join Loss Total , targeted rate-distortion optimization of structural features is performed.
[0072]
[0073] Where m represents the number of pixels. represents the predicted image and X represents the original image.
[0074] Specific implementation method 2: This implementation method is different from the specific implementation method 1 in that the remote sensing image data is input into the compression module, and the specific working process of the compression module outputting features is as follows:
[0075] The compression module includes compression block 1, compression block 2, slice dynamic pyramid module SDPM, compression block 3, and compression block 4 in sequence;
[0076] Compression block 1 includes the first convolutional layer (two-dimensional) and the first GDN in sequence. The convolution kernel size of the first convolutional layer is 7×7, the number of input channels is 3, and the number of output channels is N / 4;
[0077] Compression block 2 includes the second convolutional layer (two-dimensional) and the second GDN in sequence. The convolution kernel size of the second convolutional layer is 3×3, the number of input channels is N / 4, and the number of output channels is N / 2.
[0078] Compression block 3 includes the third convolutional layer (two-dimensional) and the third GDN in sequence. The convolution kernel size of the third convolutional layer is 3×3, the number of input channels is N / 2, and the number of output channels is 3N / 4.
[0079] Compression block 4 includes the fourth convolution layer (two-dimensional) and the fourth GDN in sequence. The convolution kernel size of the fourth convolution layer is 3×3, the number of input channels is 3N / 4, and the number of output channels is N.
[0080] GDN stands for generalized divisive normalization function; N is the number of channels;
[0081] The original remote sensing image data is sequentially input into compression block 1, compression block 2, slice dynamic pyramid module SDPM, compression block 3, and compression block 4, and the output features of compression block 4 are used as the output features of the compression module.
[0082] The other steps and parameters are the same as those in the first embodiment.
[0083] Specific implementation method three: This implementation method is different from specific implementation methods one or two in that the working process of the slice dynamic pyramid module SDPM is as follows:
[0084] Remote sensing images often contain a large number of irregular features (such as features formed by the overlap of multiple objects, features of odd shapes) and multi-scale features (such as short-range information such as ships, and long-range information such as rivers spanning the entire city). Therefore, a SDPM is designed to achieve dynamic extraction of irregular features and multi-scale features. The overall structure of SDPM is shown in the figure. Figure 4 It mainly includes three core components, namely, the channel slicing strategy SS similar to the residual connection, SEB for enhancing the ability to extract weird features, and PEB for enhancing the ability to capture multi-scale features.
[0085] When a remote sensing image data block is input into SDPM, SS will slice it along the channel dimension, then some channels will remain unchanged, while other channels will be extracted through convolution, and finally the two sets of features will be spliced in the channel dimension. This strategy uses a channel preservation method similar to residual connection, making the model more cautious and faster when adjusting parameters, significantly improving the model's ability to understand complex channel features.
[0086] The slicing dynamic pyramid module SDPM includes the slicing strategy SS, the irregular shape feature extraction block SEB and the pyramid feature enhancement block PEB in sequence;
[0087] 1) Remote sensing image data input slicing strategy SS, slicing strategy SS output feature map Output SS ; The specific process is:
[0088] Output SS =
[0089] F concat (Tensor H×W×m (Input H×W×C ),Conv 3×3 (Tensor H×W×(C-m) (Input H×W×C )))
[0090] In the formula,
[0091] Input H×W×C Represents the input feature map of the slicing strategy SS; H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels;
[0092] Input H×W×C Represents the input feature map of the slicing strategy SS; H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels;
[0093] Tensor H×W×m (Input H×W×C ) represents the input feature map of the slicing strategy SS H×W×C Feature maps corresponding to the first m channels;
[0094] Tensor H×W×(C-m) (Input H×W×C ) represents the input feature map of the slicing strategy SS H×W×C Feature maps corresponding to the remaining Cm channels;
[0095] m represents the number of channels;
[0096] Conv 3×3 Represents a convolution with a kernel size of 3×3;
[0097] F concat (,) represents the function that concatenates two tensors in the channel dimension;
[0098] Output SS Represents the output feature map of SS;
[0099] 2) Slicing strategy SS output feature map Output SS Input irregular shape feature extraction block SEB, irregular shape feature extraction block SEB output feature map Output SEB ; The specific process is:
[0100] The irregular shape feature extraction block SEB is a dynamic convolution layer, and the convolution kernel size of the dynamic convolution layer is 3×3;
[0101] Slicing strategy SS output feature map OutputSS After a layer of dynamic convolution layer, the dynamic convolution layer outputs the feature map Output SEB ;
[0102] The difference between dynamic convolution and ordinary convolution is that the position of the convolution kernel can be adaptively adjusted;
[0103] Output SS It will enter SEB to extract irregular features. In remote sensing images, many features are obtained by overlapping multiple objects, and also contain many strange-shaped features. The shapes of these features are irregular, and it is difficult for conventional square convolution kernels to effectively sample their features. Therefore, deformable convolution was introduced to construct SEB, thereby enhancing the extraction of weird features. The convolution here is added with offset learning, which allows the size and position of the convolution kernel to be dynamically adjusted according to the current image to be recognized, so that the shape of the convolution kernel can adapt to the shape and size of different objects. Here, Output SEB Represents the output feature map of SEB.
[0104] 3) Irregular shape feature extraction block SEB output feature map Output SEB Input pyramid feature enhancement block PEB, pyramid feature enhancement block PEB output feature map Output PEB ; The specific process is:
[0105] Irregular shape feature extraction block SEB output feature map Output SEB The input pyramid feature enhancement block PEB is used to extract multi-scale features and solve the problems of insufficient receptive field and loss of detail features. The tensor is subjected to pyramid-like increasing convolution through four groups of convolution kernel sizes and grouping numbers to reduce the dimension of the feature map, thereby obtaining high-quality spatial features. The four reduced-dimensional features are then concatenated in the channel dimension, and the channel features are fused through point convolution. Finally, residual connections are introduced to speed up network training.
[0106] Irregular shape feature extraction block SEB output feature map Output SEB Input 3×3 convolution layer (2D), 3×3 convolution layer outputs feature map A;
[0107] Irregular shape feature extraction block SEB output feature map Output SEB Input 5×5 convolution layer (2D), 5×5 convolution layer outputs feature map B;
[0108] Irregular shape feature extraction block SEB output feature map Output SEB Input 7×7 convolution layer (2D), 7×7 convolution layer outputs feature map C;
[0109] Irregular shape feature extraction block SEB output feature map Output SEB Input 9×9 convolution layer (two-dimensional), 9×9 convolution layer outputs feature map D;
[0110] The working process of PEB is expressed as:
[0111] Output PEB =Conv 1×1 (F concat (A,B,C,D))+Output SEB
[0112] In the formula, Conv1×1 represents point convolution;
[0113] F concat (,) represents the function that concatenates tensors in the channel dimension;
[0114] A, B, C, and D represent four groups of convolution kernel sizes and grouping numbers that increase in a pyramidal manner;
[0115] Output PEB Represents the output feature map of PEB;
[0116] The output feature map of PEB is the final output feature map of the slice dynamic pyramid module SDPM.
[0117] The other steps and parameters are the same as those in the first or second embodiment.
[0118] Specific implementation method 4: This implementation method is different from any one of the specific implementation methods 1 to 3 in that the compression module outputs features and inputs them into the encoding and decoding module, and the encoding and decoding module outputs features; the specific process is as follows:
[0119] The compression module outputs feature y and inputs it into the Q quantizer, which outputs feature feature Input AE arithmetic coding, and the output features of AE arithmetic coding are input into AD arithmetic decoding;
[0120] The output feature y of the compression module is input into the probability model Probability Model. The output of the probability model ProbabilityModel is respectively input into AE arithmetic coding and AD arithmetic decoding. The output feature of AE arithmetic coding is input into AD arithmetic decoding.
[0121] The AD arithmetic decoding output features are used as the output features of the encoding and decoding module.
[0122] The other steps and parameters are the same as those in Specific Embodiments 1 to 3.
[0123] Specific implementation method 5: This implementation method is different from any one of specific implementation methods 1 to 4 in that the compression module outputs feature y and inputs it into the probability model Probability Model, and the output of the probability model Probability Model is respectively input into AE arithmetic coding and AD arithmetic decoding; the specific process is:
[0124] The probability model includes hyper encoder, hyper decoder, latent representation space enhancement module LSM, Q quantizer, AE arithmetic coding, AD arithmetic decoding, and Factorized Entropy.
[0125] The compression module outputs the feature y which is input into the hyper encoder, and the hyper encoder outputs the feature;
[0126] The output feature of the super encoder is input into the latent representation space enhancement module LSM, and the latent representation space enhancement module LSM outputs the feature z; the output feature z of the latent representation space enhancement module LSM is input into the Q quantizer, and the Q quantizer outputs the feature
[0127] Q quantizer output characteristics Input AE arithmetic coding, AE arithmetic coding output features Input AD arithmetic decoding, AD arithmetic decoding output features
[0128] Q quantizer output characteristics Input Factorized Entropy, the features output by Factorized Entropy are input into AE arithmetic coding and AD arithmetic decoding respectively, and AD arithmetic decoding outputs features
[0129] AD arithmetic decoding output characteristics Input the latent representation space enhancement module LSM, the output features of the latent representation space enhancement module LSM are input into the hyper decoder, the output features of the hyper decoder are used as the output of the probability model Probability Model, and the output of the probability model Probability Model is respectively input into AE arithmetic coding and AD arithmetic decoding.
[0130] The other steps and parameters are the same as those in Specific Embodiments 1 to 4.
[0131] Specific implementation method 6: This implementation method is different from any one of the specific implementation methods 1 to 5 in that the compression module outputs the feature y and inputs it into the super encoder, and the super encoder outputs the feature; the specific process is:
[0132] The super encoder includes a fifth 3×3 convolution layer, a first RELU activation function layer, a sixth 3×3 convolution layer, a second RELU activation function layer, and a seventh 3×3 convolution layer in sequence;
[0133] The compression module output feature y is sequentially input into the fifth 3×3 convolution layer, the first RELU activation function layer, the sixth 3×3 convolution layer, the second RELU activation function layer, and the seventh 3×3 convolution layer. The seventh 3×3 convolution layer output feature is used as the super encoder output feature.
[0134] The latent representation space enhancement module LSM outputs features that are input into the super decoder, and the super decoder outputs features; the specific process is:
[0135] The super decoder sequentially includes an eighth 3×3 convolution layer, a third RELU activation function layer, a ninth 3×3 convolution layer, a fourth RELU activation function layer, a tenth 3×3 convolution layer, and a fifth RELU activation function layer;
[0136] The output features of the latent representation space enhancement module LSM are sequentially input into the eighth 3×3 convolution layer, the third RELU activation function layer, the ninth 3×3 convolution layer, the fourth RELU activation function layer, the tenth 3×3 convolution layer, and the fifth RELU activation function layer, and the output features of the fifth RELU activation function layer are used as the super decoder output features.
[0137] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 5.
[0138] Specific implementation method 7: This implementation method is different from any one of specific implementation methods 1 to 6 in that the super encoder outputs feature E and inputs it into the latent representation space enhancement module LSM, and the latent representation space enhancement module LSM outputs feature z; the specific process is:
[0139] The latent representation space enhancement module LSM includes large kernel band convolution LKE and point convolution in sequence;
[0140] The working process of LKE is expressed as:
[0141] Output LKE =
[0142] E+Conv 1×7 (Conv 7×1 (E))+Conv 1×11 (Conv 11×1 (E))+Conv 1×21 (Conv 21×1 (E)
[0143] In the formula, E represents the feature of the input feature after a convolution with a convolution kernel size of 5×5; Conv 7×1 Represents the input feature after a convolution with a convolution kernel size of 7×1; Conv 1×7 Represents the input feature after a convolution with a convolution kernel size of 1×7; Conv 1×11 Represents the input feature after a convolution with a convolution kernel size of 1×11; Conv 11×1 Represents the input feature after a convolution with a convolution kernel size of 11×1; Conv 1×21 Represents the input feature after a convolution with a convolution kernel size of 1×21; Conv 21×1 Represents the input feature after a convolution with a convolution kernel size of 21×1; Output LKE Represents the output features of LKE;
[0144] Output features of LKE LKE The final feature map Output is obtained by interacting channel features through a point convolution LSM :
[0145] Output LSM =Conv1×1(Output LKE )
[0146] In the formula, Conv1×1 represents 1×1 point convolution;
[0147] Feature map Output LSM That is the output feature z of the latent representation space enhancement module LSM.
[0148] The latent space representation capability of conventional probability models is insufficient. The feature maps in the probability model are deep features after multiple feature extractions, which contain a large number of spatial and channel features, while the convolution blocks in the conventional probability model cannot accommodate a large number of features, which leads to feature loss. For this, the present invention designs LSM to enhance the deep feature representation capability of the probability model, thereby greatly improving the estimation accuracy of the probability model.
[0149] The feature extraction process of LSM is divided into two parts, including large-kernel bar convolution for spatial feature enhancement (LKE) and channel feature extraction. LKE mainly adopts a four-branch structure, and the convolutions are all super-large kernel convolutions, which can greatly improve the potential space representation ability. However, super-large kernel convolutions will bring a huge computational burden, so banded convolutions are adopted to obtain the same large receptive field while halving the computational complexity. In addition, residual branches are introduced to speed up network training. It is worth noting that the convolution kernels in LKE are all single-channel modes, and there is no interaction of channel features.
[0150] In general, LSM separates spatial feature extraction from channel feature extraction and uses strip convolution with ultra-large kernels to achieve high potential space representation capabilities with low complexity, thereby greatly improving the estimation accuracy of the probability model.
[0151] The other steps and parameters are the same as those in Specific Embodiments 1 to 6.
[0152] Specific implementation eight: This implementation differs from any one of specific implementations one to seven in that the encoding and decoding module outputs features and inputs them into the reconstruction module, and the specific working process of the reconstruction module outputting the reconstructed image is as follows:
[0153] The reconstruction module includes reconstruction block 4, reconstruction block 3, slice dynamic pyramid module SDPM, reconstruction block 2, and reconstruction block 1 in sequence;
[0154] The reconstruction block 4 includes a fourteenth convolutional layer (two-dimensional) and a fourth IGDN in sequence, the convolution kernel size of the fourteenth convolutional layer is 3×3, the number of input channels is N, and the number of output channels is 3N / 4;
[0155] The reconstruction block 3 sequentially includes a thirteenth convolutional layer (two-dimensional) and a third IGDN, the convolution kernel size of the thirteenth convolutional layer is 3×3, the number of input channels is 3N / 4, and the number of output channels is N / 2;
[0156] The reconstruction block 2 includes a twelfth convolutional layer (two-dimensional) and a second IGDN in sequence. The convolution kernel size of the twelfth convolutional layer is 3×3, the number of input channels is N / 2, and the number of output channels is N / 4.
[0157] The reconstruction block 1 includes the eleventh convolutional layer (two-dimensional) and the first IGDN in sequence. The convolution kernel size of the eleventh convolutional layer is 3×3, the number of input channels is N / 4, and the number of output channels is 3;
[0158] The IGDN represents the inverse operation of the generalized split normalization function GDN; N is the number of channels;
[0159] The output features of the encoding and decoding module are sequentially input into reconstruction block 4, reconstruction block 3, slice dynamic pyramid module SDPM, reconstruction block 2, and reconstruction block 1. The output image of reconstruction block 1 is the reconstructed image output by the reconstruction module.
[0160] The other steps and parameters are the same as those in Specific Embodiments 1 to 7.
[0161] Specific embodiment 9: This embodiment is different from any one of specific embodiments 1 to 8 in that the original remote sensing image data and the reconstructed image data are respectively input into the multidimensional feature extraction module MDEI guided by the side information, and the multidimensional feature extraction module MDEI guided by the side information respectively outputs the structural features of the original part and the structural features of the reconstructed part, and calculates the structural difference Loss between the structural features of the original part and the structural features of the reconstructed part MDSE ;
[0162] The specific process is:
[0163] The side information guided multi-dimensional feature extraction module MDEI includes a horizontal attention module HA, a vertical attention module VA and an edge feature extraction module EEM;
[0164] 1) The original remote sensing image data is rotated through the permute operation (dimensional transformation) to obtain three data blocks X4, X1 and X5;
[0165] 2) Data block X4 is input into the horizontal attention module HA, and the horizontal attention module HA outputs X6 H×W×C ; expressed as:
[0166] X6 H×W×C =
[0167] Reconstruction(BarConv 1×3 (a(Avgpool(X4 H×C×W ))+b(Stdpool(X4 H×C×W ))))
[0168] Among them, BarConv 1×3 Represents a strip convolution with a convolution kernel shape of 1×3;
[0169] Avgpool stands for average pooling, and Stdpool stands for global standard deviation pooling;
[0170] H represents the height of the feature map, C represents the number of channels of the feature map, and W represents the width of the feature map;
[0171] a and b represent weight coefficients, which are 0.5 and 0.5 respectively;
[0172] Reconstruction includes Sigmoid, Expand, and Permute in turn;
[0173] X6 H×W×C Represents the horizontal structural features; mainly used to restore the data block to its original shape and perform data fusion with other branches;
[0174] Sigmoid represents the Sigmoid activation function, Expand represents dimension expansion, and Permute represents dimension transformation;
[0175] 3) Data block X5 is input into the vertical attention module VA, and the vertical attention module VA outputs X7 H×W×C ; expressed as:
[0176] X7 H×W×C =
[0177] Reconstruction(BarConv 1×3 (c(Avgpool(X5 C×W×H )+d(Stdpool(X5 C×W×H ))
[0178] Among them, BarConv 1×3 Represents a strip convolution with a convolution kernel shape of 1×3;
[0179] c and d represent weight coefficients, which are 0.5 and 0.5 respectively;
[0180] Reconstruction includes Sigmoid, Expand, and Permute in turn; X7 H×W×C Represents the vertical structural features; mainly used to restore the data block to its original shape and merge data with other branches;
[0181] 4) Data block X1 is input into edge feature extraction module EEM, and edge feature extraction module EEM outputs X3 H×W×C ; expressed as:
[0182] X3 H×W×C =
[0183] X1 H×W×C -GaussianConv2(Upsample(Downsample(GaussianConv1(X1 H×W×C ))))
[0184] Among them, GaussianConv1 represents the first Gaussian convolution;
[0185] Downsample represents interval sampling downsampling;
[0186] Upsample represents interpolation upsampling, and the insertion value is 0;
[0187] GaussianConv2 represents the second Gaussian convolution;
[0188] X3 H×W×C Represents the edge features of the image;
[0189] 5) Output X6 to the horizontal attention module HA H×W×C , vertical attention module VA output X7 H×W×C , edge feature extraction module EEM output X3 H×W×C Fusion, obtain the structural features of remote sensing images Output MDEI(Origin) :
[0190] Output MDEI(Origin) =e(X6 H×W×C )+f(X3 H×W×C )+g(X7 H×W×C )
[0191] In the formula, e, f, and g are the weight coefficients of the corresponding branches, which are 0.5, 0.1, and 0.5 respectively;
[0192] 6) Similarly, obtain the structural features of the reconstructed image Output MDEI(Reconstruction) ;
[0193] 7) Through Output MDEI(Origin) and Output MDEI(Reconstruction) Calculate the loss MDSE :
[0194] Loss MDSE =L MSE (Output MDEI(Origin) ,Output MDEI(Reconstruction) )
[0195] Where, L MSE represents the loss measured using MSE.
[0196] Loss MDSE High-quality image compression is achieved by aligning the complex structural features between the original image and the reconstructed image through rate-distortion optimization.
[0197] Remote sensing images contain rich structural features, including edge features, texture features and structural information in all directions. The loss of structural features will lead to block effects and blurring. Therefore, a three-branch MDEI is designed to extract structural features at different levels. In addition, the present invention classifies structural features into horizontal structural features, vertical structural features and edge features. Corresponding to these three types of structural features, the present invention designs HA, VA and EEM respectively, and multiplies the features of different levels obtained by weight coefficients and fuses them to obtain the Loss for calculating the structural loss. MDSE Finally, through Loss MDSE A rate-distortion optimization strategy focusing on structural feature alignment is constructed.
[0198] MDEI is essentially a module that aligns the structural features between the original image and the reconstructed image through loss. Therefore, it mainly consists of three parts: the module MDEI (Origin) for extracting the structure of the original image, the loss, and the module MDEI (Reconstruction) for extracting the structural features of the reconstructed image. Here, the loss adopts MSE. Since MDEI (Origin) and MDEI (Reconstruction) are symmetrical structures, only MDEI (Origin) is introduced here.
[0199] MDEI (Origin) is a three-branch structure, including HA for capturing horizontal architectural features, EEM for capturing edge features, and VA for capturing vertical structural features. First, the data block is input into MDEI (Origin), and it is turned through the permute operation to obtain three data blocks X4, X1, and X5. In HA, X4 is converted into a 1×1×W vector through Avgpool and Stdpool, so that only the horizontal features of the data are retained. After that, strip convolution is used to extract the horizontal structural features of the data under low complexity conditions. Then, through a series of deformation operations, the original shape of the data block is restored.
[0200] The working principle of VA is similar to that of HA, except that it extracts longitudinal structural features;
[0201] The EEM branch is mainly used to extract edge features. Its structure is shown in the figure below: Figure 3As shown. The convolution kernel of the Gaussian convolution is initialized by Gaussian to smooth and reduce noise on the image. Downsample here uses interval sampling downsampling. Upsample uses interpolation upsampling, and the insertion value is 0. After the input data has undergone the first Gaussian convolution, downsampling, upsampling and the second Gaussian convolution, the difference in the features in the data is greatly reduced, so that all features tend to be backgrounded. At this time, the edge features of the image are obtained by subtracting the data block whose features are backgrounded from the original data block (containing rich edge features).
[0202] The other steps and parameters are the same as those in Specific Embodiments 1 to 8.
[0203] Specific implementation method 10: This implementation method is different from the specific implementation methods 1 to 9 in that the DMENet network model is used to compress the original remote sensing image to be measured; the specific process is:
[0204] Based on Loss MDEI Calculate the DMENet network model loss function, expressed as:
[0205] argminProposedLoss Total =R+λ(D+ψLoss MDEI )
[0206] Where ψ represents Loss MDEI The weight coefficient of ; R represents the entropy rate, that is, the cross entropy between the potential marginal distribution and the learned entropy model; Loss Total represents the loss function of the DMENet network model; λ represents the penalty coefficient; D represents the degree of distortion between the original image and the reconstructed image;
[0207] The DMENet network model is trained based on the training set until the loss function converges to obtain a trained DMENet network model;
[0208] The training set includes original remote sensing images and reconstructed images;
[0209] The original remote sensing image to be tested is input into the compression module in the trained DMENet network model, and the compression module outputs the compressed data.
[0210] Rate-Distortion Optimization
[0211] The core training goal of the compression framework is to achieve a balance between compression rate and distortion. To achieve this goal, a rate-distortion optimization strategy is usually added to the compression framework to guide the model to train efficiently. In short, the strategy aims to ensure that while compressing data, information loss is minimized to achieve high-quality compression. The rate-distortion optimization strategy can be expressed as: argminLoss Total =R+λD
[0212] Where R represents the entropy rate, i.e., the cross entropy between the potential marginal distribution and the learned entropy model. D represents the degree of distortion between the original image and the reconstructed image, and different bit rates can be controlled by adjusting the penalty coefficient λ.
[0213]
[0214] In the formula, the bit rate is represented by the potential information and side information Together form.
[0215]
[0216] In the formula, is a learnable entropy model, It is a super encoder.
[0217] In order to further improve the efficiency and quality of image compression, this study proposes a new rate-distortion optimization strategy. Total Loss MDSE , that is, by aligning the structural features between the original image and the reconstructed image, the ability of the entire network to reconstruct structural features is improved. This new rate-distortion optimization strategy can be expressed as:
[0218] arg min ProposedLoss Total =R+λ(D+ψLoss MDEI )
[0219] Where ψ represents Loss MDSE The weight coefficient of .
[0220] The other steps and parameters are the same as those in Specific Embodiments 1 to 9.
[0221] Example:
[0222] Adequate experiments were conducted on the San Francisco, NWPU-RESISC45 and UC-Merced remote sensing image datasets. The selected datasets contain rich ground object information, which can effectively evaluate the effectiveness of the proposed DMENet method. The proposed DMENet method is compared with some excellent compression methods, including traditional codecs and compression models based on deep learning, to verify the superiority of the proposed method. Traditional image compression methods include JPEG2000, BPG and WebP. Compression models based on deep learning include Minnen ta al., Balle et al. (hyperprior), Balle et al. (factorized-relu), Tong2023 and Shi2024. The experimental results show that the proposed DMENet method has the best performance in evaluation indicators such as PSNR and MS-SSIM. In addition, this experiment also evaluates the quality of reconstructed images obtained by different compression methods from the perspective of remote sensing image classification to further verify the superiority of the DMENet method.
[0223] Dataset:
[0224] 1) Dataset San Francisco: It is a remote sensing image dataset from . It is a remote sensing image with a resolution of 17408×17408, covering various ground features such as building areas, coasts, roads, ports, lakes, etc. This paper cuts it into 256×256 pixel images and selects 3000 valid images to form a dataset. These images are divided into training set, validation set and test set in a ratio of 8:1:1. Figure 6 Some samples are shown.
[0225] 2) Dataset NWPU-RESISC45: NWPU-RESISC45 is provided by Northwestern Polytechnical University (NWPU) in China. It contains 45 categories of remote sensing image scenes, 700 images per category, and 256×256 pixels per image. The dataset has diverse scenes, including airports, deserts, churches, forests, etc. 140 images per category are selected to form a dataset of 6,300 images, which are divided into training set, validation set, and test set in a ratio of 8:1:1. Figure 7 Some samples are shown.
[0226] 3) Dataset UC-Merced: UC-Merced is provided by the University of California, Merced. It contains 21 categories of remote sensing images, 100 images per category, a total of 2100 images, with a resolution of 256×256 pixels. The images cover a variety of landforms such as farmland, airports, and forests. The dataset is divided into training set, validation set, and test set in a ratio of 8:1:1. Figure 8 Some samples are shown.
[0227] Evaluation index: In terms of image quality assessment, the present invention uses two common indexes: peak signal-to-noise ratio and multi-scale structural similarity index. In the remote sensing image scene classification part, the present invention uses overall accuracy and confusion matrix to measure the classification performance.
[0228] 1) Peak signal-to-noise ratio: PSNR compares the reconstructed image with the original image from the perspective of mean square error. The larger the PSNR value, the higher the fidelity of the reconstructed image. The peak signal-to-noise ratio can be expressed as
[0229]
[0230] Among them, MSE represents the mean square error between the original image and the reconstructed image, and max 2 (X i ) represents i The square of the maximum pixel in a band, C represents the number of bands.
[0231] 2) Multi-scale structural similarity index measurement: MS-SSIM is a multi-scale structural similarity index that measures the difference between the original image and the reconstructed image by merging image details at different resolutions. Its value range is 0 to 1, and the higher the value, the higher the similarity, that is, the higher the quality of the reconstructed image. The formula of MS-SSIM can be expressed as:
[0232]
[0233] Among them, M represents different resolutions, μ X , Represent the means of the original image and the reconstructed image, σ X , represents the standard deviation between the original image and the reconstructed image, represents the covariance between the original image and the reconstructed image, α m , m It indicates the relative importance between the two terms. C1 and C2 are constant terms used to prevent the divisor from being zero.
[0234] In order to more clearly compare the difference in MS-SSIM values, they are converted into decibel values, which can be expressed as:
[0235] MS-SSIM=-10log 10 (1-D MS-SSIM )
[0236] 3) Remote sensing scene classification index: This paper selects two widely used classification evaluation indexes to measure the quality of reconstructed images, including overall accuracy (OA) and confusion matrix (CM). The OA value is obtained by dividing the number of correctly classified images by the total number of test images, which reflects the overall performance of a classification model. CM reflects the detailed classification error and confusion degree between different scene categories. Each row in CM represents the true category, and each column represents the predicted category.
[0237] Experimental setup: In this study, the proposed DMENet method is implemented by the PyTorch framework. The optimizer used is the Adam optimizer. Two optimizers are used in this network, one is the main optimizer between the main encoder (Compression) and the main decoder (Reconstruction), and the other is the auxiliary optimizer between the super encoder and the super decoder. For the main optimizer, the initial learning rate is set at 10 -4 During network training, when the learning rate decays to 10 -6 The best training model of DMENet will be stored. For the auxiliary optimizer, its initial learning rate is set at 10 -3 . The batch size is set to 8 during training. For the fairness of the experiment, the neural network models are trained on NVIDIA GeForce RTX3090, and the traditional encoding and decoding parts are all performed on the same CPU (i9-9900KCPU@3.60GHz). The penalty coefficient λ values commonly used in the present invention are [0.660, 0.508, 0.211, 0.072, 0.033, 0.013, 0.007]. In the proposed rate-distortion optimization strategy, Loss MDSE The weight coefficient ψ is set to 0.0185. N is set to 256 in Cblock1-4, Rblock1-4, Hyper encoder and Hyperdecoder. In the remote sensing scene classification part, the benchmark model used for testing is EMTCAL (Efficient multiscale transformer and cross-level attention learning), the dataset used for training is NWPU-RESISC45, and the training-test ratio is 10%-90%. The images used for compression and the images used for remote sensing scene classification training are not intertwined, and the reconstructed images are only used for testing the classification performance, not for training the classification network.
[0238] Rate-distortion performance curve: This experiment uses PSNR and MS-SSIM to evaluate the rate-distortion performance of the model. Fig. 9 , Fig.10 and Fig.11The PSNR and MS-SSIM rate-distortion performance curves of all model experiments on the San Francisco, NWPU-RESISC45, and UC-Merced datasets are shown respectively. Among the traditional codec-based image compression methods, BPG shows excellent rate-distortion performance, which is better than WebP and JPEG2000 in most cases. This significant advantage is mainly due to the multi-channel encoding technology adopted by BPG, which allows independent encoding of different color channels and achieves fine control of image detail features. Among the image compression methods based on deep learning, Balle et al. (factorized-relu) shows relatively poor rate-distortion performance, mainly because it only uses simple convolutional layers and has limited feature extraction capabilities. Although it can be improved by increasing the number of stacked convolutional layers, this will significantly increase the number of model parameters and prolong the inference time. On the NWPU-RESISC45 and UC-Merced datasets, Shi2024 and Tong2023 methods achieved relatively good compression performance due to their excellent attention mechanism design and reasonable residual convolution modules. However, they performed poorly on the San Francisco dataset, indicating that they are less robust. The other methods performed mediocre in rate-distortion performance, mainly because they lacked a strong attention mechanism and an excellent rate-distortion optimization strategy. The DMENet method proposed in this paper achieved the best rate-distortion performance on the three datasets at the same time. This superior performance not only strongly proves the robustness of the DMENet method, but also strongly proves the effectiveness of MDEI, SDPM, LSM and the newly proposed rate-distortion optimization strategy in the DMENet method.
[0239] Visualization comparison experiment of reconstructed images: In order to further verify the visual effect of DMENet reconstructed images, the present invention conducted a visualization comparison experiment. Fig.12 These are the reconstructed images of the San Francisco dataset by various methods at 0.25bpp, as well as their local enlarged images. Fig.13 This is the reconstruction result of each method on the dataset UC-Merced at 0.28bpp. Fig.12 For example, among the traditional image compression methods based on codec, the zebra crossing of the BPG method retains more texture information compared to JPEG2000 and WebP. However, the zebra crossing in the reconstructed area of JPEG2000 and WebP has lost its clear edges and is blurred. The main reason is that the BPG method has multi-channel encoding technology, which has a stronger ability to reconstruct detail features. The image compression comparison method based on deep learning generally achieves better visualization effects than traditional image compression methods, but is still worse than the DMENet method. Fig.12In the reconstructed images of methods such as Minnen et al., Balle et al., and Balle et al. (factorized-relu), there are generally some artifacts and noise. This causes the local image to be blurred. A careful observation shows that the pixel transitions of the reconstructed images of the three comparison methods are too rough, which leads to color flattening and distortion. Finally, the Tong2023, Shi2024 and DMENet methods with better visualization effects among the comparison methods are compared. The roof of the local enlarged image in the DMENet method retains more texture features and clearer object edges. The enlarged areas of the reconstructed images of Tong2023 and Shi2024 are too smooth and lose some detail features. Therefore, DMENet achieved the best visualization effect on the San Francisco dataset. In addition, Fig.13 In the experiment, DMENet also achieved the best visualization effect on the UC-Merced dataset. This also verifies that the proposed method has strong robustness. From the perspective of visualization, the above experiments fully prove that the proposed DM ENet can reconstruct images with more complete and clearer structural features, and also prove that the new rate-distortion optimization strategy plays a key role in aligning structures.
[0240] Remote sensing image scene classification: This paper uses reconstructed images obtained by different compression methods for remote sensing scene classification to verify the effectiveness of the proposed DMENet method from an application perspective. The selected dataset is NWPU-RESISC45. The remote sensing scene classification benchmark model is EMTCAL. To ensure the fairness of the experiment, the reconstructed images of different methods are obtained at a bit rate of 0.7bpp. Fig.14 The overall accuracy (OA) obtained by reconstructing images obtained by different methods when classifying remote sensing image scenes. Among them, the proposed DMENet achieved the highest OA value and the best classification performance, exceeding 0.74% compared with Minnen taal.; exceeding 0.74% compared with Balle et al. (hyperprior); exceeding 1.31% compared with Balle et al. (factorized-relu); exceeding 0.37% compared with Tong2023 and exceeding 0.37% compared with Shi2024.
[0241] Complexity analysis: The present invention conducts complexity analysis, and the test parameters mainly include Parameter, FLOPs, GPU Memory, Compression time and Inference time. The input image format is 3×256×256. By comparison, it is found that the proposed DMENet has a great advantage in the amount of parameters. By comparing Parameter, it can be found that DMENet has achieved the smallest amount of parameters, which are only 42.40%, 53.48%, 95.32%, 19.23%, and 14.37% of Minnen ta al., Balle et al. (hyperprior), Balle et al. (factorized-relu), Tong2023, Shi2024 and other methods, respectively, which fully demonstrates the superiority of the DMENet method. By comparing FLOPs, it can be found that DMENet still has the least number of parameters, which are only 20.37%, 20.80%, 21.23%, 8.20%, and 4.18% of Minnen ta al., Balle et al. (hyperprior), Balle et al. (factorized-relu), Tong2023, and Shi2024, respectively. Compared with GPU Memory, DMENet is at a medium level. The reason is that some multi-branch structures are used in the network, which leads to more parallel computing, resulting in a relatively large GPU Memory overhead. Comparing Compression time and Inference time, it can be found that DMENet's time consumption is third, but at the same bit rate, DMENet's PSNR and MS-SSIM are much better than other comparison methods. These experiments strongly prove that DMENet can achieve very excellent rate-distortion performance at a lower complexity.
[0242] Table 3 Complexity experiments of different compression methods
[0243]
[0244] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information, characterized in that: The specific process of the method is: Build the DMENet network model; The DMENet network model represents a dynamic feature enhancement network model guided by multi-dimensional collaborative side information; The DMENet network model includes a multi-dimensional feature extraction module MDEI guided by side information, a compression module, a coding and decoding module, and a reconstruction module; The side information guided multi-dimensional feature extraction module MDEI comprises a horizontal attention module HA, a vertical attention module VA and an edge feature extraction module EEM; The compression module includes compression block 1, compression block 2, slice dynamic pyramid module SDPM, compression block 3, and compression block 4 in sequence; The reconstruction module includes reconstruction block 4, reconstruction block 3, slice dynamic pyramid module SDPM, reconstruction block 2, and reconstruction block 1 in sequence; The slicing dynamic pyramid module SDPM includes a slicing strategy SS, an irregular shape feature extraction block SEB, and a pyramid feature enhancement block PEB in sequence; The coding and decoding module includes a probability model, a Q quantizer, an AE arithmetic coding and an AD arithmetic decoding; The probability model ProbabilityModel includes a super encoder, a super decoder, a latent representation space enhancement module LSM, a Q quantizer, an AE arithmetic encoding, an AD arithmetic decoding, and a Factorized Entropy; The working process of the DMENet network model is as follows: The original remote sensing image data is input into the compression module, and the compression module outputs features; The compression module outputs features that are input into the encoding and decoding module, and the encoding and decoding module outputs features; The encoding and decoding module outputs features that are input into the reconstruction module, and the reconstruction module outputs the reconstructed image; The original remote sensing image data and the reconstructed image data are respectively input into the side information guided multidimensional feature extraction module MDEI, which outputs the structural features of the original part and the structural features of the reconstructed part respectively, and calculates the structural difference Loss between the structural features of the original part and the structural features of the reconstructed part MDEI .
2. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 1 is characterized in that: The specific working process of inputting the remote sensing image data into the compression module and outputting the features of the compression module is as follows: The compression module includes compression block 1, compression block 2, slice dynamic pyramid module SDPM, compression block 3, and compression block 4 in sequence; Compression block 1 includes the first convolutional layer and the first GDN in sequence. The convolution kernel size of the first convolutional layer is 7×7, the number of input channels is 3, and the number of output channels is N / 4. Compression block 2 includes the second convolutional layer and the second GDN in sequence. The convolution kernel size of the second convolutional layer is 3×3, the number of input channels is N / 4, and the number of output channels is N / 2. Compression block 3 includes the third convolutional layer and the third GDN in sequence. The convolution kernel size of the third convolutional layer is 3×3, the number of input channels is N / 2, and the number of output channels is 3N / 4. Compression block 4 includes a fourth convolutional layer and a fourth GDN in sequence. The convolution kernel size of the fourth convolutional layer is 3×3, the number of input channels is 3N / 4, and the number of output channels is N. GDN stands for generalized divisive normalization function; N is the number of channels; The original remote sensing image data is sequentially input into compression block 1, compression block 2, slice dynamic pyramid module SDPM, compression block 3, and compression block 4, and the output features of compression block 4 are used as the output features of the compression module.
3. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 2 is characterized in that: The working process of the slice dynamic pyramid module SDPM is as follows: The slicing dynamic pyramid module SDPM includes the slicing strategy SS, the irregular shape feature extraction block SEB and the pyramid feature enhancement block PEB in sequence; 1) Remote sensing image data input slicing strategy SS, slicing strategy SS output feature map Output SS ; The specific process is: Output SS = F concat (Tensor H×W×m (Input H×W×C ),Conv 3×3 (Tensor H×W×(C-m) (Input H×W×C ))) In the formula, Input H×W×C Represents the input feature map of the slicing strategy SS; H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels; Tensor H×W×m (Input H×W×C ) represents the input feature map of the slicing strategy SS H×W×C Feature maps corresponding to the first m channels; Tensor H×W×(C-m) (Input H×W×C ) represents the input feature map of the slicing strategy SS H×W×C Feature maps corresponding to the remaining Cm channels; m represents the number of channels; Conv 3×3 Represents a convolution with a kernel size of 3×3; F concat (,) represents the function that concatenates two tensors in the channel dimension; Output SS Represents the output feature map of SS; 2) Slicing strategy SS output feature map Output SS Input irregular shape feature extraction block SEB, irregular shape feature extraction block SEB output feature map Output SEB ; The specific process is: The irregular shape feature extraction block SEB is a dynamic convolution layer, and the convolution kernel size of the dynamic convolution layer is 3×3; Slicing strategy SS output feature map Output SS After a layer of dynamic convolution layer, the dynamic convolution layer outputs the feature map Output SEB ; 3) Irregular shape feature extraction block SEB output feature map Output SEB Input pyramid feature enhancement block PEB, pyramid feature enhancement block PEB output feature map Output PEB ; The specific process is: Irregular shape feature extraction block SEB output feature map Output SEB Input 3×3 convolution layer, 3×3 convolution layer outputs feature map A; Irregular shape feature extraction block SEB output feature map Output SEB Input 5×5 convolution layer, 5×5 convolution layer outputs feature map B; Irregular shape feature extraction block SEB output feature map Output SEB Input 7×7 convolution layer, 7×7 convolution layer outputs feature map C; Irregular shape feature extraction block SEB output feature map Output SEB Input 9×9 convolution layer, 9×9 convolution layer outputs feature map D; The working process of PEB is expressed as: Output PEB =Conv 1×1 (F concat (A,B,C,D))+Output SEB In the formula, Conv 1×1 represents point convolution; F concat (,) represents the function that concatenates tensors in the channel dimension; Output PEB Represents the output feature map of PEB; The output feature map of PEB is the final output feature map of the slice dynamic pyramid module SDPM.
4. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 3 is characterized in that: The compression module outputs features which are input into the encoding and decoding module, and the encoding and decoding module outputs features; The specific process is: The compression module outputs feature y and inputs it into the Q quantizer, which outputs feature feature Input AE arithmetic coding, and the output features of AE arithmetic coding are input into AD arithmetic decoding; The compression module output feature y is input into the probability model Probability Model, the output of the probability model Probability Model is input into AE arithmetic coding and AD arithmetic decoding respectively, and the AE arithmetic coding output feature is input into AD arithmetic decoding; The AD arithmetic decoding output features are used as the output features of the encoding and decoding module.
5. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 4 is characterized in that: The compression module outputs feature y and inputs it into the probability model. The output of the probability model is respectively input into AE arithmetic coding and AD arithmetic decoding. The specific process is as follows: The probability model includes super encoder, super decoder, latent representation space enhancement module LSM, Q quantizer, AE arithmetic coding, AD arithmetic decoding, and Factorized Entropy. The compression module outputs the feature y which is input into the super encoder, and the super encoder outputs the feature; The output feature of the super encoder is input into the latent representation space enhancement module LSM, and the latent representation space enhancement module LSM outputs the feature z; The output feature z of the latent representation space enhancement module LSM is input into the Q quantizer, and the Q quantizer outputs the feature Q quantizer output characteristics Input AE arithmetic coding, AE arithmetic coding output features Input AD arithmetic decoding, AD arithmetic decoding output features Q quantizer output characteristics Input Factorized Entropy, the features output by Factorized Entropy are input into AE arithmetic coding and AD arithmetic decoding respectively, and AD arithmetic decoding outputs features AD arithmetic decoding output characteristics The latent representation space enhancement module LSM is input, the output features of the latent representation space enhancement module LSM are input into the super decoder, the output features of the super decoder are used as the output of the probability model Probability Model, and the output of the probability model Probability Model is respectively input into AE arithmetic coding and AD arithmetic decoding.
6. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 5 is characterized in that: The compression module outputs the feature y which is input into the super encoder, and the super encoder outputs the feature; The specific process is: The super encoder includes a fifth 3×3 convolution layer, a first RELU activation function layer, a sixth 3×3 convolution layer, a second RELU activation function layer, and a seventh 3×3 convolution layer in sequence; The compression module output feature y is sequentially input into the fifth 3×3 convolution layer, the first RELU activation function layer, the sixth 3×3 convolution layer, the second RELU activation function layer, and the seventh 3×3 convolution layer. The seventh 3×3 convolution layer output feature is used as the super encoder output feature. The latent representation space enhancement module LSM outputs features that are input into the super decoder, and the super decoder outputs features; the specific process is: The super decoder sequentially includes an eighth 3×3 convolution layer, a third RELU activation function layer, a ninth 3×3 convolution layer, a fourth RELU activation function layer, a tenth 3×3 convolution layer, and a fifth RELU activation function layer; The output features of the latent representation space enhancement module LSM are sequentially input into the eighth 3×3 convolution layer, the third RELU activation function layer, the ninth 3×3 convolution layer, the fourth RELU activation function layer, the tenth 3×3 convolution layer, and the fifth RELU activation function layer, and the output features of the fifth RELU activation function layer are used as the super decoder output features.
7. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 6 is characterized in that: The super encoder outputs features that are input into a latent representation space enhancement module LSM, and the latent representation space enhancement module LSM outputs features z; The specific process is: The latent representation space enhancement module LSM includes large kernel band convolution LKE and point convolution in sequence; The working process of LKE is expressed as: Output LKE = E+Conv 1×7 (Conv 7×1 (E))+Conv 1×11 (Conv 11×1 (E))+Conv 1×21 (Conv 21×1 (AND)) In the formula, E represents the feature of the input feature after a convolution with a convolution kernel size of 5×5; Conv 7×1 Represents the input feature after a convolution with a convolution kernel size of 7×1; Conv 1×7 Represents the input feature after a convolution with a convolution kernel size of 1×7; Conv 1×11 Represents the input feature after a convolution with a convolution kernel size of 1×11; Conv 11×1 Represents the input feature after a convolution with a convolution kernel size of 11×1; Conv 1×21 Represents the input feature after a convolution with a convolution kernel size of 1×21; Conv 21×1 Represents the input feature after a convolution with a convolution kernel size of 21×1; Output LKE Represents the output features of LKE; Output features of LKE LKE The final feature map Output is obtained by interacting channel features through a point convolution LSM : Output LSM =Conv 1×1 (Output LKE ) In the formula, Conv 1×1 Represents 1×1 point convolution; Feature map Output LSM That is the output feature z of the latent representation space enhancement module LSM.
8. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 7 is characterized in that: The encoding and decoding module outputs features and inputs them into the reconstruction module. The specific working process of the reconstruction module outputting the reconstructed image is as follows: The reconstruction module includes reconstruction block 4, reconstruction block 3, slice dynamic pyramid module SDPM, reconstruction block 2, and reconstruction block 1 in sequence; The reconstruction block 4 includes a fourteenth convolutional layer and a fourth IGDN in sequence, the convolution kernel size of the fourteenth convolutional layer is 3×3, the number of input channels is N, and the number of output channels is 3N / 4; The reconstruction block 3 includes a thirteenth convolutional layer and a third IGDN in sequence, the convolution kernel size of the thirteenth convolutional layer is 3×3, the number of input channels is 3N / 4, and the number of output channels is N / 2; The reconstruction block 2 includes a twelfth convolutional layer and a second IGDN in sequence, the convolution kernel size of the twelfth convolutional layer is 3×3, the number of input channels is N / 2, and the number of output channels is N / 4; The reconstruction block 1 includes the eleventh convolutional layer and the first IGDN in sequence, the convolution kernel size of the eleventh convolutional layer is 3×3, the number of input channels is N / 4, and the number of output channels is 3; The IGDN represents the inverse operation of the generalized split normalization function GDN; N is the number of channels; The output features of the encoding and decoding module are sequentially input into reconstruction block 4, reconstruction block 3, slice dynamic pyramid module SDPM, reconstruction block 2, and reconstruction block 1. The output image of reconstruction block 1 is the reconstructed image output by the reconstruction module.
9. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 8, characterized in that: The original remote sensing image data and the reconstructed image data are respectively input into the multi-dimensional feature extraction module MDEI guided by the side information, and the multi-dimensional feature extraction module MDEI guided by the side information respectively outputs the structural features of the original part and the structural features of the reconstructed part, and calculates the structural difference Loss between the structural features of the original part and the structural features of the reconstructed part MDEI ; The specific process is: The side information guided multi-dimensional feature extraction module MDEI includes a horizontal attention module HA, a vertical attention module VA and an edge feature extraction module EEM; 1) The original remote sensing image data is redirected through the permute operation to obtain three data blocks X4, X1 and X5; 2) Data block X4 is input into the horizontal attention module HA, and the horizontal attention module HA outputs X6 H×W×C ; expressed as: X6 H×W×C = Reconstruction(BarConv 1×3 (a(Avgpool(X4 H×C×W ))+b(Stdpool(X4 H×C×W )))) Among them, BarConv 1×3 Represents a strip convolution with a convolution kernel shape of 1×3; Avgpool stands for average pooling, and Stdpool stands for global standard deviation pooling; H represents the height of the feature map, C represents the number of channels of the feature map, and W represents the width of the feature map; a and b represent weight coefficients; Reconstruction includes Sigmoid, Expand, and Permute in turn; X6 H×W×C Represents horizontal structural features; Sigmoid represents the Sigmoid activation function, Expand represents dimension expansion, and Permute represents dimension transformation; 3) Data block X5 is input into the vertical attention module VA, and the vertical attention module VA outputs X7 H×W×C ; expressed as: X7 H×W×C = Reconstruction(BarConv 1×3 (c(Avgpool(X5 C×W×H )+d(Stdpool(X5 C×W×H )) Among them, BarConv 1×k Represents a strip convolution with a convolution kernel shape of 1×3; c and d represent weight coefficients; Reconstruction includes Sigmoid, Expand, and Permute in turn; X7 H×W×C Represents the longitudinal structural characteristics; 4) Data block X1 is input into edge feature extraction module EEM, and edge feature extraction module EEM outputs X3 H×W×C ; expressed as: X3 H×W×C = X1 H×W×C -GaussianConv2(Upsample(Downsample(GaussianConv1(X1 H×W×C )))) Among them, GaussianConv1 represents the first Gaussian convolution; Downsample represents interval sampling downsampling; Upsample represents interpolation upsampling, and the insertion value is 0; GaussianConv2 represents the second Gaussian convolution; X3 H×W×C Represents the edge features of the image; 5) Output X6 to the horizontal attention module HA H×W×C , vertical attention module VA output X7 H×W×C , edge feature extraction module EEM output X3 H×W×C Fusion, obtain the structural features of remote sensing images Output MDEI (Origin): Output MDEI(Origin) =e(X6 H×W×C )+f(X3 H×W×C )+g(X7 H×W×C ) In the formula, e, f, and g are the weight coefficients of the corresponding branches; 6) Similarly, obtain the structural features of the reconstructed image Output MDEI(Reconstruction) ; 7) Through Output MDEI(Origin) and Output MDEI(Reconstruction) Calculate the loss MDSE : Loss MDSE =L MSE (Output MDEI(Origin) ,Output MDEI(Reconstruction) ) Where, L MSE represents the loss measured using MSE.
10. The remote sensing image compression method based on a dynamic feature enhancement network guided by multi-dimensional collaborative side information according to claim 9, characterized in that: The DMENet network model is used to compress the original remote sensing image to be tested; the specific process is: Based on Loss MDEI Calculate the DMENet network model loss function, expressed as: argminProposedLoss Total =R+λ(D+ψLoss MDEI ) Where ψ represents Loss MDEI The weight coefficient; R represents the entropy rate; Loss Total represents the loss function of the DMENet network model; λ represents the penalty coefficient; D represents the degree of distortion between the original image and the reconstructed image; The DMENet network model is trained based on the training set until the loss function converges to obtain a trained DMENet network model; The training set includes original remote sensing images and reconstructed images; The original remote sensing image to be tested is input into the compression module in the trained DMENet network model, and the compression module outputs the compressed data.
Citation Information
Cited By
Hyperspectral remote sensing image compression method based on attention and quantization coding optimization
CN120499378A