Hyperspectral remote sensing image compression method based on attention and quantization coding optimization

By introducing a space-spectral attention mechanism and a two-stage quantitative encoding strategy, the reconstruction quality and compression efficiency of hyperspectral remote sensing images under high compression ratio are improved, and the problems of poor reconstruction quality and high computational complexity in the existing technology are solved, and are suitable for small satellite platforms.

CN120499378AActive Publication Date: 2025-08-15WUHAN UNIV

Patent Information

Application Number
CN202510604873.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-15
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The prior art has poor reconstruction quality and high-spectral remote sensing images under high compression ratio conditions, which is difficult to meet the needs of satellite platforms with limited resources.

Method used

The method based on attention and quantized coding optimization is adopted to improve the model's perception of global information through the spatial-spectral attention mechanism, and combined with the two-stage quantitative coding strategy, it reduces storage and transmission costs.

Benefits of technology

Under high compression ratio conditions, the reconstruction performance of hyperspectral remote sensing images is significantly improved, and the computational complexity is reduced. It is suitable for small satellite platforms with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499378A_ABST
    Figure CN120499378A_ABST
Patent Text Reader

Abstract

The invention provides a hyperspectral remote sensing image compression method based on attention and quantization coding optimization, and the method comprises the steps: carrying out a network model training process: carrying out the processing, cutting and enhancement of hyperspectral remote sensing image data, and constructing a sample set for training; extracting low-dimensional feature representation of the sample data by using a lightweight encoder, wherein the encoder integrates a convolutional layer and a spectrum multi-head self-attention module; a decoder fusing a space-spectrum attention mechanism is adopted to gradually reconstruct a hyperspectral remote sensing image from low-dimensional features; optimizing coding and decoding model parameters through a combined loss function; a quantization coding two-stage compression process: performing adaptive quantization on the features, and mapping the floating point type features into discrete integers based on a logarithm mapping strategy; performing two-stage coding compression on the quantized features; and recovering feature representation through decoding and inverse quantization, and inputting a trained decoder network to reconstruct a hyperspectral remote sensing image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral remote sensing image compression, and in particular to a technical solution for realizing hyperspectral remote sensing image compression and high-quality reconstruction under high compression ratio conditions. Background Art

[0002] With significant improvements in the spectral, spatial, and temporal resolution of hyperspectral remote sensing satellites, the accuracy and real-time performance of hyperspectral remote sensing image data have been significantly improved. The daily data generated by in-orbit hyperspectral satellites worldwide has exceeded terabytes, posing significant challenges to the storage and transmission of hyperspectral data. Furthermore, efficient hyperspectral data compression and reconstruction methods are essential for building integrated real-time communication, navigation, and telecomputing service systems. Therefore, there is an urgent need to develop efficient and intelligent hyperspectral compression and reconstruction technologies to improve the acquisition and processing of massive amounts of hyperspectral image data.

[0003] Traditional frequency-domain methods, primarily based on the Nyquist-Shannon sampling theory, focus on bit-level lossy compression, typically employing a step-by-step "sampling first, then compression" approach. This approach not only increases overall compression time but also struggles to effectively compensate for degradation in reconstruction performance at high compression ratios. Compressed sensing algorithms, by contrast, can simultaneously perform data compression while sampling, theoretically overcoming the limitations of traditional sampling theorems and enabling superior image reconstruction even at low sampling rates. However, high computational complexity and long reconstruction times limit their real-time application on small satellites.

[0004] In recent years, breakthroughs in deep learning technology have sparked interest in its application in high-dimensional tensor data compression. Deep learning-based methods effectively capture data features through end-to-end learning mechanisms and, with the help of pre-trained models, enable data compression and reconstruction in a relatively short period of time. However, under extremely high compression ratios, current deep learning-based methods still fail to fully surpass traditional frequency domain methods. For example: CN118524229A discloses a multi-level encoding method and system for hyperspectral images based on channel attention. This method extracts low-rank spectral features through a channel attention embedding module and combines it with a hierarchical variational autoencoder for multi-level encoding. However, a disadvantage is that the decoding end relies on a spectral feature memory unit matrix to perform a linear mapping of the reduced-dimensional features to restore the original number of bands. This design not only introduces complex high-dimensional matrix multiplication operations, but also fails to further compress the extracted low-dimensional features, making it difficult to achieve a higher compression rate.

[0005] CN114422784A discloses a convolutional neural network-based method for compressing multispectral remote sensing images from unmanned aerial vehicles (UAVs). This method extracts image feature information through a convolutional autoencoder and combines it with multi-level quantization and a Gaussian mixture entropy coding module to remove feature redundancy. However, the proposed method has several drawbacks: its encoding and decoding network structure is simple, making it difficult to capture global spatial relationships and inter-band spectral dependencies in images. Furthermore, the use of a Gaussian mixture entropy coding module increases the complexity of the entire encoding process, making it difficult to adapt to resource-limited satellite platforms.

[0006] CN110348487A discloses a hyperspectral image compression method and device based on deep learning. This method uses an encoding network to obtain feature results, and then uses a quantization network to obtain the final compressed code stream. However, a disadvantage is that the network still follows the processing method of natural images and splits the hyperspectral image into three-channel image blocks. This makes it impossible to effectively explore the spectral correlation between bands and also destroys the trend of continuous change in the actual spectral response curve, affecting the reconstruction effect of hyperspectral remote sensing imagery.

[0007] CN115511983A discloses a remote sensing image compression algorithm based on a deep attention network and scene perception. This algorithm uses an attention feature extraction module to improve the performance of the codec network and combines scene-aware transfer learning to optimize the compression effect in specific scenarios. However, its shortcomings include: the network's spatial attention and channel attention modules are calculated independently, without collaborative fusion, and cannot effectively mine the complementary information between them; and the algorithm relies on fine-tuning the migration of different scenarios, resulting in high model redundancy and difficulty meeting the deployment requirements of spaceborne platforms.

[0008] It can be seen that the existing technology has not yet solved the pain points of poor reconstruction quality and high computational complexity under high compression ratio.

[0009] This paper proposes a hyperspectral remote sensing image compression method based on attention and quantization coding optimization, which achieves significant improvements in reconstruction performance at high compression ratios. Firstly, the attention mechanism employed in deep learning enhances the model's ability to perceive global information, thereby improving image reconstruction accuracy. Secondly, a two-stage quantization coding strategy significantly improves compression efficiency, effectively reducing storage and transmission costs. Summary of the Invention

[0010] Aiming at the shortcomings of the above-mentioned existing technologies in the compression and reconstruction of hyperspectral remote sensing images under high compression ratio conditions, the present invention proposes a hyperspectral remote sensing image compression method based on attention and quantization coding optimization.

[0011] The technical solution adopted by the present invention is a hyperspectral remote sensing image compression method based on attention and quantization coding optimization, which performs the following process: The network model training process includes: The hyperspectral remote sensing image data is normalized and clipped to generate strip data according to the push-broom imaging machine, and then enhanced to construct a sample set for training; A lightweight encoder is used to extract low-dimensional feature representations of sample data. The encoder integrates convolutional layers and a spectral multi-head self-attention module to capture long-range dependencies in the spectral dimension. A decoder that integrates spatial-spectral attention mechanism is used to gradually reconstruct hyperspectral remote sensing images from low-dimensional features. The decoder contains cascaded hybrid attention modules that combine spatial attention and spectral attention weighted features. Optimize the encoding and decoding model parameters by combining loss functions; The quantization coding two-stage compression process includes: Input the image to be compressed into the trained encoder network to extract low-dimensional feature representation; Adaptively quantize features and map floating-point features to discrete integers based on a logarithmic mapping strategy; Perform two-stage coding compression on the quantized features; The feature representation is restored by decoding and dequantization, and input into the trained decoder network to reconstruct the hyperspectral remote sensing image.

[0012] Moreover, the spatial-spectral attention mechanism is implemented as follows, Perform global average pooling and maximum pooling on the input features to generate spectral attention weights and spatial attention weights respectively; The fusion generates comprehensive attention weights.

[0013] Moreover, the adaptive quantization is implemented as follows: Normalize and logarithmically map the feature data according to the preset quantization level; Select the storage format based on the quantization level.

[0014] Moreover, a parallel branch module is added after one convolutional layer in the encoder, including Convolution branch, used to extract local spatial feature information through convolution; The Transformer branch sets up a spectral multi-head attention module and a multi-layer perceptron based on the Transformer implementation to model the dependencies between spectral channels.

[0015] Furthermore, the hybrid attention module of the decoder includes: The input features are evenly divided along the channel and input into the convolution branch and the Transformer branch respectively; The convolutional branch weights the output through spatial-spectral attention; The Transformer branch sets up a spectral multi-head self-attention module and multi-scale convolution based on Transformer to extract multi-scale features.

[0016] Moreover, the combined loss function includes an absolute error loss L1 and a spectral angle mapping loss SAM. The absolute error loss L1 is used to constrain the pixel-level reconstruction fineness, and the spectral angle mapping loss SAM is used to constrain the directional consistency of the reconstructed data in the spectral space.

[0017] Moreover, when performing two-stage coding compression on the quantized features, the integer features obtained by adaptive quantization are first converted into byte streams, and then the Brotli lossless compression algorithm is used for two-stage compression.

[0018] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the hyperspectral remote sensing image compression method based on attention and quantization coding optimization as described above is implemented.

[0019] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the hyperspectral remote sensing image compression method based on attention and quantization coding optimization as described above.

[0020] On the other hand, the present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the hyperspectral remote sensing image compression method based on attention and quantization coding optimization as described above.

[0021] This paper constructs a hyperspectral remote sensing image compression and reconstruction framework based on attention and quantization coding optimization, effectively improving the image reconstruction quality under high compression ratio conditions. By introducing a spatial-spectral attention mechanism, this solution overcomes the limitations of previous deep convolutional networks in capturing the global dependencies of hyperspectral data. It adopts a combined strategy of absolute error and spectral angle mapping loss to further ensure the consistency of the reconstructed image in the spectral space. At the same time, it adopts a two-stage quantization coding strategy to fully exploit the redundancy in the feature representation, significantly improving the compression efficiency. In addition, the encoding module of this method adopts a lightweight structural design, which is suitable for small satellite platforms with limited resources and has good practical application value.

[0022] Compared with the prior art CN118524229A, the advantages of the hyperspectral remote sensing image compression method based on attention and quantization coding optimization proposed in the present invention are: First, the spatial-spectral attention mechanism is used to effectively mine the overall dependencies between the space and bands of hyperspectral images, thereby enhancing the feature expression capability rather than single channel attention. At the same time, the quantization coding module effectively releases feature redundancy, and relying on the adaptive quantization strategy, the quantization accuracy can be flexibly adjusted according to actual needs to avoid repeated training; During the network training stage, spectral angle mapping loss is incorporated to effectively improve the spectral semantic consistency of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flow chart of a method for compressing and reconstructing hyperspectral remote sensing images according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall encoding and decoding network structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the encoder network structure according to an embodiment of the present invention; Figure 4 Schematic diagram of the network structure of a single hybrid attention module in the decoder of an embodiment of the present invention; Figure 5 Schematic diagram of the structure of the spatial-spectral attention mechanism according to an embodiment of the present invention; Figure 6 This is a flowchart of the two-stage compression part of quantization encoding according to an embodiment of the present invention; Figure 7 Schematic diagram of rate-distortion curve comparison between the embodiment of the present invention and four baseline algorithms on a test set; Figure 8 Schematic diagram of the reconstructed image effects of the embodiment of the present invention and four baseline algorithms under similar compression bit rate conditions; DETAILED DESCRIPTION In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0024] like Figure 1 As shown, the embodiment of the present invention provides a hyperspectral remote sensing image compression method based on attention and quantization coding optimization, which includes two parts: The network model training part includes: preprocessing and enhancing the training data; extracting feature representations using the encoder; reconstructing data using the decoder; iteratively training and optimizing the encoding and decoding model; and saving the hyperspectral compression reconstruction model when the evaluation accuracy reaches a predetermined threshold.

[0025] The two-stage compression part of quantization coding includes: inputting the hyperspectral remote sensing image to be compressed and performing block processing; using the trained encoder to extract feature representation; adaptively quantizing the features according to the preset quantization level, and further compressing the discretized feature results using the Brotli lossless compression method to obtain the final compressed code stream; sequentially using Brotli decoding and inverse quantization operations to restore the original feature representation; and using the trained decoder to reconstruct the hyperspectral remote sensing image data.

[0026] The specific steps of the network model training part of the embodiment are as follows: Step 1.1: Training data preprocessing and enhancement: This involves normalizing the hyperspectral remote sensing image data and cropping it to generate strip data based on the push-broom imaging machine. Random vertical and horizontal flipping operations are performed on the cropped sample data to construct a sample set for training. In the embodiment, the collected hyperspectral data is automatically normalized and clipped to generate strip data according to the push-broom imaging machine. At the same time, random vertical and horizontal flip operations are performed on the clipped sample data to enhance the data diversity of the training set. In the embodiment, the main experimental data comes from the AVIRIS hyperspectral dataset, from which multiple sizes of The sample data is obtained by removing low-quality bands from the original 224-band data. In the embodiment, an experimental data set containing 3016 samples is constructed, and the size of the cropped strip data is set to , so that the data meets the encoder network module input requirements.

[0027] Step 1.2: Feature extraction: Use a lightweight encoder model to learn features from the preprocessed sample data and extract compact and representative feature representations. In this embodiment, the encoder network is primarily a three-layer convolutional neural network, whose structural units include, but are not limited to, convolution operations, batch normalization operations, and nonlinear mapping. A spectral multi-head self-attention module is integrated after the second convolutional network module, primarily used to establish long-range dependencies across channels in the spectral dimension. A multi-layer perceptron (MLP) then introduces nonlinear transformations, enabling the encoder model to learn more complex feature representations. This attention module does not alter the feature dimensions.

[0028] Preferably, the spectral multi-head self-attention module can be implemented by referring to the multi-head self-attention mechanism in the existing Transformer technology.

[0029] As a preference, under the condition of 0.5% sampling rate, the convolutional network structure and parameter configuration of the encoder are shown in Table 1. The second to last column "input size" and the last column "output size" in the table represent the feature map size of each layer input and output (excluding batch processing dimension), respectively, in the form of three tuples. Indicates that represents the number of feature channels, and Represents the height and width of the feature map respectively. Similarly, the convolution kernel size, stride and edge padding size in Table 1 are in the form of two tuples. Indicates that Represents the vertical dimension, Indicates the horizontal dimension.

[0030] Table 1

[0031] After the second convolutional module in the encoder body, a parallel branch module is added. Its structure and parameter optimization recommendations are shown in Table 2. Specifically, the features output by the second convolutional layer are first evenly divided into two parts along the channel dimension: one part is input into the convolution branch for extracting local features; the other part is input into the Transformer branch for modeling long-range dependencies in the spectral dimension. In this branch, it is sequentially input into the Transformer spectral multi-head self-attention module and the multi-layer perceptron.

[0032] Table 2

[0033] The encoder network structure provided in the embodiment is as follows Figure 2 As shown on the left, the specific settings are as follows: Under the condition of a sampling rate of 0.5%, the encoder network consists of a total of 4 main layers, of which the 1st, 2nd and 4th layers are the main convolutional layers, and the 3rd layer is a parallel branch module. The structures of each layer are as follows: The first layer: Convolutional layer 1 + LeakyReLU, including a convolutional layer with 128 convolution kernels, a kernel size of 3×2, a stride of 2×2, and an edge padding of 1×0, and a LeakyReLU activation function layer with a negative slope set to 0.01.

[0034] The second layer: Convolutional layer 2 + LeakyReLU, including a convolutional layer with 64 convolution kernels, a kernel size of 3×1, a stride of 1×1, and an edge padding of 1×0, and a LeakyReLU activation function layer with a negative slope set to 0.01.

[0035] The third layer: parallel branch module, which divides the output feature map of the second layer into two parts according to the channel dimension, and inputs them into the convolution branch and the Transformer branch respectively.

[0036] The convolution branch includes a convolution layer with 32 convolution kernels, a kernel size of 1×1, a stride of 1×1, and an edge padding of 0×0, as well as a LeakyReLU activation function layer with a negative slope of 0.01. This branch adopts a residual connection mechanism, that is, the convolution output is added to the original input features.

[0037] In this embodiment, the Transformer branch consists of a Transformer spectral multi-head self-attention module and a multi-layer perceptron (MLP). To keep the encoder lightweight and reduce computation and memory usage, the attention module uses only two heads, each with a dimension of 16. It first projects the input features into query (Q), key (K), and value (V) using a linear projection. It then uses a scaled dot-product attention mechanism to calculate the global correlation between features. Finally, the attention weights are normalized using a softmax function. This is followed by an MLP structure consisting of two fully connected layers: the first linear layer expands the channel dimension from 32 to 128, using a GELU activation function; the second linear layer reduces the dimension to 32.

[0038] The present invention further proposes that both the attention module and the MLP are equipped with residual connections and layer normalization operations to enhance the stability and expression ability of the model. Figure 3 As shown in the figure, the main framework adopts a three-layer convolutional neural network architecture, with a dual-branch parallel processing structure added after the second convolutional layer. The upper layer is the convolution branch, which effectively extracts local spatial feature information through convolutional layers and activation functions. The lower layer is the Transformer branch, which includes a spectral multi-head attention module and an MLP to establish dependencies between spectral channels. Both branches are equipped with residual connections to enhance model training stability and feature expression capabilities.

[0039] The fourth layer is Convolutional Layer 3+LeakyReLU, which includes a convolutional layer with 14 convolution kernels, a kernel size of 3×1, a stride of 2×2, and an edge padding of 1×0, as well as a LeakyReLU activation function layer with a negative slope of 0.01.

[0040] Step 1.3: Hyperspectral remote sensing image reconstruction: The compressed feature representation extracted by the encoder is input into the decoder model that integrates the spatial-spectral attention mechanism. Through layer-by-layer feature enhancement and upsampling operations, the corresponding hyperspectral remote sensing image data is gradually restored.

[0041] In the embodiment, the decoder network can be divided into five main functional modules, namely convolution layer 1, hybrid attention module, convolution layer 2, upsampling 1 and upsampling 2. Its structural units include but are not limited to transposed convolution operation, batch normalization operation, nonlinear mapping, pooling operation and upsampling operation. The first three modules mainly perform feature enhancement, and the last two modules perform spatial-spectral dimension enhancement and reconstruction. The second module is the core component of the decoder network, which consists of 16 identical hybrid attention structures. The network structure of a single hybrid attention module is as follows: Figure 4 As shown, the structural diagram of the spatial-spectral attention mechanism is as follows Figure 5 shown.

[0042] As a preference, under the condition of a sampling rate of 0.5%, the network structure and parameter configuration of the five main functional modules of the decoder are shown in Table 3. The penultimate column "input size" and the last column "output size" in the table represent the feature map size of each layer network input and output (excluding batch processing dimension), in the form of three tuples. Indicates that represents the number of feature channels, and Represents the height and width of the feature map respectively. Similarly, the convolution kernel size, stride and padding size in the table are in the form of two tuples. Indicates that Represents the vertical dimension, Indicates the horizontal dimension.

[0043] Table 3

[0044] In the decoder network, based on the comprehensive consideration of hyperspectral image reconstruction quality and computational overhead, multiple hybrid attention modules can be set. In this embodiment, 16 hybrid attention modules are preferably cascaded as the core part of feature enhancement. The structure of a single hybrid attention module is similar to the parallel branch module of the encoder described in step 1.2. Its network structure is as follows: Figure 4 As shown: First, the input features are evenly divided along the channel dimension to obtain two One of these sub-features is fed into the convolution branch, where a spatial-spectral attention mechanism calculates a hybrid attention weight, which is then weighted on the convolution output before a residual connection is made with the input. The other sub-feature is fed into the Transformer branch, which consists of a spectral multi-head self-attention module and a multi-scale feature extraction module. This is similar to the Transformer branch in the encoder and also utilizes the Transformer-based spectral multi-head self-attention module. The network structure and parameters of a single hybrid attention module are shown in Table 4.

[0045] Table 4

[0046] In the convolutional branch of the hybrid attention module, features are enhanced by fusing the spatial-spectral attention mechanism. Figure 5 , where the spatial-spectral attention mechanism is defined as follows, In the spectral attention module (SPE), global average pooling and global maximum pooling are first applied to the input features along the spatial dimension to aggregate the features of each channel to generate a global description vector. These two vectors are then passed through a multi-layer perceptron (MLP) and feature fusion is performed using convolution operations to generate the final spectral attention weights.

[0047] Assume the input feature is ,in is the batch size, is the number of channels, and are the height and width of the feature space respectively, then the calculation process of spectral attention is as follows:

[0048] in, is the average pooling operation, is the maximum pooling operation, For multi-layer perceptron operation, For the connection operation, represents the sigmoid activation function, is the convolution operation, It represents the feature representation obtained by inputting the multi-layer perceptron after global average pooling. It represents the feature representation obtained by inputting the multi-layer perceptron after the global maximum pooling. is the concatenation of two eigenvectors, is the obtained spectral attention weight.

[0049] The spatial attention module (SPA) has a similar structure to the spectral attention module (SPE), but the main difference lies in the different dimensional directions of feature extraction. SPA performs average pooling and maximum pooling on the input features along the channel dimension to extract saliency information of spatial position. The calculation process of spatial attention weight is as follows:

[0050] in, is the average pooling operation in the channel dimension, is the maximum pooling operation in the channel dimension, For the connection operation, represents the sigmoid activation function, is the convolution operation, represents the average value in the channel dimension, Indicates the maximum value in the channel dimension, is the concatenation of two eigenvectors, is the obtained spatial attention weight.

[0051] The comprehensive attention weight calculated by the final spatial-spectral attention mechanism (SSFA) can be expressed as:

[0052] in, Represents the attention coefficient of each channel and each spatial position.

[0053] The specific settings of the decoder network are as follows: Under the condition of a sampling rate of 0.5%, the decoder network consists of a total of 5 main layers. The first to third layers are feature enhancement modules, and the fourth and fifth layers are spatial-spectral enhancement modules, which are mainly used to gradually restore the spatial and spectral resolution of hyperspectral remote sensing images. The structure of each layer is as follows: The first layer: Convolutional layer 1, which includes a convolutional layer with 64 convolution kernels, a kernel size of 3×3, a stride of 1×1, and an edge padding of 1×1. The second layer: Hybrid Attention × 16, which consists of 16 hybrid attention units with the same structure connected in series. Each module first divides the input features into two parts along the channel dimension, inputting them into the convolution branch and the Transformer branch respectively. Finally, they are concatenated and fused along the channel dimension. The fused output feature size remains the same as the input.

[0054] The convolution branch contains a convolution layer with 32 convolution kernels, a kernel size of 1×1, a stride of 1×1, and an edge padding of 0, as well as a LeakyReLU activation function layer with a negative slope set to 0.01. The hybrid attention weight is then obtained through the spatial-spectral attention mechanism, and the weight is multiplied element-wise by the convolution activation output. Finally, the weighted enhanced output is added to the original input of the convolution branch through a residual connection.

[0055] Transformer branch structure reference Figure 4, including a spectral multi-head self-attention module consistent with the encoder side, with 2 heads and a single-head dimension of 16; it also includes a multi-scale feature extraction module, which first uses a convolution layer with 128 convolution kernels, a kernel size of 1×1, a stride of 1, and a padding of 0 to expand the channel dimension from 32 to 128, and then divides the channels into four groups (32 channels each), and undergoes depthwise separable convolutions of 1×1 (padding of 0), 3×3 (padding of 1), 5×5 (padding of 2), and 7×7 (padding of 3). After the multi-scale output structure is spliced into 128 channels in the channel dimension, a convolution layer with 32 convolution kernels, a kernel size of 1×1, a stride of 1, and a padding of 0 is used to restore the number of channels to 32.

[0056] The third layer: Convolutional layer 2, including a convolutional layer with 64 convolution kernels, a kernel size of 3×3, a stride of 1×1, and an edge padding of 1×1.

[0057] The fourth layer: upsampling 1, first use pixel reordering (PixelShuffle, scale=2) to reorder the 64-channel feature map output by the previous layer into 16 channels and double the spatial resolution (64×32×1→16×64×2); secondly, through a convolution layer with 64 convolution kernels, kernel size of 1×1, stride of 1, and edge padding of 0, and a LeakyReLU activation function layer with a negative slope of 0.2, the number of channels is expanded to 64; then, through a convolution layer with 64 convolution kernels, kernel size of 3×3, stride of 1, and edge padding of 1, and a LeakyReLU activation function layer with a negative slope of 0.2, the number of channels is further expanded to 128.

[0058] The fifth layer: upsampling 2, first, pixel reordering (PixelShuffle, scale=2) is used to reorder the 128-channel feature map output by the previous layer into 32 channels and double the spatial resolution (128×64×2→32×128×4); secondly, a convolution layer with 96 convolution kernels, a kernel size of 3×3, a stride of 1, and an edge padding of 1 is used, and a LeakyReLU activation function layer with a negative slope of 0.2 is used to expand the number of channels to 96; then, a convolution layer with 172 convolution kernels, a kernel size of 3×3, a stride of 1, and an edge padding of 1 is used, and a LeakyReLU activation function layer with a negative slope of 0.2 is used to restore the feature channel dimension to 172 bands of the original hyperspectral remote sensing image.

[0059] Step 1.4: Codec model training and parameter optimization: Compare the difference between the original hyperspectral remote sensing image and the reconstructed data through the loss function, adjust the model parameters according to the error feedback, and save the model when the model evaluation accuracy reaches the preset standard or the training reaches the predetermined number of iterations. Otherwise, continue to the next iterative training.

[0060] This paper employs a combined strategy of absolute error (L1) and spectral angle mapping (SAM) loss to impose dual constraints on the network training process. The L1 loss constraint ensures pixel-level reconstruction precision, while the SAM loss ensures directional consistency of the reconstructed data in spectral space. This combined loss function strategy effectively improves the quality and fidelity of hyperspectral remote sensing image reconstruction.

[0061] The combined loss function in the embodiment is set as follows: Assume that the original hyperspectral remote sensing image The spectral vector of a pixel is , reconstruct the image The spectral vector of a pixel is , is the total number of pixels in the image, then the combined loss function Defined as:

[0062] in, express norm, represents the inner product, express norm, To avoid decimals that divide by zero; Represents the absolute error loss between the original image and the reconstructed image; Represents the spectral angle mapping loss between the original image and the reconstructed image; is a weight factor used to balance the contribution of the two losses to model training, and is preferably set to 0.05 in this embodiment.

[0063] The two-stage quantization coding compression process is as follows: Figure 6 As shown, its specific implementation includes the following sub-steps: Step 2.1: Input the hyperspectral remote sensing image data to be compressed into the trained encoder network to extract the corresponding low-dimensional feature representation; In this embodiment, the hyperspectral remote sensing image data to be compressed is first processed into blocks, where the entire image is cropped into fixed-size strips according to the encoder network input requirements. These cropped image blocks are then sequentially fed into the trained encoder network to extract the corresponding low-dimensional feature representations.

[0064] Step 2.2: Adaptive quantization: Adaptively quantize the extracted feature representation according to the preset quantization level, mapping the floating-point features into discrete integer representations; In this embodiment, the floating-point feature representation output by the encoder is adaptively quantized according to a preset quantization level and mapped to a corresponding discrete integer form. To further optimize storage efficiency, differentiated storage strategies are implemented for different quantization levels: when the quantization level is less than or equal to 256, storage is performed in uint8 format; when the quantization level is greater than 256, storage is performed in int16 format. This strategy not only effectively saves storage space but is also suitable for fast compression scenarios that do not require subsequent lossless encoding.

[0065] In this invention, adaptive quantization employs a logarithmic mapping strategy, effectively reducing the loss of low-dynamic-range feature information during the quantization process. The core of this strategy is to exploit the nonlinear nature of the logarithmic function: within a low-value (absolute value) range, smaller changes are amplified, thereby making the details of the low-value areas more apparent. The adaptive quantization process in the embodiment is as follows: Assume that the input feature data is , and its maximum and minimum values are recorded as and , then the quantization operation is defined as follows:

[0066] in, is the result after normalization to the interval [0,1]; is the result of logarithmic mapping, parameter Controls the degree of nonlinearity of the logarithmic mapping, is a sign function; is the discretized result, is the set quantization level, is a rounding function. In an embodiment, the parameter Set to 0.1, quantization level Set to 256.

[0067] Step 2.3: Perform two-stage compression on the quantized features: First, convert the integer features obtained by adaptive quantization in step 2.2 into a byte stream, and then use the Brotli lossless compression algorithm to perform two-stage compression to obtain the final compression result.

[0068] The Brotli lossless compression algorithm can be specifically implemented using existing technologies, which will not be described in detail in the present invention. In the embodiment, to achieve an optimal compression ratio, the quality parameter of the Brotli lossless compression algorithm is preferably set to the highest level (quality=11).

[0069] Step 2.4: Decoding and dequantization: Perform the inverse operations of steps 2.3 and 2.2 respectively, that is, first decompress through the Brotli algorithm, and then restore the feature representation through the dequantization operation; In the embodiment, Brotli decompression is first performed on the binary compressed data generated in step 2.3, and then the decompressed data is inversely quantized based on the metadata information during the previous quantization (including maximum and minimum values, quantization mapping parameters, quantization levels, and original feature data size) to restore the original continuous feature representation.

[0070] Step 2.5: Reconstruct the hyperspectral remote sensing image: Input the recovered feature representation into the trained decoder network to obtain the reconstructed hyperspectral remote sensing image data.

[0071] In the embodiment, the restored feature representation is input into the trained decoder network, and the reconstructed data is sequentially spliced according to the index order of the cropping and blocking operation in step 2.1 to output complete hyperspectral remote sensing image data.

[0072] During specific implementation, the above process can be automatically executed using computer software technology.

[0073] Based on the above process, this paper addresses the problem of poor reconstruction performance of existing deep learning-based technologies under high compression ratios by proposing a hyperspectral remote sensing image compression method based on attention and quantization coding optimization. This method utilizes a two-stage quantization coding strategy to further remove redundancy from the low-dimensional feature representation extracted by the encoder. At the same time, the attention mechanism on the decoder side effectively improves reconstruction accuracy, achieving good reconstruction performance even under high compression ratios. It has the following key points: 1) An attention mechanism is introduced into the end-to-end encoder-decoder network model, especially the spatial-spectral attention mechanism in the decoder, to perform weighted enhancement of the reconstructed features; and a combined loss function combining absolute error and spectral angle mapping is designed to balance pixel-level reconstruction accuracy and spectral consistency, thereby improving the hyperspectral reconstruction effect.

[0074] 2) A two-stage compression method using quantization coding is proposed to effectively improve the compression efficiency of hyperspectral remote sensing images. An adaptive quantization method using a logarithmic mapping strategy is also designed. This two-stage compression strategy achieves a compression performance improvement of over 5x while maintaining reconstruction quality.

[0075] In order to verify the effectiveness of the proposed method, four representative hyperspectral remote sensing image compression benchmark algorithms were selected as comparison methods and tested on the hyperspectral dataset constructed in this experiment. The four compared hyperspectral remote sensing image compression algorithms are: Method 1: The method proposed by Du et al., referenced in “Du Q, Fowler J E. Hyperspectral image compression using JPEG2000 and principal component analysis[J]. IEEE Geoscience and Remote sensing letters, 2007, 4(2): 201-205.” Method 2: The method proposed by Hsu et al., reference “Hsu CC, Lin CH, Kao CH, et al. DCSN: Deep compressed sensing network for efficient hyperspectral data transmission of miniaturized satellite[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 59(9): 7773-7789.” Method 3: Zhou et al., “BTC-Net: Efficient bit-level tensor data compression network for hyperspectral image[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023.” Method 4: Zhang et al., “Hyperspectral Image Compression Sensing Network with CNN-Transformer Mixture Architectures[J]. IEEE Geoscience and Remote Sensing Letters, 2024.” Method 1 is a frequency-domain method, replicated in MATLAB according to the literature principles, using the JPEG2000 codec provided by Kakadu. Methods 2, 3, and 4 are learning-based methods, with code provided by the original authors and default parameter settings. The quantization level in the quantization coding module of the present invention is set to 256. To present a fairer comparison, both the learning-based algorithm and the algorithm of the present invention were trained on the same training and validation sets, obtaining compression models at different bit rates ranging from 0.0 to 3.2.

[0076] The comparative experiment contents are as follows: Image compression is performed on the test set using method 1, method 2, method 3, method 4 and the method of the present invention respectively. Figure 7 The rate-distortion curves of the five algorithms are given. Table 5 below gives the comparison results of the average spectral angle mapping (SAM), root mean square error (RMSE) and peak signal-to-noise ratio (PSNR) of the five methods at different bit rates (compression ratios). Since some comparison methods only support discrete or fixed bit rate settings and cannot cover all bit rates in the range of 0.0 to 3.2, this experiment only selects the typical bit rate models available for each algorithm in this range for comparative analysis. The bit rate indicator is bits per pixel per band (bpppb), and the compression ratio (CR) is approximately calculated as the ratio of the original data (int16) to bpppb (usually taken to the nearest multiple of 5). In addition, in order to make a visual comparison, under the conditions of roughly the same bit rate (method 1: CR=75; method 2: CR=50; method 3: CR=70; method 4: CR=50; method of the present invention: CR=70), Figure 8 An original hyperspectral image selected from the test set and the reconstruction results of each method are shown.

[0077] Table 5

[0078] from Figure 7 As can be seen from the rate-distortion curve shown, the method of the present invention can still maintain good reconstruction performance under high compression ratio (i.e., extremely low bit rate), surpassing the frequency domain-based algorithm (method 1); and the method of the present invention is significantly better than the previous deep learning-based algorithms (methods 2, 3, and 4) overall. Figure 8 From the visualization results under high compression ratio conditions, it can be seen that the learning-based algorithms are better than the frequency domain-based algorithms (Method 1). Method 1 has an obvious ringing effect. Among the learning-based algorithms, the blocking effect of the method of the present invention is the mildest, presenting a better visual effect.

[0079] In order to further verify the effectiveness of the quantization coding module proposed in this invention, especially to highlight the flexibility and feasibility of the adaptive quantization method, the following Table 6 shows the changes in the reconstruction performance indicators (SAM, RMSE, PSNR) and the corresponding bit rate (bpppb) at different quantization levels, where the parameters of the logarithmic mapping of the quantization part are All are set to 0.1, and the feature extraction stage continues to use the network model parameter configuration of the previous invention under the condition of a compression ratio of 320.

[0080] Table 6

[0081] As can be seen from the above table, through the second stage of quantization coding compression, the present invention not only significantly reduces the storage volume without the need for repeated model training, but also effectively retains the reconstruction accuracy, and can flexibly adjust the quantization level to achieve different compression rates. The adaptive quantization strategy proposed by the present invention effectively retains the reconstruction accuracy, so that the PSNR index loss at any quantization level in the above table is controlled within 1dB. For example, when the Brotli encoding method is adopted and the quantization level is 1024, the compressed bpppb is only 1 / 4 of the original index, and the PSNR index only loses 0.016dB, and the reconstruction performance is almost unaffected. It can be seen that the present invention significantly improves the compression efficiency through the two-stage compression design, and optimizes the image reconstruction effect by using the spatial-spectral attention mechanism, effectively reducing the generation of artifacts in the image reconstruction process under low bit rate conditions, and providing a practical solution for the efficient storage and transmission of hyperspectral remote sensing image data.

[0082] In specific implementation, the method proposed in the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. System devices that implement the method, such as computer-readable storage media that store the corresponding computer program of the technical solution of the present invention and computer equipment that runs the corresponding computer program, should also be within the scope of protection of the present invention.

[0083] The following embodiment describes an electronic device for establishing a hyperspectral remote sensing image compression method based on attention and quantization coding optimization provided by the present invention. The electronic device for establishing a hyperspectral remote sensing image compression method based on attention and quantization coding optimization described below and the hyperspectral remote sensing image compression method based on attention and quantization coding optimization described above can correspond to each other.

[0084] The electronic device may include a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may invoke logic instructions in the memory to execute a hyperspectral remote sensing image compression method based on attention and quantization coding optimization, which primarily includes the software processing portion of the aforementioned steps.

[0085] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0086] On the other hand, an embodiment of the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the software processing part of the hyperspectral remote sensing image compression method based on attention and quantization coding optimization provided by the above methods.

[0087] On the other hand, an embodiment of the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the software processing part of the hyperspectral remote sensing image compression method based on attention and quantization coding optimization provided by the above-mentioned methods.

[0088] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0089] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0090] It should be understood that parts not elaborated in detail in this specification belong to the prior art.

[0091] It should be understood that the above description of the preferred embodiment is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.

Claims

1. A hyperspectral remote sensing image compression method based on attention and quantization coding optimization, characterized in that: Follow the process below: The network model training process includes: The hyperspectral remote sensing image data is normalized and clipped to generate strip data according to the push-broom imaging machine, and then enhanced to construct a sample set for training; A lightweight encoder is used to extract low-dimensional feature representations of sample data. The encoder integrates convolutional layers and a spectral multi-head self-attention module to capture long-range dependencies in the spectral dimension. A decoder that integrates spatial-spectral attention mechanism is used to gradually reconstruct hyperspectral remote sensing images from low-dimensional features. The decoder contains cascaded hybrid attention modules that combine spatial attention and spectral attention weighted features. Optimize the encoding and decoding model parameters by combining loss functions; The quantization coding two-stage compression process includes: Input the image to be compressed into the trained encoder network to extract low-dimensional feature representation; Adaptively quantize features and map floating-point features to discrete integers based on a logarithmic mapping strategy; Perform two-stage coding compression on the quantized features; The feature representation is restored by decoding and dequantization, and input into the trained decoder network to reconstruct the hyperspectral remote sensing image.

2. The method according to claim 1, wherein: The spatial-spectral attention mechanism is implemented as follows: Perform global average pooling and maximum pooling on the input features to generate spectral attention weights and spatial attention weights respectively; The fusion generates comprehensive attention weights.

3. The method according to claim 1, wherein: The adaptive quantization is implemented as follows: Normalize and logarithmically map the feature data according to the preset quantization level; Select the storage format based on the quantization level.

4. The method according to claim 1, wherein: A parallel branch module is added after a convolutional layer in the encoder, including Convolution branch, used to extract local spatial feature information through convolution; The Transformer branch sets up a spectral multi-head attention module and a multi-layer perceptron based on the Transformer implementation to model the dependencies between spectral channels.

5. The method according to claim 1, wherein: The hybrid attention module of the decoder includes: The input features are evenly divided along the channel and input into the convolution branch and the Transformer branch respectively; The convolutional branch weights the output through spatial-spectral attention; The Transformer branch sets up a spectral multi-head self-attention module and multi-scale convolution based on Transformer to extract multi-scale features.

6. The method according to claim 1, wherein: The combined loss function includes an absolute error loss L1 and a spectral angle mapping loss SAM. The absolute error loss L1 is used to constrain the pixel-level reconstruction fineness, and the spectral angle mapping loss SAM is used to constrain the directional consistency of the reconstructed data in the spectral space.

7. The method according to claim 1, wherein: When performing two-stage coding compression on the quantized features, the integer features obtained by adaptive quantization are first converted into a byte stream, and then the Brotli lossless compression algorithm is used for two-stage compression.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the hyperspectral remote sensing image compression method based on attention and quantization coding optimization as described in any one of claims 1 to 7 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the hyperspectral remote sensing image compression method based on attention and quantization coding optimization as described in any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program, characterized in that: When the computer program is executed by a processor, the hyperspectral remote sensing image compression method based on attention and quantization coding optimization as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Deep learning-based generative remote sensing image compression method

    CN111683250A

  • Hyperspectral image compression method based on spatial and spectral content importance

    CN113706641A

  • BNN-based hyperspectral image classification method, apparatus and device, and memory

    CN119360111A

  • Remote sensing image compression method of dynamic feature enhancement network based on multi-dimensional collaborative side information guidance

    CN119941880A

  • Data compression method for quantitative remote sensing application of unmanned aerial vehicle

    WO2023241188A1

Cited By

  • Optical remote sensing image defogging method based on lightweight parallel attention network

    CN121746248A

  • Hyperspectral image compression network and compression method based on 3D convolution set and causal entropy model

    CN121792734A

  • A hyperspectral data processing method and system

    CN122510746A