Remote sensing image compression method and device based on frequency domain enhancement and adaptive optimization
Through the methods of frequency domain enhancement and adaptive optimization, the fidelity problem of high-frequency details and texture edges in remote sensing image compression is solved, and efficient compression and reconstruction of remote sensing images at high compression ratios are achieved.
Patent Information
- Application Number
- CN202511215563.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing remote sensing image compression methods are not effective in ensuring image reconstruction quality, especially at high compression ratios, where it is difficult to preserve high-frequency details and texture edges.
A method based on frequency domain enhancement and adaptive optimization is adopted. Image features are extracted and reconstructed through the frequency domain enhanced selective state space modules of the encoder and decoder. The entropy model is combined for probability modeling and bitstream generation. A frequency domain-aware adaptive loss weighting mechanism is introduced to optimize the training process.
It significantly improves the compression fidelity of remote sensing images in complex texture areas and detailed information areas, and can effectively compress remote sensing images while ensuring the quality of image reconstruction.
Smart Images

Figure CN120692401A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a remote sensing image compression method and device based on frequency domain enhancement and adaptive optimization. Background Art
[0002] Remote sensing imaging, by observing objects in multiple narrow bands, enables precise identification of material composition, surface structure, and even underlying features. Compared to traditional RGB images, remote sensing images contain richer and more continuous spectral information, and are widely used in remote sensing exploration, agricultural monitoring, geological exploration, urban management, and biomedical imaging. However, the enormous volume of data generated by remote sensing images also poses significant challenges to storage, transmission, and processing efficiency.
[0003] How to effectively compress remote sensing images while ensuring the quality of image reconstruction has become an important issue of widespread concern in academia and industry. Current research on remote sensing image compression can be divided into two categories: traditional compression methods and deep learning-based compression methods. Traditional compression methods mainly rely on manually designed feature extractors and combine prediction or transformation strategies to compress image redundant information. For example, prediction-based coding methods construct correlation models between pixels to infer the encoding of spectral dimensions. Representative examples include the previously released hyperspectral compression standard CCSDS 123.0-B-2 ( Lossless Multispectral&Hyperspectral Image Compression, Issue 2, November 2019, Consultative Committee for Space Data Systems (CCSDS). ), has the advantages of simple implementation, low algorithmic complexity, and easy hardware deployment. Traditional transform-based compression methods (such as discrete cosine transform and discrete wavelet transform) map the original image to a transform domain to remove spectral and spatial redundancy, typically exemplified by JPEG2000. Although these methods have been widely used in engineering practice, their artificially designed feature extraction mechanisms lack data-driven adaptability, making it difficult to maintain spectral fidelity in the reconstructed image at high compression ratios, resulting in limited compression performance. With the advancement of deep learning, image compression methods based on convolutional neural networks have matured and are gradually being applied to remote sensing image compression tasks. A typical deep learning compression framework employs an encoder-entropy model-decoder structure, where the encoder extracts high-dimensional image features, the entropy model probabilistically models the features and performs entropy coding, and the decoder is used for image reconstruction. These methods can be trained end-to-end and possess strong representation capabilities for image feature extraction and distribution learning, making them a mainstream research direction in remote sensing image compression in recent years. However, most existing models use two-dimensional convolutional neural networks to process spatial information, ignoring the frequency domain structure of the image, especially the insufficient modeling capabilities of high-frequency details, texture edges and other areas, resulting in blurry or loss of details in the reconstructed image; more importantly, most models treat context modeling loss (such as causal context adjustment loss) as a static term, and are unable to dynamically adjust its optimization degree according to the importance of the image content, limiting the compression model's performance on different image structures.
[0004] Therefore, a new technical solution is urgently needed to solve the technical problem of how to effectively compress remote sensing images while ensuring the quality of image reconstruction. Summary of the Invention
[0005] The present invention provides a remote sensing image compression method and device based on frequency domain enhancement and adaptive optimization, which are used to solve the technical problem of how to effectively compress remote sensing images while ensuring the quality of image reconstruction.
[0006] To achieve the above objectives, the present invention provides a remote sensing image compression method based on frequency domain enhancement and adaptive optimization, which compresses and decompresses remote sensing images by using a pre-trained first model deployed simultaneously at the compression end and the decompression end, including: The first model includes an encoder, an entropy model and a decoder; both the encoder and the decoder perform frequency domain analysis on the input through a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on the state space mechanism.
[0007] The first model is trained according to a preset training set and a loss function that considers pixel distortion, structure preservation, frequency domain error, perceptual quality, bit rate constraint, contextual autoregressive regularization term and auxiliary supervision loss to obtain a pre-trained first model.
[0008] Compression includes: extracting the main features of the input image through the encoder; probabilistically modeling the distribution information of the main features through the entropy model, generating a bit stream, and sending the bit stream to the decompression end.
[0009] Decompression includes: the decompression end receives the bit stream; restores the bit stream to the main features through the entropy model; inputs the main features into the decoder to obtain the reshaped image.
[0010] Preferably, both the encoder and the decoder perform frequency domain analysis on the input through a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on the state space mechanism, including: The encoder includes three groups of first modules connected in sequence; the first module includes a 5×5 convolution, 3 bottleneck residual modules and 4 frequency domain enhanced selective state space modules connected in sequence.
[0011] The decoder includes three groups of second modules connected in sequence; the second modules include 5×5 convolution, 4 frequency domain enhanced selective state space modules and 3 bottleneck residual modules connected in sequence.
[0012] In the encoder, a frequency-domain enhanced selective state-space module is used to enhance the high-frequency details and global context expressiveness of input features through frequency-domain analysis and long-range dependency modeling.
[0013] In the decoder, a frequency-domain enhanced selective state-space module is used to enhance high-frequency details and spatial consistency in decoded features via frequency-domain attention and long-range dependency modeling.
[0014] Preferably, the frequency domain enhanced selective state space module includes: When the input feature After inputting the frequency domain enhanced selective state space module, the initial features are obtained through 1×1 convolution ; Initial features Input frequency domain analysis submodule to obtain frequency domain importance weight ; Initial features and frequency domain importance weights Input the selective state space submodule to update the state and obtain the context enhanced features that integrate frequency domain information and state perception information ; Enhance features based on context , frequency domain importance weight and initial features Get features ;feature Input nonlinear activation function free block to extract features to obtain features ;feature Output features after 1×1 convolution .
[0015] Preferably, the frequency domain analysis submodule includes: Initial features After inputting the frequency domain analysis submodule, the high frequency, medium frequency, low frequency and DC components are extracted respectively through four parallel frequency domain channels. The high frequency component is extracted by a 3×3 group convolution, and the medium frequency, low frequency and DC components are extracted by a 1×1 convolution respectively to obtain the high frequency component. , intermediate frequency components , low-frequency components and DC component ; The high frequency components , intermediate frequency components , low-frequency components and DC component The fused frequency domain features are obtained by channel splicing ; Fusion frequency domain features Enter the frequency domain importance evaluation network, and pass through average pooling, 1×1 convolution, ReLU activation function layer, 1×1 convolution and Sigmoid function layer in turn to obtain the frequency domain importance weight consistent with the original channel .
[0016] Preferably, the selective state space submodule includes: Initial features and frequency domain importance weights After inputting the selective state space submodule, the frequency domain importance weight For initial features Perform channel-by-channel weighting to obtain the frequency domain weighted feature map ; The frequency domain weighted feature map Get the global feature vector by average pooling ; Global eigenvector Input three layers of perceptron to get the selectivity coefficient used to control the state update process and ; The three-layer perceptron consists of the first linear transformation layer connected in sequence , ReLU activation function layer and second linear transformation layer ; Global eigenvector Adjust the channel dimension through linear mapping to obtain the state space representation ; According to the selectivity coefficient and Combined with lightweight gating mechanism to represent state space Perform selective update, selectivity coefficient Used to control the retention ratio of the original state, selectivity coefficient The nonlinear perturbation amplitude used to adjust the state, the selectivity coefficient It is used to control the degree of injection of global mean information into the state. The three work together to achieve dynamic state adjustment and fusion under frequency domain guidance, and obtain the state representation of adaptive update ; The status is represented By linearly mapping the original channel dimension, we can obtain the context-enhanced features that combine frequency domain information and state perception information. .
[0017] Preferably, extracting the main features of the input image by an encoder; performing probability modeling on the distribution information of the main features by an entropy model, and generating a bitstream includes: The entropy model includes a super-prior encoder, a super-prior decoder, a context modeling module, a one-stage modeling module and a two-stage modeling module; the encoder extracts the main features of the input image; the super-prior encoder models the spatial and channel statistical information of the main features to obtain a first super-prior feature; the first super-prior feature is quantized to obtain a second super-prior feature; and the super-prior decoder obtains a first probability modeling parameter based on the second super-prior feature.
[0018] After the main feature is quantified, it is divided into two sub-features based on the non-uniform channel grouping of a preset ratio, and the first probability modeling parameter is divided into two sub-parameters in the same way; the sub-features and sub-parameters with the same channel size are paired and randomly assigned as the first data and the second data.
[0019] The first-stage modeling module performs modeling based on the sub-parameters in the first data to obtain a first probability distribution; the context modeling module extracts context information in the spatial and channel dimensions based on the sub-features in the first data to obtain a second probability modeling parameter; the second-stage modeling module performs modeling based on the second probability modeling parameters and the sub-parameters in the second data to obtain a second probability distribution.
[0020] Entropy encoding is performed on sub-features in the first data according to a first probability distribution to generate a first bit stream.
[0021] Entropy encoding is performed on the sub-features in the second data according to the second probability distribution to generate a second bit stream.
[0022] The second super priori feature is modeled by a Gaussian distribution model to obtain a third probability distribution; and the second super priori feature is entropy encoded according to the third probability distribution to generate a third bit stream.
[0023] Preferably, the bit stream is restored to main features through an entropy model; the main features are input into a decoder to obtain a reconstructed image, which includes: The decompression end performs probability modeling parameters according to the third bit stream to obtain a third probability distribution.
[0024] Entropy decoding is performed on the third bit stream according to the third probability distribution to obtain a second super-prior feature.
[0025] A first probability modeling parameter is obtained based on the second super prior feature through a super prior decoder; the first probability modeling parameter is divided into two sub-parameters based on a non-uniform channel grouping of a preset ratio, and matched with the channel sizes of the first data and the second data to obtain a first sub-parameter and a second sub-parameter, respectively.
[0026] Modeling is performed according to the first sub-parameter by a one-stage modeling module to obtain a first probability distribution.
[0027] Entropy decoding is performed on the first bit stream according to the first probability distribution to obtain sub-features in the first data.
[0028] The context modeling module extracts context information in spatial and channel dimensions based on the sub-features in the first data to obtain second probability modeling parameters.
[0029] Performing modeling based on the second probability modeling parameter and the second sub-parameter by a two-stage modeling module to obtain a second probability distribution; The second bit stream is entropy decoded according to the second probability distribution to obtain sub-features in the second data.
[0030] The sub-features in the first data and the sub-features in the second data are concatenated to obtain the quantized main features.
[0031] The quantized main features are restored through the decoder to obtain the reshaped image.
[0032] Preferably, the loss function includes: ; in, Represents the distortion term, which is used to measure the reconstruction error between the compressed image and the original image; Indicates bit rate loss; Represents the causal context adjustment loss with the addition of a frequency-domain-aware adaptive loss weighting mechanism, which is used to guide the model to adaptively adjust the focus of the loss function according to the frequency-domain complexity of different images; represents the auxiliary supervision loss term.
[0033] Distortion term include: ; in, Represents the square error loss, which is used to optimize pixel-level restoration accuracy. is the total number of pixels; Reshaped image representing the model output; represents the input image; Represents perceptual loss, which enhances the perception of subjective quality by comparing the distance of images in the depth feature space. For the pre-trained network Feature maps extracted by layers; Represents edge loss, which is used to enhance the preservation of image edges and structural contours. Represents the gradient of the image; Indicates frequency domain loss, and evaluates the image restoration performance in high-frequency, medium-frequency, and low-frequency regions through spectrum analysis. , , represents the spectrum amplitude, represents Fourier transform; Represents structural similarity loss, which is used to model the structural information and local consistency of the image. represents the multi-scale structural similarity index; 、 、 and are the learnable weight coefficients for perceptual loss, edge loss, frequency domain loss, and structural similarity loss, respectively.
[0034] Bit rate loss include: ; , represents the second super prior feature in the compression stage The bit rate loss term; during model training, according to the first probability modeling parameter in the compression stage Constructing the second super-prior feature in the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get .
[0035] , represents the sub-features in the first data of the compression stage Bit rate loss term; During model training, according to the sub-parameters in the first data of the compression stage Construct sub-features in the first data of the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get .
[0036] , represents the sub-features in the second data of the compression stage Bit rate loss term; During model training, according to the sub-parameters in the second data of the compression stage The fusion result of the second probability modeling parameter is used to construct the sub-features of the second data in the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get .
[0037] Auxiliary supervision loss items include: During model training, according to the second super prior feature in the compression stage Predict sub-features in the second data of the compression stage The probability distribution of , and calculate the negative log-likelihood of the feature, we get .
[0038] Preferably, the causal context adjustment loss with the frequency-domain-aware adaptive loss weighting mechanism includes: Get the first loss ,include: ; in, represents information entropy; and These are all true distributions.
[0039] , indicating that only given Under the condition of In the probability distribution The expected value of the negative log probability under , that is, the cross entropy under context-free prediction.
[0040] , given and from context modeling Under the condition of In the second probability distribution The expected value of the negative logarithmic probability under , that is, the cross entropy under context prediction.
[0041] After grayscale processing of the input image, perform two-dimensional fast Fourier transform to obtain the image amplitude ; Calculation based on preset high frequency mask , mark the left and right 1 / 4 areas of the image as high-frequency areas, including: ; High-frequency energy , total energy ; The ratio of high-frequency energy to total energy is used as the frequency domain importance index ,include: ; in, Indicates zero micro parameters.
[0042] Frequency domain importance index As a representation of the frequency domain complexity of the current image, a learnable parameter is introduced , and the frequency domain importance index Added as a learnable weight parameter to the first loss In the causal context adjustment loss, the adaptive loss weighting mechanism with frequency domain perception is obtained ,include: .
[0043] The present invention also provides a remote sensing image compression device based on frequency domain enhancement and adaptive optimization, which is used in the method of the present invention. The device includes a first unit and a second unit.
[0044] The first unit is used to train the first model according to a preset training set and a loss function that considers pixel distortion, structure preservation, frequency domain error, perceptual quality, bit rate constraint, contextual autoregressive regularization term and auxiliary supervision loss to obtain a pre-trained first model.
[0045] The second unit is used to compress and decompress remote sensing images through a pre-trained first model deployed simultaneously at the compression and decompression ends; the first model includes an encoder, an entropy model and a decoder; both the encoder and the decoder perform frequency domain analysis on the input through a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on the state space mechanism.
[0046] Compression includes: extracting the main features of the input image through the encoder; probabilistically modeling the distribution information of the main features through the entropy model, generating a bit stream, and sending the bit stream to the decompression end.
[0047] Decompression includes: the decompression end receives the bit stream; restores the bit stream to the main features through the entropy model; inputs the main features into the decoder to obtain the reshaped image.
[0048] The present invention has the following beneficial effects: The present invention's remote sensing image compression method based on frequency domain enhancement and adaptive optimization fully considers the expressive importance of different frequency components during feature modeling by introducing a selective state-space module for frequency domain enhancement. This method achieves explicit perception of image frequency domain characteristics and efficient modeling of long-range dependencies, effectively improving the under-representation of traditional convolutional networks when compressing high-frequency information. Furthermore, by introducing a frequency-domain-aware adaptive loss weighting mechanism, the causal context-adjusted loss weights during training are dynamically adjusted based on the spectral energy distribution of the input image. This enables the model to adaptively optimize the compression strategy for image content with varying frequency domain complexity, enhancing the protection of high-frequency details and suppressing overfitting in low-frequency regions. Furthermore, by combining multiple frequency domain enhancement and rate-distortion loss functions, the present invention enhances the model's explicit frequency domain and perceptual optimization capabilities. Compared to traditional methods, the present method improves the modeling capabilities of high-frequency details and long-range dependencies. The present method can significantly improve the compression fidelity of remote sensing images in areas with complex textures and detailed information. The present method can effectively compress remote sensing images while ensuring image reconstruction quality.
[0049] The remote sensing image compression device based on frequency domain enhancement and adaptive optimization of the present invention is used in the method of the present invention and has the same beneficial effects as the method of the present invention.
[0050] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings: Figure 1 It is a schematic diagram of a method flow of a preferred embodiment of the present invention.
[0052] Figure 2 Schematic diagram of an encoder according to a preferred embodiment of the present invention.
[0053] Figure 3 Schematic diagram of a decoder according to a preferred embodiment of the present invention.
[0054] Figure 4 It is a schematic diagram of the internal flow of the selective state space module for frequency domain enhancement in a preferred embodiment of the present invention.
[0055] Figure 5 It is a schematic diagram of the internal flow of the frequency domain analysis submodule of the preferred embodiment of the present invention.
[0056] Figure 6 1 is a schematic diagram of the internal flow of the selective state space submodule of the preferred embodiment of the present invention.
[0057] Figure 7 Schematic diagram of loss function data acquisition in a preferred embodiment of the present invention.
[0058] Figure 8 Schematic diagram of the compression stage of a preferred embodiment of the present invention.
[0059] Figure 9 Schematic diagram of a super-a priori encoder according to a preferred embodiment of the present invention.
[0060] Figure 10 Schematic diagram of a super-a priori decoder according to a preferred embodiment of the present invention.
[0061] Figure 11 2 is a schematic diagram of the decompression stage of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0062] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways as defined and covered by the claims.
[0063] See also Figure 1 In a preferred embodiment of the present invention, a remote sensing image compression method based on frequency domain enhancement and adaptive optimization is provided, comprising: S1. Construct the first model; the first model includes an encoder, an entropy model and a decoder; both the encoder and the decoder perform frequency domain analysis on the input through a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on the state space mechanism.
[0064] In a preferred embodiment of the present invention, both the encoder and the decoder perform frequency domain analysis on the input through a selective state space module with frequency domain enhancement, and model long-range dependencies and frequency domain structures based on the state space mechanism, including: See also Figure 2 , the encoder includes three groups of first modules connected in sequence; the first module includes a 5×5 convolution, 3 bottleneck residual modules and 4 frequency domain enhanced selective state space modules connected in sequence.
[0065] See also Figure 3 , the decoder includes three groups of second modules connected in sequence; the second module includes 5×5 convolution, 4 frequency domain enhanced selective state space modules and 3 bottleneck residual modules connected in sequence.
[0066] In the encoder, the frequency-domain enhanced selective state-space module is used to enhance the high-frequency details and global context expression capabilities of the input features through frequency-domain analysis and long-range dependency modeling, thereby improving the structural detail preservation and semantic expression of the compressed features and reducing the reconstruction error.
[0067] In the decoder, the frequency-domain enhanced selective state-space module is used to enhance the high-frequency details and spatial consistency in the decoded features through frequency-domain attention and long-range dependency modeling, achieving high-quality recovery of structural information and detail reconstruction, significantly improving the visual quality and fidelity of the reconstructed image.
[0068] See also Figure 4 In a preferred embodiment of the present invention, the frequency domain enhanced selective state space module includes: When the input feature After inputting the frequency domain enhanced selective state space module, the initial features are obtained through 1×1 convolution ; Initial features Input frequency domain analysis submodule to obtain frequency domain importance weight ; Initial features and frequency domain importance weights Input the selective state space submodule to update the state and obtain the context enhanced features that integrate frequency domain information and state perception information ; Enhance features based on context , frequency domain importance weight and initial features Get features ;feature Input nonlinear activation function free block to extract features to obtain features ;feature Output features after 1×1 convolution .
[0069] See also Figure 5 In a preferred embodiment of the present invention, the frequency domain analysis submodule includes: Initial features After inputting the frequency domain analysis submodule, the high frequency, medium frequency, low frequency and DC components are extracted respectively through four parallel frequency domain channels. The high frequency component is extracted by a 3×3 group convolution, and the medium frequency, low frequency and DC components are extracted by a 1×1 convolution respectively to obtain the high frequency component. , intermediate frequency components , low-frequency components and DC component ; The high frequency components , intermediate frequency components , low-frequency components and DC component The fused frequency domain features are obtained by channel splicing ; Fusion frequency domain features Enter the frequency domain importance evaluation network, and pass through average pooling, 1×1 convolution, ReLU activation function layer, 1×1 convolution and Sigmoid function layer in turn to obtain the frequency domain importance weight consistent with the original channel .
[0070] In a preferred embodiment of the present invention, extracting high frequency, intermediate frequency, low frequency and DC components respectively through four parallel frequency domain channels includes: ; See also Figure 6 In a preferred embodiment of the present invention, the selective state space submodule includes: Initial features and frequency domain importance weights After inputting the selective state space submodule, the frequency domain importance weight For initial features Perform channel-by-channel weighting to obtain the frequency domain weighted feature map , get the spectrum of higher importance and realize the feature recalibration driven by frequency domain. Get the global feature vector by average pooling ; Global eigenvector Input three layers of perceptron to get the selectivity coefficient used to control the state update process and ; The three-layer perceptron consists of the first linear transformation layer connected in sequence , ReLU activation function layer and second linear transformation layer ; Global eigenvector Adjust the channel dimension through linear mapping to obtain the state space representation ; According to the selectivity coefficient and Combined with lightweight gating mechanism to represent state space Perform selective update, selectivity coefficient Used to control the retention ratio of the original state, selectivity coefficient The nonlinear perturbation amplitude used to adjust the state, the selectivity coefficient It is used to control the degree of injection of global mean information into the state. The three work together to achieve dynamic state adjustment and fusion under frequency domain guidance, and obtain the state representation of adaptive update ; The status is represented By linearly mapping the original channel dimension, we can obtain the context-enhanced features that combine frequency domain information and state perception information. .
[0071] In a preferred embodiment of the present invention, the state representation of the adaptive update include: ; in, Represents the Sigmoid function; represents the hyperbolic tangent function; ,right Averaging over the channel latitude, S for The number of channels.
[0072] S2. Train the first model according to a preset training set and a loss function that considers pixel distortion, structure preservation, frequency domain error, perceptual quality, bit rate constraint, contextual autoregressive regularization term, and auxiliary supervision loss to obtain a pre-trained first model.
[0073] In a preferred embodiment of the present invention, pixel distortion includes mean square error (MSE) and multi-scale structural similarity (MS-SSIM); and structure preservation includes edge information and structural similarity (SSIM).
[0074] See also Figure 7 In the preferred embodiment of the present invention, the loss function obtains the input image from the compression stage of the first model , the second super prior feature , sub-features in the first data , sub-features in the second data and the second probability distribution Get the reshaped image of the model output from the decompression stage of the first model The specific sources of the parameters are introduced in step S3.
[0075] In a preferred embodiment of the present invention, the loss function includes: ; in, Represents the distortion term, which is used to measure the reconstruction error between the compressed image and the original image; Indicates bit rate loss; Represents the causal context adjustment loss with the addition of a frequency-domain-aware adaptive loss weighting mechanism, which is used to guide the model to adaptively adjust the focus of the loss function according to the frequency-domain complexity of different images; represents the auxiliary supervision loss term.
[0076] Distortion term include: ; in, Represents the square error loss, which is used to optimize pixel-level restoration accuracy. is the total number of pixels; Reshaped image representing the model output; represents the input image; Represents perceptual loss, which enhances the perception of subjective quality by comparing the distance of images in the depth feature space. For the pre-trained network Feature maps extracted by layers; Represents edge loss, which is used to enhance the preservation of image edges and structural contours. Represents the gradient of the image; Indicates frequency domain loss, and evaluates the image restoration performance in high-frequency, medium-frequency, and low-frequency regions through spectrum analysis. , , represents the spectrum amplitude, represents Fourier transform; Represents structural similarity loss, which is used to model the structural information and local consistency of the image. represents the multi-scale structural similarity index; 、 、 and are the learnable weight coefficients for perceptual loss, edge loss, frequency domain loss, and structural similarity loss, respectively.
[0077] Bit rate loss include: ; , represents the second super prior feature in the compression stage The bit rate loss term; during model training, according to the first probability modeling parameter in the compression stage Constructing the second super-prior feature in the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get .
[0078] , represents the sub-features in the first data of the compression stage Bit rate loss term; During model training, according to the sub-parameters in the first data of the compression stage Construct sub-features in the first data of the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get .
[0079] , represents the sub-features in the second data of the compression stage Bit rate loss term; During model training, according to the sub-parameters in the second data of the compression stage and the second probability modeling parameter of the compression stage The fusion result constructs the sub-features in the second data of the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get .
[0080] Auxiliary supervision loss items include: During model training, according to the second super prior feature in the compression stage Predict sub-features in the second data of the compression stage The probability distribution of , and calculate the negative log-likelihood of the feature, we get .
[0081] In a preferred embodiment of the present invention, according to the second super prior feature of the compression stage Predict sub-features in the second data of the compression stage The probability distribution of When using the noisy Gaussian distribution form, the output parameters are , is the mean, is the standard deviation, including: ; in, represents a Gaussian distribution, represents Gaussian noise.
[0082] In a preferred embodiment of the present invention, the causal context adjustment loss with the frequency-domain-aware adaptive loss weighting mechanism includes: Get the first loss ,include: ; in, represents information entropy; and These are all true distributions.
[0083] , indicating that only given Under the condition of In the probability distribution The expected value of the negative log probability under , that is, the cross entropy under context-free prediction.
[0084] , and from context modeling Under the condition of In the second probability distribution The expected value of the negative logarithmic probability under , that is, the cross entropy under context prediction.
[0085] After grayscale processing of the input image, perform two-dimensional fast Fourier transform to obtain the image amplitude ; Calculation based on preset high frequency mask , mark the left and right 1 / 4 areas of the image as high-frequency areas, including: ; High-frequency energy , total energy ; The ratio of high-frequency energy to total energy is used as the frequency domain importance index ,include: ; in, Indicates zero micro parameters.
[0086] Frequency domain importance index As a representation of the frequency domain complexity of the current image, a learnable parameter is introduced , and the frequency domain importance index Added as a learnable weight parameter to the first loss In the causal context adjustment loss, the adaptive loss weighting mechanism with frequency domain perception is obtained ,include: ; In a preferred embodiment of the present invention, during model training, a frequency-domain-aware adaptive loss weighting mechanism assists the context modeling module in performing causal autoregressive modeling of the main features; a frequency-domain enhancement rate-distortion loss function with multiple combinations is introduced into the loss function for joint optimization.
[0087] In a preferred embodiment of the present invention, several remote sensing images are selected as a training set, and the remaining images are selected as a test set. The training set images are expanded using a sliding window cropping method to obtain multi-scale image block samples; the first model is constructed under a deep learning framework; the training set is input into the first model for compression and decompression, and the Adam optimizer is used to minimize the above-mentioned loss function, parameter optimization training is performed, and the training results are verified using the test set, finally obtaining a pre-trained first model.
[0088] In a preferred embodiment of the present invention, a frequency-domain-aware adaptive loss weighting mechanism is implemented to dynamically adjust the contextual loss weights based on the frequency energy distribution, enabling content-adaptive rate-distortion optimization. By analyzing the frequency domain characteristics of the input image, the training loss weight distribution is dynamically adjusted. For images rich in high-frequency detail, the compression loss's focus on detail is automatically increased, thereby improving compression performance. For images dominated by low frequencies, the impact of the auxiliary loss is automatically reduced to prevent over-optimization. This enables the model to adaptively adjust the loss function's focus based on the frequency domain complexity of different images.
[0089] In a preferred embodiment of the present invention, during the model training process, a method of adding uniform noise is used to simulate the quantization effect, which is expressed as: Indicates After training, the model is used to compress and decompress the input image using the following hard quantization method: , Represents the rounding function.
[0090] S3. Deploy the pre-trained first model to the compression end and the decompression end at the same time, and compress and decompress the remote sensing image through the pre-trained first model of the compression end and the decompression end.
[0091] Compression includes: extracting the main features of the input image through the encoder; probabilistically modeling the distribution information of the main features through the entropy model, generating a bit stream, and sending the bit stream to the decompression end.
[0092] Decompression includes: the decompression end receives the bit stream; restores the bit stream to the main features through the entropy model; inputs the main features into the decoder to obtain the reshaped image.
[0093] See also Figure 8 In a preferred embodiment of the present invention, extracting the main features of the input image by an encoder; probabilistically modeling the distribution information of the main features by an entropy model, and generating a bitstream include: In a preferred embodiment of the present invention, the entropy model includes a super-a priori encoder, a super-a priori decoder, a context modeling module, a one-stage modeling module and a two-stage modeling module.
[0094] See also Figure 9 In a preferred embodiment of the present invention, the super prior encoder is composed of three consecutive convolutional layers and two activation function layers in cascade. The convolution kernel size of the first convolutional layer is 3×3, followed by a GELU activation function layer; the convolution kernel size of the second convolutional layer is 5×5, followed by a GELU activation function layer; the convolution kernel size of the third convolutional layer is 5×5.
[0095] See also Figure 10 In a preferred embodiment of the present invention, the super prior decoder is composed of three consecutive convolutional layers and two activation function layers in cascade. The convolution kernel size of the first convolutional layer is 5×5, followed by a GELU activation function layer; the convolution kernel size of the second convolutional layer is 5×5, followed by a GELU activation function layer; the convolution kernel size of the third convolutional layer is 3×3.
[0096] Extract the input image through the encoder Main features y ; The main features are encoded by the super prior y The spatial and channel statistical information is modeled to obtain the first super prior feature z ; The first super prior feature z Quantize and get the second super prior feature ; Based on the second super-prior feature through the super-prior decoder Get the first probability modeling parameter , is the mean, is the standard deviation.
[0097] Main feature y Quantify and obtain the main features after quantization ,Will Based on the preset ratio of non-uniform channel grouping to divide into two sub-features and (For example, Channel as ,back Channel as ), and the first probability modeling parameter Perform the same division to obtain two sub-parameters and ; Pair the sub-features and sub-parameters with the same channel size and randomly assign them as the first data and the second data. In the preferred embodiment of the present invention, it is assumed that the first data includes the sub-features and sub-parameters , the second data includes sub-features and sub-parameters .
[0098] Through the one-stage modeling module according to the sub-parameters in the first data Modeling is performed to obtain the first probability distribution , using the noisy Gaussian distribution form: ; in, represents a Gaussian distribution, represents Gaussian noise.
[0099] The context modeling module is used to model the sub-features in the first data Extract context information in spatial and channel dimensions to obtain the second probability modeling parameters .
[0100] Modeling is performed by a two-stage modeling module based on the second probability modeling parameter and the sub-parameters in the second data to obtain a second probability distribution, including: The second probability modeling parameter and the sub-parameters in the second data By fusion, we get: ; The second probability distribution is obtained by modeling based on the fusion results through the two-stage modeling module , using the noisy Gaussian distribution form: ; According to the first probability distribution For the sub-features in the first data Perform entropy coding to generate the first bit stream .
[0101] According to the second probability distribution For the sub-features in the second data Perform entropy coding to generate the second bit stream .
[0102] The second hyper-prior feature is modeled by Gaussian distribution Modeling is performed to obtain the third probability distribution ; According to the third probability distribution The second super prior feature Perform entropy coding to generate the third bit stream .
[0103] See also Figure 11 In a preferred embodiment of the present invention, the bit stream is restored to the main features through the entropy model; the main features are input into the decoder to obtain the reconstructed image, which includes: The decompression end uses the third bit stream Perform probability modeling parameters to obtain the third probability distribution .
[0104] According to the third probability distribution For the third bitstream Perform entropy decoding to obtain the second super prior feature .
[0105] Based on the second super-prior feature Get the first probability modeling parameter ; The first probability modeling parameter Based on the same preset ratio as above, the non-uniform channel grouping is divided into two sub-parameters and , and match them with the channel size of the first data and the second data to obtain the first sub-parameter and the second subparameter .
[0106] Through the one-stage modeling module according to the first sub-parameter Modeling is performed to obtain the first probability distribution .
[0107] According to the first probability distribution For the first bitstream Perform entropy decoding to obtain the sub-features in the first data .
[0108] The context modeling module is used to model the sub-features in the first data Extract context information in spatial and channel dimensions to obtain the second probability modeling parameters .
[0109] The second probability modeling parameter is modeled by the two-stage modeling module and the second subparameter Modeling is performed to obtain the second probability distribution ; According to the second probability distribution For the second bitstream Perform entropy decoding to obtain the sub-features in the second data .
[0110] The sub-features in the first data and the sub-features in the second data Splice to get the quantized main features ; The quantized main features are restored through the decoder to obtain the reshaped image .
[0111] The present invention's remote sensing image compression method based on frequency domain enhancement and adaptive optimization fully considers the expressive importance of different frequency components during feature modeling by introducing a selective state-space module for frequency domain enhancement. This method achieves explicit perception of image frequency domain characteristics and efficient modeling of long-range dependencies, effectively improving the under-representation of traditional convolutional networks when compressing high-frequency information. Furthermore, by introducing a frequency-domain-aware adaptive loss weighting mechanism, the causal context-adjusted loss weights during training are dynamically adjusted based on the spectral energy distribution of the input image. This enables the model to adaptively optimize the compression strategy for image content with varying frequency domain complexity, enhancing the preservation of high-frequency details and suppressing overfitting in low-frequency regions. Furthermore, by combining multiple frequency domain enhancement and rate-distortion loss functions, the present invention enhances the model's explicit frequency domain and perceptual optimization capabilities. Compared to traditional methods, this method improves the modeling capabilities of high-frequency details and long-range dependencies. It can significantly improve the compression fidelity of remote sensing images in areas with complex textures and detailed information. The present method effectively compresses remote sensing images while ensuring image reconstruction quality.
[0112] In a preferred embodiment of the present invention, a remote sensing image compression device based on frequency domain enhancement and adaptive optimization is provided, which is used in the method of the present invention. The device includes a first unit and a second unit.
[0113] The first unit is used to train the first model according to a preset training set and a loss function that considers pixel distortion, structure preservation, frequency domain error, perceptual quality, bit rate constraint, contextual autoregressive regularization term and auxiliary supervision loss to obtain a pre-trained first model.
[0114] The second unit is used to compress and decompress remote sensing images through a pre-trained first model deployed simultaneously at the compression and decompression ends; the first model includes an encoder, an entropy model and a decoder; both the encoder and the decoder perform frequency domain analysis on the input through a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on the state space mechanism.
[0115] Compression includes: extracting the main features of the input image through the encoder; probabilistically modeling the distribution information of the main features through the entropy model, generating a bit stream, and sending the bit stream to the decompression end.
[0116] Decompression includes: the decompression end receives the bit stream; restores the bit stream to the main features through the entropy model; inputs the main features into the decoder to obtain the reshaped image.
[0117] The remote sensing image compression device based on frequency domain enhancement and adaptive optimization of the present invention is used in the method of the present invention and has the same beneficial effects as the method of the present invention.
[0118] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A remote sensing image compression method based on frequency domain enhancement and adaptive optimization, characterized in that: The remote sensing image is compressed and decompressed by using the pre-trained first model deployed simultaneously on the compression and decompression ends, including: The first model includes an encoder, an entropy model, and a decoder; the encoder and decoder both perform frequency domain analysis on the input through a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on a state space mechanism; The first model is trained according to a preset training set and a loss function that considers pixel distortion, structure preservation, frequency domain error, perceptual quality, bit rate constraint, contextual autoregressive regularization term, and auxiliary supervision loss to obtain a pre-trained first model; The compression includes: extracting the main features of the input image by the encoder; performing probability modeling on the distribution information of the main features by the entropy model to generate a bit stream, and sending the bit stream to the decompression end; The decompression includes: the decompression end receives the bit stream; restores the bit stream to the main features through the entropy model; and inputs the main features into the decoder to obtain a reshaped image.
2. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 1, characterized in that: The encoder and decoder both perform frequency domain analysis on the input through a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on the state space mechanism, including: The encoder includes three groups of first modules connected in sequence; the first modules include 5×5 convolution, 3 bottleneck residual modules and 4 frequency domain enhanced selective state space modules connected in sequence; The decoder includes three groups of second modules connected in sequence; the second modules include 5×5 convolution, 4 frequency domain enhanced selective state space modules and 3 bottleneck residual modules connected in sequence; In the encoder, the frequency domain enhanced selective state space module is used to enhance the high frequency details and global context expression ability of input features through frequency domain analysis and long-range dependency modeling; In the decoder, the frequency-domain enhanced selective state-space module is used to enhance high-frequency details and spatial consistency in decoded features through frequency-domain attention and long-range dependency modeling.
3. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 2, characterized in that: The frequency domain enhanced selective state space module includes: When the input feature After inputting the frequency domain enhanced selective state space module, the initial features are obtained through 1×1 convolution ; The initial features Input frequency domain analysis submodule to obtain frequency domain importance weight ; The initial features and the frequency domain importance weight Input the selective state space submodule to update the state and obtain the context enhanced features that integrate frequency domain information and state perception information ; Enhance features based on the context , the frequency domain importance weight and the initial features Get features ; The characteristics Input nonlinear activation function free block to extract features to obtain features ; The characteristics Output features after 1×1 convolution .
4. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 3, characterized in that: The frequency domain analysis submodule includes: The initial features After inputting the frequency domain analysis submodule, the high frequency, medium frequency, low frequency and DC components are extracted respectively through four parallel frequency domain channels. The high frequency component is extracted by a 3×3 group convolution, and the medium frequency, low frequency and DC components are extracted by a 1×1 convolution respectively to obtain the high frequency component. , intermediate frequency components , low-frequency components and DC component ; The high frequency component , intermediate frequency components , low-frequency components and DC component The fused frequency domain features are obtained by channel splicing ; The fused frequency domain features Enter the frequency domain importance evaluation network, and pass through average pooling, 1×1 convolution, ReLU activation function layer, 1×1 convolution and Sigmoid function layer in turn to obtain the frequency domain importance weight consistent with the original channel .
5. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 4, characterized in that: The selective state space submodule includes: The initial features and the frequency domain importance weight After inputting the selective state space submodule, the frequency domain importance weight For the initial features Perform channel-by-channel weighting to obtain the frequency domain weighted feature map ; The frequency domain weighted feature map Get the global feature vector by average pooling ; The global feature vector Input three layers of perceptron to get the selectivity coefficient used to control the state update process and The three-layer perceptron includes a first linear transformation layer connected in sequence , ReLU activation function layer and second linear transformation layer ; The global feature vector Adjust the channel dimension through linear mapping to obtain the state space representation According to the selectivity coefficient and Combined with a lightweight gating mechanism to represent the state space Perform selective update, selectivity coefficient Used to control the retention ratio of the original state, selectivity coefficient The nonlinear perturbation amplitude used to adjust the state, the selectivity coefficient It is used to control the degree of injection of global mean information into the state. The three work together to achieve dynamic state adjustment and fusion under frequency domain guidance, and obtain the state representation of adaptive update ; The status is represented By linearly mapping the original channel dimension, we can obtain the context-enhanced features that combine frequency domain information and state perception information. .
6. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 5, characterized in that: Extracting the main features of the input image by the encoder; Probabilistically modeling the distribution information of the main features using the entropy model to generate a bitstream includes: The entropy model includes a super-prior encoder, a super-prior decoder, a context modeling module, a one-stage modeling module, and a two-stage modeling module; the encoder extracts the main features of the input image; the super-prior encoder models the spatial and channel statistical information of the main features to obtain a first super-prior feature; the first super-prior feature is quantized to obtain a second super-prior feature; and the super-prior decoder obtains a first probability modeling parameter based on the second super-prior feature. After quantizing the main feature, the main feature is divided into two sub-features based on the non-uniform channel grouping of a preset ratio, and the first probability modeling parameter is divided into two sub-parameters in the same manner; the sub-features and sub-parameters with the same channel size are paired and randomly assigned as the first data and the second data; The first-stage modeling module performs modeling based on the sub-parameters in the first data to obtain a first probability distribution; the context modeling module extracts context information in spatial and channel dimensions based on the sub-features in the first data to obtain second probability modeling parameters; the second-stage modeling module performs modeling based on the second probability modeling parameters and the sub-parameters in the second data to obtain a second probability distribution; Performing entropy encoding on sub-features in the first data according to the first probability distribution to generate a first bitstream; Performing entropy encoding on the sub-features in the second data according to the second probability distribution to generate a second bitstream; The second super a priori feature is modeled by a Gaussian distribution model to obtain a third probability distribution; and the second super a priori feature is entropy encoded according to the third probability distribution to generate a third bit stream.
7. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 6, characterized in that: Restoring the bit stream to the main features by using the entropy model; inputting the main features into the decoder to obtain the reconstructed image includes: The decompression end performs probability modeling parameters according to the third bit stream to obtain the third probability distribution; Performing entropy decoding on the third bit stream according to the third probability distribution to obtain the second super-prior feature; Obtaining the first probability modeling parameter based on the second super-prior feature by the super-prior decoder; dividing the first probability modeling parameter into two sub-parameters based on the non-uniform channel grouping of the preset ratio, and matching them with the channel sizes of the first data and the second data to obtain a first sub-parameter and a second sub-parameter, respectively; Performing modeling according to the first sub-parameters by the first-stage modeling module to obtain the first probability distribution; Performing entropy decoding on the first bit stream according to the first probability distribution to obtain sub-features in the first data; Extracting context information in spatial and channel dimensions based on the sub-features in the first data by the context modeling module to obtain the second probability modeling parameters; Performing modeling according to the second probability modeling parameter and the second sub-parameter by the two-stage modeling module to obtain the second probability distribution; Performing entropy decoding on the second bit stream according to the second probability distribution to obtain sub-features in the second data; concatenating the sub-features in the first data and the sub-features in the second data to obtain a quantized main feature; The quantized main features are restored by the decoder to obtain a reshaped image.
8. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 7, characterized in that: The loss function includes: ; in, Represents the distortion term, which is used to measure the reconstruction error between the compressed image and the original image; Indicates bit rate loss; Represents the causal context adjustment loss with the addition of a frequency-domain-aware adaptive loss weighting mechanism, which is used to guide the model to adaptively adjust the focus of the loss function according to the frequency-domain complexity of different images; represents the auxiliary supervision loss term; The distortion term include: ; in, Represents the square error loss, which is used to optimize pixel-level restoration accuracy. is the total number of pixels; Reshaped image representing the model output; represents the input image; Represents perceptual loss, which enhances the perception of subjective quality by comparing the distance of images in the depth feature space. For the pre-trained network Feature maps extracted by layers; Represents edge loss, which is used to enhance the preservation of image edges and structural contours. Represents the gradient of the image; Indicates frequency domain loss, and evaluates the image restoration performance in high-frequency, medium-frequency, and low-frequency regions through spectrum analysis. , , represents the spectrum amplitude, represents Fourier transform; Represents structural similarity loss, which is used to model the structural information and local consistency of the image. represents the multi-scale structural similarity index; 、 、 and are the learnable weight coefficients for perceptual loss, edge loss, frequency domain loss, and structural similarity loss respectively; The bit rate loss include: ; , represents the second super prior feature in the compression stage The bit rate loss term; During model training, according to the first probability modeling parameter in the compression stage Construct the second super prior feature in the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get ; , represents the sub-features of the first data in the compression stage Bit rate loss term; During model training, according to the sub-parameters in the first data in the compression stage Construct sub-features of the first data in the compression stage The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get ; , represents the sub-features of the second data in the compression stage Bit rate loss term; During model training, according to the sub-parameters in the second data in the compression stage The fusion result of the second probability modeling parameter and the second data in the compression stage is used to construct the sub-features The probability distribution of , and estimate the negative log likelihood value of each feature position, and accumulate to get ; The auxiliary supervision loss term include: During model training, according to the second super prior feature in the compression stage Predict the sub-features of the second data in the compression stage The probability distribution of , and calculate the negative log-likelihood of the feature, we get .
9. The remote sensing image compression method based on frequency domain enhancement and adaptive optimization according to claim 8, characterized in that: The causal context adjustment loss with the frequency-domain-aware adaptive loss weighting mechanism includes: Get the first loss ,include: ; in, represents information entropy; and All are true distributions; , indicating that only given Under the condition of In the probability distribution The expected value of the negative logarithmic probability under , that is, the cross entropy under context-free prediction; , given and from context modeling Under the condition of In the second probability distribution The expected value of the negative logarithmic probability under , that is, the cross entropy under context prediction; After grayscale processing of the input image, perform two-dimensional fast Fourier transform to obtain the image amplitude ; Calculation based on preset high frequency mask , mark the left and right 1 / 4 areas of the image as high-frequency areas, including: ; High-frequency energy , total energy ; The ratio of high-frequency energy to total energy is used as the frequency domain importance index ,include: ; in, Indicates zero micro parameters; Frequency domain importance index As a representation of the frequency domain complexity of the current image, a learnable parameter is introduced , and the frequency domain importance index Added as a learnable weight parameter to the first loss The causal context adjustment loss of the adaptive loss weighting mechanism with frequency domain perception is obtained ,include: 。 10. A remote sensing image compression device based on frequency domain enhancement and adaptive optimization, used in the method according to any one of claims 1 to 9, characterized in that: The device comprises a first unit and a second unit; The first unit is used to train the first model according to a preset training set and a loss function that considers pixel distortion, structure preservation, frequency domain error, perceptual quality, bit rate constraint, contextual autoregressive regularization term and auxiliary supervision loss to obtain a pre-trained first model; The second unit is configured to compress and decompress the remote sensing image using a pre-trained first model deployed simultaneously at the compression end and the decompression end; the first model includes an encoder, an entropy model, and a decoder; the encoder and decoder both perform frequency domain analysis on the input using a frequency domain enhanced selective state space module, and model long-range dependencies and frequency domain structures based on a state space mechanism; The compression includes: extracting the main features of the input image by the encoder; performing probability modeling on the distribution information of the main features by the entropy model to generate a bit stream, and sending the bit stream to the decompression end; The decompression includes: the decompression end receives the bit stream; restores the bit stream to the main features through the entropy model; and inputs the main features into the decoder to obtain a reshaped image.
Citation Information
Patent Citations
Remote sensing image compression method based on multi-scale asymmetric coding and decoding network
CN118608799A
Video compression intelligent preprocessing method, system and device based on frequency domain perception optimization and medium
CN120111245A
Image compression method and apparatus
US20220286696A1
Cited By
Image compression method based on traffic scene
CN121661160A