A space-ground integrated multispectral remote sensing image compression method and system
The star-ground joint multi-spectral remote sensing image compression framework using a variational autoencoder addresses computational resource limitations on satellites, achieving high-compression ratios and fidelity through a novel encoding and decoding architecture.
Patent Information
- Application Number
- CN202211639456.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-19
AI Technical Summary
The prior art is difficult to achieve high-magnification compression of multi-spectral remote sensing images on remote sensing satellites, resulting in data backlog and satellite-ground transmission delay, affecting the satellite's immediate service.
The compression method of satellite-ground combined multispectral remote sensing image is adopted, and the compression model is constructed using a variational autoencoder, including encoding layer, super-priori encoding, super-priori decoding and decoding layer. Combined with the semantic information reconstruction module and a multi-head decoder, efficient compression is carried out through the satellite-ground joint.
High-fidelity image reconstruction under high compression ratio is realized, the efficiency of on-star data storage and satellite bandwidth utilization is improved, the calculation process is simplified, and the multi-spectral remote sensing image compression task is adapted to any band.
Smart Images

Figure CN115955563B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-spectral remote sensing image compression, and particularly relates to a satellite-ground joint multi-spectral remote sensing image compression scheme, and proposes a multi-spectral remote sensing image compression scheme based on a variational autoencoder to achieve high magnification and high-fidelity image compression tasks. Background Art
[0002] With the development of remote sensing technology, high-resolution remote sensing satellites are faced with problems such as large amounts of data, weak on-board storage capabilities, and small satellite-ground transmission bandwidths, resulting in backlogs of on-board data and untimely downlink of satellite-ground data. These will all lead to large delays in remote sensing satellite services, seriously affecting the immediate services of the satellites. Remote sensing image compression technology can encode image data into bitstream data with less storage space, which can not only reduce the required storage space, but also improve the satellite-ground data transmission efficiency.
[0003] Optical remote sensing satellites usually carry cameras with multiple spectral bands and can obtain multi-spectral images with rich information. For example, China's GF7 satellite has four spectral band images, the HS2 satellite has 10 spectral bands, and the GFDM satellite has eight spectral band images. Due to the rich spectral and spatial information of multi-spectral images, they have been widely used in many fields such as land management, urban construction, and military reconnaissance. However, the rich information of multi-spectral remote sensing images requires more storage space and higher bandwidth. Given an 8-spectral band image with a size of 1,000×1,000 pixels and a data type of 16 bits / pixel, it occupies approximately 15 MB of storage space. For RGB images, it only requires less than 3 MB of storage space. Therefore, using image compression to reduce the data volume of remote sensing images is a very effective method. Currently, traditional image compression methods include JPEG, JPEG2000, JPEG-LS, CCSDS, etc. However, their compression ratios are usually relatively low, generally less than 10:1. Although these methods can achieve higher compression ratios, they may cause serious distortion of the images, affecting normal use.
[0004] With the development of deep learning technology, the image compression method based on variational autoencoder has better compression performance than standard compression methods such as JPEG and JPEG2000. The compression model based on variational autoencoder consists of an encoder and a decoder. The purpose of the encoder is to map the input image to the coding space through parametric non-linear transformation, and then use the quantization function and entropy coding to obtain the bitstream; during decoding, it is necessary to use entropy decoding to decode the bitstream into the latent representation feature map, and then use the decoder to reconstruct the latent representation map back into the image. Although the current compression method based on variational autoencoder has achieved good results in natural images, there are still some problems in the compression coding of multi-spectral remote sensing images, especially in the limited computing resources on satellites, it is difficult to provide sufficient computing power for the deep learning-based compression model. There are no relevant papers in domestic and foreign journals that propose a high magnification multi-spectral remote sensing image compression method for on-orbit environments. At present, there are no relevant solutions and authorized patents in China either. Summary of the Invention
[0005] Aiming at the problem of efficient compression of on-orbit multi-spectral remote sensing images, the present invention provides a satellite-ground joint multi-spectral remote sensing image compression scheme to achieve on-orbit high magnification compression tasks.
[0006] The technical solution provided by the present invention is a satellite-ground joint multi-spectral remote sensing image compression method, including the following steps:
[0007] Step 1, construct a multi-spectral remote sensing image data set;
[0008] Step 2, set up a satellite-ground joint multi-spectral remote sensing image compression model, the satellite-ground joint multi-spectral remote sensing image compression model is implemented based on a variational autoencoder, including an encoding layer g e , a hyperprior encoder g he , a hyperprior decoder g hd and a decoding layer g d ; the encoding layer is used to map the input image to the latent representation space to support entropy coding and entropy decoding of the quantized representation information; the hyperprior encoding and decoding modules are used to learn the latent representation information output by the encoding layer and use the Gaussian mixture model to realize the modeling of the latent representation; the decoding layer is used to reconstruct the original image according to the entropy decoded representation feature map;
[0009] Step 3, set up a semantic information reconstruction module, the semantic information reconstruction module is used to generate the mean, variance and bias feature map of the Gaussian distribution model, the mean and variance of the Gaussian distribution model are used to construct the Gaussian distribution model to provide a probability model for entropy coding and decoding, and the bias feature map is used to reconstruct the lost semantic information after quantization and fuse it with the quantized representation information for image reconstruction by the decoding layer;
[0010] Step 4, construct a multi-head decoder, including setting a corresponding decoder according to the number of bands of the multi-spectral image to be reconstructed.
[0011] Step 5, train the satellite-ground joint multi-spectral remote sensing image compression model.
[0012] Step 6, perform satellite-ground joint multi-spectral remote sensing image compression. The compression is divided into two parts. The first part is the on-orbit compression part, including an encoding layer, a hyperprior encoding layer, a hyperprior decoding layer, and entropy encoding. The second part is the ground reconstruction part, including entropy decoding, semantic information reconstruction, a decoding layer, and a multi-head decoder.
[0013] Moreover, in Step 1, when constructing the multi-spectral remote sensing image dataset, a high-resolution multi-mode satellite data source is used to construct an 8-band high-resolution multi-mode satellite dataset.
[0014] Moreover, in the satellite-ground joint multi-spectral remote sensing image compression model, the implementation method of the hyperprior encoding and decoding module is to use a convolutional neural network to actively learn the Gaussian distribution information of each pixel in the feature map, providing an accurate probability distribution for each element for entropy encoding and decoding, so as to better encode the quantized feature map.
[0015] Moreover, in Step 3, the semantic information reconstruction is implemented as follows.
[0016] 1) Use the feature information generated by the hyperprior decoding layer to generate the mean, variance, and bias information of the Gaussian model through a convolutional neural network.
[0017] 2) Use the generated mean and variance to construct a Gaussian distribution model, and use entropy decoding to decode the bitstream to recover the quantized representation information.
[0018] 3) Use the bias information obtained in 1) and the quantized representation information to perform a difference to obtain a difference feature map, and use a convolutional neural network to generate a semantic feature s.
[0019] 4) Use the semantic feature s generated in 3) and the quantized representation information to fuse and generate a new representation.
[0020] Moreover, in Step 4, the construction of the multi-head decoder is implemented as follows.
[0021] 1) Use the decoding layer g d to reconstruct a feature map with the same size as the input image, denoted as
[0022] 2) Use the channel separation method to divide into n groups, and the i-th group of feature maps is denoted as where n represents the number of bands of the multispectral image;
[0023] 3) Use n decoding heads to map n groups of feature maps to the space of 1 channel respectively, for generating an n-band image.
[0024] Moreover, the on-board compression implementation process includes the following steps
[0025] 1) Use the encoding layer g e to map the multispectral image x to the latent representation layer, obtaining the latent representation y;
[0026] 2) Quantize the y obtained in 1) to obtain the quantized latent representation
[0027] 3) Use the hyperprior encoding layer g he to further encode and quantize the latent representation y in 1), obtaining the representation z and the quantized representation
[0028] 4) Perform entropy encoding on the quantized representation to generate the bitstream z_bits-stream. The probability model of its entropy encoding is the probability density function obtained through offline learning. Entropy encode each element in the quantized representation ;
[0029] 5) Use the hyperprior decoding layer g hd to decode the quantized representation obtained in 3), and output two parameters, the mean Mean and the variance Variance of the Gaussian model;
[0030] 6) Construct the Gaussian distribution probability model according to the Mean and Variance generated in 5), and perform entropy encoding on the quantized latent representation in 2) to generate the bitstream y_bits-stream;
[0031] 7) Send the bitstream z_bits-stream and the bitstream y_bits-stream in 4) to the space-ground transmission link and download them to the ground.
[0032] Moreover, the ground decoding includes the following steps
[0033] 1) The ground receives the bitstream z_bits-stream and the bitstream y_bits-stream, and performs entropy decoding on z_bits-stream to obtain
[0034] 2) Input the obtained in 1) into the hyperprior decoding layer ghd Generate
[0035] 3) Input into the semantic feature reconstruction module to generate the mean Mean and variance Variance of the Gaussian distribution for the bias feature map Offset; construct a Gaussian distribution model based on the generated Mean and Variance, and perform entropy decoding on the bitstream y_bits - stream to obtain the quantized representation
[0036] 4) Then use together with Offset for semantic feature reconstruction and generate
[0037] 5) Input the generated in 4) into g d for decoding to obtain
[0038] 6) Input the generated in 5) into the multi - head decoder to reconstruct the multi - spectral image.
[0039] On the other hand, the present invention provides a space - to - ground joint multi - spectral remote sensing image compression system for implementing a space - to - ground joint multi - spectral remote sensing image compression method as described in any one of the above.
[0040] Moreover, it includes the following modules,
[0041] The first module is used for constructing a multi - spectral remote sensing image data set;
[0042] The second module is used for setting a space - to - ground joint multi - spectral remote sensing image compression model, which is implemented based on a variational auto - encoder and includes an encoding layer g e , a hyper - prior encoding g he , a hyper - prior decoding g hd and a decoding layer g d ; the encoding layer is used to map the input image to the latent representation space to support entropy encoding and entropy decoding of the quantized representation information; the hyper - prior encoding and decoding modules are used to model the latent representation by using a Gaussian mixture model through learning the latent representation information output by the encoding layer; the decoding layer is used to reconstruct the original image according to the representation feature map after entropy decoding;
[0043] The third module is used to set up a semantic information reconstruction module, which is used to generate the mean, variance, and bias feature map of the Gaussian distribution model. The mean and variance of the Gaussian distribution model are used to construct a Gaussian distribution model to provide a probability model for entropy coding and decoding. The bias feature map is used to reconstruct the lost semantic information after quantization and fuse it with the quantized representation information for image reconstruction in the decoding layer;
[0044] The fourth module is used to construct a multi-head decoder, including setting corresponding decoders according to the number of bands of the multi-spectral image to be reconstructed;
[0045] The fifth module is used to train the satellite-ground joint multi-spectral remote sensing image compression model;
[0046] The sixth module is used to perform satellite-ground joint multi-spectral remote sensing image compression. The compression is divided into two parts. The first part is the on-orbit compression part, including an encoding layer, a hyperprior encoding layer, a hyperprior decoding layer, and entropy coding. The second part is the ground reconstruction part, including entropy decoding, semantic information reconstruction, a decoding layer, and a multi-head decoder.
[0047] Alternatively, it includes a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a satellite-ground joint multi-spectral remote sensing image compression method as described above.
[0048] The present invention provides a satellite-ground joint multi-spectral remote sensing image compression framework, which solves the problem of data transmission delay caused by limited on-board data storage and satellite-ground bandwidth, and effectively ensures high-fidelity image reconstruction at a high compression ratio. This solution has the characteristics of simplicity, effectiveness, high precision, and easy implementation.
[0049] Compared with the prior art, the present invention has the following advantages:
[0050] (1) The semantic feature reconstruction module and the multi-head decoder proposed by the present invention are used at the ground decoding end and do not affect the computational efficiency of the on-board compression encoding end.
[0051] (2) The satellite-ground joint multi-spectral remote sensing image compression designed based on the convolutional neural network model by the present invention has a high compression ratio and high-fidelity compression effect, and improves the compression ratio of the existing compression method from 2-10 times to 40 times.
[0052] (3) It has strong practicability and universality. The present invention fully considers the correlation between bands and can adapt to the compression tasks of multi-spectral remote sensing images of any band.
[0053] The solution of the present invention is simple and convenient to implement, has strong practicability, solves the problems of low practicability and inconvenience in actual application existing in the related technology, can improve the user experience, and has important market value. Brief Description of the Drawings
[0054] Figure 1 Schematic diagram for configuring the network model of the embodiment of the present invention.
[0055] Figure 2 Schematic diagram for reconstructing semantic representation of the embodiment of the present invention.
[0056] Figure 3 Schematic diagram of the multi-head decoder of the embodiment of the present invention.
[0057] Figure 4 Schematic diagram of the on-satellite compression encoding end of the embodiment of the present invention.
[0058] Figure 5 Schematic diagram of the ground decoding end of the embodiment of the present invention.
[0059] Figure 6 Schematic diagram for comparing and displaying compression results of the embodiment of the present invention.
[0060] Figure 7 Schematic diagram for comparing and displaying rate-distortion curves of different compression methods of the embodiment of the present invention. Detailed implementation manners
[0061] The technical solution of the present invention will be specifically described below in conjunction with the accompanying drawings and embodiments.
[0062] The technical solution of the present invention provides a design for a space-ground joint multi-spectral remote sensing image compression framework aiming at the technical problems to be solved. Specifically in implementation, the technology of the present invention can be used for experiments with the Python programming language or for engineering applications with the C / C++ programming language.
[0063] The embodiment of the present invention provides a space-ground joint multi-spectral remote sensing image compression method, including the following steps:
[0064] Step 1, constructing a multi-spectral remote sensing image data set, including constructing an 8-band high-resolution multi-mode satellite data set, and the 8-band data set is used for model training, validation, and testing;
[0065] In this embodiment, an 8-band data set is constructed to train the compression model. This data set uses the data source of the high-resolution multi-mode remote sensing satellite in China. First, the acquired image data has a spatial resolution of 2-meter multi-spectral images, with a size of 10000×10000, and a total of 40 multi-spectral images. Limited by memory, the images in the embodiment of the present invention are divided into blocks of 500×500 in size, with a total of 16000 images, among which 10000 images are used as the training set, 3000 images are used as the validation set, and 3000 images are used as the test set.
[0066] Step 2, setting up the space-ground joint multi-spectral remote sensing image compression model, which is implemented based on a variational autoencoder and includes an encoding layer g e , a hyperprior encoder g he , a hyperprior decoder g hd , and a decoding layer g d .
[0067] The encoding layer mainly maps the input image to the latent representation space to facilitate entropy encoding AE and entropy decoding AD of the quantized representation information; among them, quantization Q is used to eliminate the information with high redundancy in the latent representation layer; entropy encoding and decoding encode the quantized representation information into a binary bitstream or decode the binary bitstream into a representation feature map;
[0068] The hyperprior encoding and decoding modules learn the latent representation information output by the encoding layer and use a Gaussian mixture model to model the latent representation. The main implementation method is to use a convolutional neural network to actively learn the Gaussian distribution information of each pixel in the representation feature map, providing an accurate probability distribution for each element for entropy encoding and decoding, so as to better encode the quantized representation feature map;
[0069] The decoding layer reconstructs the original image according to the representation feature map after entropy decoding.
[0070] Further implementation suggestions are as follows:
[0071] In the space-ground joint multi-spectral remote sensing image compression model, convolutional layers are used for feature extraction;
[0072] The encoding layer includes several convolutional layers and activation layers. The input of the encoding layer is the original image, and the output is the latent representation feature map;
[0073] Quantization is in the form of adding uniform noise or using a rounding function. Uniform noise is used during model training, and the rounding function is used during model inference;
[0074] Entropy encoding and decoding include entropy encoding and entropy decoding of the quantized latent representation;
[0075] Hyperprior encoding includes encoding the latent representation information and using quantization and entropy encoding to generate the bitstream of the edge information; the hyperprior decoding layer learns the two parameters of the mean and variance of the latent representation obeying the Gaussian distribution according to the edge information and uses them to construct a Gaussian distribution probability model;
[0076] The decoding layer includes several transposed convolutional layers and activation function layers, and the output is the compressed image.
[0077] In Step 2 of the embodiment, the encoding layer, decoding layer, hyperprior encoding layer, and hyperprior decoding layer all use convolutional layers and transposed convolutional layers for feature extraction and reconstruction. SeeFigure 1 。
[0078] Encoding layer g e comprises four convolutional layers (Conv) and three generalized divisive normalization layers (GDN). Conv is used to extract effective features, and GDN is used to increase the nonlinearity of the model; the encoding layer g e functions to map the input multi-spectral image into a high-dimensional feature space and generate a latent representation y; Figure 1 In Figure 1 , the arrow pointing to the right indicates that the output of the left module is the input of the right module; following the arrow direction, the input of the first Conv layer is the multi-spectral image of B bands, and the output is the feature map of N channels. k represents the size of the convolutional kernel, and s represents the stride of the convolutional layer. s = 2 indicates that the size of the feature map output by this convolutional layer is half of the size of the input feature map. Then, the output of the first Conv layer is input into the first GDN; the output of the first GDN is connected to the second Conv. The input of the second Conv is the feature map of N channels, the convolutional kernel is 5, and the stride is 2. The size of the output feature map is half of the size of the input feature map. Then, the output of the second Conv layer is input into the second GDN, and its output is the input of the third Conv. The convolutional kernel of the third Conv, k = 5, and the stride s = 2. The number of input and output feature maps is N. Then, the output of the third Conv layer is input into the third GDN, and the output of the third GDN is used as the input of the fourth Conv. The fourth convolutional kernel, k = 5, and the stride s = 2. The number of output feature maps is M.
[0079] Decoding layer g d comprises four transposed convolutional layers (TConv) and three inverse generalized divisive normalization layers (IGDN), Figure 1 in Figure 1 , the arrow pointing to the left indicates that the output of the right module is the input of the left module; the decoding layer g d functions to reconstruct the quantized latent representation back into the remote sensing image; following the arrow direction, the input of the first TConv is the quantized representation with M channels, the convolutional kernel k = 3, and the stride s = 2. The output feature map is twice the size of the input feature map. Then, the output of the first TConv is input into the first IGDN, and its output is used as the input of the second TConv. The convolutional kernel of the TConv is 3, and the stride is 2. The input feature map has N channels. Then, the output of the second TConv is input into the second IGDN, and its output is used as the input of the third TConv. The convolutional kernel of the third TConv is 3, and the stride is 2. The number of output feature maps is N. Finally, the output of the third TConv is input into the third IGDN, and the output of the third IGDN enters the fourth TConv to reconstruct the image of B bands using the fourth TConv.
[0080] Hyper - prior encoding layer g he It includes three convolutional layers and two activation layers (RELU); The hyper - prior encoding layer g he is used to encode the marginal information of the latent representation y; From left to right in the direction of the arrow, the input of the first Conv is the output of the encoding layer g e The output of the fourth Conv. The output of the first Conv is M feature maps, with a convolutional kernel of 3 and a stride of 1; Then it is input into the first RELU; The input of the second Conv is the output of the first RELU, and the output of the second Conv is N feature maps, with a convolutional kernel size of 3 and a stride of 2; Then it is input into the second RELU, and the output feature maps are used as the input of the third Conv. The convolutional kernel of the third Conv is 3 and the stride is 2, and N feature maps are input.
[0081] Hyper - prior decoding layer g hd It includes three transposed convolutional layers and two activation layers; The hyper - prior decoding layer g hd is used to recover through entropy decoding using the bit - stream z_bits - stream generated by the hyper - prior encoding layer Then use g hd to decode to generate the mean Mean and variance Variance of the Gaussian distribution model, and the Offset for semantic feature reconstruction; From right to left in the direction of the arrow, the input of the first TConv is the quantized representation with a convolutional kernel of 3, a stride of 2, and N feature maps; Then it is input into the first RELU, and the output feature maps are input into the second TConv, with a convolutional kernel of 3 and a stride of 2, and the output feature maps are N; Then it is input into the second RELU, and its output is the input of the third TConv, M feature maps, with a convolutional kernel of 3 and a stride of 1.
[0082] Step 3, set up a semantic information reconstruction module. The semantic information reconstruction module includes generating the mean and variance of the Gaussian distribution model and the bias feature map. The mean and variance of the Gaussian distribution model are used to construct a Gaussian distribution model to provide a probability model for entropy encoding and decoding. The bias feature map is used to reconstruct the lost semantic information after quantization and fuse it with the quantized representation information for image reconstruction in the decoding layer;
[0083] Step 3 of the embodiment includes generating the mean and variance of the Gaussian distribution model and reconstructing semantic information. See Figure 2 ; Conv1×1 represents a 1×1 convolutional layer, RELU represents an activation function, Abs represents an absolute - value module, D represents a difference, and ⊕ represents a sum; The input of the three Conv1×1 above is the quantized representation There is a RELU module after each Conv; the output feature map is divided into three parts, namely Mean, Variance, and Offset. Mean and Variance are the two parameters of the Gaussian model, and Offset is used to perform differencing with the quantized latent representation to generate a difference feature map; in the figure, there is an Abs module and three Conv1×1 / RELU modules at the bottom for reconstructing semantic information, and the inputs are the quantized representations and Offset; first, perform differencing on the quantized representation and Offset to obtain a difference feature map, take the absolute value, and then sequentially input it into the three Conv1×1 / RELU modules. The output feature map s is summed with the quantized representation to obtain the reconstructed semantic feature map
[0084] The features output by the hyperprior decoding module in step 2 are input into the semantic information reconstruction module, which is divided into two parts
[0085] The first part consists of three 1×1 convolutional layers and three activation layers, and is used to generate the mean Mean and variance Variance of the Gaussian distribution, as well as the bias feature Offset for reconstructing the semantic representation
[0086] The second part includes a differencing module (D), an absolute value module (Abs), and three 1×1 convolutional layers and three activation layers. First, use the Gaussian distribution module constructed in the first part to decode the bitstream y_bits_stream to obtain Then The difference feature map obtained by differencing and taking the absolute value with Offset is input into the subsequent convolutional layers. Through three convolutions and activations, the reconstructed semantic information s is obtained. Finally, s is fused with using an addition operation to obtain
[0087] Specifically, the implementation of semantic information reconstruction can adopt the following sub-steps
[0088] 1) Use the feature information generated by the hyperprior decoding layer to generate the mean, variance, and bias information of the Gaussian model through a convolutional neural network
[0089] 2) Use the generated mean and variance to construct a Gaussian distribution model, and use entropy decoding to decode the bitstream to recover the quantized latent representation
[0090] 3) Use the offset information obtained in 1) and the quantized representation information to perform a difference operation to obtain a difference feature map, and use a convolutional neural network to generate a semantic feature s;
[0091] 4) Use the semantic feature s generated in 3) and the quantized representation information to fuse and generate a new representation.
[0092] Step 4: Construct a multi-head decoder. The multi-head decoder is configured with decoders according to the number of bands of the multi-spectral image to be reconstructed. For example, for 8-band data, 8 decoders need to be set;
[0093] Step 4 of the embodiment includes channel separation and a multi-head decoder. Refer to Figure 3 ; In the figure, s represents the channel separation operation. Each decoder head has the same module structure. For example, decoder head 1 is composed of three Conv1×1 / RELU; the decoding layer g d outputs a feature map Use channel separation to divide it into n groups, and each group has the same number of feature maps; then input the n groups of feature maps into n decoder heads respectively to generate corresponding band images. For example, decoder head 1 generates the first band image, and decoder head n generates the nth band image.
[0094] Channel separation is to divide the feature map output by the decoding layer into n groups.
[0095] The multi-head decoder is composed of three convolutional layers and two activation layers. In order to capture features under different fields of view, the embodiments of the present invention preferably use three types of convolutional kernels, namely 1×1, 3×3, and 5×5; each group of feature maps is respectively input into its own decoder head to generate each spectral image.
[0096] Specifically, the construction of the multi-head decoder can adopt the following sub-steps:
[0097] 1) Use the decoding layer g d to reconstruct a feature map with the same size as the input image, denoted as
[0098] 2) Use the channel separation method to divide it into n groups, and the i-th group of feature maps is denoted as where n represents the number of bands of the multi-spectral image;
[0099] 3) Use n decoder heads to map the n groups of feature maps into a 1-channel space respectively to generate an n-band image.
[0100] Step 5: Set and train an image compression model, including training the compression model parameters on an 8-band dataset;
[0101] Set up and train an image compression model. The model training is to train the model, update parameters, and converge the model on an 8-spectrum dataset.
[0102] In specific implementation, the space-ground joint multi-spectral remote sensing image compression model set in step 2 is trained in the training set of the multi-spectral dataset constructed in step 1, and the rate-distortion accuracy of the model is verified on the validation set. A total of 300 epochs are trained, the initial learning rate is set to 0.0001, and the learning rate is reduced once every 100 epochs, and the learning rate is reduced to 1 / 10 of the previous one. When training the target model, an Ubuntu 20.04 LTS system is required, the software environment requires Pytorch 1.7 and Python 3.6, and the main computing platform is an NVIDIA RTX2080Ti graphics card. The loss function of the compression model adopts rate-distortion optimization, and the formula is as follows:
[0103] Loss = R + λ × D (1)
[0104] Among them, R represents the estimated bit rate, D represents the distortion degree, and λ represents the balance coefficient, which is used to balance the influence of R and D. The larger λ is, the greater the influence of D on model training, and vice versa, the greater the influence of R on model training.
[0105] The calculation formula of R is as follows:
[0106]
[0107] Among them represents the estimation of the bitstream y_bits_stream, represents the Gaussian distribution probability model of represents the estimation of the bitstream z_bits_stream, represents the probability distribution model of represents the quantized latent representation, represents the z quantized representation output by the hyperprior coding layer g he and p x is the probability density model of the latent representation of the input image x.
[0108] The calculation formula of D is as follows:
[0109]
[0110] Among them, x represents the input multi-spectral image, represents the compressed multi-spectral image.
[0111] Step 6, space-ground joint multi-spectral remote sensing image compression. The space-ground joint multi-spectral remote sensing image compression divides the compression model into two parts. The first part is the on-orbit compression part, including an encoding layer, a hyperprior encoding layer, a hyperprior decoding layer, and entropy encoding. The second part is the ground reconstruction part, including entropy decoding, semantic information reconstruction, a decoding layer, and a multi-head decoder.
[0112] Specifically, according to the task requirements of space-ground joint compression, in the embodiment, the space-ground joint compression is divided into two parts as follows:
[0113] See Figure 4 , where Q in the figure represents the quantization operation, and AE represents the entropy encoding operation. The first part is on-board encoding, which is divided into the following steps:
[0114] 1) Use the encoding layer g e to map the multi-spectral image x to the latent representation layer, obtaining the latent representation y;
[0115] 2) Quantize the y obtained in 1) to obtain the quantized latent representation
[0116] 3) Use the hyperprior encoding layer g he to further encode and quantize the latent representation y in 1), obtaining the representation z and the quantized representation
[0117] 4) Perform entropy encoding on the quantized representation to generate the code stream z_bits-stream. The probability model of its entropy encoding is the probability density function obtained through offline learning. Each element in the quantized representation is entropy encoded according to the probability density function;
[0118] 5) Use the hyperprior decoding layer g hd to decode the quantized representation obtained in 3) and output two parameters, the mean Mean and the variance Variance of the Gaussian model;
[0119] 6) Construct the Gaussian distribution probability model according to the Mean and Variance generated in 5) to perform entropy encoding on the quantized latent representation in 2) and generate the code stream y_bits-stream;
[0120] 7) Send the code stream z_bits-stream and the code stream y_bits-stream in 4) to the space-ground transmission link and download them to the ground.
[0121] See Figure 5, in the figure, AD represents entropy decoding, and bitstreams represent the binary data streams transmitted from the satellite to the ground. represents the quantized latent representation. represents the hyperprior encoding layer g he the quantized representation of the output z. is the reconstructed semantic feature map. is the decoding layer g d the output feature map; The second part is ground decoding, which is divided into the following steps:
[0122] 1) The ground receives the code stream z_bits-stream and the code stream y_bits-stream, and performs entropy decoding on the z_bits-stream to obtain
[0123] 2) Input the obtained in 1) into the hyperprior decoding layer g hd to generate
[0124] 3) Input into the semantic feature reconstruction module to generate the mean Mean and variance Variance of the Gaussian distribution for the bias feature map Offset; construct a Gaussian distribution model according to the generated Mean and Variance, and perform entropy decoding on the code stream y_bits-stream to obtain the quantized representation
[0125] 4) Then use together with Offset for semantic feature reconstruction and generate
[0126] 5) Input the generated in 4) into g d for decoding to obtain
[0127] 6) Input the generated in 5) into the multi-head decoder to reconstruct the multi-spectral image.
[0128] Through the above step 6, high-magnification satellite-ground joint image compression can be effectively achieved.
[0129] To facilitate understanding the technical effects of the present invention, the application comparison between the present invention and the traditional method is provided for reference in Figure 6 and Figure 7 . Figure 6It shows that the method proposed by the present invention can achieve higher peak signal-to-noise ratio (PSNR) and better multi-scale similarity (MS-SSIM) at lower bit rates (bpp). According to the results of manual visual inspection, it can also be seen that a satellite-ground joint multi-spectral remote sensing image compression method proposed by the present invention has better compression performance. Figure 7 The rate-distortion curves of three methods are shown. The more to the upper left the curve is, the better the performance. It can be seen that the method proposed by the present invention has stronger compression performance compared with the conventional JPEG2000 compression method.
[0130] In specific implementation, the method proposed by the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. The system device for implementing the method, such as a computer-readable storage medium storing the corresponding computer program of the technical solution of the present invention and a computer device including running the corresponding computer program, should also be within the protection scope of the present invention.
[0131] In some possible embodiments, a satellite-ground joint multi-spectral remote sensing image compression system is provided, including the following modules.
[0132] The first module is used to construct a multi-spectral remote sensing image data set.
[0133] The second module is used to set up a satellite-ground joint multi-spectral remote sensing image compression model. The satellite-ground joint multi-spectral remote sensing image compression model is implemented based on a variational autoencoder and includes an encoding layer g e , a hyperprior encoder g he , a hyperprior decoder g hd and a decoding layer g d ; the encoding layer is used to map the input image to a latent representation space to support entropy encoding and entropy decoding of the quantized representation information; the hyperprior encoding and decoding modules are used to model the latent representation by using a Gaussian mixture model through learning the latent representation information output by the encoding layer; the decoding layer is used to reconstruct the original image according to the representation feature map after entropy decoding.
[0134] The third module is used to set up a semantic information reconstruction module. The semantic information reconstruction module is used to generate the mean, variance and bias feature map of a Gaussian distribution model. The mean and variance of the Gaussian distribution model are used to construct a Gaussian distribution model to provide a probability model for entropy encoding and decoding. The bias feature map is used to reconstruct the lost semantic information after quantization and fuse it with the quantized representation information for image reconstruction by the decoding layer.
[0135] The fourth module is used to construct a multi-head decoder, including setting up corresponding decoders according to the number of bands of the multi-spectral image to be reconstructed.
[0136] The fifth module is used to train the space-ground joint multi-spectral remote sensing image compression model;
[0137] The sixth module is used to perform space-ground joint multi-spectral remote sensing image compression. The compression is divided into two parts. The first part is the on-orbit compression part, including an encoding layer, a hyperprior encoding layer, a hyperprior decoding layer, and entropy encoding. The second part is the ground reconstruction part, including entropy decoding, semantic information reconstruction, a decoding layer, and a multi-head decoder.
[0138] In some possible embodiments, a space-ground joint multi-spectral remote sensing image compression system is provided, including a readable storage medium. A computer program is stored on the readable storage medium. When the computer program is executed, the above-mentioned task-oriented remote sensing image compression method is implemented.
[0139] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A space-ground collaborative multi-spectral remote sensing image compression method, characterized in that It includes the following steps: Step 1: Construct a multi-spectral remote sensing image dataset; Step 2, set up a satellite-ground joint multi-spectral remote sensing image compression model, which is based on a variational autoencoder and includes an encoding layer g e , a hyperprior encoder g he , a hyperprior decoder g hd and a decoding layer g d ; the encoding layer is used to map the input image to a latent representation space to support entropy encoding and entropy decoding of the quantized representation information; the hyperprior encoding and decoding modules are used to model the latent representation by learning the latent representation information output by the encoding layer and using a Gaussian mixture model; the decoding layer is used to reconstruct the original image according to the representation feature map after entropy decoding; Step 3: Set a semantic information reconstruction module, which is used to generate the mean, variance and bias feature map of the Gaussian distribution model. The mean and variance of the Gaussian distribution model are used to construct the Gaussian distribution model to provide a probability model for entropy coding and decoding. The bias feature map is used to reconstruct the lost semantic information after quantization and fuse it with the quantized representation information for image reconstruction in the decoding layer; Step 4: Construct a multi-head decoder, including setting corresponding decoders according to the number of bands of the multi-spectral image to be reconstructed; Step 5: Train a satellite-ground joint multi-spectral remote sensing image compression model; Step 6: Perform satellite-ground joint multi-spectral remote sensing image compression. The compression is divided into two parts. The first part is the on-orbit compression part, including an encoding layer, a hyperprior encoding layer, a hyperprior decoding layer and entropy coding; The second part is the ground reconstruction part, including entropy decoding, semantic information reconstruction, a decoding layer and a multi-head decoder.
2. The method for compressing satellite-ground integrated multi-spectral remote sensing images according to claim 1, wherein: In Step 1, when constructing the multi-spectral remote sensing image dataset, a high-resolution multi-mode satellite data source is used to construct an 8-band high-resolution multi-mode satellite dataset.
3. The method for compressing satellite-ground integrated multi-spectral remote sensing images according to claim 1, wherein: In the satellite-ground joint multi-spectral remote sensing image compression model, the hyperprior encoding and decoding modules are implemented by using a convolutional neural network to actively learn the Gaussian distribution information of each pixel in the representation feature map, providing an accurate probability distribution for each element for entropy coding and decoding, so as to better encode the quantized representation feature map.
4. The method for compressing satellite-ground combined multi-spectral remote sensing images according to claim 1, wherein: In Step 3, the semantic information reconstruction is implemented as follows 1) Feature information generated by the hyperprior decoding layer Generate the mean, variance, and bias information of the Gaussian model through a convolutional neural network; 2) Construct a Gaussian distribution model using the generated mean and variance, and use entropy decoding to decode the bitstream to recover the quantized representation information 3) Using the bias information obtained in 1) and the quantized representation information to obtain a differential feature map through difference, and generating a semantic feature s using a convolutional neural network; 4) Generate a new representation by fusing the semantic feature s generated in 3) with the quantized representation information 5. The method for compressing satellite-ground integrated multi-spectral remote sensing images according to claim 1, wherein: In Step 4, the construction of the multi-head decoder is implemented as follows 1) Use the decoding layer g d to reconstruct a feature map with the same size as the input image, denoted as 2) Use the channel separation method to separate into n groups, and the feature map of the i-th group is denoted as where n represents the number of bands of the multispectral image; 3) Use n decoder heads to map n groups of feature maps to the space of 1 channel respectively to generate an n-band image.
6. The method for compressing satellite-ground integrated multi-spectral remote sensing images according to claim 1 or 2 or 3 or 4 or 5, characterized in that: The on-board compression implementation process includes the following steps 1) Use the encoding layer g e Map the multi-spectral image x to the latent representation layer to obtain the latent representation y; 2) Quantize the y obtained in 1) to obtain a quantized latent representation 3) Using the hyperprior coding layer g he to further encode and quantize the latent representation y in 1) to obtain the representation z and the quantized representation 4) Entropy encoding is performed on the quantized representation to generate a bitstream z_bits-stream. The probability model for entropy encoding is a probability density function obtained through offline learning. Entropy encoding is performed on each element in the quantized representation according to the probability density function; 5) Use the hyperprior decoding layer g hd to decode the quantized representation obtained in 3) and output two parameters, the mean Mean and variance Variance of the Gaussian model; 6) Construct a Gaussian distribution probability model based on the Mean and Variance generated in 5), and perform entropy coding on the quantized latent representation in 2) to generate a bitstream y_bits-stream; 7) Send the code stream z_bits-stream and the code stream y_bits-stream in 4) to the satellite-ground transmission link and download them to the ground.
7. A satellite-ground integrated multispectral remote sensing image compression method according to claim 6, characterized in that: The ground decoding includes the following steps 1) The ground receives the bitstreams z_bits-stream and y_bits-stream, and performs entropy decoding on the z_bits-stream to obtain 2) Input the result obtained in 1) into the hyperprior decoding layer g hd to generate 3) Input into the semantic feature reconstruction module to generate the mean Mean and variance Variance of the Gaussian distribution for the bias feature map Offset; construct a Gaussian distribution model based on the generated Mean and Variance, and perform entropy decoding on the bitstream y_bits-stream to obtain the quantized representation 4) Then, utilize to perform semantic feature reconstruction together with Offset and generate 5) Input the generated in 4) into g d for decoding to obtain 6) Input the generated in 5) into the multi-head decoder to reconstruct the multi-spectral image.
8. A space-ground integrated multi-spectral remote sensing image compression system, characterized in that: It is used to implement a satellite-ground joint multi-spectral remote sensing image compression method as described in any one of claims 1-7.
9. The space-ground integrated multi-spectral remote sensing image compression system according to claim 8, characterized in that: It includes the following modules. The first module is used to construct a multi-spectral remote sensing image dataset; The second module is used to set up a satellite-ground joint multi-spectral remote sensing image compression model, which is implemented based on a variational autoencoder and includes an encoding layer g e , a hyperprior encoder g he , a hyperprior decoder g hd and a decoding layer g d ; the encoding layer is used to map the input image to a latent representation space to support entropy encoding and entropy decoding of the quantized representation information; the hyperprior encoding and decoding modules are used to model the latent representation by using a Gaussian mixture model through learning the latent representation information output by the encoding layer; the decoding layer is used to reconstruct the original image according to the representation feature map after entropy decoding; The third module is used to set a semantic information reconstruction module, which is used to generate the mean, variance and bias feature map of the Gaussian distribution model. The mean and variance of the Gaussian distribution model are used to construct the Gaussian distribution model to provide a probability model for entropy coding and decoding. The bias feature map is used to reconstruct the lost semantic information after quantization and fuse it with the quantized representation information for image reconstruction in the decoding layer; The fourth module is used to construct a multi-head decoder, including setting corresponding decoders according to the number of bands of the multi-spectral image to be reconstructed; The fifth module is used to train a satellite-ground joint multi-spectral remote sensing image compression model; The sixth module is used to perform satellite-ground joint multi-spectral remote sensing image compression. The compression is divided into two parts. The first part is the on-orbit compression part, including an encoding layer, a hyperprior encoding layer, a hyperprior decoding layer and entropy coding; The second part is the ground reconstruction part, including entropy decoding, semantic information reconstruction, a decoding layer and a multi-head decoder.
10. The space-ground integrated multi-spectral remote sensing image compression system according to claim 8, characterized in that: It includes a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a space-ground integrated multi-spectral remote sensing image compression method according to any one of claims 1-7.
Citation Information
Patent Citations
Method and device for image compression and decompression based on variational auto-encoder
CN113497938A
Task-oriented remote sensing image compression method and system
CN115131673A