Semantic scalable image coding method, system, device and storage medium
By building a scalable image encoder containing a compressed model and a cross-layer context model, the joint compression of images and features is realized, solving the problems of inconsistent image compression and semantic analysis accuracy and information redundancy in the prior art, and improving coding efficiency.
Patent Information
- Application Number
- CN202111134977.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-09-27
AI Technical Summary
While ensuring the quality of human visual images, the existing image compression scheme fails to effectively consider the accuracy of semantic analysis tasks, and there is information redundancy when compressing images and features separately, resulting in low encoding efficiency.
A scalable image encoder containing a compressed model and a cross-layer context model is constructed. By combining the compression of the input image and the image features output by the feature extractor, the compressed features are used as a priori to estimate the hidden representation probability distribution of the input image, and encoded by the entropy encoding model.
Taking into account image quality and semantic analysis task accuracy for human vision, reduce information redundancy and improve coding efficiency.
Smart Images

Figure CN115880379B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image compression technology, and in particular to a semantic scalable image coding method, system, device and storage medium. Background Art
[0002] Image compression aims to represent the original pixel space in a more compact form to reduce the storage and transmission overhead of the image. Currently, the widely used traditional image compression frameworks include JPEG, JPEG2000, BPG, etc. In recent years, with the rapid development of deep learning technology, end-to-end image compression frameworks based on deep learning have also achieved great success. The end-to-end image compression framework mainly consists of three parts: encoder, decoder and entropy coding module. The encoder transforms the image into a more compact latent representation. After the latent representation is quantized, it is losslessly compressed by the entropy coding module to obtain a bit stream for transmission. At the receiving end, the bit stream is decompressed to obtain the quantized latent representation. The decoder transforms the quantized latent representation back to the pixel space to reconstruct the input image. The goal of image compression is to minimize the bit rate required after encoding while ensuring the quality of the reconstructed image.
[0003] In end-to-end image compression, in order to solve the problem of non-differentiable quantization, a common method is to use additive uniformly distributed noise to simulate the quantization process during training. An existing model framework is Figure 1 As shown. Among them, the encoder g a It is composed of multiple layers of convolution, which transforms the image signal x into the latent representation y to be encoded. In order to improve the efficiency of entropy coding, existing work introduces an additional variable z as a super prior, and assumes that when z is given, each element of y is conditionally independent, which can reduce the spatial redundancy when encoding y. For convenience, the conditional probability of y given z is usually modeled as a Gaussian distribution. In specific implementation, a transformation h is used a Based on y, we get z, and then through another transformation h s , based on z, generate the relevant parameters of the Gaussian distribution P(y|z), namely the mean and variance; Figure 1 In the above figure, U|Q have the same meaning. Q stands for quantization, which is used during testing, and U stands for adding uniform noise, which is used to simulate quantization during training. represents x, y, and z after adding uniform noise. represents x, y, z after adding uniform noise; g s Denotes the decoder. The final optimized loss is R+λD, where R is the sum of the bit rates of variables y and z, D is the error of image reconstruction, and λ is the weighting coefficient.
[0004] However, existing image compression schemes only consider the image quality for human vision, and do not consider the accuracy of semantic analysis tasks for compressed images. Distortion at low bit rates will seriously reduce the accuracy of semantic tasks. Moreover, since the image-based semantic analysis process (i.e., feature extraction process) often requires a lot of computing resources, simply compressing the image means that the cloud needs to bear a large computing load, which will hinder the deployment of deep learning models in practice.
[0005] Deep neural networks can be viewed as stacked feature extractors. The features extracted from images have rich semantic information and can be used for different visual task analysis, such as image classification, object detection, etc. In cloud-based visual analysis systems, image and video data are collected at the front end, and the semantic analysis process is completed in the cloud. Traditional data communication systems usually adopt the "compress first, then analyze" mode, that is, the front end completes the image acquisition, compression and encoding, and the cloud is responsible for semantic analysis of the decoded image. However, since feature extraction based on deep learning usually requires a large amount of computing resources, the cloud is difficult to bear the huge computing overhead brought by the increasing scale of data. In order to reduce the bandwidth required for transmission and reduce the computing overhead of the cloud, another "analysis first, then compression" mode can be adopted. The front end completes the image acquisition and feature extraction, and the extracted features are compressed and transmitted to the cloud. The cloud directly performs semantic analysis based on the decoded features. Unlike image compression, the technology related to feature compression is not very mature and needs further research.
[0006] In the prior art, most methods for compressing deep neural network feature maps are based on image or video compression methods. Existing work regards a feature map of dimension C as a video sequence of C frames, and compresses it frame by frame using a video encoder (such as HEVC, etc.). Since HEVC requires the length and width of the input to be an integer multiple of 8, and can only process inputs of 8 bits or higher, the feature map needs to be preprocessed before compression, including zero padding on the edges and scaling the value range to [0, 255] before quantizing it to an integer.
[0007] However, since the size of feature maps is usually relatively small relative to the input image, and the spatial correlation of feature maps is also weak, the compression efficiency of using video encoders such as HEVC for each dimension of feature maps alone is low. On the other hand, in many cases, machine vision cannot completely replace human reasoning and decision-making. However, due to the loss of a large amount of information in the feature extraction process, it is difficult to reconstruct the input image with high quality based on the feature map, which also limits the application scenarios of feature compression. In addition, if feature maps and images of different levels are compressed at the same time, the existing technology does not consider the features of adjacent layers and the correlation between features and images, so there is a lot of information redundancy, resulting in insufficient coding efficiency. Summary of the invention
[0008] The purpose of the present invention is to provide a semantically scalable image coding method, system, device and storage medium, which can take into account both the image quality for human vision and the accuracy of semantic analysis tasks, and improve the coding efficiency.
[0009] The objective of the present invention is achieved through the following technical solutions:
[0010] A semantically scalable image coding method, comprising:
[0011] Build a scalable image encoder that includes a compression model and a cross-layer context model;
[0012] The scalable image encoder is used to jointly compress the input image and the image features output by the feature extractor, including: using the scalable image encoder to compress the image features, using the compressed image features as a priori, estimating the probability distribution of the implicit representation of the input image output by the feature encoder in the compression model, inputting the compressed image features into a cross-layer context model, and estimating the parameters of the probability distribution; using the parameters of the probability distribution by the entropy coding model in the compression model to entropy encode the input image, and then obtaining the encoded image through the decoder in the compression model.
[0013] A semantically scalable image coding system, comprising:
[0014] A model building unit for a scalable image encoder including a compression model and a cross-layer context model;
[0015] The joint compression unit is used to use the scalable image encoder to jointly compress the input image and the image features output by the feature extractor, including: using the scalable image encoder to compress the image features, using the compressed image features as a priori, estimating the probability distribution of the implicit representation of the input image output by the feature encoder in the compression model, inputting the compressed image features into the cross-layer context model, and estimating the parameters of the probability distribution; using the entropy coding model in the compression model to entropy encode the input image using the parameters of the probability distribution, and then obtaining the encoded image through the decoder in the compression model.
[0016] A processing device, comprising: one or more processors; a memory for storing one or more programs;
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0018] A readable storage medium stores a computer program, which implements the above method when the computer program is executed by a processor.
[0019] It can be seen from the technical solution provided by the present invention that by jointly performing feature and image compression, the problems existing in the prior art of separately compressing images and compressing features are solved, which not only takes into account the image quality for human vision and the accuracy of semantic analysis tasks, but also reduces information redundancy and bit rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0021] Figure 1 A schematic diagram of an existing image coding framework provided as the background technology of the present invention;
[0022] Figure 2 A schematic diagram of a semantically scalable image coding method provided by an embodiment of the present invention;
[0023] Figure 3 A flowchart of the solution of embodiment 1 provided for the embodiment of the present invention;
[0024] Figure 4 A schematic diagram of image compression performance comparison results provided by an embodiment of the present invention;
[0025] Figure 5 A schematic diagram of a semantically scalable image coding system provided by an embodiment of the present invention;
[0026] Figure 6 A schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0028] First, the terms that may be used in this article are explained as follows:
[0029] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.
[0030] The following is a detailed description of a semantic scalable image coding method provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professionals in the field. If no specific conditions are specified in the embodiments of the present invention, the conventional conditions in the field or the conditions recommended by the manufacturer shall be followed.
[0031] The semantic scalable image coding method provided by the embodiment of the present invention mainly comprises the following steps:
[0032] 1. Build a scalable image encoder consisting of a compression model and a cross-layer context model.
[0033] In an embodiment of the present invention, the cross-layer context model may include three convolutional layers arranged in sequence, and an activation function Leaky ReLU is provided between adjacent convolutional layers.
[0034] In an embodiment of the present invention, the compression model includes: a feature encoder, a feature decoder and an entropy coding model; wherein the encoder may include three convolutional layers arranged in sequence, and a nonlinear normalization layer GDN (generalized divisive normalization) is provided between adjacent convolutional layers; arithmetic coding and arithmetic decoding are performed in sequence in the entropy coding model; the decoder may include three convolutional layers arranged in sequence, and an IGDN layer is provided between adjacent convolutional layers, and the IGDN layer performs the inverse process of the GDN layer; the entropy coding model can use an existing model structure, for example, a decomposable probability model.
[0035] 2. Use the scalable image encoder to jointly compress the input image and the image features output by the feature extractor, including: using the scalable image encoder to compress the image features, using the compressed image features as a priori, estimating the probability distribution of the implicit representation of the input image output by the feature encoder in the compression model, inputting the compressed image features into a cross-layer context model, and estimating the parameters of the probability distribution; using the entropy coding model in the compression model to entropy encode the input image using the parameters of the probability distribution, and then obtaining the encoded image through a decoder.
[0036] In the embodiment of the present invention, when the input image and image features are jointly compressed, the image features involved can be image features of one level or image features of multiple levels. 1) When the image features only contain image features of one level, the image features are directly compressed using the compression model, and the compressed image features are used as a priori for encoding the input image; the image feature compression process is as follows: the input image features are passed through the encoder and the entropy encoder to obtain a quantized latent representation, which is reconstructed through the feature decoder to obtain the compressed image features. 2) When the image features contain image features of multiple levels, the compression model is used to compress the image features of the highest level N (in the same way as in the aforementioned case 1); then the images are compressed from high to low. When compressing the image features of the i-th level, the compressed image features of the j+1th level are used as a priori to estimate the probability distribution of the latent representation of the image features of the j-th level output by the feature encoder in the compression model, and the compressed image features of the j+1th level are input into the cross-layer context model to estimate the parameters of the probability distribution; the entropy coding model in the compression model entropy codes the image feature input image of the j-th level using the parameters of the probability distribution, and then the compressed image features of the j-th level are reconstructed by the decoder; i=1,...,N-1; the compressed feature image of the 1st level will be used as a priori for encoding the input image.
[0037] Based on the above scheme, during training, the compression model and the cross-layer context model are jointly trained to optimize the rate-distortion loss function R+λD, where R is the sum of the bit rates of the input image and image features, D is the reconstruction error of the input image and image features, and λ is the weighting coefficient.
[0038] like Figure 2As shown, it is a schematic diagram of the semantic scalable image coding method provided by the present invention. The coding method adopts a new image coding method, which solves the problems existing in the existing single compression of images and compression features by jointly performing image features and image compression. As mentioned above, the image features involved in the joint compression may be multiple levels. Since the low-level features are extracted from the image signal, and the high-level features are extracted from the low-level features, there is a strong correlation between the features of adjacent layers and the image. There is a lot of information redundancy in the single compression. In the above scheme of the present invention, the high-level features, low-level features and input images are progressively jointly compressed to reduce information redundancy and improve coding efficiency. That is, the high-level image features are used as input, as a cross-layer prior, and the conditional probability distribution of the low-level features is predicted based on the prior. When encoding the low-level features, the entropy encoder is designed based on the predicted conditional probability distribution to reduce information redundancy between layers. Similarly, when compressing the image signal, the cross-layer context model uses the low-level features as input, predicts the conditional probability distribution of the image signal, and then performs entropy coding based on the predicted probability distribution.
[0039] like Figure 2 As shown, the scenario in which the joint coding result is applicable is shown. Figure 2 The various semantic tasks in can be selected according to actual conditions. For example, semantic task C can be image recognition, semantic task B can be target detection, and semantic task A can be target segmentation. Through the joint compression provided by the present invention, a scalable bit stream can be obtained, which is the output of the entropy encoder model and is a result obtained after encoding in the manner described above. Among them, the bottom layer is the basic layer, which is a bit stream of high-level image features. It is the result of high-level image features obtained by the entropy coding model and contains coarse-grained semantic information; the penultimate layer to the top layer are the first, second, and third enhancement layers, respectively. The first enhancement layer is a higher-level feature, the second enhancement layer is a low-level feature, and the third enhancement layer is an image layer. Similarly, the corresponding image features or the input image are obtained by the entropy coding model. According to different task requirements, the corresponding bit stream can be selected to reconstruct the corresponding features after passing through the decoder of the compression model to complete the corresponding task.
[0040] The high level, higher level, and low level here are relative concepts. High and low correspond to the depth of the feature extractor in the deep neural network. The specific division method can be adjusted by the user according to actual conditions or experience. For example, high-level features can correspond to image features output by the last feature extractor, low-level features can correspond to image features output by the first feature extractor, and image features output by the remaining feature extractors can be called higher-level features.
[0041] It should be noted that: 1) The cross-layer context model provided in the embodiment of the present invention is a universal compression model that can be used in different feature extraction networks to reduce information redundancy between layers, such as image classification networks, target detection networks, image segmentation networks, etc. 2) The cross-layer context model provided in the embodiment of the present invention is combined with other end-to-end compression or entropy coding models, such as hyper-prior models, autoregressive context models, etc.
[0042] For ease of understanding, three embodiments are described below.
[0043] Embodiment 1
[0044] In this embodiment, the proposed cross-layer context model is combined with an image classification network based on a residual network (ResNet), and the multi-layer image features of the convolutional neural network for image classification and the input image are compressed at the same time. In this embodiment, the residual network consists of four stages, so N is taken as 4, such as Figure 3 As shown, the main steps are as follows:
[0045] Step 1: Train the image classification network and perform the highest level image feature f N The feature f after pooling u , plus the bit rate constraint, the classification accuracy and the compression performance of the highest layer features are jointly optimized during training to obtain the compressed features This feature can be directly used to complete the classification task through a linear classifier. The bit rate constraint can be implemented through a decomposable probabilistic model.
[0046] Step 2: Compress the highest-level image feature f4 in the intermediate layer features. The compressed image feature is recorded as In the compression branch, En N Denotes the encoder, De N represents a decoder, N is a serial number (in this embodiment, N=4), Q represents quantization, AE represents formula encoding, and AD represents formula decoding. The feature encoder and decoder structures used are as follows Figure 3 As shown on the right, the entropy coding model uses a decomposable probability model.
[0047] Step 3: Compress the lower image features f in the intermediate layer features from high to low j , j = 1, 2, 3. When compressing the j-th layer image features, the compressed j+1-th layer image features As a priori, estimate the latent representation z of the j-th layer image features j After compression, the j+1th layer image features Assuming that the conditional probability distribution conforms to the Gaussian distribution, the cross-layer context model (CLCM) converts the compressed j+1 layer image features As input, the mean μ and variance σ of the Gaussian distribution are predicted, and the predicted parameters are used in the entropy coding process of the j-th layer image features; the cross-layer context model structure is as follows Figure 3 As shown on the right; in the three dotted boxes on the right, the parameters of the convolutional layer (Conv) are A×B×C: A is the number of channels output by the convolutional layer, B and C are the sizes of the convolutional kernels of the convolutional layer. For example, the parameters of the first convolutional layer of the cross-layer context model, K1×3×3, indicate that the number of output channels of the convolutional layer is K1 (i.e., there are K1 convolutional kernels), and the size of the convolutional kernel is 3×3.
[0048] Step 4: Compress the shallowest image features compressed in step 3 As a priori, estimate the probability distribution of the input image latent representation y, which is used to encode the input image x. Similarly, assuming that the compressed image feature The conditional probability of the time-hidden representation y conforms to the Gaussian distribution, and the compressed image features Input the cross-layer context model to predict the mean μ and variance σ of the Gaussian distribution, which is used for entropy coding of the input image. Finally, the decoder completes the coding, and the coding result is recorded as
[0049] Step 5: Jointly train the compression model and cross-layer context model from step 2 to step 4 to optimize the rate-distortion loss function R+λD, where R is the sum of the bit rates of the image signal and each layer feature, D is the reconstruction error of the image signal and each layer feature, and λ is the weighting coefficient. After the training is completed, the final scalable image encoder is obtained.
[0050] Embodiment 2
[0051] This embodiment combines the proposed cross-layer context model with the image classification network. Different from the first embodiment, this embodiment considers compressing only the features and image signals of the middle single layer.
[0052] The main steps of this embodiment are as follows:
[0053] Step 1: Compress the intermediate layer image feature f2, and the compression result is recorded as The feature encoder and decoder structures used are as follows Figure 3 As shown on the right, the entropy coding model uses a decomposable probability model.
[0054] Step 2: Compress the image features compressed in the first step As a priori, the probability distribution of the image signal hidden representation y is estimated to encode the image signal. Assuming that the compressed image feature The conditional probability of the time-hidden representation y conforms to the Gaussian distribution, which compresses the image features The cross-layer context model is input to predict the mean μ and variance σ of the Gaussian distribution, which is used for entropy coding of the input image and finally the encoding is completed by the decoder.
[0055] Step 3: Jointly train the compression model and cross-layer context model of the first two steps to optimize the rate-distortion loss function R+λD, where R is the sum of the bit rates of the image signal and each layer feature, D is the reconstruction error of the image signal and each layer feature, and λ is the weighting coefficient. After the training is completed, the final scalable image encoder is obtained.
[0056] Embodiment 3
[0057] This embodiment combines the proposed cross-layer context model with a feature pyramid-based object detection network to simultaneously compress image features of different scales used for object detection and the input image.
[0058] The specific steps of this embodiment are as follows:
[0059] Step 1: Compress the highest level image feature p5 in the feature pyramid. The compression result is recorded as The feature encoder and decoder structures used are as follows Figure 2 As shown on the right, the entropy coding model uses a decomposable probability model.
[0060] Step 2: Compress the lower features p in the feature pyramid from high to low j , j = 2, 3, 4. When compressing the jth layer image features, the compressed j+1th layer image features As a priori, estimate the latent representation z of the j-th layer image features j Given the j+1th layer image features after compression Assuming that the conditional probability distribution conforms to the Gaussian distribution, the cross-layer context model (CLCM) converts the compressed j+1 layer image features As input, the mean μ and variance σ of the Gaussian distribution are predicted, and the predicted parameters are used in the entropy encoding process of the j-th layer image features.
[0061] Step 3: Compress the shallowest features in the third step As a priori, the probability distribution of the latent representation y of the image signal is estimated and used to encode the input image. Similarly, assuming that the conditional probability of the latent representation y given the feature p2 conforms to the Gaussian distribution, the compressed image feature The cross-layer context model is input to predict the mean μ and variance σ of the Gaussian distribution, which is used for entropy coding of the input image and finally the encoding is completed by the decoder.
[0062] Step 4: Jointly train the compression model and cross-layer context model from step 2 to step 4 to optimize the rate-distortion loss function R+λD, where R is the sum of the bit rates of the image signal and each layer feature, D is the reconstruction error of the image signal and each layer feature, and λ is the weighting coefficient. After the training is completed, the final scalable image encoder is obtained.
[0063] In order to illustrate the effect of the above scheme of the present invention, the following two comparative experiments are used for illustration.
[0064] 1) Compared with the method of not using the cross-layer context model and compressing each layer feature separately, the scheme of embodiment 1 reduces the bit rate of image features f1 and f2 by 23.6% and 11.0% respectively on the CUB-200-2011 dataset. On the FGVC-Aircraft dataset, the bit rate of image features f1 and f2 is reduced by 21.9% and 10.7% respectively.
[0065] 2) The compression performance of the input image is significantly improved compared to the method without using the cross-layer context model. The results are as follows Figure 4 As shown in the figure, the left part is the experimental results of the CUB-200-2011 dataset; the right part is the experimental results on the FGVC-Aircraft dataset. The f1 and f2 in the brackets refer to the cross-layer prior when the features f1 and f2 are used as compressed images.
[0066] Another embodiment of the present invention further provides a semantically scalable image coding system, which is mainly used to implement the method provided in the above embodiment, such as Figure 5 As shown, the system mainly includes:
[0067] A model building unit for a scalable image encoder including a compression model and a cross-layer context model;
[0068] The joint compression unit is used to use the scalable image encoder to jointly compress the input image and the image features output by the feature extractor, including: using the scalable image encoder to compress the image features, using the compressed image features as a priori, estimating the probability distribution of the implicit representation of the input image output by the feature encoder in the compression model, inputting the compressed image features into the cross-layer context model, and estimating the parameters of the probability distribution; using the entropy coding model in the compression model to entropy encode the input image using the parameters of the probability distribution, and then obtaining the encoded image through the decoder in the compression model.
[0069] Another embodiment of the present invention further provides a processing device, such as Figure 6 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the aforementioned embodiments.
[0070] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0071] In the embodiment of the present invention, the specific types of the memory, input device and output device are not limited; for example:
[0072] The input device may be a touch screen, an image acquisition device, a physical button or a mouse, etc.;
[0073] The output device may be a display terminal;
[0074] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0075] Another embodiment of the present invention further provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0076] In the embodiment of the present invention, the readable storage medium is a computer-readable storage medium and can be set in the aforementioned processing device, for example, as a memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk, etc., which can store program codes.
[0077] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A semantically scalable image coding method, characterized in that: include: Build a scalable image encoder that includes a compression model and a cross-layer context model; The compression model includes: a feature encoder, a feature decoder and an entropy coding model; during training, the compression model and the cross-layer context model are jointly trained to optimize the rate-distortion loss function R+λD, where R is the sum of the bit rates of the input image and the image features, D is the reconstruction error of the input image and the image features, and λ is a weighting coefficient; The scalable image encoder is used to jointly compress the input image and the image features output by the feature extractor, including: using the scalable image encoder to compress the image features, using the compressed image features as a priori, estimating the probability distribution of the implicit representation of the input image output by the feature encoder in the compression model, inputting the compressed image features into a cross-layer context model, and estimating the parameters of the probability distribution; using the parameters of the probability distribution by the entropy coding model in the compression model to entropy encode the input image, and then obtaining the encoded image through the feature decoder in the compression model.
2. A semantically scalable image coding method according to claim 1, characterized in that: The compressing the image features by using a scalable image encoder comprises: When the image features only contain one level of image features, the compression model is used to compress the image features.
3. A semantically scalable image coding method according to claim 1, characterized in that: The compressing the image features by using a scalable image encoder comprises: When the image features contain multiple levels of image features, the compression model is used to compress the image features of the highest layer N; then the images are compressed from high to low. When compressing the image features of the jth layer, the compressed image features of the j+1th layer are used as priors to estimate the probability distribution of the latent representation of the image features of the jth layer output by the feature encoder in the compression model, and the compressed image features of the j+1th layer are input into the cross-layer context model to estimate the parameters of the probability distribution; the entropy coding model in the compression model uses the parameters of the probability distribution to entropy encode the image feature input image of the jth layer, and then the compressed image features of the jth layer are reconstructed by the feature decoder; j=1,…,N-1; Among them, the compressed feature image of the first layer will be used as a priori for encoding the input image.
4. A semantically scalable image coding method according to claim 1, characterized in that: The cross-layer context model includes three convolutional layers arranged in sequence, and an activation function Leaky ReLU is arranged between adjacent convolutional layers.
5. A semantically scalable image coding method according to claim 1, characterized in that: The feature encoder in the compression model includes three convolutional layers arranged in sequence, and a nonlinear normalization layer GDN is arranged between adjacent convolutional layers.
6. A semantically scalable image coding method according to claim 1, characterized in that: The feature decoder in the compression model includes three convolutional layers arranged in sequence, and an IGDN layer is arranged between adjacent convolutional layers.
7. A semantically scalable image coding system, characterized in that include: A model building unit for a scalable image encoder including a compression model and a cross-layer context model; The compression model includes: a feature encoder, a feature decoder and an entropy coding model; during training, the compression model and the cross-layer context model are jointly trained to optimize the rate-distortion loss function R+λD, where R is the sum of the bit rates of the input image and the image features, D is the reconstruction error of the input image and the image features, and λ is a weighting coefficient; A joint compression unit is used to use the scalable image encoder to jointly compress the input image and the image features output by the feature extractor, including: using the scalable image encoder to compress the image features, using the compressed image features as a priori, estimating the probability distribution of the implicit representation of the input image output by the feature encoder in the compression model, inputting the compressed image features into the cross-layer context model, and estimating the parameters of the probability distribution; using the entropy coding model in the compression model to entropy encode the input image using the parameters of the probability distribution, and then obtaining the encoded image through the feature decoder in the compression model.
8. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
End-to-end binocular image joint compression method and device, equipment and medium
CN112702592A
Image compression processing method and device, computer equipment and storage medium
CN113313777A