Coding method, system, device and storage medium

By utilizing the correlation between the reconstructed image and the residual image in the enhancement layer encoding, the probability distribution of the residual image is predicted, thus solving the problem of low compression efficiency of the enhancement layer and improving the overall image compression efficiency.

CN119520804BActive Publication Date: 2026-03-20HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In existing technologies, the compression efficiency of the enhancement layer is poor, resulting in low resolution and quality of video images in channel-limited or complex environments.

Method used

During the enhancement layer coding process, the correlation between the reconstructed image of the base layer and the residual image of the enhancement layer is utilized. The reconstructed image of the base layer is reused as prior data to predict the probability distribution of the residual image and entropy coding is performed to improve compression efficiency.

Benefits of technology

Without the need to store or transmit additional prior data, the compression efficiency of the image in the enhancement layer is improved, thereby enhancing the overall compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119520804B_ABST
    Figure CN119520804B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a coding method, system, device and storage medium. The method comprises the following steps: determining a residual image of an image in an enhancement layer according to the image and a reconstructed image of the image in a base layer; predicting a probability distribution corresponding to the residual image by using a trained probability model according to the reconstructed image; and performing entropy coding on the residual image according to the probability distribution corresponding to the residual image to obtain a code stream of the image in the enhancement layer. The coding scheme provided by the embodiments of the present application can improve the compression efficiency of the image in the enhancement layer, and further improve the overall compression efficiency of the image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of coding and decoding technology, and in particular to a coding and decoding method, system, device and storage medium. BACKGROUND

[0002] Video coding technology is an effective means to reduce video redundancy information and reduce video transmission bandwidth, and is widely used in cloud desktop, video live broadcast, video conference and other scenarios.

[0003] In order to further enhance the user's viewing experience and recognition ability, a common solution in a video coding system is to introduce layered coding technology. Layered coding is to encode in the base layer and the enhancement layer respectively to obtain the base layer code stream and the enhancement layer code stream. The base layer code stream can make the decoding end completely decode the basic video content, but the video image obtained by the base layer code stream may have low resolution or poor quality. In a limited channel or complex channel environment, the base layer code stream can ensure that the decoding end can receive a smooth video image that can be watched. When the channel environment is good or the channel resources are abundant, the enhancement layer code stream can be additionally transmitted to improve the resolution or video quality. However, in the prior art, the compression efficiency of the enhancement layer is poor. SUMMARY

[0004] In view of the above problems, the present application is proposed to provide a coding and decoding method, system, device and storage medium which solve the above problems or at least partially solve the above problems.

[0005] Therefore, in an embodiment of the present application, a coding method is provided, comprising:

[0006] determining a residual image of the image in the enhancement layer according to the image and the reconstructed image of the image in the base layer;

[0007] predicting a probability distribution corresponding to the residual image by using a trained probability model according to the reconstructed image;

[0008] performing entropy coding on the residual image according to the probability distribution corresponding to the residual image to obtain a code stream of the image in the enhancement layer.

[0009] In another embodiment of the present application, a decoding method is provided, comprising:

[0010] obtaining a code stream of the image in the base layer and a code stream of the image in the enhancement layer;

[0011] decoding the code stream of the image in the base layer to obtain a reconstructed image;

[0012] According to the reconstructed image, a trained probability model is used to predict a probability distribution corresponding to the residual image between the image and the reconstructed image;

[0013] According to the probability distribution corresponding to the residual image, the code stream of the image in the enhancement layer is entropy decoded to obtain a reconstructed residual image;

[0014] According to the reconstructed residual image, the reconstructed image is corrected to obtain a corrected reconstructed image.

[0015] In another embodiment of the present application, a coding system is provided, comprising: an encoding end and a decoding end;

[0016] The encoding end is configured to: determine a residual image of the image in the enhancement layer according to the image and a reconstructed image of the image in the base layer; use a trained probability model to predict a probability distribution corresponding to the residual image according to the reconstructed image; entropy encode the residual image according to the probability distribution corresponding to the residual image to obtain a code stream of the image in the enhancement layer; and send the code stream of the image in the base layer and the code stream of the image in the enhancement layer to the decoding end device.

[0017] The decoding end is configured to: receive the code stream of the image in the base layer and the code stream of the image in the enhancement layer; decode the code stream of the image in the base layer to obtain a reconstructed image; use a trained probability model to predict a probability distribution corresponding to the residual image between the image and the reconstructed image according to the reconstructed image; entropy decode the code stream of the image in the enhancement layer according to the probability distribution corresponding to the residual image to obtain a reconstructed residual image; and correct the reconstructed image according to the reconstructed residual image to obtain a corrected reconstructed image.

[0018] In another embodiment of the present application, an electronic device is provided. The electronic device comprises: a memory and a processor, wherein,

[0019] The memory is configured to store a program.

[0020] The processor is coupled to the memory and is configured to execute the program stored in the memory to implement any of the above methods.

[0021] In another embodiment of the present application, a computer readable storage medium storing a computer program is provided, and the computer program is executable by a computer to implement any of the above methods.

[0022] In the technical scheme provided by the embodiment of the present application, in the process of enhancement layer coding, the correlation between the reconstructed image of the base layer and the residual image of the enhancement layer is utilized, the reconstructed image of the base layer is multiplexed, and the reconstructed image of the base layer is used as prior data to predict the probability distribution of the residual image feature. In this way, additional prior data does not need to be stored or transmitted, the compression efficiency of the image in the enhancement layer can be improved, and the overall compression efficiency of the image is further improved. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical scheme in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0024] Figure 1 The structural block diagram of the coding system provided by an embodiment of the present application is provided.

[0025] Figure 2a The flowchart of the coding method provided by an embodiment of the present application is provided.

[0026] Figure 2b The prior data preprocessing flowchart provided by an embodiment of the present application is provided.

[0027] Figure 3 The flowchart of the decoding method provided by an embodiment of the present application is provided.

[0028] Figure 4 The flowchart of the model training method provided by an embodiment of the present application is provided.

[0029] Figure 5 The structural block diagram of the cloud desktop system provided by an embodiment of the present application is provided.

[0030] Figure 6 The flowchart of the cloud desktop image coding method provided by an embodiment of the present application is provided.

[0031] Figure 7 The experimental control chart related to the PSNR index provided by an embodiment of the present application is provided.

[0032] Figure 8 The experimental control chart related to the SSIM index provided by an embodiment of the present application is provided.

[0033] Figure 9 The experimental control chart related to the LPIPS index provided by an embodiment of the present application is provided.

[0034] Figure 10An experimental control diagram related to the DISTS index provided by an embodiment of the present application is shown in the following table:

[0035] Figure 11 A structural block diagram of an electronic device provided by an embodiment of the present application is shown in the following table. DETAILED DESCRIPTION

[0036] Generally, an image is composed of two chrominance maps and one luminance map. Taking a YCbCr format image as an example, YCbCr = YCbCr (Y, Cb, Cr), wherein Y represents luminance (Luminance or luma), that is, a gray value; and Cb and Cr represent chrominance, which is used to reflect the concentration offset of blue and red. YCbCr is often used in the video field.

[0037] In order to reduce the code rate consumption, in the base layer, the original image needs to be converted from YCbCr4:4:4 to YCbCr4:2:0 in format through chrominance downsampling, and then the image in YCbCr4:2:0 format is encoded to obtain the code stream of the base layer. The chrominance downsampling process of converting the original data from YCbCr4:4:4 to YCbCr4:2:0 will cause information loss, resulting in obvious artifacts and blurring in the image picture or video picture obtained by subsequent decoding. In the enhancement layer, the residual image between the original image and the reconstructed image of the base layer is encoded to obtain the code stream of the enhancement layer. At the decoding end, the code stream of the base layer and the code stream of the enhancement layer are combined to improve the objective quality and subjective quality of the image picture or video picture at the decoding end.

[0038] How to efficiently and high-quality compress the residual image has high research significance and application value.

[0039] Embodiments of the present application propose a residual image encoding scheme. That is, in the process of encoding the enhancement layer, the correlation between the reconstructed image of the base layer and the residual image of the enhancement layer is used to reuse the reconstructed image of the base layer as prior data to predict the probability distribution of the residual image feature. In this way, additional prior data does not need to be stored or transmitted, which can improve the compression efficiency of the image in the enhancement layer, and further improve the overall compression efficiency of the image.

[0040] In order to enable personnel in the technical field to better understand the present application scheme, the technical scheme in the embodiments of the present application will be described clearly and completely according to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor belong to the scope of protection of the present application.

[0041] Furthermore, some processes described in the specification, claims, and accompanying drawings of this application include multiple operations that appear in a specific order. These operations may be performed out of order or in parallel. Operation numbers such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be performed sequentially or in parallel. It should be noted that the terms "first," "second," etc., used herein are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0042] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0043] Before introducing the encoding and decoding methods provided in the embodiments of this application, the system architecture involved in this solution will be described. For example... Figure 1 As shown, the encoding / decoding system includes: an encoding end and a decoding end; wherein,

[0044] The encoding end is used to perform the following steps:

[0045] 11. Determine the residual image of the image in the enhancement layer based on the image and its reconstructed image in the base layer.

[0046] 12. Based on the reconstructed image, predict the probability distribution corresponding to the residual image using a trained probability model.

[0047] 13. Based on the probability distribution corresponding to the residual image, entropy coding is performed on the residual image to obtain the bitstream of the image in the enhancement layer.

[0048] 14. Send the bitstream of the image in the base layer and the bitstream of the image in the enhancement layer to the decoding end.

[0049] The decoding end is used to perform the following steps:

[0050] 21. Receive the bitstream of the image in the base layer and the bitstream of the image in the enhancement layer.

[0051] 22. Decode the bitstream of the image at the base layer to obtain the reconstructed image.

[0052] 23. predicting, according to the reconstructed image, a probability distribution corresponding to the residual image between the image and the reconstructed image by using the probability model.

[0053] 24. entropy decoding, according to the probability distribution corresponding to the residual image, a code stream of the image in the enhancement layer to obtain a reconstructed residual image.

[0054] 25. correcting, according to the reconstructed residual image, the reconstructed image to obtain a corrected reconstructed image.

[0055] The image can be a common image or a video frame in a video.

[0056] The coding system provided by the embodiments of the present application can be applied in cloud desktop, video live broadcast, video conference and the like. Taking the cloud desktop as an example, the image can be a video frame in a cloud desktop video stream, and the encoding end is a cloud desktop server; the decoding end is a cloud desktop client.

[0057] In the technical solution provided by the embodiments of the present application, in the process of enhancement layer coding, the correlation between the reconstructed image of the base layer and the residual image of the enhancement layer is used, the reconstructed image of the base layer is reused, and the reconstructed image of the base layer is used as prior data to predict the probability distribution of the residual image feature. In this way, additional prior data does not need to be stored or transmitted, the compression efficiency of the image in the enhancement layer can be improved, and the overall compression efficiency of the image is improved.

[0058] The specific processing process of the encoding end device and the decoding end device in the coding system and the interaction process between the two will be described in detail in the following embodiments.

[0059] Figure 2a A flowchart of an encoding method provided by the embodiments of the present application is shown. The execution subject of the method can be the encoding end in the coding system. As shown in the figure, the encoding method comprises the following steps. Figure 2a

[0060] 201. determining, according to the image and the reconstructed image of the image in the base layer, a residual image of the image in the enhancement layer;

[0061] 202. predicting, according to the reconstructed image, a probability distribution corresponding to the residual image by using the trained probability model;

[0062] 203. entropy encoding, according to the probability distribution corresponding to the residual image, the residual image to obtain a code stream of the image in the enhancement layer.

[0063] In 201, the image is a video frame in a cloud desktop video stream.

[0064] ​A residual image between the image and the reconstructed image can be calculated as the residual image of the image in the enhancement layer. The residual image is used to represent the difference between the image and the reconstructed image.

[0065] When the resolution of the reconstructed image is different from the resolution of the image, the reconstructed image can be resampled to obtain a resampled reconstructed image; the resolution of the resampled reconstructed image is the same as the resolution of the image. Then, a residual image between the image and the resampled reconstructed image can be calculated as the residual image of the image in the enhancement layer.

[0066] In a coding scenario, generally, the resolution of the reconstructed image is less than the resolution of the image, and therefore, the resampling is specifically up-sampling.

[0067] In 202, the probability model can also be referred to as an entropy model. The probability model can be selected according to actual needs, and embodiments of the present application do not make specific limitations thereon. The probability model can be a Gaussian probability model, a Laplace probability model, etc. The probability model is based on a neural network and can be trained according to training samples.

[0068] In actual applications, the reconstructed image can be input into the trained probability model, and the probability model outputs a probability distribution corresponding to the predicted residual image.

[0069] In an example, the probability distribution corresponding to the residual image can include a probability distribution of data values of each data bit in the residual image. Generally, any probability distribution can be represented by distribution parameters. Taking a Gaussian probability distribution as an example, the Gaussian probability distribution can be represented by two distribution parameters, i.e., a mean value and a variance. Therefore, the probability model predicts the distribution parameters of the probability distribution corresponding to the residual image.

[0070] In 203, entropy coding is a coding process in which no information is lost according to the entropy principle. The information entropy is the average amount of information of a signal source. Common entropy coding includes Shannon coding, Huffman coding, and arithmetic coding. Entropy coding is a lossless coding / lossless compression method.

[0071] In an example, the distribution parameters of the probability distribution of the data values of each data bit in the residual image and the residual image can be input into an entropy encoder to perform entropy coding, to obtain the code stream of the image in the enhancement layer.

[0072] The residual image is obtained by subtracting the reconstructed image from the image, and thus the residual image is related to the reconstructed image. Generally, low-frequency signals and high-frequency signals in the image are also called low-frequency components and high-frequency components. The high-frequency component in the image refers to a place where the image intensity (brightness / gray level) changes sharply, i.e., an edge (outline, texture); and the low-frequency component in the image refers to a place where the image intensity (brightness / gray level) changes gently, i.e., a large color block. Compared with the image, the reconstructed image has more high-frequency signal distortion and less low-frequency signal distortion.

[0073] Therefore, the position of the high-frequency information in the reconstructed image can reveal the position of the high-frequency signal in the residual image, the high-frequency signal in the residual image belongs to important signals, and a relatively large code rate can be consumed for compression; and the position of the low-frequency information in the reconstructed image can reveal the position of the low-frequency signal in the residual image, the low-frequency signal in the residual image belongs to unimportant information, and a relatively small code rate can be consumed for compression. It can be seen that, when the residual image is entropy encoded, the reconstructed image related to the residual image can be used to more accurately predict the probability distribution corresponding to the residual image, so as to reduce the overall code rate consumption.

[0074] In the technical scheme provided by the embodiments of the present application, in the process of enhancement layer coding, the correlation between the reconstructed image of the base layer and the residual image of the enhancement layer is used, and the reconstructed image of the base layer is reused as prior data to predict the probability distribution of the residual image feature. In this way, additional prior data does not need to be stored or transmitted, the compression efficiency of the image in the enhancement layer can be improved, and the overall compression efficiency of the image is further improved.

[0075] In actual application, the luminance component is very important information, and generally, resolution reduction is not allowed, and only resolution reduction is allowed on two chroma components to improve the compression efficiency. Resolution is used to represent the number of pixel points of an image in the width direction and the height direction. For example, the resolution of image A is 100*50, which indicates that image A has 100 pixel points in the width direction and 50 pixel points in the height direction.

[0076] In an example, the resolution of the luminance map of the reconstructed image is the same as the resolution of the luminance map of the image; and the resolution of the chroma map of the reconstructed image is smaller than the resolution of the chroma map of the image. That is, the resolution is not reduced on the luminance component, and the resolution is reduced on the chroma component. Therefore, the coding loss on the luminance component can be ignored. Therefore, in the enhancement layer, only the loss information on the chroma component needs to be encoded. The above "determining the residual image of the image in the enhancement layer according to the image and the reconstructed image of the image in the base layer" in 201 can be implemented by the following steps:

[0077] 2011, calculate a chroma residual image between a chroma map of the image and a chroma map of the reconstructed image.

[0078] 2012, determine the chroma residual image as a residual image of the image at an enhancement layer.

[0079] In 2011, the chroma residual image is used to represent a difference between a chroma map of the image and a chroma map of the reconstructed image.

[0080] Since the resolution of the chroma map of the reconstructed image is less than the resolution of the chroma map of the image, the chroma map of the reconstructed image needs to be up-sampled to obtain an up-sampled chroma map. The up-sampling algorithm used for the up-sampling can include, but is not limited to, nearest neighbor interpolation, bilinear interpolation. The resolution of the up-sampled chroma map is consistent with the resolution of the chroma map of the image.

[0081] The difference between the chroma map of the image and the corresponding pixel points in the up-sampled chroma map is calculated to obtain a chroma residual image.

[0082] In 2012, the chroma residual image is used as a residual image of the image at an enhancement layer.

[0083] Since the resolution of the chroma map of the reconstructed image is different from the resolution of the luminance map of the reconstructed image, a scale alignment process is needed to be able to be used as an input of the probability model. In 202, the step of "predicting a probability distribution corresponding to the residual image using a trained probability model according to the reconstructed image" can be implemented as follows:

[0084] 2021, up-sample the chroma map of the reconstructed image to obtain an up-sampled chroma map.

[0085] The resolution of the up-sampled chroma map is the same as the resolution of the luminance map of the reconstructed image.

[0086] 2022, determine input data of the probability model according to the up-sampled chroma map and the luminance map of the reconstructed image.

[0087] In 2021, the up-sampling process can refer to the corresponding content in the above embodiments, which will not be described here.

[0088] In an example, in 2022, the up-sampled chroma map and the luminance map of the reconstructed image can be used as different channel maps in the input data. In this embodiment, the input data includes two channel maps, one is the up-sampled chroma map, and the other is the luminance map of the reconstructed image.

[0089] In the embodiment, not only the chroma component of the reconstructed image is used, but also the luminance component of the reconstructed image is used to preset the probability distribution corresponding to the chroma residual image, and the prediction accuracy is good.

[0090] In another example, the luminance component of the reconstructed image has less distortion compared with the image, and the chroma component of the reconstructed image has more distortion. This is because, when the image is encoded, the chroma downsampling operation is performed on the image, and the luminance downsampling operation is not performed. The chroma downsampling operation is the most important reason for the large chroma distortion. In order to improve the prediction accuracy of the probability distribution of the entropy model, the same downsampling operation is performed on the luminance image of the reconstructed image to simulate the chroma downsampling process, and the luminance residual image between the luminance image of the reconstructed image and the luminance image after downsampling is used to improve the prediction accuracy of the probability distribution of the entropy model. The luminance residual image carries the luminance downsampling distortion. Since the downsampling algorithm used for the luminance downsampling is the same as the downsampling algorithm used for the chroma downsampling, the luminance downsampling distortion is similar to the chroma downsampling distortion. Therefore, the luminance residual image carrying the luminance downsampling distortion can improve the prediction accuracy of the probability distribution of the entropy model for the chroma residual image. Specifically, the step of “determining the input data of the probability model according to the upsampled chroma image and the luminance image of the reconstructed image” in the above 2022 can be implemented by the following steps:

[0091] S11, determining a downsampling algorithm according to which the chroma downsampling operation performed on the image by the base layer is performed.

[0092] S12, performing luminance downsampling on the luminance image of the reconstructed image according to the downsampling algorithm to obtain a luminance image after downsampling.

[0093] S13, determining a luminance residual image between the luminance image of the reconstructed image and the luminance image after downsampling.

[0094] The resolution of the luminance residual image is the same as that of the luminance image of the reconstructed image.

[0095] S14, taking the upsampled chroma image, the luminance image of the reconstructed image, and the luminance residual image as different channel images in the input data of the probability model, respectively.

[0096] In the above S11, the downsampling algorithm according to which the chroma downsampling operation performed on the image by the base layer is selected according to actual needs, and the embodiment of the present application does not make specific limitation. The downsampling algorithm can include but is not limited to: average pooling algorithm, maximum value pooling algorithm.

[0097] In S12, the luminance of the reconstructed image is down-sampled according to the down-sampling algorithm to obtain a down-sampled luminance.

[0098] For example, if the down-sampling algorithm for the chroma down-sampling operation is the average pooling algorithm, the luminance of the reconstructed image is down-sampled according to the average pooling algorithm to obtain a down-sampled luminance.

[0099] In S13, the resolution of the down-sampled luminance is smaller than that of the luminance of the reconstructed image, and thus the down-sampled luminance needs to be up-sampled to obtain an up-sampled luminance before the luminance residual image is calculated. The resolution of the up-sampled luminance is consistent with that of the luminance of the reconstructed image. The luminance residual image is obtained by calculating the difference between the corresponding pixels in the luminance of the reconstructed image and the up-sampled luminance. The up-sampling algorithm for the down-sampled luminance can be selected according to actual needs, and embodiments of the present application do not make specific limitations thereon. In a specific example, the up-sampling algorithm for the down-sampled luminance can be consistent with the up-sampling algorithm for the chroma up-sampling operation in step 201.

[0100] In S14, the up-sampled chroma, the luminance of the reconstructed image, and the luminance residual image are respectively taken as different channel images in the input data of the probability model. That is, the input data includes three channel images, namely, the up-sampled chroma, the luminance of the reconstructed image, and the luminance residual image. The input data is the final required prior data.

[0101] In the present scheme, the joint input of multi-channel information can improve the accuracy of probability distribution estimation.

[0102] Figure 2b A processing flow example from the reconstructed image to the prior data is shown. As shown in Figure 2b The luminance of the reconstructed image is sequentially subjected to the average pooling 2001 and the nearest neighbor interpolation 2002, and then the difference between the luminance of the reconstructed image and the luminance residual image is calculated to obtain the luminance residual image. The chroma of the reconstructed image is subjected to the nearest neighbor interpolation 2003 to obtain the up-sampled chroma. The luminance of the reconstructed image, the luminance residual image, and the up-sampled chroma are taken as the final prior data.

[0103] In an implementable scheme, the probability distribution of the data value of each data bit in the residual image can be directly predicted by the probability model, and then the residual image is directly input into the entropy encoder for encoding.

[0104] In another implementation, the probability distribution corresponding to the residual image comprises a probability distribution of a residual image feature of the residual image. The "predicting, according to the reconstructed image, the probability distribution corresponding to the residual image by using the trained probability model" in 202 can comprise:

[0105] predicting, according to the reconstructed image, the probability distribution of the residual image feature by using the trained probability model.

[0106] Correspondingly, the "entropy encoding, according to the probability distribution corresponding to the residual image, the residual image to obtain the code stream of the image in the enhancement layer" in 203 can comprise:

[0107] entropy encoding, according to the probability distribution of the residual image feature, the residual image feature to obtain the code stream of the image in the enhancement layer.

[0108] In this embodiment, instead of directly entropy encoding the residual image, the residual image is first feature extracted to obtain the residual image feature of the residual image, and then the residual image feature is entropy encoded. Since there is a large amount of redundant information in the residual image, by feature extraction, the important information is extracted and then entropy encoded, which can effectively improve the overall compression efficiency of the residual image.

[0109] Optionally, the method can further comprise:

[0110] 204. feature extracting, by using the trained encoding model, the residual image to obtain the residual image feature.

[0111] The encoding model can be based on a neural network. The encoding model can be trained based on training samples in advance.

[0112] In an example, the residual image is the chroma residual image, which generally comprises two chroma components. The two chroma components can be input into the encoding model as two channel images.

[0113] In actual application, the data type of the residual image feature output by the encoding model is floating point type; and the entropy encoding requires the data type of the input data to be integer type. Therefore, the residual image can be input into the trained encoding model to obtain the initial residual image feature output by the encoding model; and the initial residual image feature is quantized to obtain the residual image feature.

[0114] The steps 201 to 204 are all performed in the enhancement layer, and the processing flow of the base layer will be introduced below. Specifically, the method can further comprise:

[0115] 205. In the base layer, chroma down-sampling the image to obtain a down-sampled image.

[0116] 206. Compressing the down-sampled image to obtain a bitstream of the image in the base layer.

[0117] 207. Decoding the bitstream of the image in the base layer to obtain the reconstructed image.

[0118] The specific implementation of step 205 can be referred to the corresponding content in the above-mentioned embodiments.

[0119] In 206, the compression process performed on the down-sampled image can include feature extraction, entropy encoding, etc., which are not specifically limited in the embodiments of the present application.

[0120] In 207, the decoding process performed on the bitstream of the image in the base layer can include entropy decoding, feature decoding, etc., which are not specifically limited in the embodiments of the present application.

[0121] Figure 3 A flowchart of a decoding method provided by an embodiment of the present application is shown. The execution subject of the decoding method can be the decoding end in the coding and decoding system. As shown in the figure, the method can include: Figure 3

[0122] 301. Obtaining a bitstream of an image in a base layer and a bitstream of the image in an enhancement layer.

[0123] 302. Decoding the bitstream of the image in the base layer to obtain a reconstructed image.

[0124] 303. According to the reconstructed image, predicting a probability distribution corresponding to a residual image between the image and the reconstructed image by using a trained probability model.

[0125] 304. According to the probability distribution corresponding to the residual image, performing entropy decoding on the bitstream of the image in the enhancement layer to obtain a reconstructed residual image.

[0126] 305. According to the reconstructed residual image, correcting the reconstructed image to obtain a corrected reconstructed image.

[0127] In 302, the decoding process performed on the bitstream of the image in the base layer can include entropy decoding, feature decoding, etc., which are not specifically limited in the embodiments of the present application.

[0128] In 303, the probability model here is the same as the probability model used in the encoding end. The specific implementation process of 303 can be referred to the corresponding content in the above-mentioned embodiments, which will not be described herein.

[0129] ​The probability model outputs a distribution parameter of a probability distribution corresponding to the residual image.

[0130] In 304, the distribution parameter of the probability distribution corresponding to the residual image and the code stream of the image in the enhancement layer are input into an entropy decoder; and a reconstructed residual image is determined according to an output result of the entropy decoder.

[0131] In an example, the probability distribution corresponding to the residual image includes a probability distribution of data values of each data bit in the residual image, and the entropy decoder outputs the reconstructed residual image.

[0132] In another example, the probability distribution corresponding to the residual image includes a probability distribution of residual image features of the residual image, that is, a probability distribution of data values of each data bit in the residual image features, and the entropy decoder outputs the residual image features. Then, the reconstructed residual image can be determined according to the residual image features. In an implementable solution, the residual image features can be decoded by using a trained decoding model to obtain the reconstructed residual image. The encoding model can be based on a neural network. The decoding model and the encoding model can be trained together to improve the training effects of the two models. The specific training process will be described in the following embodiments.

[0133] In 305, the reconstructed residual image can be superimposed on the reconstructed image to obtain a corrected reconstructed image.

[0134] In an example, the reconstructed residual image is a reconstructed chroma residual image, and the resolution of the reconstructed chroma residual image is greater than the resolution of a chroma map of the reconstructed image. Therefore, the reconstructed image can be chroma up-sampled to obtain an up-sampled reconstructed image; the resolution of a chroma map of the up-sampled reconstructed image is the same as the resolution of the reconstructed chroma residual image; and the reconstructed chroma residual image is superimposed on the chroma map of the up-sampled reconstructed image to obtain the corrected reconstructed image.

[0135] In the technical solution provided by the embodiments of the present application, in the process of enhancement layer encoding, the correlation between the reconstructed image of the base layer and the residual image of the enhancement layer is utilized, the reconstructed image of the base layer is multiplexed, and the reconstructed image of the base layer is used as prior data to predict the probability distribution of the residual image features. In this way, additional prior data does not need to be stored or transmitted, the compression efficiency of the image in the enhancement layer can be improved, and the overall compression efficiency of the image is improved.

[0136] The technical solutions provided by the embodiments of the present application will be described below with reference to the accompanying drawings. Figure 4 The training method of the model involved in the embodiments of the present application will be described. As shown in FIG. 1, the method can include the following steps. Figure 4 ​

[0137] 401、determine a sample residual image of the sample image in the enhancement layer according to the sample image and the reconstructed sample image of the sample image in the base layer.

[0138] 402、extract features of the sample residual image by using an encoding model to obtain sample residual image features.

[0139] 403、decode the sample residual image features by using a decoding model to obtain a reconstructed sample residual image.

[0140] 404、correct the reconstructed sample image according to the reconstructed sample residual image to obtain a corrected reconstructed sample image.

[0141] 405、optimize parameters of the encoding model and the decoding model with an optimization objective of minimizing a first loss function.

[0142] The first loss function is determined according to a difference between the sample image and the corrected reconstructed sample image.

[0143] In 401, a sample residual image between the sample image and the reconstructed sample image is determined as the sample residual image of the sample image in the enhancement layer.

[0144] The specific calculation process can refer to the calculation process of the “residual image” in the above embodiments, which is not described in detail here.

[0145] In 402 and 403, the internal structure of the encoding model and the decoding model can be designed according to actual needs, and the embodiments of the present application do not make specific limitations.

[0146] In 404, the correction process of the reconstructed sample image can refer to the correction process of the “reconstructed image” in the above embodiments, which is not described in detail here.

[0147] In 405, the gradient descent algorithm can be used to optimize the parameters of the encoding model and the decoding model. The difference between the sample image and the corrected reconstructed sample image can be represented by mean square error (MSE), structural similarity (SSIM), etc.

[0148] Optionally, the training method can further include:

[0149] 406、determine a prediction probability distribution of the sample residual image features according to the reconstructed sample image by using the probability model.

[0150] 407. optimize the probability model according to a gradient descent algorithm.

[0151] wherein the second loss function is determined according to a difference between the predicted probability distribution and a true probability distribution of the sample residual image feature.

[0152] The implementation of 406 can refer to the implementation of 202, which will not be repeated here.

[0153] In 407, the probability model can be optimized according to a gradient descent algorithm. The second loss function can be specifically a Shannon cross entropy.

[0154] In actual applications, the encoding model, the decoding model and the probability model can be jointly trained. Specifically, a total loss function can be determined according to the first loss function and the second loss function; and the encoding model, the decoding model and the probability model can be optimized according to a minimization of the total loss function. The total loss function is a sum of the first loss function and the second loss function.

[0155] The scheme provided by the embodiments of the present application improves the feature extraction of the residual image in a data-driven manner, and is more targeted for processing of chroma residual images with different contents.

[0156] The embodiments of the present application will be described in detail below taking a cloud desktop scenario as an example:

[0157] As shown in FIG. 5, a cloud desktop server 501 and a cloud desktop client 502 are connected through a network. Figure 5 As shown in FIG. 5, a cloud desktop server 501 and a cloud desktop client 502 are connected through a network.

[0158] The cloud desktop server 501 sends a basic layer code stream of a cloud desktop image and an enhancement layer code stream of the cloud desktop image to the cloud desktop client 502. The cloud desktop image is a video frame in a cloud desktop video stream.

[0159] As shown in FIG. 5, a cloud desktop server 501 and a cloud desktop client 502 are connected through a network. Figure 6 As shown in FIG. 5, a cloud desktop server 501 and a cloud desktop client 502 are connected through a network.

[0160] The encoding process of the cloud desktop server: input the chroma residual image 1 between the cloud desktop image and the reconstructed cloud desktop image 9 into the trained encoding model 2 to perform feature extraction, to obtain initial chroma residual image features; input the initial chroma residual image features into a quantizer to perform quantization processing, to obtain chroma residual image features; input the chroma residual image features into an entropy encoder 5, and the entropy encoder 5 performs entropy encoding on the chroma residual image features according to a probability distribution N(μ, θ) of the chroma residual image features output by the probability model, to obtain the enhancement layer code stream of the cloud desktop image. μ is a mean of a Gaussian distribution, and θ is a variance of the Gaussian distribution.

[0161] Note: Since the dynamic range of residual image is relatively smaller, mostly concentrated around 0, it is not normalized to avoid the dynamic range of the learned latent variable feature (i.e. the initial chroma residual image feature in the above) being too small.

[0162] Decoding process of cloud desktop client: input the code stream of cloud desktop image in enhancement layer into entropy decoder 6, entropy decoder 6 performs entropy decoding on the code stream of cloud desktop image in enhancement layer according to the probability distribution N(μ, θ) of chroma residual image feature output by the probability model, to obtain the chroma residual image feature; input the chroma residual image feature into the decoding model 7 for feature decoding, to obtain the reconstructed chroma residual image.

[0163] In addition, the cloud desktop server can also perform chroma down-sampling on the cloud desktop image in the basic layer to obtain a down-sampled image; encode the down-sampled image to obtain the code stream of the cloud desktop image in the basic layer; and decode the code stream of the cloud desktop image in the basic layer to obtain the reconstructed cloud desktop image.

[0164] The cloud desktop client further comprises: decoding the code stream of the cloud desktop image in the basic layer to obtain the reconstructed cloud desktop image; and correcting the reconstructed cloud desktop image in the basic layer according to the reconstructed chroma residual image to obtain the corrected reconstructed cloud desktop image.

[0165] Probability distribution process performed by both the cloud desktop server and the cloud desktop client: the prior preprocessing model 10 pre-processes the reconstructed cloud desktop image to obtain prior data. The preprocessing can include the size alignment and the acquisition of the luminance residual image in the above. The preprocessing process can refer to the corresponding content in the above embodiments, which will not be described here. Input the prior data into the prior feature extraction module 1001 of the probability model 11 for feature extraction to obtain the prior feature. The prior feature extraction module also adopts the structure of neural network, retains a relatively large receptive field, and extracts the global characteristics of the prior data. The prior data also contains high-frequency information such as edges and textures in the original image, which needs to be decomposed and extracted by the prior encoder, and the high-frequency features close to the chroma residual are screened out to improve the probability prediction accuracy. Input the prior feature into the probability prediction module 1002 of the probability model 11 to predict the probability distribution N(μ, θ) of the chroma residual image feature.

[0166] The encoding module can extract effective features of the chroma residual image to obtain the latent variable feature. The encoding module and the decoding module are responsible for the conversion between the image space and the latent variable space of the chroma residual image. The encoding module and the decoding module based on the neural network can more effectively adapt to the content characteristics of the chroma residual, and the representation ability of the latent variable feature is more powerful. The probability model takes the reconstructed image as input, extracts the global prior information of the chroma residual image, establishes an independent Gaussian distribution for each position of the latent variable feature, and predicts the Gaussian distribution parameters by using the global prior. The independent Gaussian model combined with the prior information can more effectively adapt to the residual data distribution with high sparsity, and improve the prediction accuracy of the probability model. Entropy encoder / decoder: The arithmetic encoder / decoder can be used to encode / decode the quantized latent variable feature by using the Gaussian probability distribution parameters provided by the probability model, to realize the conversion between the feature and the code stream. The scheme provided in the embodiment of the application can effectively improve the compression efficiency of the chroma residual image and achieve better compression quality under the same code rate.

[0167] Experimental comparison:

[0168] The following respectively uses subjective and objective performance indicators, for example: peak signal-to-noise ratio (Peak signal-to-noise ratio, commonly abbreviated as PSNR), SSIM, learned perceptual image patch similarity (Learned Perceptual Image Patch Similarity, LPIPS), DISTS (Differentiable Image Saliency Transform for Improved Scalability and Portability of Image Quality Assessment, image quality evaluation index based on differentiable image saliency transform) to show the technical effect of each scheme. The chroma residual image obtained by the present scheme and several existing compression schemes is superimposed with the reconstructed image of the base layer, and the final performance result is compared. It can be seen that, compared with the existing compression method, the present scheme can achieve better results in each index.

[0169] As shown in Figure 7 : the present scheme performs better in the PSNR index. Note: the higher the PSNR index, the better.

[0170] As shown in Figure 8 : the present scheme performs better in the SSIM index. Note: the higher the SSIM index, the better.

[0171] As shown in Figure 9 : the present scheme performs better in the LPIPS index. Note: the lower the LPIPS index, the better.

[0172] As shown inFigure 10 As shown, the present scheme performs better in the DISTS index. Note: The lower the DISTS index, the better.

[0173] It should be noted that, Figure 7 , Figure 8 , Figure 9 and Figure 10 The multiple unlabeled curves in the multiple existing compression schemes correspond to multiple existing compression schemes.

[0174] To sum up, the compression scheme based on the neural network provided by the present scheme supports the dual chroma channel residual as the input of the encoding model, and inputs the dual chroma information as the different channel images of the encoding model into the encoding model for processing. Moreover, the encoding model supports the input of pixels with various bit depths, for example, 16-bit depth. For the probability / entropy estimation problem, a probability model based on the reconstructed image is introduced to predict the probability distribution of the residual image. The reconstructed data of the basic layer shared by the encoding / decoding end is used to estimate the distribution of the residual information, which omits the encoding and transmission of the hyper-prior information compared with the traditional compression algorithm. Moreover, the probability model adopts the joint input of multiple channel information, which can improve the accuracy of probability estimation. For the problem that the size of the luminance / chroma component of the reconstructed image is inconsistent, a method including nearest neighbor sampling is introduced to align the size and then input into the probability model.

[0175] Figure 11 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. As shown in Figure 11 The electronic device includes a memory 1101 and a processor 1102. The memory 1101 can be configured to store various other data to support operations on the electronic device. Examples of the data include instructions for operating any application or method on the electronic device. The memory 1101 can be realized by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read only memory (EEPROM), an erasable programmable read only memory (EPROM), a programmable read only memory (PROM), a read only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.

[0176] The memory 1101 is configured to store programs;

[0177] The processor 1102 is coupled to the memory 1101 and is used to execute the program stored in the memory 1101 to implement the methods provided in the above-described method embodiments.

[0178] Furthermore, such as Figure 11 As shown, the electronic device also includes: communication component 1103, display 1104, power supply component 1105, audio component 1106, and other components. Figure 11 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 11 The components shown.

[0179] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a computer, can implement the steps or functions of the methods provided in the above-described method embodiments.

[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM (Read Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0182] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the same; although the present application has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An encoding method, characterized in that, include: Based on the image and its reconstructed image in the base layer, a residual image of the image in the enhancement layer is determined, the residual image including the chromaticity residual image between the chromaticity map of the image and the chromaticity map of the reconstructed image; Based on the reconstructed image, the probability distribution corresponding to the residual image is predicted using a trained probability model; Based on the probability distribution corresponding to the residual image, entropy coding is performed on the residual image to obtain the bitstream of the image in the enhancement layer; The step of predicting the probability distribution corresponding to the residual image using a trained probability model based on the reconstructed image includes: upsampling the chroma map of the reconstructed image to obtain an upsampled chroma map; the resolution of the upsampled chroma map is the same as the resolution of the luminance map of the reconstructed image; and determining the input data of the probability model based on the upsampled chroma map and the luminance map of the reconstructed image.

2. The method according to claim 1, characterized in that, The resolution of the luminance map of the reconstructed image is the same as the resolution of the luminance map of the image. Based on the upsampled chromaticity map and the luminance map of the reconstructed image, the input data for the probability model is determined, including: Determine the downsampling algorithm used for the chroma downsampling operation performed on the image at the base layer; According to the downsampling algorithm, the brightness map of the reconstructed image is downsampled to obtain the downsampled brightness map; Determine a brightness residual map between the brightness map of the reconstructed image and the brightness map after downsampling; the resolution of the brightness residual map is the same as the resolution of the brightness map of the reconstructed image; The upsampled chroma map, the luminance map of the reconstructed image, and the luminance residual map are used as different channel maps in the input data of the probability model.

3. The method according to claim 1 or 2, characterized in that, The probability distribution corresponding to the residual image includes the probability distribution of the residual image features of the residual image; Based on the reconstructed image, predicting the probability distribution corresponding to the residual image using a trained probability model includes: Based on the reconstructed image, the probability distribution of the residual image features is predicted using a trained probability model; Based on the probability distribution corresponding to the residual image, entropy coding is performed on the residual image to obtain the bitstream of the image in the enhancement layer, including: Based on the probability distribution of the residual image features, entropy coding is performed on the residual image features to obtain the bitstream of the image in the enhancement layer.

4. The method according to claim 3, characterized in that, Also includes: The residual image features are obtained by using a trained encoding model to extract features from the residual image.

5. The method according to claim 4, characterized in that, The residual image is feature extracted using a trained encoding model to obtain residual image features, including: The residual image is input into the trained encoding model to obtain the initial residual image features output by the encoding model; The initial residual image features are quantized to obtain the residual image features.

6. The method according to claim 1 or 2, characterized in that, The image is a video frame from a cloud desktop video stream.

7. The method according to claim 1 or 2, characterized in that, Also includes: At the base layer, the image is chroma downsampled to obtain the downsampled image; The downsampled image is encoded to obtain the bitstream of the image at the base layer; The reconstructed image is obtained by decoding the bitstream of the image at the base layer.

8. A decoding method, characterized in that, include: Obtain the bitstream of the image at the base layer and the bitstream of the image at the enhancement layer; The image is decoded at the base layer bitstream to obtain the reconstructed image; Based on the reconstructed image, a trained probability model is used to predict the probability distribution corresponding to the residual image between the image and the reconstructed image. The residual image includes the chromaticity residual image between the chromaticity map of the image and the chromaticity map of the reconstructed image. Based on the probability distribution corresponding to the residual image, entropy decoding is performed on the bitstream of the image in the enhancement layer to obtain the reconstructed residual image; Based on the reconstructed residual image, the reconstructed image is corrected to obtain the corrected reconstructed image; The step of predicting the probability distribution corresponding to the residual image between the reconstructed image and the reconstructed image using a trained probability model, based on the reconstructed image, includes: upsampling the chroma map of the reconstructed image to obtain an upsampled chroma map; the resolution of the upsampled chroma map is the same as the resolution of the luminance map of the reconstructed image; and determining the input data of the probability model based on the upsampled chroma map and the luminance map of the reconstructed image.

9. A codec system, characterized in that, include: Encoding end and decoding end; The encoding end is configured to: determine the residual image of the image in the enhancement layer based on the image and its reconstructed image in the base layer, wherein the residual image includes a chroma residual image between the chroma map of the image and the chroma map of the reconstructed image; predict the probability distribution corresponding to the residual image using a trained probability model based on the reconstructed image; perform entropy encoding on the residual image based on the probability distribution corresponding to the residual image to obtain the bitstream of the image in the enhancement layer; and send the bitstream of the image in the base layer and the bitstream of the image in the enhancement layer to the decoding end. The decoding end is configured to: receive the bitstream of the image at the base layer and the bitstream of the image at the enhancement layer; decode the bitstream of the image at the base layer to obtain the reconstructed image; predict the probability distribution corresponding to the residual image between the image and the reconstructed image using the probability model based on the reconstructed image; perform entropy decoding on the bitstream of the image at the enhancement layer based on the probability distribution corresponding to the residual image to obtain the reconstructed residual image; and correct the reconstructed image based on the reconstructed residual image to obtain the corrected reconstructed image. The step of predicting the probability distribution of the residual image between the reconstructed image and the reconstructed image using the probability model includes: upsampling the chroma map of the reconstructed image to obtain an upsampled chroma map; the resolution of the upsampled chroma map is the same as the resolution of the luminance map of the reconstructed image; and determining the input data of the probability model based on the upsampled chroma map and the luminance map of the reconstructed image.

10. An electronic device, characterized in that, include: Memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program stored in the memory to implement the method of any one of claims 1 to 8.

11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a computer, it can implement the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Entropy coding / decoding method and device

    CN114339262A

  • Scalable encoding and decoding method of color video,and apparatus thereof

    KR1020060108253A