Coding, decoding methods and related devices
By encoding the original image and generating a second code stream of fidelity map, the problem that the prior art cannot obtain the encoded image distortion intensity information on the decoding end is solved, and detailed distortion intensity information is obtained on the decoding end, which improves the evaluation and recovery performance of image compression quality.
Patent Information
- Application Number
- CN202110170984.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-02-08
AI Technical Summary
Existing video or image encoding and decoding schemes cannot obtain distortion intensity information of the encoded image at the decoding end, including the inability to obtain distortion intensity information and overall distortion intensity information of each area.
By encoding the original image to obtain a first code stream, and encoding the fidelity map to obtain a second code stream, the fidelity map is used to represent the distortion between the original image and the reconstructed image. The decoding terminal decodes the first code stream and the second code stream, and can obtain distortion intensity information of the encoded image.
It realizes the acquisition of distortion intensity information of the encoded image on the decoding side, which can more effectively evaluate the compression quality and recovery performance of the image.
Smart Images

Figure CN114913249B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application are directed to the field of video or image compression technology based on artificial intelligence (AI), and in particular to a coding and decoding method and related equipment. Background Art
[0002] Video coding (video encoding and decoding) is widely used in digital video applications such as broadcast digital television, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVD and Blu-ray Discs, video content acquisition and editing systems, and security applications in camcorders.
[0003] Even in the case of short films, large amounts of video data are required to describe them, which can create difficulties when the data is to be sent or otherwise transmitted across networks with limited bandwidth capacity. Therefore, video data is often compressed before being transmitted across modern telecommunications networks. The size of the video can also be an issue when storing the video on a storage device, as memory resources may be limited. Video compression devices typically use software and / or hardware on the source side to encode the video data prior to transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received by a video decompression device on the destination side. With limited network resources and the growing demand for higher video quality, there is a need for improved compression and decompression techniques that can increase compression rates with little to no impact on image quality.
[0004] In recent years, the application of deep learning in the field of image and video coding has gradually become a trend. However, various existing neural network-based video or image coding solutions cannot obtain the distortion strength information of the encoded image at the decoding end, for example, the distortion strength information of each area in an encoded image and the overall distortion strength information of an encoded image cannot be obtained at the decoding end. Summary of the invention
[0005] The present application provides an encoding and decoding method and related devices, which can obtain the distortion intensity information of the encoded image at the decoding end.
[0006] The above and other objects are achieved by the subject matter of the independent claims. Other implementations are apparent from the dependent claims, the detailed description and the drawings.
[0007] Specific embodiments are outlined in the attached independent claims, further embodiments are outlined in the dependent claims.
[0008] According to a first aspect, the present application relates to a coding method. The method is performed by a coding device. The method comprises: encoding an original image to obtain a first code stream; encoding a fidelity map to obtain a second code stream, wherein the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of a reconstructed image, and the reconstructed image is obtained after decoding the first code stream. In an embodiment of the present application, an original image is encoded to obtain a first code stream, and a fidelity map is encoded to obtain a second code stream, and the fidelity map is used to represent the distortion between at least a portion of the original image and at least a portion of the reconstructed image, wherein the distortion includes a difference; a decoding end decodes the first code stream to obtain a reconstructed image of the original image, and a decoding end decodes the second code stream to obtain a reconstructed map of the fidelity map (also referred to as a reconstructed fidelity map); and if the encoding is lossless encoding, the reconstructed map of the fidelity map is the same as the fidelity map; if the encoding is lossy encoding, the reconstructed map of the fidelity map includes the encoding distortion generated by encoding the fidelity map; therefore, the fidelity map can be used to represent the distortion between at least a portion of the original image and at least a portion of the reconstructed image; therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0009] In a possible design, the method also includes: dividing the original image into multiple first image blocks, and dividing the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the original image is the same as the division strategy for dividing the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; or dividing a preset area of the original image into multiple first image blocks, and dividing the preset area of the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the preset area of the original image is the same as the division strategy for dividing the preset area of the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; calculating the fidelity value of any second image block according to any second image block among the multiple second image blocks and the first image block corresponding to the any second image block, the fidelity map including the fidelity value of any second image block, and the fidelity value of any second image block is used to represent the distortion between any second image block and the first image block corresponding to the any second image block. Among them, the position of any second image block among the multiple second image blocks in the reconstructed image is the same as the position of the first image block corresponding to the any second image block in the original image; the position of the preset area of the original image in the original image is the same as the position of the preset area of the reconstructed image in the reconstructed image, and the position of any second image block among the multiple second image blocks in the preset area of the reconstructed image is the same as the position of the first image block corresponding to the any second image block in the preset area of the original image.In an embodiment of the present application, the size of the original image is the same as the size of the reconstructed image, and the size and position of the preset area in the original image are the same as those in the reconstructed image; the original image is divided into a plurality of first image blocks according to the same division strategy, and the reconstructed image is divided into a plurality of second image blocks; or the preset area of the original image is divided into a plurality of first image blocks according to the same division strategy, and the preset area of the reconstructed image is divided into a plurality of second image blocks; there is a one-to-one correspondence between the plurality of first image blocks obtained by division and the second image blocks obtained by division, wherein the size of any first image block is the same, the size of any second image block is also the same, and the size of the first image block is also the same as that of the second image block; therefore, the first image block and the second image block can be used as The basic unit of fidelity calculation, that is, the fidelity value of any second image block can be calculated according to any second image block among the multiple second image blocks and its corresponding first image block, and the fidelity values of the multiple second image blocks are also the fidelity values of each area of the reconstructed image, and a fidelity map can be obtained according to the fidelity values of the multiple second image blocks; wherein, when the first image block is obtained by dividing the original image and the second image block is obtained by dividing the reconstructed image, the fidelity map is used to characterize the fidelity of the reconstructed image; when the first image block is obtained by dividing the preset area of the original image and the second image block is obtained by dividing the preset area of the reconstructed image, the fidelity map is used to characterize the fidelity of the preset area of the reconstructed image; thereby facilitating the acquisition of a fidelity map for characterizing the distortion intensity information of the encoded image.
[0010] In a possible design, the fidelity map includes a plurality of first elements, the plurality of second image blocks correspond to the plurality of first elements one by one, the value of any first element among the plurality of first elements is the fidelity value of the second image block corresponding to the any first element, the position of the any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image. The first element may also be referred to as a pixel point of the fidelity map. When calculating the fidelity map, the image is divided into basic units for fidelity calculation, so there are as many first elements in the fidelity map as there are basic units; any first element has two attributes, namely, the fidelity value and the position of the fidelity value in the fidelity map. In an embodiment of the present application, the fidelity map is a two-dimensional array, and the reconstructed image is divided into multiple second image blocks. The fidelity map can be obtained according to the fidelity values of the multiple second image blocks, that is, multiple first elements can be determined according to the multiple second image blocks, and the multiple second image blocks correspond one-to-one to the multiple first elements. The value of any first element among the multiple first elements is the fidelity value of the second image block corresponding thereto; and the position of any first element among the multiple first elements in the fidelity map is determined according to the position of the second image block corresponding thereto in the reconstructed image. Specifically, the position of any first element among the multiple first elements in the fidelity map is the same as the position of the second image block corresponding thereto in the reconstructed image or in a preset area of the reconstructed image, so that the elements at each position of the fidelity map represent the fidelity of the area corresponding to its position in the reconstructed image or in the preset area of the reconstructed image, which is beneficial for the fidelity map to be used to represent the distortion intensity information of the encoded image.
[0011] In a possible design, the second image block includes three color components, the fidelity map is a three-dimensional array including three dimensions of color component, width and height, the two-dimensional array under any color component A in the fidelity map includes multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element, the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image. Wherein, when the fidelity map is a three-dimensional array including three dimensions of color component, width and height, the height indicates that the two-dimensional array under any color component A includes multiple row first elements, the width indicates that the two-dimensional array under any color component A includes multiple column first elements, the number of the multiple first elements is equal to the product of the width and the height, and the color component A is any one of the three color components. In an embodiment of the present application, the original image or the reconstructed image includes three color components. When calculating the fidelity map, a two-dimensional array fidelity map is calculated for any color component. The two-dimensional arrays for the three color components constitute a fidelity map for a three-dimensional array. The first element in the two-dimensional array for any color component A in the fidelity map of the three-dimensional array represents the fidelity of the reconstructed image or the area corresponding to its position in a preset area of the reconstructed image for any color component A. This is beneficial for the fidelity map of the three-dimensional array to represent the distortion intensity information of the three color components of the encoded image.
[0012] In a possible design, encoding the fidelity map to obtain the second code stream includes: performing entropy encoding on any of the first elements to obtain the second code stream, the entropy encoding of the any of the first elements being independent of the entropy encoding of other first elements; or, determining a probability distribution of a value of the any of the first elements or a predicted value of the any of the first elements according to a value of at least one of the encoded first elements, and performing entropy encoding on the any of the first elements according to the probability distribution of the value of the any of the first elements or the predicted value of the any of the first elements to obtain the second code stream; wherein the second code stream includes code streams of the multiple first elements. Specifically, for the entropy coding process of any first element in the fidelity map, if there is no encoded first element, entropy coding is directly performed on the any first element to obtain a code stream of the any first element; if there is an encoded first element, the probability distribution of the value of the any first element or the predicted value of the any first element is determined according to the value of at least one first element in the encoded first elements, and entropy coding is performed on the any first element according to the probability distribution of the value of the any first element or the predicted value of the any first element to obtain a code stream of the any first element; wherein the second code stream includes the code streams of the multiple first elements. In an embodiment of the present application, a fidelity map is encoded to obtain a second code stream, that is, any first element in the fidelity map is encoded to obtain a code stream of any first element in the fidelity map, and the second code stream includes a code stream of any first element in the fidelity map; in the entropy coding process, the value of the encoded first element can be used to determine the probability distribution of the value of the currently encoded first element or the predicted value of the currently encoded first element, for example, the value of the adjacent first element to the left, above, upper left, etc. of the currently encoded first element is used to determine the probability distribution of the value of the currently encoded first element, and then the currently encoded first element is entropy coded according to the probability distribution of the value of the currently encoded first element or the predicted value of the currently encoded first element, so as to assist in improving the entropy coding efficiency.
[0013] In a possible design, encoding the fidelity map to obtain a second code stream includes: quantizing any of the first elements to obtain a quantized first element; encoding the quantized first element to obtain the second code stream; wherein the second code stream includes a code stream of the multiple first elements. The quantization step size for quantizing any of the multiple first elements may be the same or different. In an embodiment of the present application, the fidelity map is encoded to obtain a second code stream, that is, any first element in the fidelity map is encoded to obtain a code stream of any first element in the fidelity map, and the second code stream includes the code stream of any first element in the fidelity map; in the process of encoding any first element in the fidelity map, any first element can be quantized, and then any quantized first element can be encoded to obtain a code stream of any first element; quantizing any first element is to quantize the value of any first element, or to scale any fidelity value in the fidelity map; the purpose of quantization is to reduce the dynamic range of the fidelity value in the fidelity map to reduce the encoding overhead of the fidelity map.
[0014] According to a second aspect, the present application relates to a decoding method. The method is performed by a decoding device. The method comprises: decoding a first code stream to obtain a reconstructed image of an original image; decoding a second code stream to obtain a reconstructed image of a fidelity map, wherein the second code stream is obtained by encoding the fidelity map, and the reconstructed image of the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image. In an embodiment of the present application, an original image is encoded to obtain a first code stream, and a fidelity map is encoded to obtain a second code stream, wherein the fidelity map is used to represent the distortion between at least a portion of an original image and at least a portion of a reconstructed image, wherein the distortion includes a difference; a decoding end decodes the first code stream to obtain a reconstructed image of the original image, and a decoding end decodes the second code stream to obtain a reconstructed map of the fidelity map; and if the encoding is lossless encoding, the reconstructed map of the fidelity map is the same as the fidelity map; if the encoding is lossy encoding, the reconstructed map of the fidelity map includes the encoding distortion generated by encoding the fidelity map; therefore, the reconstructed map of the fidelity map can be used to represent the distortion between at least a portion of an original image and at least a portion of a reconstructed image; therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0015] In a possible design, the fidelity map includes a fidelity value of any second image block among a plurality of second image blocks, and the fidelity value of any second image block is used to represent the distortion between the any second image block and the original image block corresponding to the any second image block. The plurality of second image blocks are obtained by dividing the reconstructed image, the plurality of second image blocks correspond one-to-one to the plurality of original image blocks, the original image blocks are image blocks in the original image, for example, the original image blocks are the aforementioned first image blocks; the plurality of original image blocks are obtained by dividing the original image, the plurality of second image blocks are obtained by dividing the reconstructed image, and the division strategy of dividing the original image is the same as the division strategy of dividing the reconstructed image; or the plurality of original image blocks are obtained by dividing a preset area of the original image, the plurality of second image blocks are obtained by dividing a preset area of the reconstructed image, and the division strategy of dividing the preset area of the original image is the same as the division strategy of dividing the preset area of the reconstructed image.
[0016] In one possible design, the fidelity map includes multiple first elements, the multiple second image blocks correspond one-to-one to the multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0017] In one possible design, the second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions: color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0018] In one possible design, the decoding of the second code stream to obtain a reconstructed image of the fidelity map includes: decoding the second code stream to obtain a reconstruction fidelity value of any of the first elements; obtaining a reconstructed image of the fidelity map according to the reconstruction fidelity value of any of the first elements. The reconstruction fidelity value of the first element is also a reconstruction of the value of the first element. The position of the reconstruction fidelity value of any of the first elements in the reconstructed image of the fidelity map is determined according to the position of the second image block corresponding to any of the first elements in the reconstructed image. Alternatively, the second code stream includes the position of any of the first elements in the fidelity map; the position of the reconstruction fidelity value of any of the first elements in the reconstructed image of the fidelity map is determined according to the position of any of the first elements in the fidelity map. In the embodiment of the present application, the second code stream includes the code stream of any first element in the fidelity map, so the reconstruction fidelity value of any first element can be obtained by decoding the second code stream; it should be understood that if the encoding is lossless encoding, the reconstruction fidelity value of the first element is the value of the first element; if the encoding is lossy encoding, the reconstruction fidelity value of the first element includes the encoding distortion generated by encoding the first element, and the reconstruction fidelity value of the first element is the sum of the value of the first element and the encoding distortion; thus, according to the reconstruction fidelity value of any first element, a reconstruction map of the fidelity map can be obtained, so the reconstruction map of the fidelity map can be used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image; therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0019] In one possible design, the second code stream is obtained by encoding the quantized first element; and the decoding of the second code stream to obtain a reconstructed image of the fidelity map includes: decoding the second code stream to obtain a reconstruction fidelity value of the quantized first element; dequantizing the reconstruction fidelity value of the quantized first element to obtain a reconstruction fidelity value of any first element; and obtaining the reconstructed image of the fidelity map according to the reconstruction fidelity value of any first element. In an embodiment of the present application, in order to reduce the coding overhead, the encoding end may quantize the first element and then encode it to obtain a code stream of the first element, so the code stream of the first element obtained by the decoding end may be obtained by encoding the quantized first element; in this case, the second code stream is decoded to obtain the reconstruction fidelity value of the quantized first element, and the reconstruction fidelity value of the quantized first element needs to be dequantized to obtain the reconstruction fidelity value of any first element; then, a reconstruction map of the fidelity map can be obtained based on the reconstruction fidelity value of any first element; in this way, the distortion intensity information of the encoded image can be obtained at the decoding end, and the coding overhead can be reduced.
[0020] In a possible design, the method further includes: processing the reconstructed image or the preset area of the reconstructed image according to the reconstruction map of the fidelity map to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determining whether to apply the reconstructed image according to the reconstruction map of the fidelity map. In an embodiment of the present application, the decoding end may process the reconstructed image or the preset area of the reconstructed image according to the reconstruction map of the fidelity map to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determine whether to apply the reconstructed image according to the reconstruction map of the fidelity map; thereby facilitating the application of the reconstructed image.
[0021] According to a third aspect, the present application relates to a decoding method. The method is performed by a decoding device. The method comprises: decoding a first code stream to obtain a reconstructed image of an original image and target quantization parameter information, wherein the target quantization parameter information comprises quantization parameter values of all or part of the second image blocks in a plurality of second image blocks of the reconstructed image; constructing a quantization parameter map of the reconstructed image according to the target quantization parameter information, wherein the quantization parameter map of the reconstructed image is used to represent the distortion between at least a portion of the area of the original image and at least a portion of the area of the reconstructed image. It should be understood that at the decoding end, the main purpose of the quantization parameter is to perform an inverse quantization operation; of course, the quantization parameter itself represents the signal distortion (fidelity), so the quantization parameter map of the reconstructed image constructed according to the quantization parameter can be used to represent the distortion between at least a portion of the area of the original image and at least a portion of the area of the reconstructed image. In the prior art, when the decoding end performs a decoding operation on the first code stream obtained by encoding the original image, the quantization parameter values of each area (second image block) of the reconstructed image will not be obtained. In an embodiment of the present application, the decoding end can obtain the quantization parameter values of each area (second image block) of the reconstructed image by decoding the first code stream obtained by encoding the original image; and a quantization parameter map of the reconstructed image can be constructed according to the quantization parameter values of all or part of the multiple second image blocks of the reconstructed image, and the quantization parameter map of the reconstructed image can be used to represent the distortion between at least a part of the area of the original image and at least a part of the area of the reconstructed image. Therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0022] In a possible design, the second image block is a coding unit. In an embodiment of the present application, the encoding end divides the original image into multiple coding units, and encodes the multiple coding units obtained by dividing the original image to obtain a first code stream; the decoding end decodes the first code stream to obtain a reconstructed image and target quantization parameter information of the original image, and the target quantization parameter information includes the quantization parameter values of all or part of the coding units in the multiple coding units; according to the target quantization parameter information, a quantization parameter map of the reconstructed image can be constructed; and the quantization parameter map of the reconstructed image is a form of fidelity map, when the target quantization parameter information includes the quantization parameter values of all the coding units in the multiple coding units, the quantization parameter map of the reconstructed image is a fidelity map of the entire reconstructed image; when the target quantization parameter information includes the quantization parameter values of part of the coding units in the multiple coding units, the quantization parameter map of the reconstructed image is a fidelity map of a preset area of the reconstructed image; therefore, the quantization parameter map of the reconstructed image can be used to characterize the fidelity of the reconstructed image or to characterize the fidelity of a preset area of the reconstructed image; therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0023] In a possible design, the quantization parameter map of the reconstructed image includes a plurality of second elements, the plurality of second image blocks correspond one-to-one to the plurality of second elements, the value of any second element among the plurality of second elements is the quantization parameter value of the second image block corresponding to the any second element, the position of the any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of the any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image. The second element may also be referred to as a pixel point of the quantization parameter map.
[0024] In one possible design, the second image block includes three color components, the quantization parameter map of the reconstructed image is a three-dimensional array including three dimensions of color component, width and height, the two-dimensional array under any color component A in the quantization parameter map of the reconstructed image includes multiple second elements, the value of any second element among the multiple second elements is the quantization parameter value of the color component A of the second image block corresponding to the any second element, the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image.
[0025] In one possible design, constructing the quantization parameter map of the reconstructed image according to the target quantization parameter information includes: when the target quantization parameter information includes quantization parameter values of some coding units among the multiple coding units, obtaining the quantization parameter value of the target coding unit according to the quantization parameter value of the some coding units and / or the reference quantization parameter map, wherein the reference quantization parameter map is the quantization parameter map of the reference image of the reconstructed image, and the target coding unit is the coding unit among the multiple coding units except the some coding units; obtaining the quantization parameter map of the reconstructed image according to the quantization parameter value of the some coding units and the quantization parameter value of the target coding unit. Specifically, when the target quantization parameter information includes the quantization parameter values of all the coding units among the multiple coding units, the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of all the coding units; when the target quantization parameter information includes the quantization parameter values of some of the coding units among the multiple coding units, the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of the some of the coding units, or the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of the some of the coding units and a reference quantization parameter map, wherein the reference quantization parameter map is a quantization parameter map of a reference image of the reconstructed image. In an embodiment of the present application, when the target quantization parameter information includes the quantization parameter values of all coding units among a plurality of coding units, a quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of all coding units, and the quantization parameter map of the reconstructed image obtained in this case can be used to characterize the fidelity of the entire reconstructed image; when the target quantization parameter information includes the quantization parameter values of some coding units among a plurality of coding units, the quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some coding units, and the quantization parameter map of the reconstructed image obtained in this case can be used to characterize the fidelity of a preset area of the reconstructed image; when the target quantization parameter information includes the quantization parameter values of some coding units among a plurality of coding units, the quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some coding units and the reference quantization parameter map A quantization parameter map of a reconstructed image. Since the reference quantization parameter map is a quantization parameter map of a reference image of the reconstructed image, the quantization parameter value of any one of the multiple coding units except the part of the coding units can be obtained according to the reference quantization parameter map. In this way, the quantization parameter of any one of the multiple coding units can be obtained. The quantization parameter map of the reconstructed image obtained in this case can be used to characterize the fidelity of the entire reconstructed image or to characterize the fidelity of a preset area of the reconstructed image. Therefore, in any case in which the target quantization parameter information obtained by decoding includes the quantization parameter values of all or part of the coding units in the multiple coding units, the embodiment of the present application can obtain a quantization parameter map of the reconstructed image for characterizing the fidelity of the reconstructed image or for characterizing the fidelity of the preset area of the reconstructed image.
[0026] In one possible design, obtaining the quantization parameter value of the target coding unit according to the quantization parameter values of the partial coding units includes: determining the quantization parameter value of the target coding unit according to the quantization parameter value of at least one coding unit in the partial coding units. In an embodiment of the present application, when the quantization parameter value of a certain coding unit cannot be obtained from the first bitstream, the quantization parameter of the spatial neighborhood of the coding unit can be used for filling. Specifically, when the target quantization parameter information includes quantization parameter values of some coding units among multiple coding units, the quantization parameter value of any one coding unit among the multiple coding units except the some coding units can also be determined according to the quantization parameter value of at least one coding unit among the some coding units, so that the quantization parameter value of any one coding unit among the multiple coding units can be ensured; and, a quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some or all coding units among the multiple coding units; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of all coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of the entire reconstructed image; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of some coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of a preset area of the reconstructed image.
[0027] In a possible design, the reference quantization parameter map includes multiple reference elements, and the value of any reference element among the multiple reference elements is the quantization parameter value of the coding unit in the reference image; the quantization parameter value of the target coding unit is obtained according to the reference quantization parameter map, including: taking the value of the target element as the quantization parameter value of any target coding unit in the target coding unit, wherein the target element is a reference element in the reference quantization parameter map, and the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image, or the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image and the motion vector of any target coding unit. The reference element is another name for the second element. In an embodiment of the present application, when the quantization parameter value of a certain coding unit cannot be obtained from the first code stream, the quantization parameter of the temporal neighborhood of the coding unit can be used for filling. Specifically, for any coding unit whose quantization parameter value cannot be obtained from the first bitstream, the value of the target element in the reference quantization parameter map can be used as the quantization parameter value of the coding unit, and the position of the target element in the reference quantization parameter map is determined according to the position of the coding unit in the reconstructed image, or the position of the target element in the reference quantization parameter map is determined according to the position of the coding unit in the reconstructed image and the motion vector of the coding unit. In this way, the quantization parameter value of any coding unit among the multiple coding units can be obtained; and the quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some or all coding units among the multiple coding units; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of all coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of the entire reconstructed image; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of some coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of the preset area of the reconstructed image.
[0028] In one possible design, the method further includes: storing the reconstructed image and the quantization parameter map of the reconstructed image in association with each other, so as to use the reconstructed image as a reference image and the quantization parameter map of the reconstructed image as a reference quantization parameter map. In an embodiment of the present application, the quantization parameter map of the reconstructed image can be stored as a reference quantization parameter map for constructing a quantization parameter map of a subsequent decoded image, thereby facilitating the construction of a quantization parameter map of a subsequent decoded image.
[0029] In a possible design, the method further includes: processing the reconstructed image or a preset area of the reconstructed image according to the quantization parameter map of the reconstructed image to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determining whether to apply the reconstructed image according to the quantization parameter map of the reconstructed image. In an embodiment of the present application, the decoding end may process the reconstructed image or the preset area of the reconstructed image according to the reconstruction map of the fidelity map to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determine whether to apply the reconstructed image according to the reconstruction map of the fidelity map; thereby facilitating the application of the reconstructed image.
[0030] According to a fourth aspect, the present application relates to an encoding device, and the beneficial effects can be found in the description of the first aspect, which will not be repeated here. The encoding device includes: a video encoder, which is used to encode an original image to obtain a first code stream; a fidelity map encoder, which is used to encode a fidelity map to obtain a second code stream, wherein the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of a reconstructed image, and the reconstructed image is obtained after decoding the first code stream.
[0031] In a possible design, the encoding device also includes a fidelity map calculator, which is used to: divide the original image into multiple first image blocks, and divide the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the original image is the same as the division strategy for dividing the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; or divide a preset area of the original image into multiple first image blocks, and divide the preset area of the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the preset area of the original image is the same as the division strategy for dividing the preset area of the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; calculate the fidelity value of any second image block according to any second image block among the multiple second image blocks and the first image block corresponding to the any second image block, and the fidelity map includes the fidelity value of any second image block, and the fidelity value of any second image block is used to represent the distortion between any second image block and the first image block corresponding to the any second image block.
[0032] In one possible design, the fidelity map includes multiple first elements, the multiple second image blocks correspond one-to-one to the multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0033] In one possible design, the second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions: color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0034] In a possible design, the fidelity map encoder is specifically used to: perform entropy encoding on any of the first elements to obtain the second code stream, and the entropy encoding of the any of the first elements is independent of the entropy encoding of other first elements; or, determine the probability distribution of the value of any of the first elements or the predicted value of any of the first elements according to the value of at least one first element among the encoded first elements, and perform entropy encoding on any of the first elements according to the probability distribution of the value of any of the first elements or the predicted value of any of the first elements to obtain the second code stream; wherein the second code stream includes the code streams of the multiple first elements.
[0035] In one possible design, the fidelity map encoder is specifically used to: quantize any one of the first elements to obtain a quantized first element; encode the quantized first element to obtain the second code stream; wherein the second code stream includes the code streams of the multiple first elements.
[0036] According to a fifth aspect, the present application relates to a decoding device, and the beneficial effects can be found in the description of the second aspect, which will not be repeated here. The decoding device includes: a video decoder, which is used to decode a first bitstream to obtain a reconstructed image of the original image; a fidelity map decoder, which is used to decode a second bitstream to obtain a reconstructed image of a fidelity map, wherein the second bitstream is obtained by encoding the fidelity map, and the reconstructed image of the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image.
[0037] In one possible design, the fidelity map includes a fidelity value of any second image block among a plurality of second image blocks, and the fidelity value of any second image block is used to represent the distortion between the any second image block and an original image block corresponding to the any second image block.
[0038] In one possible design, the fidelity map includes multiple first elements, the multiple second image blocks correspond one-to-one to the multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0039] In one possible design, the second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions: color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0040] In one possible design, the fidelity map decoder is specifically used to: decode the second code stream to obtain a reconstruction fidelity value of any of the first elements; and obtain a reconstruction map of the fidelity map based on the reconstruction fidelity value of any of the first elements.
[0041] In one possible design, the second code stream is obtained by encoding the quantized first element; the fidelity map decoder is specifically used to: decode the second code stream to obtain the reconstruction fidelity value of the quantized first element; dequantize the reconstruction fidelity value of the quantized first element to obtain the reconstruction fidelity value of any first element; and obtain a reconstruction map of the fidelity map according to the reconstruction fidelity value of any first element.
[0042] According to the sixth aspect, the present application relates to a decoding device, and the beneficial effects can be found in the description of the third aspect, which will not be repeated here. The decoding device includes: a video decoder, which is used to decode a first code stream to obtain a reconstructed image and target quantization parameter information of the original image, wherein the target quantization parameter information includes the quantization parameter values of all or part of the second image blocks in the multiple second image blocks of the reconstructed image; a quantization parameter map builder, which is used to construct a quantization parameter map of the reconstructed image according to the target quantization parameter information, wherein the quantization parameter map of the reconstructed image is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image.
[0043] In one possible design, the second image block is a coding unit.
[0044] In one possible design, the quantization parameter map of the reconstructed image includes multiple second elements, the multiple second image blocks correspond one-to-one to the multiple second elements, the value of any second element among the multiple second elements is the quantization parameter value of the second image block corresponding to the any second element, the position of any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image.
[0045] In one possible design, the second image block includes three color components, the quantization parameter map of the reconstructed image is a three-dimensional array including three dimensions of color component, width and height, the two-dimensional array under any color component A in the quantization parameter map of the reconstructed image includes multiple second elements, the value of any second element among the multiple second elements is the quantization parameter value of the color component A of the second image block corresponding to the any second element, the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image.
[0046] In one possible design, the quantization parameter map builder is specifically used to: when the target quantization parameter information includes quantization parameter values of some coding units among the multiple coding units, obtain the quantization parameter value of the target coding unit according to the quantization parameter value of the some coding units and / or the reference quantization parameter map, wherein the reference quantization parameter map is the quantization parameter map of the reference image of the reconstructed image, and the target coding unit is the coding unit other than the some coding units among the multiple coding units; obtain the quantization parameter map of the reconstructed image according to the quantization parameter value of the some coding units and the quantization parameter value of the target coding unit.
[0047] In one possible design, the reference quantization parameter map includes multiple reference elements, and the value of any reference element among the multiple reference elements is the quantization parameter value of the coding unit in the reference image; the quantization parameter map builder is specifically used to: use the value of the target element as the quantization parameter value of any target coding unit among the target coding units, wherein the target element is a reference element in the reference quantization parameter map, and the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image, or the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image and the motion vector of any target coding unit.
[0048] According to the seventh aspect, the present application relates to a coding device, and the beneficial effects can be found in the description of the first aspect, which will not be repeated here. The coding device has the function of implementing the behavior in the method instance of the first aspect above. The function can be implemented by hardware, or by hardware executing corresponding software implementation. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the decoding device includes a processing unit, and the processing unit is used to: encode the original image to obtain a first code stream; encode the fidelity map to obtain a second code stream, wherein the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image, and the reconstructed image is obtained after decoding the first code stream.
[0049] In a possible design, the processing unit is further used to: divide the original image into multiple first image blocks, and divide the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the original image is the same as the division strategy for dividing the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; or divide a preset area of the original image into multiple first image blocks, and divide the preset area of the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the preset area of the original image is the same as the division strategy for dividing the preset area of the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; calculate the fidelity value of any second image block according to any second image block among the multiple second image blocks and the first image block corresponding to the any second image block, and the fidelity map includes the fidelity value of any second image block, and the fidelity value of any second image block is used to represent the distortion between any second image block and the first image block corresponding to the any second image block.
[0050] In one possible design, the fidelity map includes multiple first elements, the multiple second image blocks correspond one-to-one to the multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0051] In one possible design, the second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions: color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0052] In a possible design, the processing unit is specifically used to: perform entropy encoding on any of the first elements to obtain the second code stream, and the entropy encoding of the any of the first elements is independent of the entropy encoding of other first elements; or, determine the probability distribution of the value of any of the first elements or the predicted value of any of the first elements according to the value of at least one first element among the encoded first elements, and perform entropy encoding on any of the first elements according to the probability distribution of the value of any of the first elements or the predicted value of any of the first elements to obtain the second code stream; wherein the second code stream includes the code streams of the multiple first elements.
[0053] In one possible design, the processing unit is specifically used to: quantize any one of the first elements to obtain a quantized first element; encode the quantized first element to obtain the second code stream; wherein the second code stream includes the code streams of the multiple first elements.
[0054] According to an eighth aspect, the present application relates to an encoding device, and the beneficial effects can be found in the description of the second aspect, which will not be repeated here. The encoding device has the function of implementing the behavior in the method instance of the second aspect. The function can be implemented by hardware, or by hardware executing corresponding software implementation. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the decoding device includes a processing unit, and the processing unit is used to: decode a first code stream to obtain a reconstructed image of the original image; decode a second code stream to obtain a reconstructed image of a fidelity map, wherein the second code stream is obtained by encoding the fidelity map, and the reconstructed image of the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image.
[0055] In one possible design, the fidelity map includes a fidelity value of any second image block among a plurality of second image blocks, and the fidelity value of any second image block is used to represent the distortion between the any second image block and an original image block corresponding to the any second image block.
[0056] In one possible design, the fidelity map includes multiple first elements, the multiple second image blocks correspond one-to-one to the multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0057] In one possible design, the second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions: color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0058] In a possible design, the processing unit is specifically used to: decode the second code stream to obtain a reconstruction fidelity value of any of the first elements; and obtain a reconstruction map of the fidelity map according to the reconstruction fidelity value of any of the first elements.
[0059] In a possible design, the second code stream is obtained by encoding the quantized first element; the processing unit is specifically used to: decode the second code stream to obtain the reconstruction fidelity value of the quantized first element; dequantize the reconstruction fidelity value of the quantized first element to obtain the reconstruction fidelity value of any first element; and obtain a reconstruction map of the fidelity map according to the reconstruction fidelity value of any first element.
[0060] According to the ninth aspect, the present application relates to a coding device, and the beneficial effects can be found in the description of the third aspect, which will not be repeated here. The coding device has the function of implementing the behavior in the method instance of the third aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the decoding device includes a processing unit, which is used to: decode the first code stream to obtain a reconstructed image and target quantization parameter information of the original image, and the target quantization parameter information includes the quantization parameter values of all or part of the second image blocks in the multiple second image blocks of the reconstructed image; construct a quantization parameter map of the reconstructed image according to the target quantization parameter information, wherein the quantization parameter map of the reconstructed image is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image.
[0061] In one possible design, the second image block is a coding unit.
[0062] In one possible design, the quantization parameter map of the reconstructed image includes multiple second elements, the multiple second image blocks correspond one-to-one to the multiple second elements, the value of any second element among the multiple second elements is the quantization parameter value of the second image block corresponding to the any second element, the position of any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image.
[0063] In one possible design, the second image block includes three color components, the quantization parameter map of the reconstructed image is a three-dimensional array including three dimensions of color component, width and height, the two-dimensional array under any color component A in the quantization parameter map of the reconstructed image includes multiple second elements, the value of any second element among the multiple second elements is the quantization parameter value of the color component A of the second image block corresponding to the any second element, the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image.
[0064] In one possible design, the processing unit is specifically used to: when the target quantization parameter information includes quantization parameter values of some coding units among the multiple coding units, obtain the quantization parameter value of the target coding unit according to the quantization parameter values of the some coding units and / or a reference quantization parameter map, wherein the reference quantization parameter map is a quantization parameter map of a reference image of the reconstructed image, and the target coding unit is a coding unit among the multiple coding units other than the some coding units; obtain the quantization parameter map of the reconstructed image according to the quantization parameter values of the some coding units and the quantization parameter value of the target coding unit.
[0065] In one possible design, the reference quantization parameter map includes multiple reference elements, and the value of any reference element among the multiple reference elements is the quantization parameter value of the coding unit in the reference image; the processing unit is specifically used to: use the value of the target element as the quantization parameter value of any target coding unit among the target coding units, wherein the target element is a reference element in the reference quantization parameter map, and the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image, or the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image and the motion vector of any target coding unit.
[0066] The method described in the first aspect of the present application can be performed by the device described in the seventh aspect of the present application. Other features and implementations of the method described in the first aspect of the present application directly depend on the functionality and implementation of the device described in the seventh aspect of the present application.
[0067] The method described in the second aspect of the present application can be performed by the device described in the eighth aspect of the present application. Other features and implementations of the method described in the second aspect of the present application directly depend on the functionality and implementation of the device described in the eighth aspect of the present application.
[0068] The method described in the third aspect of the present application can be performed by the device described in the ninth aspect of the present application. Other features and implementations of the method described in the second aspect of the present application directly depend on the functionality and implementation of the device described in the ninth aspect of the present application.
[0069] According to a tenth aspect, the present application relates to a device for encoding a video stream, comprising a processor and a memory, wherein the memory stores instructions, and the instructions enable the processor to execute the method described in the first aspect.
[0070] According to an eleventh aspect, the present application relates to a device for decoding a video stream, comprising a processor and a memory, wherein the memory stores instructions, and the instructions enable the processor to execute the method described in the second aspect.
[0071] According to a twelfth aspect, the present application relates to a device for decoding a video stream, comprising a processor and a memory, wherein the memory stores instructions, and the instructions enable the processor to execute the method described in the third aspect.
[0072] According to a thirteenth aspect, the present application provides a computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to encode video data. The instructions cause the one or more processors to perform the method of the first, second, or third aspect or any possible embodiment of the first, second, or third aspect.
[0073] According to a fourteenth aspect, the present application relates to a computer program product comprising a program code, which, when running, executes the method in the first, second or third aspect or any possible embodiment of the first, second or third aspect.
[0074] According to a fifteenth aspect, the present application relates to an encoder (20), comprising a processing circuit for executing the method in the first aspect or any possible embodiment of the first aspect.
[0075] According to a sixteenth aspect, the present application relates to a decoder (30) comprising a processing circuit for executing the method of the second or third aspect or any possible embodiment of the second or third aspect.
[0076] According to the seventeenth aspect, the present application relates to an encoder, comprising: one or more processors; a non-transitory computer-readable storage medium, coupled to the processor and storing a program executed by the processor, wherein when the program is executed by the processor, the encoder executes the method in the first aspect or any possible embodiment of the first aspect.
[0077] According to the eighteenth aspect, the present application relates to a decoder, comprising: one or more processors; a non-transitory computer-readable storage medium, coupled to the processor and storing a program executed by the processor, wherein the program, when executed by the processor, causes the decoder to execute the method in the second or third aspect or any possible embodiment of the second or third aspect.
[0078] According to the nineteenth aspect, the present application relates to a non-transitory computer-readable storage medium, comprising program code, which, when executed by a computer device, is used to execute the method in the first, second or third aspect or any possible embodiment of the first, second or third aspect.
[0079] According to a twentieth aspect, the present application relates to a non-transitory storage medium, comprising a bit stream encoded according to the method in the first aspect or any possible embodiment of the first aspect.
[0080] According to the twenty-first aspect, the present application relates to an electronic device, which includes the encoding device described in the fourth aspect and / or the decoding device described in the fifth aspect or the sixth aspect.
[0081] One or more embodiments will be described in detail in the drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The following is an introduction to the drawings used in the embodiments of the present application.
[0083] Figure 1A A block diagram of an example of a video decoding system for implementing an embodiment of the present application, wherein the system uses a neural network to encode or decode a video image;
[0084] Figure 1B A block diagram of another example of a video decoding system for implementing an embodiment of the present application, wherein the video encoder and / or the video decoder uses a neural network to encode or decode a video image;
[0085] Figure 2 A block diagram of an example of a video encoder for implementing an embodiment of the present application, wherein the video encoder 20 uses a neural network to encode video images;
[0086] Figure 3 A block diagram of an example of a video decoder for implementing an embodiment of the present application, wherein the video decoder 30 uses a neural network to decode video images;
[0087] Figure 4 is a schematic block diagram of a video decoding device for implementing an embodiment of the present application;
[0088] Figure 5 is a schematic block diagram of a video decoding device for implementing an embodiment of the present application;
[0089] Figure 6 A schematic diagram of an image codec based on a deep neural network for implementing an embodiment of the present application;
[0090] Figure 7 A schematic block diagram of an encoding method provided in an embodiment of the present application;
[0091] Figure 8 A schematic diagram of dividing an original image or a preset area of an original image provided in an embodiment of the present application;
[0092] Fig. 9 A schematic diagram of dividing a reconstructed image or a preset area of a reconstructed image provided in an embodiment of the present application;
[0093] Fig.10A schematic diagram of a fidelity map provided in an embodiment of the present application;
[0094] Fig.11 A schematic block diagram of a decoding method provided in an embodiment of the present application;
[0095] Fig.12 A schematic diagram of a reconstruction image of a fidelity map provided in an embodiment of the present application;
[0096] Fig.13 A schematic block diagram of another decoding method provided in an embodiment of the present application;
[0097] Fig.14 A schematic diagram of a quantization parameter map provided in an embodiment of the present application;
[0098] Fig.15 A schematic diagram of another quantization parameter map provided in an embodiment of the present application;
[0099] Fig.16 A schematic block diagram of an encoding device provided in an embodiment of the present application;
[0100] Fig.17 A schematic block diagram of a decoding device provided in an embodiment of the present application;
[0101] Fig.18 A schematic block diagram of a decoding device provided in an embodiment of the present application;
[0102] Fig.19 A schematic block diagram of a coding device provided in an embodiment of the present application;
[0103] Fig. 20 A schematic block diagram of a decoding device provided in an embodiment of the present application;
[0104] Fig.21 A schematic block diagram of a decoding device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0105] The embodiments of the present application provide a video image compression technology based on artificial intelligence, and in particular, provide a video compression technology based on a neural network (NN), and specifically provide an inter-frame prediction technology / intra-frame prediction technology / filtering technology based on a neural network to improve the traditional hybrid video encoding and decoding system.
[0106] Video coding generally refers to processing a sequence of images that form a video or video sequence. In the field of video coding, the terms "picture", "frame" or "image" can be used as synonyms. Video coding (or commonly referred to as coding) includes two parts: video encoding and video decoding. Video coding is performed on the source side and generally includes processing (e.g., compressing) the original video image to reduce the amount of data required to represent the video image (thereby more efficiently storing and / or transmitting). Video decoding is performed on the destination side and generally includes inverse processing relative to the encoder to reconstruct the video image. The "encoding" of the video image (or commonly referred to as the image) involved in the embodiment should be understood as the "encoding" or "decoding" of the video image or video sequence. The encoding part and the decoding part are also collectively referred to as codec (encoding and decoding, CODEC).
[0107] In the case of lossless video coding, the original video image can be reconstructed, that is, the reconstructed video image has the same quality as the original video image (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy video coding, further compression is performed by quantization, etc. to reduce the amount of data required to represent the video image, but the decoder side cannot fully reconstruct the video image, that is, the quality of the reconstructed video image is lower or worse than the quality of the original video image.
[0108] Several video coding standards belong to the category of "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the pixel domain with 2D transform coding in the transform domain for applying quantization). Any image in a video sequence is usually divided into a set of non-overlapping blocks, which are usually encoded at the block level. In other words, the encoder usually processes, i.e., encodes the video at the block (video block) level, for example, by generating a prediction block through spatial (intra-frame) prediction and temporal (inter-frame) prediction; subtracting the prediction block from the current block (currently processed / to be processed block) to obtain a residual block; transforming the residual block in the transform domain and quantizing the residual block to reduce the amount of data to be transmitted (compressed), while the decoder side applies the inverse processing part relative to the encoder to the encoded or compressed block to reconstruct the current block for representation. In addition, the encoder needs to repeat the processing steps of the decoder so that the encoder and the decoder generate the same predictions (e.g., intra-frame predictions and inter-frame predictions) and / or reconstructed pixels for processing, i.e., encoding subsequent blocks.
[0109] In the following embodiment of the decoding system 10, the encoder 20 and the decoder 30 are based on Figures 1A to 3 Give a description.
[0110] Figure 1A1 is a schematic block diagram of an exemplary decoding system 10, such as a video decoding system 10 (or simply a decoding system 10) that can utilize the techniques of the present application. The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) in the video decoding system 10 represent devices that can be used to perform various techniques according to various examples described in the present application.
[0111] like Figure 1A As shown, the decoding system 10 includes a source device 12, and the source device 12 is used to provide coded image data 21 such as coded images to a destination device 14 for decoding the coded image data 21. The coded image data is also a bit stream, a compressed code stream or a code stream, so the coded image data 21 can also be called a bit stream 21, a compressed code stream 21 or a code stream 21.
[0112] The source device 12 includes an encoder 20 , and may additionally or optionally include an image source 16 , a preprocessor (or a preprocessing unit) 18 such as an image preprocessor, and a communication interface (or a communication unit) 22 .
[0113] The image source 16 may include or may be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer-animated images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage that stores any of the above images.
[0114] In order to distinguish the processing performed by the pre-processor (or pre-processing unit) 18 , the image (or image data) 17 may also be referred to as a raw image (or raw image data) 17 .
[0115] The preprocessor 18 is used to receive (raw) image data 17 and preprocess the image data 17 to obtain a preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color adjustment, or denoising. It is understood that the preprocessing unit 18 may be an optional component.
[0116] The video encoder (or encoder) 20 is used to receive the pre-processed image data 19 and provide the encoded image data 21 (hereinafter referred to as Figure 2 etc. for further description).
[0117] The communication interface 22 in the source device 12 can be used to receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device through the communication channel 13 for storage or direct reconstruction.
[0118] The destination device 14 includes a decoder 30 and, in addition or alternatively, may include a communication interface (or communication unit) 28 , a post-processor (or post-processing unit) 32 , and a display device 34 .
[0119] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or any other processed version) directly from the source device 12 or from any other source device such as a storage device, for example, the storage device is a encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0120] The communication interface 22 and the communication interface 28 can be used to send or receive encoded image data (or encoded data) 21 through a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any type of combination thereof.
[0121] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission coding or processing for transmission over a communication link or network.
[0122] The communication interface 28 corresponds to the communication interface 22 , for example, and can be used to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain the encoded image data 21 .
[0123] Both the communication interface 22 and the communication interface 28 can be configured as follows Figure 1A A unidirectional communication interface, or a bidirectional communication interface, indicated by an arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14, and can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded image data transmission, etc.
[0124] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image, reconstructed image) 31 (hereinafter referred to as Figure 3 etc. for further description).
[0125] The post-processor 32 is used to post-process the decoded image data 31 (also called reconstructed image data) such as the decoded image to obtain the post-processed image data 33 such as the post-processed image. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color adjustment, cropping or resampling, or any other processing for generating the decoded image data 31 for display by the display device 34 or the like.
[0126] The display device 34 is used to receive the post-processed image data 33 to display the image to a user or viewer, etc. The display device 34 can be or include any type of display for displaying the reconstructed image, such as an integrated or external display screen or display. For example, the display screen can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.
[0127] The decoding system 10 also includes a training engine 25, which is used to train the encoder 20 (especially the mode selection unit 260 in the encoder 20) or the decoder 30 (especially the mode application unit 360 in the decoder 30) to process an input image or image region or image block to generate a prediction value of the input image or image region or image block.
[0128] In the embodiment of the present application, the training data used to train the encoder 20 or the decoder 30 can be stored in a database (not shown), and the training engine 25 obtains the target model (for example, a neural network for inter-frame prediction, intra-frame prediction, or loop filtering) based on the training data. It should be noted that the embodiment of the present application does not limit the source of the training data, for example, the training data can be obtained from the cloud or other places for model training.
[0129] The target model trained by the training engine 25 can be applied to the decoding system 10, 40, for example, Figure 1AThe source device 12 (e.g., encoder 20) or the destination device 14 (e.g., decoder 30) shown. The training engine 25 can be trained in the cloud to obtain a target model, and then the decoding system 10 downloads and uses the target model from the cloud; or, the training engine 25 can be trained in the cloud to obtain a target model and use the target model, and the decoding system 10 directly obtains the processing result from the cloud. For example, the training engine 25 trains to obtain a target model with a filtering function, and the decoding system 10 downloads the target model from the cloud, and then the loop filter 220 in the encoder 20 or the loop filter 320 in the decoder 30 can filter the input reconstructed image or image block according to the target model to obtain a filtered image or image block. For another example, the training engine 25 trains to obtain a target model with a filtering function, and the decoding system 10 does not need to download the target model from the cloud. The encoder 20 or the decoder 30 transmits the reconstructed image or image block to the cloud, and the cloud performs filtering on the reconstructed image or image block through the target model to obtain a filtered image or image block and transmit it to the encoder 20 or the decoder 30.
[0130] although Figure 1A The source device 12 and the destination device 14 are shown as independent devices, but the device embodiment may also include the source device 12 and the destination device 14 or the functions of the source device 12 and the destination device 14 at the same time, that is, the source device 12 or the corresponding function and the destination device 14 or the corresponding function at the same time. In these embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function can be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0131] According to the description, Figure 1A The presence and (exact) division of different units or functions in the source device 12 and / or the destination device 14 shown may vary depending on the actual device and application, which will be obvious to the skilled person.
[0132] The encoder 20 (eg, video encoder 20) or the decoder 30 (eg, video decoder 30) or both may be configured as follows: Figure 1B The processing circuitry shown may be implemented, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding dedicated processors, or any combination thereof. The encoder 20 may be implemented by processing circuitry 46 to include reference Figure 2The various modules discussed in encoder 20 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented by processing circuit 46 to include reference Figure 3 The processing circuit 46 may be used to perform the various operations discussed below. Figure 5 As shown, if part of the technology is implemented in software, the device can store the instructions of the software in a suitable non-transitory computer-readable storage medium, and use one or more processors to execute the instructions in hardware, thereby performing the technology of the present application. One of the video encoder 20 and the video decoder 30 can be integrated into a single device as part of a combined codec (encoder / decoder, CODEC), such as Figure 1B shown.
[0133] Source device 12 and destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smart phone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, etc., and may not use or use any type of operating system. In some cases, source device 12 and destination device 14 may be equipped with components for wireless communication. Therefore, source device 12 and destination device 14 may be wireless communication devices.
[0134] In some cases, Figure 1A The video decoding system 10 shown is merely exemplary, and the techniques provided herein may be applicable to video encoding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from a local memory, sent over a network, and so on. A video encoding device may encode data and store the data in a memory, and / or a video decoding device may retrieve data from a memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to a memory and / or retrieve and decode data from a memory.
[0135] Figure 1B According to an exemplary embodiment, Figure 2 The video encoder 20 and / or Figure 330. The video decoding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video codec implemented by a processing circuit 46), an antenna 42, one or more processors 43, one or more memory storages 44, and / or a display device 45.
[0136] like Figure 1B As shown, imaging device 41, antenna 42, processing circuit 46, video encoder 20, video decoder 30, processor 43, memory storage 44 and / or display device 45 can communicate with each other. In different examples, video decoding system 40 can include only video encoder 20 or only video decoder 30.
[0137] In some instances, antenna 42 can be used to transmit or receive a coded bit stream of video data. In addition, in some instances, display device 45 can be used to present video data. Processing circuit 46 can include application-specific integrated circuit (ASIC) logic, graphics processor, general processor, etc. Video decoding system 40 can also include optional processor 43, which can similarly include application-specific integrated circuit (ASIC) logic, graphics processor, general processor, etc. In addition, memory storage 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory storage 44 can be implemented by cache memory. In other instances, processing circuit 46 can include memory (e.g., cache, etc.) for implementing image buffer, etc.
[0138] In some examples, video encoder 20 implemented by logic circuits may include an image buffer (e.g., implemented by processing circuit 46 or memory storage 44) and a graphics processing unit (e.g., implemented by processing circuit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video encoder 20 implemented by processing circuit 46 to implement reference Figure 2 The various modules discussed in connection with the encoder and / or any other encoder system or subsystem described herein. Logic circuits may be used to perform the various operations discussed herein.
[0139] In some examples, the video decoder 30 may be implemented by the processing circuit 46 in a similar manner to implement the various modules discussed with reference to the video decoder 30 of FIG. 3 and / or any other decoder systems or subsystems described herein. In some examples, the logic circuit-implemented video decoder 30 may include an image buffer (implemented by the processing circuit 46 or the memory storage 44) and a graphics processing unit (implemented by the processing circuit 46, for example). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 30 implemented by the processing circuit 46 to implement the video decoder 30 of FIG. 3 and / or any other decoder systems or subsystems described herein. Figure 3 and / or the various modules discussed above for any other decoder system or subsystem described herein.
[0140] In some examples, antenna 42 may be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data related to the encoded video frames discussed herein, indicators, index values, mode selection data, etc., such as data related to the encoded partitions (e.g., transform coefficients or quantized transform coefficients, (as discussed) optional indicators, and / or data defining the encoded partitions). Video decoding system 40 may also include video decoder 30 coupled to antenna 42 and used to decode the encoded bitstream. Display device 45 is used to present the video frames.
[0141] It should be understood that for the examples described with reference to the video encoder 20 in the embodiments of the present application, the video decoder 30 can be used to perform the reverse process. With respect to the signaling syntax elements, the video decoder 30 can be used to receive and parse such syntax elements and decode the related video data accordingly. In some examples, the video encoder 20 can entropy encode the syntax elements into an encoded video bitstream. In such examples, the video decoder 30 can parse such syntax elements and decode the related video data accordingly.
[0142] For ease of description, the embodiments of the present application are described with reference to the Versatile video coding (VVC) reference software or the High-Efficiency Video Coding (HEVC) developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those of ordinary skill in the art understand that the embodiments of the present application are not limited to HEVC or VVC.
[0143] The following is an introduction to the encoder and encoding method as well as the decoder and decoding method.
[0144] 1. Encoder and encoding method
[0145] Figure 2 FIG. 2 is a schematic block diagram of an example of a video encoder 20 for implementing the technology of the present application. Figure 2 In the example of , the video encoder 20 includes an input terminal (or input interface) 201, a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (decoded picture buffer, DPB) 230, a mode selection unit 260, an entropy coding unit 270 and an output terminal (or output interface) 272. The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254 and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The illustrated video encoder 20 may also be referred to as a hybrid video encoder or a hybrid video codec-based video encoder.
[0146] See also Figure 2 The inter-frame prediction module / intra-frame prediction module / loop filtering module includes (is) a trained target model (also called a neural network) that is used to process an input image or an image region or an image block to generate a prediction value of the input image block. For example, the neural network for inter-frame prediction / intra-frame prediction / loop filtering is used to receive an input image or an image region or an image block.
[0147] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208 and the mode selection unit 260 constitute the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244 and the intra-frame prediction unit 254 constitute the backward signal path of the encoder, wherein the backward signal path of the encoder 20 corresponds to the signal path of the decoder (see Figure 3 The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded image buffer 230, the inter-frame prediction unit 244 and the intra-frame prediction unit 254 also constitute a “built-in decoder” of the video encoder 20.
[0148] (1) Images and image segmentation (images and blocks)
[0149] The encoder 20 may be configured to receive, via an input 201 or the like, an image (or image data) 17, e.g., an image in a sequence of images forming a video or a video sequence. The received image or image data may also be a pre-processed image (or pre-processed image data) 19. For simplicity, the following description uses image 17. Image 17 may also be referred to as a current image, an original image, or an image to be encoded (particularly when distinguishing the current image from other images in video encoding, e.g., previously encoded images and / or decoded images in the same video sequence, i.e., a video sequence also including the current image).
[0150] A (digital) image is or can be considered as a two-dimensional array or matrix of pixels with intensity values. The pixels in the array can also be called pixels (pixel or pel) (short for picture elements). The number of pixels in the array or image in the horizontal and vertical directions (or axes) determines the size and / or resolution of the image. In order to represent color, three color components are usually used, that is, the image can be represented as or include three pixel arrays. In the RBG format or color space, the image includes corresponding red, green and blue pixel arrays. However, in video coding, any pixel is usually represented in a brightness / chrominance format or color space, such as YCbCr, including a brightness component indicated by Y (sometimes also represented by L) and two chrominance components represented by Cb and Cr. The brightness (luma) component Y represents the brightness or grayscale level intensity (for example, the two are the same in grayscale images), while the two chrominance (abbreviated as chroma) components Cb and Cr represent the chrominance or color information components. Accordingly, an image in YCbCr format includes a brightness pixel array of brightness pixel values (Y) and two chrominance pixel arrays of chrominance values (Cb and Cr). An image in RGB format can be converted or transformed into YCbCr format, and vice versa, a process also referred to as color conversion or transformation. If the image is black and white, the image may include only a brightness pixel array. Accordingly, the image may be, for example, a brightness pixel array in a monochrome format or a brightness pixel array and two corresponding chrominance pixel arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0151] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit ( Figure 2), is used to divide the image 17 into multiple (usually non-overlapping) image blocks 203, where the image block 203 is a collection of pixels. The image block 203 is sometimes also called a current block 203, an original block 203, or a partitioned block 203. For example, an image 17 to be encoded is first divided into non-overlapping image blocks 203, and each image block 203 is processed in turn in a given order (e.g., a row scan order). These image blocks 203 may also be called root blocks, macroblocks (H.264 / AVC) or coding tree blocks (CTB), or coding tree units (CTU) in the H.265 / HEVC and VVC standards. The partitioning unit can be used to use the same block size for all images in a video sequence and a corresponding grid with a limited block size, or to change the block size between images or image subsets or image groups, and to partition any image into corresponding blocks.
[0152] In other embodiments, the video encoder may be configured to directly receive an image block 203 of the image 17, for example, one, several or all blocks constituting the image 17. The image block 203 may also be referred to as a current image block or an image block to be encoded.
[0153] Like image 17, image block 203 is also or can be considered as a two-dimensional array or matrix composed of pixels with intensity values (pixel values), but image block 203 is smaller than image 17. In other words, image block 203 may include one pixel array (e.g., a brightness array in the case of monochrome image 17 or a brightness array or chrominance array in the case of color image) or three pixel arrays (e.g., one brightness array and two chrominance arrays in the case of color image 17) or any other number and / or type of arrays according to the adopted color format. The number of pixels in the horizontal and vertical directions (or axes) of image block 203 defines the size of image block 203. Accordingly, image block 203 may be an array of M×N (M columns×N rows) pixels, or an array of M×N transform coefficients, etc. For example, if the size of image block 203 is N×N, it means that image block 203 is a two-dimensional pixel array, and its horizontal and vertical sizes are both N.
[0154] In one embodiment, Figure 2 The video encoder 20 shown is used to encode the image 17 block by block, for example, to perform encoding and prediction on any image block 203.
[0155] In one embodiment, Figure 2The video encoder 20 shown may also be used to partition and / or encode an image using slices (also referred to as video slices), where an image may be partitioned or encoded using one or more slices (usually non-overlapping). Any slice may include one or more blocks (e.g., coding tree units CTUs) or one or more block groups (e.g., coding tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).
[0156] In one embodiment, Figure 2 The video encoder 20 shown can also be used to partition and / or encode an image using slices / coding block groups (also called video coding block groups) and / or coding blocks (also called video coding blocks), wherein an image can be partitioned or encoded using one or more slices / coding block groups (usually non-overlapping), any slice / coding block group may include one or more blocks (e.g., CTU) or one or more coding blocks, etc., wherein any coding block may be in a rectangular shape, etc., and may include one or more complete or partial blocks (e.g., CTU).
[0157] (2) Residual calculation
[0158] The residual calculation unit 204 is used to calculate the residual block 205 (the prediction block 265 is described in detail later) according to the image block 203 and the prediction block 265 in the following manner: for example, the pixel value of the prediction block 265 is subtracted from the pixel value of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain. Among them, the encoder 20 performs intra-frame prediction or inter-frame prediction on the image block to obtain the predicted value of the pixel therein, and the set of predicted values of the pixels in the image block is called the prediction of the image block, also known as the prediction block. The difference between the original value of the pixel in the image block and the predicted value of the pixel in the image block is further calculated, and the set of the difference between the original value of the pixel in the image block and the predicted value of the pixel in the image block is called the residual of the image block, also known as the residual block.
[0159] (3) Transformation
[0160] The transform processing unit 206 is used to perform discrete cosine transform (DCT) or discrete sine transform (DST) on the pixel values of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 may also be called transform residual coefficients, representing the residual block 205 in the transform domain.
[0161] The transform processing unit 206 may be used to apply an integerized approximation of the DCT / DST, such as the transform specified for H.265 / HEVC. This integerized approximation is typically scaled by a factor compared to an orthogonal DCT transform. In order to maintain the norm of the residual block processed by the forward transform and the inverse transform, other scaling factors are used as part of the transform process. The scaling factor is typically selected based on certain constraints, such as whether the scaling factor is a power of 2 for the shift operation, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. For example, a specific scaling factor is specified for the inverse transform on the encoder 20 side by the inverse transform processing unit 212 (and for the corresponding inverse transform on the decoder 30 side by, for example, the inverse transform processing unit 312), and correspondingly, a corresponding scaling factor may be specified for the forward transform on the encoder 20 side by the transform processing unit 206.
[0162] In one embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) may be used to output transform parameters such as one or more transform types, for example, directly output or output after encoding or compression by the entropy coding unit 270, for example, so that the video decoder 30 may receive and use the transform parameters for decoding.
[0163] (4) Quantification
[0164] The quantization unit 208 is used to quantize the transform coefficients 207 by, for example, scalar quantization or vector quantization to obtain quantized transform coefficients 209, also known as quantized transform coefficients 209, referred to as quantized coefficients 209. The quantized transform coefficients 209 may also be referred to as quantized residual coefficients 209.
[0165] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, different degrees of scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to a finer quantization, while a larger quantization step size corresponds to a coarser quantization. A suitable quantization step size may be indicated by a quantization parameter (QP). For example, a quantization parameter may be an index of a predefined set of suitable quantization step sizes. For example, a smaller quantization parameter may correspond to fine quantization (a smaller quantization step size), a larger quantization parameter may correspond to coarse quantization (a larger quantization step size), and vice versa. Quantization may include dividing by a quantization step size, while a corresponding or inverse dequantization performed by an inverse quantization unit 210 or the like may include multiplying by a quantization step size. Embodiments according to some standards, such as HEVC, may be used to determine a quantization step size using a quantization parameter. In general, the quantization step size may be calculated using a fixed-point approximation of an equation containing division according to the quantization parameter. Other scaling factors may be introduced for quantization and dequantization to recover the norm of the residual block that may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, a custom quantization table may be used and indicated from the encoder to the decoder in the bitstream, etc. Quantization is a lossy operation, where the larger the quantization step size, the greater the loss.
[0166] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) may be configured to output a quantization parameter (QP), for example, directly output or output after being encoded or compressed by the entropy coding unit 270, for example, so that the video decoder 30 may receive and use the quantization parameter for decoding.
[0167] (5) Dequantization
[0168] The inverse quantization unit 210 is used to perform the inverse quantization of the quantization unit 208 on the quantized coefficients 209 to obtain dequantized coefficients 211, for example, performing the inverse quantization scheme of the quantization scheme performed by the quantization unit 208 according to or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211 may also be called dequantized residual coefficients 211 or inverse quantized coefficients 211, corresponding to the transform coefficients 207, but due to the loss caused by quantization, the inverse quantized coefficients 211 are usually not completely the same as the transform coefficients 207.
[0169] (6) Inverse transformation
[0170] The inverse transform processing unit 212 is used to perform an inverse transform of the transform performed by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 in the pixel domain. The reconstructed residual block 213 may also be referred to as a transform block 213.
[0171] (7) Reconstruction
[0172] The reconstruction unit 214 (e.g., the summer 214) is used to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the pixel domain, for example, by adding the pixel point values of the reconstructed residual block 213 and the pixel point values of the prediction block 265.
[0173] (8) Filtering
[0174] The loop filter unit 220 (or simply "loop filter" 220) is used to filter the reconstruction block 215 to obtain the filter block 221, or is generally used to filter the reconstructed pixel points to obtain the filtered pixel point value. For example, the loop filter unit is used to smoothly perform pixel conversion or improve video quality, and coding distortions such as block effects and ringing effects can be removed through loop filtering. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Figure 2 2 is shown as a loop filter, but in other configurations, the loop filter unit 220 can be implemented as a loop filter. The filtering block 221 can also be called a filtering reconstruction block 221.
[0175] In one embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) may be used to output loop filter parameters (e.g., SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, directly or after being entropy encoded by the entropy encoding unit 270, such that the decoder 30 may receive and use the same or different loop filter parameters for decoding.
[0176] (9) Decoded Image Buffer
[0177] The decoded picture buffer (DPB) 230 may be a reference picture memory for storing reference pictures (or reference picture data) for use by the video encoder 20 when encoding video data. The DPB 230 may be formed by any of a variety of memory devices, such as a dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The decoded picture buffer 230 may be used to store one or more filter blocks 221. The decoded picture buffer 230 may also be used to store other previous filter blocks of the same current picture or a different picture such as a previous reconstructed picture, such as a previously reconstructed and filtered block 221, and may provide a complete previously reconstructed, i.e., decoded picture (and corresponding reference blocks and pixels) and / or a partially reconstructed current picture (and corresponding reference blocks and pixels), such as for inter-frame prediction. The decoded image buffer 230 may also be used to store one or more unfiltered reconstructed blocks 215, or generally to store unfiltered reconstructed pixels, for example, reconstructed blocks 215 that have not been filtered by the loop filter unit 220, or reconstructed blocks or reconstructed pixels that have not undergone any other processing.
[0178] (10) Mode selection (segmentation and prediction)
[0179] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254, which are used to receive or obtain original image data such as an original block 203 (current block 203 of the current image 17) and reconstructed image data from the decoded image buffer 230 or other buffers (e.g., a column buffer, not shown in the figure), such as filtered and / or unfiltered reconstructed pixels or reconstructed blocks of the same (current) image and / or one or more previously decoded images. The reconstructed image data is used as reference image data required for prediction such as inter-frame prediction or intra-frame prediction to obtain a prediction block 265 or a prediction value 265.
[0180] The mode selection unit 260 can be used to determine or select a partition for the current block prediction mode (including non-partition) and prediction mode (such as intra-frame or inter-frame prediction mode), generate a corresponding prediction block 265, so as to calculate the residual block 205 and reconstruct the reconstruction block 215.
[0181] In one embodiment, the mode selection unit 260 may be used to select a segmentation and prediction mode (e.g., from prediction modes supported or available by the mode selection unit 260) that provides the best match or minimum residual (minimum residual means better compression in transmission or storage), or provides minimum signaling overhead (minimum signaling overhead means better compression in transmission or storage), or considers or balances both of the above. The mode selection unit 260 may be used to determine the segmentation and prediction mode based on rate distortion optimization (RDO), i.e., select the prediction mode that provides minimum rate distortion optimization. The terms "best", "lowest", "optimal", etc. herein do not necessarily mean "best", "lowest", "optimal" in general, but may also refer to situations where termination or selection criteria are met, for example, values exceeding or falling below a threshold or other restrictions may result in a "suboptimal selection" but reduce complexity and processing time.
[0182] In other words, the partitioning unit 262 may be used to partition an image in a video sequence into a sequence of coding tree units (CTUs) 203, which may be further partitioned into smaller block portions or sub-blocks (again forming blocks), for example by iteratively using quad-tree partitioning (QT), binary-tree partitioning (BT) or triple-tree partitioning (TT) or any combination thereof, and for performing prediction on each of the block portions or sub-blocks, for example, wherein the mode selection comprises selecting a tree structure of the partitioned blocks 203 and selecting a prediction mode to be applied to each of the block portions or sub-blocks.
[0183] The segmentation (eg, performed by segmentation unit 262) and prediction processes (eg, performed by inter-prediction unit 244 and intra-prediction unit 254) performed by video encoder 20 are described in detail below.
[0184] (11)Split
[0185] The partitioning unit 262 may partition (or divide) a coding tree unit 203 into smaller parts, such as small blocks in a square or rectangular shape. For an image with three pixel arrays, a CTU consists of N×N luminance pixel blocks and two corresponding chrominance pixel blocks. The maximum allowed size of a luminance block in a CTU is specified as 128×128 in the Versatile Video Coding (VVC) standard under development, but may be specified as a value different from 128×128 in the future, such as 256×256. The CTUs of an image may be concentrated / grouped into slices / coding block groups, coding blocks, or bricks. A coding block covers a rectangular area of an image, and a coding block may be divided into one or more bricks. A brick consists of multiple CTU rows within a coding block. A coding block that is not partitioned into multiple bricks may be called a brick. However, a brick is a true subset of a coding block and is therefore not called a coding block. VVC supports two coding block group modes, namely a raster scan slice / coding block group mode and a rectangular slice mode. In raster scan coded block group mode, a slice / coded block group contains a sequence of coded blocks in a raster scan of the coded blocks of an image. In rectangular slice mode, a slice contains multiple bricks of an image, which together form a rectangular area of the image. The bricks within a rectangular slice are arranged in the order of the brick raster scan of the slice. These smaller blocks (also called sub-blocks) can be further split into smaller parts. This is also called tree partitioning or hierarchical tree partitioning, where a root block at root tree level 0 (hierarchy level 0, depth 0), etc., can be recursively split into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchy level 1, depth 1). These blocks can be split into two or more blocks at the next lower level, such as tree level 2 (hierarchy level 2, depth 2), etc., until the partitioning ends (because the end criteria are met, such as reaching the maximum tree depth or minimum block size). Blocks that are not further split are also called leaf blocks or leaf nodes of the tree. A tree divided into two parts is called a binary tree (binary-tree, BT), a tree divided into three parts is called a ternary tree (ternary-tree, TT), and a tree divided into four parts is called a quadtree (quad-tree, QT).
[0186] For example, a coding tree unit (CTU) may be or include a CTB of a luminance pixel, two corresponding CTBs of a chrominance pixel of an image with three pixel arrays, or a CTB of a pixel of a monochrome image, or a CTB of a pixel of an image encoded using three independent color planes and a syntax structure (for encoding pixels). Accordingly, a coding tree block (CTB) may be an N×N pixel block, where N may be set to a certain value so that the component is divided into CTBs, which is segmentation. A coding unit (coding unit, CU) may be or include a coding block of luminance pixels, two corresponding coding blocks of chrominance pixels of an image with three pixel arrays, or a coding block of a pixel of a monochrome image, or a coding block of a pixel of an image encoded using three independent color planes and a syntax structure (for encoding pixels). Accordingly, a coding block (CB) may be an M×N pixel block, where M and N may be set to a certain value so that the CTB is divided into coding blocks, which is segmentation.
[0187] For example, in an embodiment, according to HEVC, a coding tree unit (CTU) may be divided into multiple CUs using a quadtree structure represented as a coding tree. A decision is made at the leaf-CU level whether to use inter-frame (temporal) prediction or intra-frame (spatial) prediction to encode an image area. Any leaf-CU may be further divided into one, two, or four PUs according to the PU partition type. The same prediction process is used within a PU, and relevant information is transmitted to the decoder in units of PUs. After the residual block is obtained by applying the prediction process according to the PU partition type, the leaf-CU may be partitioned into transform units (TUs) according to other quadtree structures similar to the coding tree for the CU.
[0188] For example, in an embodiment, according to the latest video coding standard currently under development (called Versatile Video Coding (VVC), a combined quadtree of nested multi-type trees (such as binary trees and ternary trees) is used to divide the segment structure for segmenting the coding tree unit. In the coding tree structure within the coding tree unit, the CU can be square or rectangular. For example, the coding tree unit (CTU) is first segmented by a quadtree structure. The quadtree leaf nodes are further segmented by a multi-type tree structure. The multi-type tree structure has four types of partitioning: vertical binary tree partitioning (SPLIT_BT_VER), horizontal binary tree partitioning (SPLIT_BT_HOR), vertical ternary tree partitioning (SPLIT_TT_VER), and horizontal ternary tree partitioning (SPLIT_TT_HOR). The multi-type leaf nodes are called coding units (CUs), unless the CU is too large for the maximum transform length, such segmentation is used for prediction and transform processing without any other partitioning. In most cases, this means that the block sizes of CU, PU, and TU in the coding block structure of the quadtree nested multi-type tree are the same. This exception occurs when the maximum supported transform length is less than the width or height of the color component of the CU. VVC has developed a unique signaling mechanism for split partition information in a coding structure with a quadtree nested multi-type tree. In the signaling mechanism, the coding tree unit (CTU) is first split by the quadtree structure as the root of the quadtree. Then any quadtree leaf node (when large enough) is further split into a multi-type tree structure. In the multi-type tree structure, the first flag (mtt_split_cu_flag) indicates whether the node is further split. When the node is further split, the second flag (mtt_split_cu_vertical_flag) is used to indicate the split direction, and the third flag (mtt_split_cu_binary_flag) is used to indicate whether the split is a binary tree split or a ternary tree split. According to the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the decoder can derive the multi-type tree split mode (MttSplitMode) of the CU based on predefined rules or tables. It should be noted that for certain designs, such as the 64×64 luma block and 32×32 chroma pipeline design in the VVC hardware decoder, TT splitting is not allowed when the width or height of the luma coding block is greater than 64. TT splitting is also not allowed when the width or height of the chroma coding block is greater than 32. The pipeline design divides the image into multiple virtual pipeline data units (VPDUs), and any VPDU is defined as a non-overlapping unit in the image. In the hardware decoder, consecutive VPDUs are processed simultaneously in multiple pipeline stages.In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is desirable to keep the VPDU small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the ternary tree (TT) and binary tree (BT) partitioning may increase the VPDU size.
[0189] In addition, it should be noted that when a part of the tree node block exceeds the bottom or the right boundary of the image, the tree node block is forcibly divided until all the pixels of any coding CU are located within the image boundary.
[0190] For example, the intra sub-partitions (ISP) tool may divide the luma intra prediction block vertically or horizontally into two or four sub-partitions according to the block size.
[0191] In one example, mode select unit 260 of video encoder 20 may be used to perform any combination of the segmentation techniques described above.
[0192] As described above, the video encoder 20 is operable to determine or select the best or optimal prediction mode from a (predetermined) set of prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0193] (12) Intra-frame prediction
[0194] The intra prediction mode set may include 35 different intra prediction modes, for example, non-directional modes like DC (or mean) mode and plane mode, or directional modes as defined in HEVC, or may include 67 different intra prediction modes, for example, non-directional modes like DC (or mean) mode and plane mode, or directional modes as defined in VVC. For example, several traditional angle intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks defined in VVC. For another example, in order to avoid the division operation of DC prediction, only the longer side is used to calculate the average value of non-square blocks. In addition, the intra prediction result of the plane mode can also be modified using the position-dependent intra prediction combination (PDPC) method.
[0195] The intra prediction unit 254 is configured to generate an intra prediction block 265 using reconstructed pixels of adjacent blocks of the same current image according to an intra prediction mode in the intra prediction mode set.
[0196] The intra-frame prediction unit 254 (or generally the mode selection unit 260) is also used to output intra-frame prediction parameters (or generally information indicating the selected intra-frame prediction mode of the block) in the form of syntax elements 266 to the entropy coding unit 270 for inclusion in the encoded image data 21, so that the video decoder 30 can perform operations, such as receiving and using the prediction parameters for decoding.
[0197] (13) Inter-frame prediction
[0198] In a possible implementation, the set of inter-prediction modes depends on the available reference picture (i.e., at least part of the previously decoded picture stored in the DBP 230 as mentioned above) and other inter-prediction parameters, for example, on whether to use the entire reference picture or only a part of the reference picture, such as a search window area around the area of the current block, to search for the best matching reference block, and / or on whether to perform pixel interpolation such as half-pixel, quarter-pixel and / or 1 / 16 interpolation, for example.
[0199] In addition to the above prediction modes, skip mode and / or direct mode may also be employed.
[0200] For example, extended merge prediction, the merge candidate list of this mode consists of the following five candidate types in order: spatial MVP from spatially adjacent CUs, temporal MVP from collocated CUs, history-based MVP from FIFO tables, pairwise average MVP, and zero MV. Decoder side motion vector refinement (DMVR) based on bilateral matching can be used to increase the accuracy of the MV of the merge mode. Merge mode with MVD (MMVD) is derived from the merge mode with motion vector difference. The MMVD flag is sent immediately after the skip flag and merge flag are sent to specify whether the CU uses the MMVD mode. The CU-level adaptive motion vector resolution (AMVR) scheme can be used. AMVR supports encoding the MVD of the CU with different precisions. The MVD of the current CU is adaptively selected according to the prediction mode of the current CU. When the CU is encoded in merge mode, the combined inter / intra prediction (CIIP) mode can be applied to the current CU. The CIIP prediction is obtained by weighted averaging the inter-frame and intra-frame prediction signals. For affine motion compensated prediction, the affine motion field of the block is described by the motion information of 2 control points (4 parameters) or 3 control points (6 parameters) motion vectors. Subblock-based temporal motion vector prediction (SbTMVP) is similar to the temporal motion vector prediction (TMVP) in HEVC, but the motion vector of the sub-CU within the current CU is predicted. Bidirectional optical flow (BDOF), formerly known as BIO, is a simplified version that reduces calculations, especially in terms of the number of multiplications and the size of the multipliers. In triangle partitioning mode, the CU is evenly divided into two triangular parts in two ways: diagonal partitioning and anti-diagonal partitioning. In addition, the bidirectional prediction mode is extended on the basis of simple averaging to support weighted averaging of two prediction signals.
[0201] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in Figure 217). The motion estimation unit may be configured to receive or acquire the image block 203 (the current image block 203 of the current image 17) and the decoded image 231, or at least one or more previously reconstructed blocks, for example, reconstructed blocks of one or more other / different previously decoded images 231, for motion estimation. For example, a video sequence may include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 may be part of or form a sequence of images forming the video sequence.
[0202] For example, the encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different images in a plurality of other images, and provide the reference image (or reference image index) and / or the offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. The offset is also referred to as a motion vector (MV).
[0203] The motion compensation unit is used to obtain, for example, receive, inter-frame prediction parameters, and perform inter-frame prediction based on or using the inter-frame prediction parameters to obtain an inter-frame prediction block 246. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block based on a motion / block vector determined by motion estimation, and may also include performing interpolation with sub-pixel accuracy. Interpolation filtering can generate pixel points of other pixels from pixel points of known pixels, thereby potentially increasing the number of candidate prediction blocks that can be used to encode the image block. Once a motion vector corresponding to a PU of a current image block is received, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists.
[0204] The motion compensation unit may also generate syntax elements associated with blocks and video slices for use by the video decoder 30 when decoding image blocks of the video slice. In addition, or as an alternative to slices and corresponding syntax elements, coding block groups and / or coding blocks and corresponding syntax elements may be generated or used.
[0205] (14) Entropy Coding
[0206] The entropy coding unit 270 is used to apply an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC scheme (CALVC), an arithmetic coding scheme, a binarization algorithm, a context adaptive binary arithmetic coding (CABAC), a syntax-based context-adaptive binary arithmetic coding (SBAC), a probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements to obtain coded image data 21 that can be output in the form of a coded bit stream 21 through an output terminal 272, so that the video decoder 30, etc. can receive and use the parameters for decoding. The coded bit stream 21 can be transmitted to the video decoder 30, or it can be stored in a memory for later transmission or retrieval by the video decoder 30.
[0207] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform based encoder 20 may directly quantize the residual signal without a transform processing unit 206 for certain blocks or frames. In another implementation, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.
[0208] 2. Decoder and Decoding Method
[0209] Figure 3 An exemplary video decoder 30 for implementing the technology of the present application is shown. The video decoder 30 is used to receive the encoded image data 21 (e.g., the encoded bit stream 21) encoded by the encoder 20, for example, to obtain a decoded image 331. The encoded image data 21 or the bit stream includes information for decoding the encoded image data 21, such as data representing image blocks of an encoded video slice (and / or a coding block group or coding block) and related syntax elements.
[0210] exist Figure 3In the example of FIG. 3 , decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter-frame prediction unit 344, and an intra-frame prediction unit 354. Inter-frame prediction unit 344 may be or include a motion compensation unit. In some examples, video decoder 30 may perform substantially the same as reference Figure 2 The decoding process of the video encoder 100 is the reverse of the encoding process described above.
[0211] See also Figure 3 The inter-frame prediction module / intra-frame prediction module / loop filtering module includes (is) a trained target model (also called a neural network) that is used to process an input image or an image region or an image block to generate a prediction value of the input image block. For example, the neural network for inter-frame prediction / intra-frame prediction / loop filtering is used to receive an input image or an image region or an image block.
[0212] As described for encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer 230, inter prediction unit 344, and intra prediction unit 354 also constitute a “built-in decoder” of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 122, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Therefore, the explanation of the corresponding units and functions of video encoder 20 applies accordingly to the corresponding units and functions of video decoder 30.
[0213] (1) Entropy decoding
[0214] The entropy decoding unit 304 is used to parse the bit stream 21 (or generally the encoded image data 21) and perform entropy decoding on the encoded image data 21 to obtain the quantization coefficients 309 and / or the decoded encoding parameters ( Figure 3The entropy decoding unit 304 may be used to provide the inter-frame prediction parameters, intra-frame prediction parameters and / or other syntax elements to the mode application unit 360, and to provide other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice and / or video block level. In addition, or as an alternative to slices and corresponding syntax elements, coding block groups and / or coding blocks and corresponding syntax elements may be received or used.
[0215] (2) Dequantization
[0216] The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally information related to inverse quantization) and a quantization coefficient from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse quantize the decoded quantization coefficient 309 based on the quantization parameter to obtain an inverse quantization coefficient 311, which may also be referred to as a transform coefficient 311 or a dequantized coefficient 311. The inverse quantization process may include using the quantization parameter calculated by the video encoder 20 for any video block in the video slice to determine a degree of quantization, and also determine a degree of inverse quantization that needs to be performed.
[0217] (3) Inverse transformation
[0218] The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain a reconstructed residual block 313 in the pixel domain. The reconstructed residual block 313 may also be referred to as a transform block 313. The transform may be an inverse transform, such as an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may also be configured to receive transform parameters or corresponding information from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304) to determine the transform applied to the dequantized coefficients 311.
[0219] (4) Reconstruction
[0220] The reconstruction unit 314 (eg, the summer 314 ) is used to add the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the pixel domain, for example, by adding the pixel point values of the reconstructed residual block 313 and the pixel point values of the prediction block 365 .
[0221] (5) Filtering
[0222] The loop filter unit 320 (in or after the encoding loop) is used to filter the reconstruction block 315 to obtain a filter block 321, so as to smoothly perform pixel conversion or improve video quality, etc. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 320 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chromascaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Figure 3 Although shown as a loop filter in FIG. 3 , in other configurations, the loop filter unit 320 may be implemented as a loop filter.
[0223] (6) Decoded Image Buffer
[0224] The decoded video blocks 321 in one picture are then stored in a decoded picture buffer 330 which stores the decoded pictures 331 as reference pictures for subsequent motion compensation of other pictures and / or for respective output display.
[0225] The decoder 30 is used to output the decoded image 331 through the output terminal 312, etc., for display to the user or for the user to view.
[0226] (7) Prediction
[0227] The inter-frame prediction unit 344 may be functionally the same as the inter-frame prediction unit 244 (particularly the motion compensation unit), and the intra-frame prediction unit 354 may be functionally the same as the intra-frame prediction unit 254, and may determine the division or segmentation and perform prediction based on the segmentation and / or prediction parameters or corresponding information received from the coded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304). The mode application unit 360 may be used to perform prediction (intra-frame or inter-frame prediction) of any block based on the reconstructed image, block or corresponding pixel point (filtered or unfiltered), and obtain a prediction block 365.
[0228] When the video slice is encoded as an intra coded (I) slice, the intra prediction unit 354 in the mode application unit 360 is used to generate a prediction block 365 for the image block of the current video slice based on the indicated intra prediction mode and data from a previously decoded block of the current image. When the video image is encoded as an inter coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) in the mode application unit 360 is used to generate a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter prediction, these prediction blocks can be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 can construct reference frame list 0 and list 1 using a default construction technique based on the reference pictures stored in the decoded picture buffer 330. The same or similar processes may be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks) in addition to or as an alternative to slices (e.g., video slices), e.g., a video may be encoded using I, P or B coding block groups and / or coding blocks.
[0229] The mode application unit 360 is used to determine prediction information for a video block of a current video slice by parsing motion vectors and other syntax elements, and use the prediction information to generate a prediction block for the current video block being decoded. For example, the mode application unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for encoding the video block of the video slice, the inter-frame prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information of one or more reference picture lists for the slice, the motion vector for any inter-frame coded video block of the slice, the inter-frame prediction state for any inter-frame coded video block of the slice, and other information to decode the video block in the current video slice. In addition to slices (e.g., video slices) or as an alternative to slices, the same or similar process can be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks), for example, the video can be encoded using I, P, or B coding block groups and / or coding blocks.
[0230] In one embodiment, Figure 3 The video encoder 30 shown may also be used to partition and / or decode an image using slices (also referred to as video slices), where an image may be partitioned or decoded using one or more slices (usually non-overlapping). Any slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., coding blocks in the H.265 / HEVC / VVC standard and bricks in the VVC standard).
[0231] In one embodiment, Figure 3 The video decoder 30 shown can also be used to segment and / or decode an image using slices / coding block groups (also called video coding block groups) and / or coding blocks (also called video coding blocks), wherein an image can be segmented or decoded using one or more slices / coding block groups (usually non-overlapping), and any slice / coding block group may include one or more blocks (e.g., CTUs) or one or more coding blocks, etc., wherein any coding block may be in a rectangular shape, etc., and may include one or more complete or partial blocks (e.g., CTUs).
[0232] Other variations of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-transform based decoder 30 may directly dequantize the residual signal without the inverse transform processing unit 312 for certain blocks or frames. In another implementation, the video decoder 30 may have the dequantization unit 310 and the inverse transform processing unit 312 combined into a single unit.
[0233] It should be understood that the processing result of the current step can be further processed in the encoder 20 and the decoder 30 and then output to the next step. For example, after interpolation filtering, motion vector derivation or loop filtering, the processing result of interpolation filtering, motion vector derivation or loop filtering can be further operated, such as clipping or shifting operation.
[0234] It should be noted that the derived motion vector of the current block (including but not limited to the control point motion vector of the affine mode, the sub-block motion vector of the affine, plane, ATMVP mode, the temporal motion vector, etc.) can be further operated. For example, the value of the motion vector is limited to a predefined range according to the representation bit of the motion vector. If the representation bit of the motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" represents the power. For example, if bitDepth is set to 16, the range is -32768 to 32767; if bitDepth is set to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of 4 4×4 sub-blocks in an 8×8 block) is limited so that the maximum difference between the integer parts of the 4 4×4 sub-block MVs does not exceed N pixels, for example, not more than 1 pixel. Two methods of limiting the motion vector according to bitDepth are provided here.
[0235] Although the above embodiments mainly describe video coding and decoding, it should be noted that the embodiments of the decoding system 10, the encoder 20 and the decoder 30 and other embodiments described herein can also be used for still image processing or coding and decoding, that is, the processing or coding and decoding of a single image in video coding and decoding that is independent of any previous or consecutive images. In general, if the image processing is limited to a single image 17, the inter-frame prediction unit 244 (encoder) and the inter-frame prediction unit 344 (decoder) may not be available. All other functions (also called tools or techniques) of the video encoder 20 and the video decoder 30 can also be used for still image processing, such as residual calculation 204 / 304, transform processing unit 206, quantization unit 208, inverse quantization unit 210 / 310, inverse transform processing unit 212 / 312, segmentation unit 262, intra-frame prediction unit 254 / 354, loop filter 220 / 320, entropy encoding unit 270 and entropy decoding unit 304, etc.
[0236] It should be noted that the operations of the modules in the encoder 20 and the decoder 30 have a corresponding relationship. For example, in the encoder 20 and the decoder 30, the operations of the inter-frame prediction unit 244 and the inter-frame prediction unit 344, and the intra-frame prediction unit 254 and the intra-frame prediction unit 354 are exactly the same; while the entropy coding unit 270 and the entropy decoding unit 304, the transform processing unit 206 and the inverse transform processing unit 212 / 312, the quantization unit 208 and the inverse quantization unit 210 / 310, etc., are all paired inverse operations. Therefore, after the prediction, transform, quantization, entropy coding and other operations of the encoder 20 are specified, the prediction, inverse transform, inverse quantization, entropy decoding and other operations of the decoder 30 are also determined accordingly.
[0237] Figure 4Schematic diagram of a video decoding device 400 provided in an embodiment of the present application. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 400 may be a decoder, such as Figure 1A The video decoder 30 in the embodiment may also be an encoder, for example Figure 1A The video encoder 20 in.
[0238] The video decoding device 400 includes: an input port 410 (or input port 410) and a receiving unit (Rx) 420 for receiving data; a processor, a logic unit or a central processing unit (CPU) 430 for processing data; for example, the processor 430 here can be a neural network processor 430; a transmitter unit (Tx) 440 and an output port 450 (or output port 450) for transmitting data; and a memory 460 for storing data. The video decoding device 400 may also include an optical-to-electrical (OE) component and an electrical-to-optical (EO) component coupled to the input port 410, the receiving unit 420, the transmitting unit 440 and the output port 450 for the output or output of optical or electrical signals.
[0239] The processor 430 is implemented by hardware and software. The processor 430 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the input port 410, the receiving unit 420, the sending unit 440, the output port 450, and the memory 460. The processor 430 includes a decoding module 470 (e.g., a decoding module 470 based on a neural network). The decoding module 470 implements the embodiments disclosed above. For example, the decoding module 470 performs, processes, prepares, or provides various encoding operations. Therefore, the decoding module 470 provides substantial improvements to the functions of the video decoding device 400 and affects the switching of the video decoding device 400 to different states. Alternatively, the decoding module 470 is implemented with instructions stored in the memory 460 and executed by the processor 430.
[0240] The memory 460 includes one or more disks, tape drives, and solid-state hard disks, and can be used as an overflow data storage device for storing such programs when such programs are selected for execution, and for storing instructions and data read during program execution. The memory 460 can be volatile and / or non-volatile, and can be a read-only memory (ROM), a random access memory (RAM), a ternary content-addressable memory (TCAM), and / or a static random-access memory (SRAM).
[0241] Figure 5 A simplified block diagram of an apparatus 500 is provided for an exemplary embodiment, and the apparatus 500 may be used as Figure 1A Either or both of the source device 12 and the destination device 14 in .
[0242] The processor 502 in the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices that are currently available or will be developed in the future and are capable of manipulating or processing information. Although a single processor such as the processor 502 shown in the figure may be used to implement the disclosed implementation, using more than one processor is faster and more efficient.
[0243] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 via a bus 512. The memory 504 may also include an operating system 508 and an application 510, which includes at least one program that allows the processor 502 to perform the methods described herein. For example, the application 510 may include applications 1 to N, and also include a video decoding application that performs the methods described herein.
[0244] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element that may be used to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.
[0245] Although bus 512 in device 500 is described herein as a single bus, bus 512 may include multiple buses. In addition, auxiliary storage may be directly coupled to other components of device 500 or accessed through a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, device 500 may have a variety of configurations.
[0246] In order to facilitate those skilled in the art to understand the present application, some terms in the embodiments of the present application are explained below, and relevant technical knowledge involved in the embodiments of the present application is introduced.
[0247] 1. Terminology
[0248] Auto-Encoder (AE): A specific neural network structure.
[0249] End-to-end (E2E): usually refers to the technical solution of a learning task being implemented entirely using neural networks, so all parameters of the neural network can be optimized simultaneously through gradient backpropagation.
[0250] Coding picture: usually contains three matrices of an image, which respectively store the reconstruction of the intensity values of the three color components of the image, YUV or RGB, and also contains the coding information such as block division, coding mode, quantization parameters, etc. of all coding units of the image.
[0251] Decoded picture: usually contains three matrices of an image, which store the reconstruction of the intensity values of the three color components of the image, YUV or RGB.
[0252] Coding Unit (CU): usually contains three matrices of an image block, which respectively store the reconstruction of the intensity values of the three color components of the image block, YUV or RGB, and also contains the coding information of the image block, such as block division, coding mode, quantization parameter, etc.
[0253] Reconstructed image: refers to the image containing coding distortion after the encoding operation.
[0254] 2. Related technical knowledge
[0255] Video compression coding technology is widely used in multimedia services, broadcasting, video communication and storage. In recent years, the two major standard organizations, ITU-T and ISO / IEC, have jointly formulated and released three video codec standards, H.264 / AVC, H.265 / HEVC and H.266 / VVC, in 2003, 2013 and 2020 respectively. At the same time, the AVS standard group also formulated and output a series of video image codec standards such as AVS1, AVS2, and AVS3. In addition, the AOM Alliance also released the AV1 video codec solution in 2018. The above video coding and decoding technologies all adopt a hybrid coding and decoding solution based on block division and transform quantization, and continuously iterate the technology in specific modules such as block division, prediction, transformation, entropy coding, and loop filtering, so as to continuously improve the compression efficiency of video images.
[0256] (1) Related technical knowledge
[0257] In recent years, the academic community has begun to study image encoding and decoding solutions based on deep neural networks (DNN). Figure 6 A typical image encoding and decoding scheme based on deep neural network is shown. The scheme adopts the autoencoder network structure. The input x is the original image to be encoded, which can be expressed as a w x ×h x ×c x Array of x 、h x , c x Respectively represent the width, height, and number of color components of the input image. The analyzer and synthesizer are usually deep convolutional neural networks (CNNs). The analyzer performs dimensionality reduction on the input image x to obtain its latent representation y, which can usually be represented as a w y ×h y ×c y Array, w y 、h y , c yare respectively the width, height, and number of latent channels of the latent expression. Each element in the latent expression y output by the analyzer can be a floating point number or an integer, and can be quantized or normally rounded to obtain a more compact integer expression y'. Then entropy coding is performed on y' to obtain a compressed bitstream. Entropy decoding is the inverse operation of entropy coding, and its purpose is to parse y' from the compressed bitstream. Entropy coding and entropy decoding use the same probability model, and the state of the probability model can be updated synchronously in the encoder and decoder to ensure encoding and decoding matching. Optionally, the encoder can transmit the probability model parameters to the decoder to ensure that entropy coding and entropy decoding use the same probability model. The synthesizer obtains the encoded image based on y' and reconstructs x'. Among them, the input image can be divided into blocks, and each image block is sent to Figure 3 The encoder shown performs encoding operation to output a compressed code stream, and decodes the compressed code stream to obtain reconstruction of the image block.
[0258] Figure 2 The quantization unit 208 is used to quantize the transform coefficient 207 output by the transform processing unit 206 by applying scalar quantization or vector quantization to obtain a quantization level 209 of the transform coefficient 207, wherein the quantization level 209 is the output after quantization of the transform coefficient, also called the quantized transform coefficient 209, or the quantized coefficient 209.
[0259] For example, for scalar quantization, different quantization step sizes (QS) may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, and a larger quantization step size corresponds to coarser quantization. The quantization step size used may be indicated by a quantization parameter (QP). For example, a quantization parameter may be an index to a predefined set of suitable quantization step sizes. For example, a smaller quantization parameter may correspond to fine quantization (smaller quantization step size), and a larger quantization parameter may correspond to coarse quantization (larger quantization step size), and vice versa. Figure 2 The inverse quantization unit 210 in Figure 3 The inverse quantization unit 310 in performs the same inverse quantization operation, its input is the quantization coefficient, and its output is the dequantization coefficient.
[0260] An embodiment of the encoder 20 (or quantization unit 208) can be used to output a quantization scheme and a quantization step size, for example, directly output or output after entropy encoding by the entropy coding unit 270 or any other entropy coding unit, so that the decoder 30 can receive and perform corresponding inverse quantization operations.
[0261] The quantization step size, or equivalent quantization parameter, used in the quantization unit 208 and the inverse quantization unit 210 / 310 must be the same. The quantization parameter is specified by the encoder 20 and is directly output or output after entropy encoding by the entropy encoding unit 270 or any other entropy encoding unit. The decoder 30 can receive the quantization parameter information output by the encoder 20, and obtain the quantization parameter specified by the encoder 20 after entropy decoding by the entropy decoding unit 304 or any other entropy decoding unit. The following takes the HEVC standard scheme as an example to illustrate how the encoder passes the quantization parameter to the decoder, as shown in Table 1.
[0262] Table 1
[0263]
[0264] As can be seen from Table 1, the initial QP value is passed in the Picture Parameter Set (PPS), which is init_qp_minus26+26. In addition, the control flag cu_qp_delta_enabled_flag is used to determine whether different QPs can be specified for different CUs. If the control flag indicates false, all CUs in the entire image use the same QP, so it is impossible to specify a different QP for each CU in the image. If the control flag indicates true, a different QP can be specified for each CU in the image, and the QP information can be written into the video bitstream when encoding a CU.
[0265] Note that in the HEVC standard, CUs can have different sizes, from 64×64 to 8×8. In extreme cases, a coded image is divided into 8×8 CUs. At this time, a QP needs to be transmitted for each 8×8 CU, which may bring huge QP coding overhead and a significant increase in the coded video rate. To avoid this extreme situation, the HEVC standard specifies the quantization group (QG) size through the syntax diff_cu_qp_delta_depth in the PPS. When the size of the coding tree unit is 64×64, the mapping relationship between the two is shown in Table 2.
[0266] Table 2
[0267] diff_cu_qp_delta_depth value 0 1 2 3 QG size 64×64 32×32 16×16 8×8
[0268] QG is the basic transmission unit for transmitting QP. In other words, each QG can only transmit at most one QP. If the CU size is smaller than the QG size, that is, a QG contains multiple CUs, the QP is only transmitted in the first CU containing a non-zero quantization level, and the QP will be used for the dequantization operation of all CUs in the QG. If the CU size is greater than or equal to the QG size, that is, a CU contains one or more QGs, whether to transmit the QP information of the CU is determined based on whether the CU contains a non-zero quantization level.
[0269] The initial QP value transmitted in the PPS is applied to all coded images within the scope of the PPS. When processing each coded image, stripe, sub-image, or slice, the initial QP value can be further adjusted to obtain the QP reference value of the processed coded image, stripe, sub-image, or slice. For example, the HEVC standard transmits the syntax slice_qp_delta in the slice header (SH), which means that a differential value is superimposed on the initial QP value transmitted in the PPS to obtain the QP reference value of the slice, as shown in Table 3.
[0270] Table 3
[0271] slice_segment_header(){ Descriptor slice_qp_delta se(v)
[0272] When processing each CU, HEVC determines whether each transform unit (TU) is the first TU in the QG where the CU is located that contains a non-zero quantization level. If so, the QP differential information of the CU is transmitted. Specifically, the QP differential information of the CU includes the QP differential absolute value cu_qp_delta_abs and the CU differential sign cu_qp_delta_sign_flag, as shown in Table 4; wherein, the QP differential value of the CU is cu_qp_delta_abs×(1-2×cu_qp_delta_sign_flag). Because a CU only transmits at most one QP differential information, when a CU contains multiple TUs, the QP differential information is only transmitted when processing the first TU containing a non-zero quantization level.
[0273] Table 4
[0274]
[0275] Even under the constraint of QG, if the QP value is transmitted for at most one CU, the overhead caused by QP value encoding will still significantly reduce the efficiency of video compression. Therefore, QP values are generally predictively encoded. Taking the HEVC standard as an example, the QP prediction value of a CU is derived based on the QP values of the left adjacent QG, the upper adjacent QG, and the previous encoded QG, that is, the QP value of the processed QG in the neighborhood of the current QG is used to generate a prediction of the QP value of the current QG. After the encoder determines the QP value of a CU based on the content complexity and the code control strategy, it only needs to transmit the difference between the QP value of the CU and the QP prediction value of the CU; after the decoder obtains the QP differential value of a CU through decoding, it obtains the QP prediction value through the same prediction operation, and the QP value of the CU can be obtained by superimposing it with the QP differential value.
[0276] Using the specified quantization step size Qstep to quantize the specified signal, the intensity of quantization distortion can be obtained by theoretical analysis. For example, assuming a uniformly distributed signal source, the mean square error of the quantization distortion of uniform scalar quantization is Qstep 2 / 12.
[0277] However, the QP mechanism in the existing HEVC standard or any other hybrid coding scheme, although it also uses uniform scalar quantization based on Qstep, cannot accurately indicate the actual distortion of each image block for the following reasons:
[0278] First, if the residual value of any sampling point in the residual block of an image block is less than the quantization step size Qstep, the encoder will quantize the residual of the image block to all 0s. This situation is more common in P and B frame encoding, and it occurs when the reference image quality is high and the Qstep setting of the current encoded image block is large. At this time, although the decoder can obtain the effective QP information of the current block, the QP cannot reflect the actual distortion of the image block.
[0279] Second, if the residual quantization of an image block is all 0, that is, no residual will be transmitted for the image block, the decoder will skip the inverse quantization operation. To avoid transmitting useless QP information, the encoder will not pass the QP information of the image block to the decoder. At this time, the decoder cannot obtain the valid QP information of the current block at all.
[0280] This is determined by the design purpose of the QP mechanism in contemporary video codecs. The purpose of the current QP mechanism is to perform correct inverse quantization at the decoding end, rather than to obtain accurate distortion intensity information at the decoding end.
[0281] (2) Related technical knowledge 2
[0282] The academic community has also proposed a learning-based deep image coding and decoding scheme. In this learning-based deep image coding and decoding scheme, the image to be encoded is passed through the encoder subnetwork to obtain the latent expression of the input image, and then processed by the quantization module to obtain multiple quantized code blocks. On the other hand, the encoder calculates the importance map based on the input image. The encoder uses the importance map to crop the quantized code blocks, and entropy encodes the cropped quantized code blocks together with the importance map and transmits them to the decoder. In this scheme, the importance map can be used to control the number of quantized code blocks that need to be transmitted, thereby realizing the function of bit rate control. Therefore, the importance map is essentially equivalent to the quantization parameter (QP), which makes regional bit allocation according to the content of the image.
[0283] The role of the importance map in this scheme is to tell the decoding end which coding blocks are included in the bitstream, guide the decoder to decode and obtain these coding blocks, and set the coding blocks not included in the bitstream to 0, so as to obtain a complete latent expression and input it into the decoder subnetwork for subsequent decoding operations.
[0284] In the current mainstream image compression coding schemes based on deep neural networks, the image to be encoded is first input into an analyzer subnetwork (such as an encoder subnetwork) to extract the latent expression of the input image. On the one hand, since the latent expression is a dimensionality reduction representation of the input image, information loss has been introduced, that is, the distortion of the image signal has been introduced in the latent expression, and this distortion cannot be represented by the importance map of the above scheme. On the other hand, the importance map is only used to determine which coding blocks do not need to be encoded and transmitted, and cannot indicate the degree of signal distortion caused by discarding these coding blocks. Therefore, the above importance map cannot indicate the distortion intensity information of each area in the encoded image.
[0285] In summary, the existing video image coding and decoding schemes based on deep neural networks use the trained models to perform video or image coding and decoding operations. Such methods usually optimize subjective visual quality based on human eye perception, which is essentially to make flexible bit allocation between different coded images and different regions within an image. On the one hand, these coding and decoding schemes do not transmit the signal quality (or distortion intensity) of each coded image or each region within a coded image to the decoding end. On the other hand, in many cases, the true distortion intensity of the coded image cannot be obtained by human eye observation; for example, in some coded images obtained by processing the codec network trained based on the Generative Adverarial Network (GAN) method, false texture information that cannot be distinguished by the human eye will be included. Therefore, the existing coding and decoding schemes based on deep neural networks cannot obtain the distortion intensity information of the current coded image at the decoding end. Specifically, the distortion intensity information of each region in an image and the overall distortion intensity information of an image cannot be obtained at the decoding end. The distortion intensity information of the coded image can be used to assist in judging whether the content of a certain area in the image is credible, so it is very important for applications such as video surveillance. Therefore, if you want to use a video image encoding and decoding solution based on a deep neural network in an actual product or service, you need to inform the decoder of this information. It should be noted that even if you use a traditional hybrid-based encoding and decoding solution, you cannot obtain the accurate distortion intensity information of the current encoded image through the transmitted quantization parameters. The specific reasons have been explained above. Therefore, for some video products or services that use traditional hybrid encoding and decoding solutions, it is also necessary to use some method to enable the decoder to obtain the accurate distortion intensity information of the encoded image.
[0286] In view of the technical problems introduced above, the embodiments of the present application provide a coding, decoding method and related equipment. In the embodiments of the present application, no matter which video image coding and decoding scheme is used, the encoding end can compare the original image and the reconstructed image, calculate the fidelity information of the reconstructed image (including the fidelity information of each image area in the reconstructed image), and carry the fidelity information in the compressed code stream to inform the decoding end; wherein, the reconstructed image is a reconstructed image of the original image, that is, the encoded output image. When calculating the fidelity of the reconstructed image, any quality evaluation method with reference, such as MSE (Mean Squared Error), SAD (Sum of Absolute Difference), SSIM (Structural Similarity), etc., can be used to calculate the distortion intensity value of the reconstructed image relative to the original image, or to obtain an indication mark of whether there is synthesized false image content in the reconstructed image content. Fidelity can be calculated at any granularity, such as the entire image, or any M×N image block in an image, and so on. The fidelity can be calculated for any color component separately, for example, three fidelities can be calculated for the three color components of RGB or YUV of an image, or the distortion strength of the three color components can be combined to obtain one fidelity. Even for the traditional hybrid video image encoding and decoding scheme, the implementation of the present application can also use the QP information obtained by the decoding end and the fidelity of the image area referenced by the current image block in the inter-frame prediction operation to derive the fidelity of the current image block.
[0287] The technical solution provided by this application is introduced in detail below in conjunction with specific implementation methods.
[0288] Figure 7 700 is a flowchart showing a process 700 of an encoding method according to an embodiment of the present application. The process 700 may be performed by an encoding device, such as the video encoder 20. The process 700 is described as a series of steps or operations, and it should be understood that the process 700 may be performed in various orders and / or occur simultaneously, not limited to Figure 7 The execution order shown. Process 700 includes but is not limited to the following steps or operations:
[0289] 701. Encode an original image to obtain a first code stream.
[0290] It should be understood that the original image is also the image 17 , so the original image is an encoded image; the first code stream is also the encoded image data 21 .
[0291] 702. Encode a fidelity map to obtain a second code stream, wherein the fidelity map is used to represent distortion between at least a partial area of the original image and at least a partial area of a reconstructed image, and the reconstructed image is obtained after decoding the first code stream.
[0292] It should be understood that the reconstructed image is also the decoded image 231, so the reconstructed image is a decoded image. Since the first code stream is obtained by encoding the original image, the reconstructed image obtained by decoding the first code stream is a reconstructed image of the original image, and the reconstructed image has the same image size as the original image.
[0293] The fidelity map can be calculated based on the original image and the reconstructed image, or the fidelity map can be calculated based on the preset area of the original image and the preset area of the reconstructed image. When the fidelity map is calculated based on the original image and the reconstructed image, the fidelity map is used to characterize the fidelity of the entire reconstructed image; when the fidelity map is calculated based on the preset area of the original image and the preset area of the reconstructed image, the fidelity map is used to characterize the fidelity of the preset area of the reconstructed image. The preset area of the original image is also an image of a certain area in the original image, which can be an image block; the preset area of the reconstructed image is also an image of a certain area in the reconstructed image, which can also be an image block.
[0294] In order to enable the decoding end to obtain the distortion intensity information of the encoded image, the first code stream and the second code stream need to be transmitted from the encoding end to the decoding end. The first code stream and the second code stream can be combined and transmitted from the encoding end to the decoding end, or the first code stream and the second code stream can be separately transmitted from the encoding end to the decoding end. In addition, if the fidelity map is calculated based on the preset area of the original image and the preset area of the reconstructed image, the position of the preset area in the original image and / or the position of the preset area in the reconstructed image needs to be transmitted from the encoding end to the decoding end, wherein when the preset area is a rectangle, the position of the preset area can be represented by the coordinates of the preset area, and the coordinates of the preset area are usually represented as the upper left corner brightness pixel coordinates of the preset area; and the position of the preset area in the original image and / or the position of the preset area in the reconstructed image can be combined with at least one of the first code stream and the second code stream and transmitted from the encoding end to the decoding end; the position of the preset area in the original image and / or the position of the preset area in the reconstructed image can also be not combined with at least one of the first code stream and the second code stream, and transmitted from the encoding end to the decoding end separately. The encoding end and decoding end described in the embodiments of the present application may be different electronic devices or different hardware units of the same electronic device, and the present application does not make any specific limitation on this.
[0295] It should be understood that since the second bitstream is obtained by encoding the fidelity map, the first bitstream can be decoded to obtain a reconstructed image of the fidelity map. Depending on the encoding method, the reconstructed image of the decoded fidelity map can be the same as the fidelity map or different from the fidelity map. Specifically, if the encoding is lossless, the reconstructed image of the fidelity map is the same as the fidelity map; if the encoding is lossy, the reconstructed image of the fidelity map includes the encoding distortion generated by encoding the fidelity map.
[0296] In an embodiment of the present application, an original image is encoded to obtain a first code stream, and a fidelity map is encoded to obtain a second code stream, wherein the fidelity map is used to represent the distortion between at least a portion of an original image and at least a portion of a reconstructed image; the distortion includes a difference; a decoding end decodes the first code stream to obtain a reconstructed image of the original image, and a decoding end decodes the second code stream to obtain a reconstructed map of the fidelity map; and if the encoding is lossless encoding, the reconstructed map of the fidelity map is the same as the fidelity map; if the encoding is lossy encoding, the reconstructed map of the fidelity map includes the encoding distortion generated by encoding the fidelity map; therefore, the reconstructed map of the fidelity map can be used to represent the distortion between at least a portion of an original image and at least a portion of a reconstructed image; therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0297] In a possible design, the method also includes: dividing the original image into multiple first image blocks, and dividing the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the original image is the same as the division strategy for dividing the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; or dividing a preset area of the original image into multiple first image blocks, and dividing the preset area of the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the preset area of the original image is the same as the division strategy for dividing the preset area of the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; calculating the fidelity value of any second image block according to any second image block among the multiple second image blocks and the first image block corresponding to the any second image block, the fidelity map including the fidelity value of any second image block, and the fidelity value of any second image block is used to represent the distortion between any second image block and the first image block corresponding to the any second image block.
[0298] Among them, the position of any second image block among the multiple second image blocks in the reconstructed image is the same as the position of the first image block corresponding to the any second image block in the original image; the position of the preset area of the original image in the original image is the same as the position of the preset area of the reconstructed image in the reconstructed image, and the position of any second image block among the multiple second image blocks in the preset area of the reconstructed image is the same as the position of the first image block corresponding to the any second image block in the preset area of the original image.
[0299] Among them, the division strategy for dividing the original image is the same as the division strategy for dividing the reconstructed image, and the division strategy for dividing the preset area of the original image is the same as the division strategy for dividing the preset area of the reconstructed image, which means that the size of any first image block among the multiple first image blocks obtained by division is the same as the size of the corresponding second image block, and the position of any first image block among the multiple first image blocks obtained by division in the original image is the same as the position of the corresponding second image block in the reconstructed image; when the image block is rectangular, the position of the image block can be represented by the coordinates of the image block, and the coordinates of the image block are usually represented as the upper left corner brightness pixel coordinates of the image block. When dividing the original image or the reconstructed image, it can be divided according to basic units, that is, divided according to the same size of any image block, so that the size of any first image block obtained by division is the same, the size of any second image block obtained by division is also the same, and the size of the first image block and the size of the second image block are the size of the basic unit. It should be understood that the size of the basic unit can be determined according to the size of the original image or the reconstructed image or the number of the first image blocks or the second image blocks to be divided, and the minimum size of the basic unit can be 1×1 pixels, and the maximum size of the basic unit can be the size of the original image or the reconstructed image.
[0300] Among them, the fidelity can be calculated separately for any color component, for example, three fidelities are calculated separately for the three color components of RGB or YUV of an image, or the distortion intensity of the three color components can be fused to obtain a fidelity. When a fidelity is calculated for any color component, the fidelity map is a three-dimensional array; when a fidelity is calculated by fusion of three color components, the fidelity map is a two-dimensional array. Therefore, the fidelity map calculated for a reconstructed image or a preset area of a reconstructed image is a two-dimensional array of (W / M)×(H / N), or a three-dimensional array of (W / M)×(H / N)×C, wherein W and H represent the width and height of the original image or the reconstructed image, W and H represent the width and height of the preset area of the original image or the preset area of the reconstructed image, M and N represent the width and height of the basic unit used for the fidelity calculation, and C represents the number of color components of the original image or the reconstructed image. Without loss of generality, to simplify the description, the following assumes that C is 1 to describe the specific implementation method of the embodiment of the present application.
[0301] Specifically, the size of the basic unit can be set to M×N, and the original image and the reconstructed image can be divided according to the basic unit. The original image can be divided into R rows and S columns, with a total of R×S first image blocks, where R=(W / M), S=(H / N); and the reconstructed image can be divided into R rows and S columns, with a total of R×S second image blocks; and the size of the first image blocks and the second image blocks obtained by division is M×N. For example, the original image is divided into 4 rows and 7 columns, with a total of 4×7=28 first image blocks, as shown in FIG. Figure 8 As shown; and the reconstructed image is divided into 4 rows and 7 columns, a total of 4×7=28 second image blocks, such as Fig. 9 As shown; and the size of the first image block and the second image block obtained by division is M×N. The same method is also used to divide the preset area of the original image and the preset area of the reconstructed image. The preset area of the original image can be divided into R rows and S columns, with a total of R×S first image blocks; and the preset area of the reconstructed image can be divided into R rows and S columns, with a total of R×S second image blocks; and the size of the first image block and the second image block obtained by division is M×N.
[0302] Among them, after dividing to obtain R×S first image blocks and R×S second image blocks, the fidelity of any second image block relative to the corresponding first image block can be calculated, that is, the fidelity of the second image block in the i-th row and j-th column in the reconstructed image relative to the first image block in the i-th row and j-th column in the original image is calculated, wherein 1≤i≤R, 1≤j≤S, thereby obtaining the fidelity of any second image block among the R×S second image blocks and the corresponding first image block, that is, obtaining the fidelity values of the R×S second image blocks; and then obtaining a fidelity map based on the fidelity values of these R×S second image blocks.
[0303] In one example, the decoding is performed by a decoder or a decoding device; the decoder or the decoding device stores the size of the first image block and / or the second image block; or, the decoder or the decoding device stores the number of the first image blocks and / or the number of the second image blocks; or, the size of the first image block and / or the second image block is the input for the decoding; or, the number of the first image blocks and / or the number of the second image blocks is the input for the decoding; or, the decoder or the decoding device stores the size of the basic unit and / or the number of the basic units; or, the size of the basic unit and / or the number of the basic units is the input for the decoding.
[0304] Specifically, the decoding end stores the values of M and N or the values of R and S, or the values of M and N or the values of R and S as the input of decoding; in this way, it can be ensured that after the decoding end decodes and obtains the reconstructed image and the reconstructed image of the fidelity map, it knows that the value of any element in the reconstructed image of the fidelity map represents the fidelity of which image block in the reconstructed image. When the values of M and N or the values of R and S are the input of decoding, they can be combined with at least one of the first code stream and the second code stream and then transmitted to the decoding end, or they can be not combined with at least one of the first code stream and the second code stream, but transmitted to the decoding end alone.
[0305] It should be noted that Figure 8 Divide the original image and Fig. 9 The reconstructed image is divided in a uniform manner, and the sizes of any first image block or second image block obtained are the same. It is understandable that the original image or the reconstructed image can also be divided in a non-uniform manner, in which case the sizes of any first image block or second image block obtained by division are the same. When non-uniform division is used, the decoding end needs to know the size and position of any first image block and / or the second image block.
[0306] In an embodiment of the present application, the size of the original image is the same as the size of the reconstructed image, and the size and position of the preset area in the original image are the same as the size and position in the reconstructed image; the original image is divided into a plurality of first image blocks according to the same division strategy, and the reconstructed image is divided into a plurality of second image blocks; or the preset area of the original image is divided into a plurality of first image blocks according to the same division strategy, and the preset area of the reconstructed image is divided into a plurality of second image blocks; there is a one-to-one correspondence between the plurality of first image blocks obtained by division and the second image blocks obtained by division, wherein the size of any first image block is the same, the size of any second image block is also the same, and the size of the first image block is also the same as the size of the second image block; therefore, the first image block and the second image block can be As a basic unit for fidelity calculation, the fidelity value of any second image block can be calculated based on any second image block among the multiple second image blocks and its corresponding first image block, and the fidelity values of the multiple second image blocks are also the fidelity values of each area of the reconstructed image. A fidelity map can be obtained according to the fidelity values of the multiple second image blocks; wherein, when the first image block is obtained by dividing the original image and the second image block is obtained by dividing the reconstructed image, the fidelity map is used to characterize the fidelity of the reconstructed image; when the first image block is obtained by dividing a preset area of the original image and the second image block is obtained by dividing a preset area of the reconstructed image, the fidelity map is used to characterize the fidelity of the preset area of the reconstructed image; thereby facilitating the acquisition of a fidelity map for characterizing the distortion intensity information of the encoded image.
[0307] In a possible design, the fidelity map includes a plurality of first elements, the plurality of second image blocks correspond to the plurality of first elements one-to-one, the value of any first element among the plurality of first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image. The first element may also be referred to as a pixel point of the fidelity map.
[0308] The fidelity map construction process is as follows: determining a plurality of first elements based on the plurality of second image blocks, wherein the plurality of second image blocks correspond one-to-one to the plurality of first elements, and the value of any first element among the plurality of first elements is the fidelity value of the second image block corresponding thereto; obtaining the fidelity map based on the plurality of first elements, wherein the position of any first element among the plurality of first elements in the fidelity map is determined based on the position of the second image block corresponding thereto in the reconstructed image, or the position of any first element among the plurality of first elements in the fidelity map is determined based on the position of the second image block corresponding thereto in a preset area of the reconstructed image.
[0309] When calculating the fidelity map, the image is divided into basic units for fidelity calculation. Therefore, there are as many first elements in the fidelity map as there are basic units; that is, there are as many first elements in the fidelity map as there are first image blocks or second image blocks; any first element has two attributes, namely, the fidelity value and the position of the fidelity value in the fidelity map. Therefore, any first element in the fidelity map is used to represent the fidelity value of the corresponding second image block, and the value of any first element in the fidelity map is the fidelity value of the corresponding second image block.
[0310] Specifically, the reconstructed image or the preset area of the reconstructed image is divided into R×S second image blocks, and the fidelity map is a two-dimensional array of R rows and S columns, with a total of R×S first elements, and the R×S second image blocks correspond to the R×S first elements one by one, and the value of any first element among the R×S first elements is the fidelity value of the corresponding second image block. If the fidelity map of the entire reconstructed image is calculated, the position of any first element among the R×S first elements in the fidelity map is the same as the position of the corresponding second image block in the reconstructed image; if the fidelity map of the preset area of the reconstructed image is calculated, the position of any first element among the R×S first elements in the fidelity map is the same as the position of the corresponding second image block in the preset area of the reconstructed image. That is, the first element of the i-th row and j-th column in the fidelity map corresponds to the second image block of the i-th row and j-th column in the reconstructed image or the preset area of the reconstructed image, and the value of the first element of the i-th row and j-th column in the fidelity map is the fidelity value of the second image block of the i-th row and j-th column in the reconstructed image or the preset area of the reconstructed image, where 1≤i≤R, 1≤j≤S. For example, Fig.10 As shown, Fig.10 yes Fig. 9 The fidelity map of the reconstructed image shown, Fig.10 The grid in represents a first element, and the number in any grid represents the value of the first element, that is, the fidelity value of the second image block corresponding to the first element; Fig. 9 In , the reconstructed image is divided into 4 rows and 7 columns, with a total of 4×7=28 second image blocks; Fig.10 In , the fidelity map is a two-dimensional array with 4 rows and 7 columns, with a total of 4×7=28 first elements; Fig.10 The first element of the i-th row and j-th column in Fig. 9 The second image block in the i-th row and j-th column corresponds to, Fig.10 The value of the first element in row i and column j is Fig. 9 The fidelity value of the second image block in the i-th row and j-th column in , where 1≤i≤4, 1≤j≤7.
[0311] In an embodiment of the present application, the fidelity map is a two-dimensional array, and the reconstructed image is divided into multiple second image blocks. The fidelity map can be obtained according to the fidelity values of the multiple second image blocks, that is, multiple first elements can be determined according to the multiple second image blocks, and the multiple second image blocks correspond one-to-one to the multiple first elements. The value of any first element among the multiple first elements is the fidelity value of the second image block corresponding thereto; and the position of any first element among the multiple first elements in the fidelity map is determined according to the position of the second image block corresponding thereto in the reconstructed image. Specifically, the position of any first element among the multiple first elements in the fidelity map is the same as the position of the second image block corresponding thereto in the reconstructed image or in a preset area of the reconstructed image, so that the elements at each position of the fidelity map represent the fidelity of the area corresponding to its position in the reconstructed image or in the preset area of the reconstructed image, which is beneficial for the fidelity map to be used to represent the distortion intensity information of the encoded image.
[0312] In a possible design, the second image block includes three color components, the fidelity map is a three-dimensional array including three dimensions of color component, width and height, the two-dimensional array under any color component A in the fidelity map includes multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element, the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image. Wherein, when the fidelity map is a three-dimensional array including three dimensions of color component, width and height, the height indicates that the two-dimensional array under any color component A includes multiple row first elements, the width indicates that the two-dimensional array under any color component A includes multiple column first elements, the number of the multiple first elements is equal to the product of the width and the height, and the color component A is any one of the three color components.
[0313] In an embodiment of the present application, the original image or the reconstructed image includes three color components. When calculating the fidelity map, a two-dimensional array fidelity map is calculated for any color component. The two-dimensional arrays for the three color components constitute a fidelity map for a three-dimensional array. The first element in the two-dimensional array for any color component A in the fidelity map of the three-dimensional array represents the fidelity of the reconstructed image or the area corresponding to its position in a preset area of the reconstructed image for any color component A. This is beneficial for the fidelity map of the three-dimensional array to represent the distortion intensity information of the three color components of the encoded image.
[0314] In a possible design, encoding the fidelity map to obtain the second code stream includes: performing entropy encoding on any of the first elements to obtain the second code stream, the entropy encoding of the any of the first elements being independent of the entropy encoding of other first elements; or, determining a probability distribution of a value of the any of the first elements or a predicted value of the any of the first elements according to a value of at least one of the encoded first elements, and performing entropy encoding on the any of the first elements according to the probability distribution of the value of the any of the first elements or the predicted value of the any of the first elements to obtain the second code stream; wherein the second code stream includes code streams of the multiple first elements.
[0315] Specifically, for the entropy coding process of any first element in the fidelity map, if there is no encoded first element, entropy coding is directly performed on the any first element to obtain a code stream of the any first element; if there is an encoded first element, the probability distribution of the value of the any first element or the predicted value of the any first element is determined according to the value of at least one first element in the encoded first elements, and entropy coding is performed on the any first element according to the probability distribution of the value of the any first element or the predicted value of the any first element to obtain a code stream of the any first element; wherein the second code stream includes the code streams of the multiple first elements.
[0316] Among them, in the entropy coding process of any first element, the value of the encoded first element can be used to determine the probability distribution of the value of the first element, and the value of the first element is also the fidelity value of the second image block, that is, the probability distribution of the fidelity value of the current code is determined by using the encoded fidelity value, so as to assist in improving the entropy coding efficiency. Among them, the input for arithmetic coding of symbols is the symbol probability distribution. Among them, any encoded fidelity value can be used to determine the probability distribution of the fidelity value of the current code. For example, the encoded fidelity values at the left, top, top left, etc. of the fidelity value of the current code in the fidelity map determine the probability distribution of the fidelity value of the current code, to assist in improving the entropy coding efficiency. For example, different Huffman code tables can be selected according to the probability distribution of the fidelity value of the current code, or the sub-interval division method of the arithmetic coding can be determined.
[0317] Alternatively, during the entropy coding process, for any first element, the value of the encoded first element can be used to determine the predicted value of the first element, that is, the encoded fidelity value can be used to determine the predicted value of the currently encoded fidelity, and then the difference between the value of the first element and the predicted value of the first element is entropy coded to obtain the code stream of the first element.
[0318] In an embodiment of the present application, the fidelity map is encoded to obtain a second code stream, that is, any first element in the fidelity map is encoded to obtain a code stream of any first element, and the second code stream includes the code stream of any first element in the fidelity map; in the entropy coding process, the value of the encoded first element can be used to determine the probability distribution of the value of the currently encoded first element, for example, the value of the adjacent first element to the left, above, upper left, etc. of the currently encoded first element is used to determine the probability distribution of the value of the currently encoded first element or the predicted value of the currently encoded first element, and then the currently encoded first element is encoded according to the probability distribution of the value of the currently encoded first element or the predicted value of the currently encoded first element, so as to assist in improving the entropy coding efficiency.
[0319] It should be understood that the above entropy coding is only an exemplary entropy coding method; in the entropy coding process, the present application can adopt various existing entropy coding techniques, such as Huffman coding, arithmetic coding, context modeling arithmetic coding (a context modeling method is described here to assist arithmetic coding), binary arithmetic coding, etc.; the present application does not make specific limitations on this.
[0320] In one possible design, encoding the fidelity map to obtain a second code stream includes: quantizing any of the first elements to obtain a quantized first element; encoding the quantized first element to obtain the second code stream; wherein the second code stream includes a code stream of the multiple first elements.
[0321] The quantization step sizes for quantizing any first element among the multiple first elements may be the same or different.
[0322] In one example, the decoding is performed by a decoder or a decoding device; the decoder or the decoding device stores a quantization step for quantizing any first element of the plurality of first elements; or the quantization step for quantizing any first element of the plurality of first elements is input to the decoding, so that the decoding end can restore the fidelity value of the original order of magnitude.
[0323] In an embodiment of the present application, the fidelity map is encoded to obtain a second code stream, that is, any first element in the fidelity map is encoded to obtain a code stream of any first element in the fidelity map, and the second code stream includes the code stream of any first element in the fidelity map; in the process of encoding any first element in the fidelity map, any first element can be quantized, and then any quantized first element can be encoded to obtain a code stream of any first element; quantizing any first element is to quantize the value of any first element, or to scale any fidelity value in the fidelity map; the purpose of quantization is to reduce the dynamic range of the fidelity value in the fidelity map to reduce the encoding overhead of the fidelity map.
[0324] Fig.11 1 is a flowchart showing a process 1100 of a decoding method according to an embodiment of the present application. The process 1100 may be performed by a decoding device, such as a video decoder 30. The process 1100 is described as a series of steps or operations. It should be understood that the process 1100 may be performed in various orders and / or may occur simultaneously, not limited to Fig.11 The execution order shown. Process 1100 includes but is not limited to the following steps or operations:
[0325] 1101. Decode a first code stream to obtain a reconstructed image of an original image.
[0326] The first code stream is a code stream obtained by encoding the original image, that is, the encoded image data 21; the reconstructed image of the original image is also the decoded image 331, hereinafter referred to as the reconstructed image.
[0327] 1102. Decode a second code stream to obtain a reconstructed image of a fidelity map, wherein the second code stream is obtained by encoding the fidelity map, and the reconstructed image of the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image.
[0328] Among them, since the reconstruction map of the fidelity map is obtained by decoding the bitstream of the fidelity map, the reconstruction map of the fidelity map has the same size and properties as the fidelity map; when the fidelity map is used to characterize the fidelity of the entire reconstructed image, the reconstruction map of the fidelity map is also used to characterize the fidelity of the entire reconstructed image; when the fidelity map is used to characterize the fidelity of a preset area of the reconstructed image, the reconstruction map of the fidelity map is also used to characterize the fidelity of the preset area of the reconstructed image. If the fidelity map is used to characterize the fidelity of a preset area of the reconstructed image, the decoding end needs to obtain the position of the preset area in the original image and / or the position of the preset area in the reconstructed image from the encoding end, so that after the decoding end decodes and obtains the reconstruction map of the fidelity map, it knows which specific area in the reconstructed image the reconstruction map of the fidelity map is used to characterize the fidelity.
[0329] It should be understood that since the second bitstream is obtained by encoding the fidelity map, the first bitstream can be decoded to obtain a reconstructed image of the fidelity map. Depending on the encoding method, the reconstructed image of the decoded fidelity map can be the same as the fidelity map or different from the fidelity map. Specifically, if the encoding is lossless, the reconstructed image of the fidelity map is the same as the fidelity map; if the encoding is lossy, the reconstructed image of the fidelity map includes the encoding distortion generated by encoding the fidelity map.
[0330] In an embodiment of the present application, an original image is encoded to obtain a first code stream, and a fidelity map is encoded to obtain a second code stream, and the fidelity map is used to represent the distortion between at least a portion of the original image and at least a portion of the reconstructed image, wherein the distortion includes a difference; a decoding end decodes the first code stream to obtain a reconstructed image of the original image, and a decoding end decodes the second code stream to obtain a reconstructed map of the fidelity map; and if the encoding is lossless encoding, the reconstructed map of the fidelity map is the same as the fidelity map; if the encoding is lossy encoding, the reconstructed map of the fidelity map includes the encoding distortion generated by encoding the fidelity map; therefore, the reconstructed map of the fidelity map can be used to represent the distortion between at least a portion of the original image and at least a portion of the reconstructed image; therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0331] In one possible design, the fidelity map includes a fidelity value of any second image block among a plurality of second image blocks, and the fidelity value of any second image block is used to represent the distortion between the any second image block and an original image block corresponding to the any second image block.
[0332] Among them, the multiple second image blocks are obtained by dividing the reconstructed image, the multiple second image blocks correspond one-to-one to the multiple original image blocks, the original image blocks are image blocks in the original image, for example, the original image blocks are the aforementioned first image blocks; the multiple original image blocks are obtained by dividing the original image, the multiple second image blocks are obtained by dividing the reconstructed image, and the division strategy for dividing the original image is the same as the division strategy for dividing the reconstructed image; or the multiple original image blocks are obtained by dividing a preset area of the original image, the multiple second image blocks are obtained by dividing a preset area of the reconstructed image, and the division strategy for dividing the preset area of the original image is the same as the division strategy for dividing the preset area of the reconstructed image.
[0333] In one possible design, the fidelity map includes multiple first elements, the multiple second image blocks correspond one-to-one to the multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0334] In one possible design, the second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions: color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0335] In one possible design, decoding the second code stream to obtain a reconstructed image of the fidelity map includes: decoding the second code stream to obtain a reconstruction fidelity value of any of the first elements; and obtaining the reconstructed image of the fidelity map according to the reconstruction fidelity value of any of the first elements.
[0336] The reconstruction fidelity value of the first element is also the reconstruction of the value of the first element. The position of the reconstruction fidelity value of any first element in the reconstruction map of the fidelity map is determined according to the position of the second image block corresponding to any first element in the reconstructed image. Alternatively, the second code stream includes the position of any first element in the fidelity map; the position of the reconstruction fidelity value of any first element in the reconstruction map of the fidelity map is determined according to the position of any first element in the fidelity map.
[0337] Among them, any first element in the reconstruction map of the fidelity map has two attributes, namely, the reconstruction fidelity value of the any first element and the position of the reconstruction fidelity value in the reconstruction map of the fidelity map.
[0338] It should be understood that if the encoding is lossless encoding, the reconstruction fidelity value of any first element is the value of the first element. At this time, a, b, c, d, e, f, g, and h in the reconstruction image of the fidelity map are 15, 67, 99, 134, 16, 76, 123, and 187, respectively. Fig.10 and Fig.12 As shown, Fig.12 A grid in represents a first element, and the number in any grid represents the reconstruction fidelity value of the first element. If the encoding is lossy encoding, the reconstruction fidelity value of any first element is the sum of the value of the first element and the encoding distortion.
[0339] Among them, if there are multiple first elements in the fidelity map, then the reconstructed image or the preset area of the reconstructed image can also be divided into multiple second image blocks, and the multiple first elements correspond to the multiple second image blocks one-to-one, and the value of any first element among the multiple first elements is the fidelity value of the corresponding second image block; since there are multiple reconstruction fidelity values of the first elements in the reconstruction map of the fidelity map, the reconstruction fidelity values of the multiple first elements correspond to the multiple first elements one-to-one, so the reconstruction fidelity values of the multiple first elements correspond to the multiple second image blocks one-to-one, and the reconstruction fidelity value of any first element among the reconstruction fidelity values of the multiple first elements is the fidelity value of the corresponding second image block.
[0340] Specifically, Fig.12 As shown, since the reconstructed image or the preset area of the reconstructed image is divided into R×S second image blocks, the reconstructed image of the fidelity map is a two-dimensional array of R rows and S columns, with a total of R×S first elements, and the R×S second image blocks correspond to the R×S first elements one by one, and the reconstruction fidelity value of any first element among the R×S first elements is the fidelity value of the second image block corresponding thereto. If the reconstructed image of the fidelity map of the entire reconstructed image is calculated, the position of any first element among the R×S first elements in the reconstructed image of the fidelity map is the same as the position of the second image block corresponding thereto in the reconstructed image; if the reconstructed image of the fidelity map of the preset area of the reconstructed image is calculated, the position of any first element among the R×S first elements in the reconstructed image of the fidelity map is the same as the position of the second image block corresponding thereto in the preset area of the reconstructed image. That is, the first element of the i-th row and j-th column in the reconstruction map of the fidelity map corresponds to the second image block of the i-th row and j-th column in the reconstructed image or the preset area of the reconstructed image, and the reconstruction fidelity value of the first element of the i-th row and j-th column in the reconstruction map of the fidelity map is the fidelity value of the second image block of the i-th row and j-th column in the reconstructed image or the preset area of the reconstructed image, where 1≤i≤R, 1≤j≤S.
[0341] In the embodiment of the present application, the second code stream includes the code stream of any first element in the fidelity map, so the reconstruction fidelity value of any first element can be obtained by decoding the second code stream; it should be understood that if the encoding is lossless encoding, the reconstruction fidelity value of the first element is the value of the first element; if the encoding is lossy encoding, the reconstruction fidelity value of the first element includes the encoding distortion generated by encoding the first element, and the reconstruction fidelity value of the first element is the sum of the value of the first element and the encoding distortion; thus, according to the reconstruction fidelity value of any first element, a reconstruction map of the fidelity map can be obtained, so the reconstruction map of the fidelity map can be used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image; therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0342] In one possible design, the second code stream is obtained by encoding the quantized first element; and the decoding of the second code stream to obtain a reconstructed image of the fidelity map includes: decoding the second code stream to obtain a reconstruction fidelity value of the quantized first element; dequantizing the reconstruction fidelity value of the quantized first element to obtain a reconstruction fidelity value of any first element; and obtaining the reconstructed image of the fidelity map according to the reconstruction fidelity value of any first element.
[0343] Among them, since the quantization step size for quantizing any first element among the multiple first elements can be the same or different, the quantization step size for inverse quantization of any first element among the multiple first elements can be the same or different; and the quantization step size for quantizing any first element among the multiple first elements is the same as the quantization step size when quantizing the corresponding first element.
[0344] In one example, the decoding is performed by a decoder or a decoding device; the decoder or the decoding device stores a quantization step for quantizing any first element of the multiple first elements; or the decoder or the decoding device stores a quantization step for dequantizing any first element of the multiple first elements; or the quantization step for quantizing any first element of the multiple first elements is the input of the decoding; or the quantization step for dequantizing any first element of the multiple first elements is the input of the decoding. In this way, the decoding end can restore the fidelity value of the original order of magnitude.
[0345] In an embodiment of the present application, in order to reduce the coding overhead, the encoding end may quantize the first element and then encode it to obtain a code stream of the first element, so the code stream of the first element obtained by the decoding end may be obtained by encoding the quantized first element; in this case, the second code stream is decoded to obtain the reconstruction fidelity value of the quantized first element, and the reconstruction fidelity value of the quantized first element needs to be dequantized to obtain the reconstruction fidelity value of any first element; then, a reconstruction map of the fidelity map can be obtained based on the reconstruction fidelity value of any first element; in this way, the distortion intensity information of the encoded image can be obtained at the decoding end, and the coding overhead can be reduced.
[0346] In one possible design, the method also includes: processing the reconstructed image or a preset area of the reconstructed image according to a reconstruction map of the fidelity map to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determining whether to apply the reconstructed image according to the reconstruction map of the fidelity map.
[0347] Specifically, the image quality of the reconstructed image or the preset area of the reconstructed image can be determined based on the reconstruction map of the fidelity map. When the image quality of the reconstructed image or the preset area of the reconstructed image is poor, the reconstructed image or the preset area of the reconstructed image is processed by an image quality enhancement algorithm to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or the fidelity of the reconstructed image or the preset area of the reconstructed image can be determined based on the reconstruction map of the fidelity map, and whether to apply the reconstructed image can be determined based on the fidelity of the reconstructed image or the preset area of the reconstructed image. For example, when the fidelity value of the reconstructed image or the preset area of the reconstructed image is lower than a preset fidelity threshold, it is determined that the reconstructed image should not be applied.
[0348] Among them, when processing the reconstructed image or the preset area of the reconstructed image, in the learning-based post-processing enhancement algorithm, the image can be divided into B distortion ranges according to the degree of distortion and trained separately to obtain multiple image enhancement models, where B is an integer greater than 1. The decoding end can determine the degree of distortion of different areas of the reconstructed image based on the fidelity map, and select different models for image enhancement in different areas. Using a training model that better matches the distortion distribution can achieve better image quality improvement effects. For another example, a single-model image enhancement algorithm can use a quantization parameter map as additional input information for the network during training and use. Under the guidance of the quantization parameter map, the network can output an image with a better enhancement effect.
[0349] In an embodiment of the present application, the decoding end can process the reconstructed image or a preset area of the reconstructed image according to the reconstruction map of the fidelity map to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determine whether to apply the reconstructed image according to the reconstruction map of the fidelity map; thereby facilitating the application of the reconstructed image.
[0350] It should be noted that Fig.11 The decoding method described is Figure 7 The inverse process of the encoding method described above, Fig.11 The steps or operations described can be found in Figure 7 A description of the steps or actions being described.
[0351] The following is a video image encoding and decoding process. Figures 7 to 12 The technical solution provided is further introduced.
[0352] 1. Encoding implementation method
[0353] (1) Original image encoding
[0354] For any video image to be encoded in the video sequence (that is, the original image), it is input into the video encoder for encoding operation, and after the encoding operation, the first code stream of the video image and the reconstructed image of the video image are output, that is, the first code stream of the original image and the reconstructed image of the original image are output; wherein, the video image to be encoded is often referred to as the encoded image when the encoding operation is performed; the specific encoding operation can refer to the description in the previous text, which will not be repeated here. The above encoding operation can be any video image encoding method, such as the existing H.264, H.265, H.266, AVS2, AVS3, AV1 and other domestic and foreign standard solutions or industry solutions, or the video image encoding solution based on deep neural network being studied by academia and industry, or other solutions that can compress the input video image. The reconstructed image refers to an image containing encoding distortion after the encoding operation. Typical encoding distortions include blocking effect, ringing effect, blur, etc. Some codec schemes based on deep learning technology, such as those based on generative adversarial networks (GANs), can also generate new types of coding distortions such as false image content or texture details.
[0355] The above method can be used to process each video image in a video sequence, thereby generating a code stream of the entire video sequence and a reconstructed image of each video image in the video sequence, wherein the code stream of the entire video sequence includes the first code stream of each video image in the video sequence.
[0356] (2) Calculation and representation of fidelity graph
[0357] For a video sequence, a fidelity map can be calculated for each video image in the video sequence, thereby obtaining a fidelity map for each video image in the video sequence; or a portion of the video images in the video sequence can be selected to calculate the fidelity map, thereby obtaining a fidelity map for a portion of the video images in the video sequence. For example, only intra-frame coded images can be selected to calculate their fidelity maps; or only scene switching frames, i.e., the first frame image after the scene switching in the video sequence, can be selected to calculate their fidelity maps; or only key frames, i.e., the lowest temporal layer images when encoded according to the hierarchical B frame structure reference structure, can be selected to calculate their fidelity maps; or other selection rules and combinations of multiple selection rules can be selected.
[0358] Among them, the fidelity map of a coded image can be calculated according to the preset basic unit size. For example, the coded image can be divided into multiple image blocks of size M×N, and each image block is used as a basic unit for calculating fidelity; M and N are the width and height of the image block respectively, and the values of M and N can be equal or unequal. The size of the basic unit can be preset and stored in the encoder and decoder; the size of the basic unit can also be flexibly set by the encoder according to the richness of the content of all video images or part of the video images in the video sequence, and the size of the set basic unit is notified to the decoder. It should be understood that in addition to directly specifying the width and height of the basic unit, the number of basic units in a coded image can also be specified, such as specifying the number of basic units in the horizontal and vertical directions of the coded image. Because a coded image is divided into a group of basic units (i.e., multiple basic units) in a uniform division manner, the above two basic unit representation methods are equivalent.
[0359] When calculating the fidelity, any quality evaluation method with a reference, such as Mean Squared Error (MSE), Sum of Absolute Difference (SAD), Structural Similarity Index Measurement (SSIM), etc., can be used to calculate the distortion intensity value of the reconstructed image relative to the original image, or to calculate whether there is a synthetic false image content in the reconstructed image content. Specifically, both the original image and the reconstructed image are evenly divided according to the size of the basic unit or the number of basic units to divide the original image into a plurality of first image blocks, and the reconstructed image is divided into a plurality of second image blocks; the fidelity of any second image block in the plurality of second image blocks relative to the corresponding first image block is calculated, and the fidelity of the second image block relative to the corresponding first image block can be called a fidelity value, so that the fidelity values of the plurality of second image blocks can be obtained; and the fidelity map of the reconstructed image can be obtained according to the fidelity values of the plurality of second image blocks. The fidelity map is an array, which includes multiple first elements, and the multiple first elements correspond one-to-one to multiple second image blocks obtained by dividing the reconstructed image. Any element among the multiple first elements is used to represent the fidelity value of the second image block corresponding to it, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to it, and the position of any first element among the multiple first elements in the array is the same as the position of the corresponding second image block in the reconstructed image.
[0360] After the fidelity map is calculated, it can be quantized. Specifically, the value of the first element in the fidelity map can be scaled using a specified quantization step size, that is, the fidelity value of the second image block corresponding to the first element in the fidelity map can be scaled using a specified quantization step size. The purpose of quantization is to reduce the dynamic range of the fidelity value in the fidelity map to reduce the encoding overhead of the fidelity map. Quantization will cause distortion to the fidelity value, so when setting the quantization step size, the encoder can set the corresponding quantization step size according to the accuracy requirements of the decoder for the fidelity of the reconstructed image. For example, the same quantization step size can be set for all coded images of a video sequence, that is, the fidelity maps of all coded images are quantized using the same quantization step size; or a quantization step size can be set for a part of the coded images of a video sequence, and another quantization step size can be set for another part of the coded images of the video sequence, that is, the fidelity maps of some coded images in all coded images are quantized using the same quantization step size; or different quantization step sizes can be set for different coded images of a video sequence, that is, the fidelity maps of any coded image in all coded images are quantized using different quantization step sizes; or even different quantization step sizes can be set for different image blocks of each coded image of a video sequence, that is, the fidelity values of all coded blocks are quantized using different quantization step sizes. Among them, the quantization step size set by the encoder needs to be informed to the decoder; that is, the encoder needs to inform the decoder which quantization step size is used for the fidelity map of which coded image to be quantized or which quantization step size is used for the fidelity value of which coded block to be quantized. In addition, in addition to the flexible specification of the quantization step size at the encoder, the preset quantization step size can also be stored in the encoder and decoder for use; that is, which quantization step size is used when the fidelity map of the encoded image is quantized or which quantization step size is used when the fidelity value of the encoded block is quantized, and the encoder and decoder are pre-stored. In this way, the encoder and decoder can obtain the same quantization step size, which enables the decoder to restore the fidelity value of the original order of magnitude in the encoder.
[0361] The above-mentioned fidelity graph can be represented by the syntax structure of Table 5. In Table 5, fidelity_metric_idc is an indication of the quality evaluation method used to calculate the fidelity, which can be used to indicate which evaluation method is used in the preset quality evaluation method list to calculate the fidelity graph. The quality evaluation method list is preset and can include quality evaluation indicators such as MSE, SSE, and SSIM; base_unit_width and base_unit_height are respectively the width and height of the basic unit used when calculating the fidelity; quantization_step is the quantization step; and fidelity_value is the fidelity value calculated in one basic unit.
[0362] Table 5
[0363]
[0364] It should be understood that the encoding end can calculate the fidelity map for the entire encoded image, or it can calculate the fidelity map for a preset area in the encoded image, wherein the process of calculating the fidelity map for the preset area in the encoded image by the encoding end is the same as the process of calculating the fidelity map for the entire encoded image, which is not repeated here. If the encoding end calculates the fidelity map for the preset area in the encoded image, the encoding end needs to inform the decoding end of the position of the preset area in the encoded image or the coordinates of the preset area in the encoded image.
[0365] (3) Fidelity Graph Coding
[0366] As shown in Table 5, when encoding the fidelity map, each fidelity_value in the fidelity map of the current encoded image can be traversed and compressed to obtain the bitstream of all fidelity_values. The second bitstream obtained by encoding the fidelity map includes the bitstream of all fidelity_values. Specifically, the value can be binarized by fixed-length coding, exponential Golomb coding, etc. to obtain a binary string, and then each binary character in the binary string is entropy encoded; or the value can be directly entropy encoded by Huffman coding, multi-value arithmetic coding, etc. In the entropy coding process, the fidelity value encoded in the fidelity map can be used to determine the probability distribution of the fidelity value of the current code or the predicted value of the fidelity value of the current code, for example, the fidelity value encoded on the left, above or above the left of the fidelity value of the current code is used to determine the probability distribution of the fidelity value of the current code or the predicted value of the fidelity value of the current code, so as to assist in improving the entropy coding efficiency. For example, different Huffman code tables may be selected or the sub-interval division method of the arithmetic coding may be determined according to the determined probability distribution of the fidelity value of the current coding or the predicted value of the fidelity value of the current coding.
[0367] Since the fidelity map is represented as a two-dimensional or three-dimensional array, any existing encoding method for monochrome images or color images can be used to compress and encode the fidelity map. In this case, it is not necessary to traverse all the fidelity values calculated by the basic units in Table 5, but directly embed the fidelity map using the existing encoding method to encode the fidelity map to obtain the second code stream.
[0368] It should be noted that the second bitstream output by the fidelity map encoding can be embedded in the first bitstream output by the original image encoding, or it can be independently managed for operations such as transmission or storage.
[0369] Through the above encoding operation, the encoding end can obtain the first code stream of the original image and the second code stream of the fidelity map. After obtaining the first code stream and the second code stream, the first code stream and the second code stream are transmitted to the decoding end; wherein the code streams transmitted from the encoding end to the decoding end can be collectively referred to as compressed code streams.
[0370] 2. Decoding end implementation method
[0371] (1) Decoding to obtain reconstructed image
[0372] After receiving the compressed code stream from the encoder, the decoder obtains the first code stream of the original image from the compressed code stream, and inputs the first code stream into the video decoder to obtain the reconstructed image through the decoding operation. The decoding operation of the first code stream by the decoder is the inverse operation of the encoding operation of the original image by the encoder, and the reconstructed image obtained by the decoder is the same as the reconstructed image obtained by the encoder. The specific decoding operation can be referred to the description above, which will not be repeated here.
[0373] (2) Decoding to obtain the reconstructed image of the fidelity image
[0374] After receiving the compressed bitstream from the encoder, the decoder obtains the second bitstream of the fidelity map from the compressed bitstream, and performs a decoding operation on the second bitstream to obtain a reconstructed image of the fidelity map. The decoding operation of the decoder on the second bitstream is the inverse operation of the encoding operation of the encoder on the fidelity map. If the encoding operation of the encoder on the fidelity map adopts a lossless encoding method, as shown in Table 5, traverses each fidelity value in the fidelity map, and performs lossless entropy encoding on each fidelity value, then the reconstructed image of the fidelity map obtained by the decoding operation of the decoder is the same as the fidelity map calculated by the encoder. If the encoding operation of the fidelity map by the encoder adopts a lossy encoding method, such as using a monochrome image encoding method to process the fidelity map, then the reconstructed image of the fidelity map obtained by the decoding operation of the decoder is different from the fidelity map calculated by the encoder, because the reconstructed image of the fidelity map obtained by the decoding operation of the decoder will contain the encoding distortion introduced by the encoding operation of the fidelity map by the encoder.
[0375] (3) Application Fidelity Diagram
[0376] The decoding end determines the signal distortion of a reconstructed image or a preset area in the reconstructed image based on the reconstruction map of the fidelity map, and applies it to different business environments. For example, in a video surveillance scenario, the degree of distortion of a certain image area in the reconstructed image is determined based on the reconstruction map of the fidelity map. If it is greater than a preset threshold, the reconstructed image will not be applied; in a video conferencing scenario, the fidelity map can be used to determine whether the content seen is real or false. For another example, based on the degree of distortion of a certain image area in the reconstructed image, one of the image enhancement methods can be selected from a group of image enhancement methods according to preset rules and applied to the image area to improve the image quality.
[0377] Fig.13 1 is a flowchart showing a process 1300 of a decoding method according to another embodiment of the present application. The process 1300 may be performed by a decoding device, such as a video decoder 30. The process 1300 is described as a series of steps or operations. It should be understood that the process 1300 may be performed in various orders and / or may occur simultaneously, not limited to Fig.13 The execution order shown. Process 1300 includes but is not limited to the following steps or operations:
[0378] 1301. Decode a first bitstream to obtain a reconstructed image and target quantization parameter information of an original image, where the target quantization parameter information includes quantization parameter values of all or part of a plurality of second image blocks of the reconstructed image.
[0379] It should be understood that the first code stream is the encoded image data 21; the reconstructed image is the decoded image 331, so the reconstructed image is a decoded image; the original image is the image 17, so the original image is an encoded image.
[0380] In one possible design, the second image block is a coding unit.
[0381] The first code stream can be obtained by encoding the original image by an encoder specified by any existing mainstream video image coding standard (H.264, H.265, H.266, AVS3, etc.); wherein, in the original image coding process, the original image is divided into multiple coding units (i.e., coding blocks), and any one of the multiple coding units is encoded, thereby obtaining the code stream of any one of the multiple coding units, and the first code stream includes the code stream of the multiple coding units. It should be understood that the original image is divided into multiple coding units, and how to divide them is determined by the encoder. Correspondingly, the decoder that decodes the first code stream is any decoder that can perform the decoding operation specified by the aforementioned video image coding standard. The decoder performs the standard decoding operation on the input first code stream and outputs a reconstructed image of the original image.
[0382] The embodiment of the present application decodes the first code stream, and in addition to outputting the reconstructed image of the original image, it also outputs the target quantization parameter information obtained during the decoding process, which is also the quantization parameter information used in the original image encoding process. The quantization parameter information includes the quantization parameter values of all or part of the multiple coding units obtained by dividing any original image in the encoding process, the positions of all or part of the multiple coding units obtained by dividing any original image in the original image, and the sizes of all or part of the multiple coding units obtained by dividing any original image; wherein the quantization parameter value is also the value of the quantization parameter, and the quantization parameter is also the quantization parameter used by the quantization unit 208. In the quantization unit 208, any coding unit uses the corresponding quantization parameter for quantization. Therefore, the target quantization parameter information includes the quantization parameter values of all or part of the multiple coding units obtained by dividing the above original image, the position of any coding unit in the original image obtained by dividing the above original image, and the size of any coding unit in the multiple coding units obtained by dividing the above original image. The position of a coding unit in the original image can be represented by the coordinates of the coding unit, which are usually represented by the coordinates of the brightness pixel of the upper left corner of the coding unit. Since a coding unit is usually rectangular, its size can be represented by its width and height, usually in terms of the number of brightness pixels; if a coding unit is square, its size can be represented by only the side length or area.
[0383] 1302. Construct a quantization parameter map of the reconstructed image according to the target quantization parameter information, wherein the quantization parameter map of the reconstructed image is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image.
[0384] Among them, the quantization parameter map is also the quantization parameter map of the reconstructed image, which is a data structure indicating the quantization parameter value of the coding unit in the reconstructed image; and the quantization parameter value of the coding unit in the reconstructed image is also the quantization parameter value of the coding unit in the original image. It should be understood that the main purpose of the quantization parameter is to perform an inverse quantization operation; of course, the quantization parameter itself represents the signal distortion (fidelity), so the quantization parameter map can be used to characterize the fidelity of the reconstructed image, that is, the quantization parameter map is a form of fidelity map. Therefore, the quantization parameter map of the reconstructed image constructed according to the quantization parameter can be used to characterize the fidelity of the reconstructed image or can be used to characterize the fidelity of a preset area of the reconstructed image. The quantization parameter map is a two-dimensional array or a three-dimensional array, and the value of the second element in the quantization parameter map is the quantization parameter value of the coding unit; wherein any second element in the quantization parameter map has two attributes, namely, the quantization parameter value and the position of the quantization parameter value in the quantization parameter map. It should be understood that each quantization parameter value in the quantization parameter map represents the degree of distortion of a certain image block in the corresponding reconstructed image, and it is obvious that the quantization parameter map is also a form of fidelity map. If a quantization parameter map is constructed by the quantization parameter values of the three color components RGB or YUV of a reconstructed image, the quantization parameter map is a three-dimensional array; if a quantization parameter map is constructed by the fused quantization parameter values of the three color components of a reconstructed image, the quantization parameter map is a two-dimensional array. Without loss of generality, to simplify the description, the following describes the specific implementation method by taking the quantization parameter map as a two-dimensional array.
[0385] The quantization parameter map can be represented in a variety of ways, some of which are listed below as examples.
[0386] The first one is that the original image is divided into multiple coding units, so that the reconstructed image includes multiple coding units, and the coding units are used as basic units for constructing the quantization parameter map. At this time, the coding units in the reconstructed image are also the second image blocks, so the multiple coding units are multiple basic units for constructing the quantization parameter map; these multiple basic units correspond one-to-one to the multiple second elements in the quantization parameter map, and these multiple basic units are these multiple coding units, so the multiple coding units correspond one-to-one to the multiple second elements in the quantization parameter map, and the value of any second element in the quantization parameter map is set to the quantization parameter value of the corresponding coding unit.
[0387] In this representation, the position of any second element in the quantization parameter map is the same as the position of the corresponding coding unit in the reconstructed image. Furthermore, the quantization parameter map can also have the same spatial resolution as the corresponding original image or reconstructed image, that is, the quantization parameter map and the original image or reconstructed image have the same size. Fig.14As shown, the original image or the reconstructed image includes 6 coding units, that is, the original image or the reconstructed image includes 6 coding units, then there are 6 grids in the quantization parameter map, that is, there are 6 second elements in the quantization parameter map, wherein the 6 coding units correspond to the 6 second elements one by one, and the number in any grid is the quantization parameter corresponding to the corresponding coding unit; the size of the original image is W×H, the size of the reconstructed image is also W×H, then the size of the quantization parameter map is also W×H. Fig.14 The case where the quantization parameter map and the corresponding original image or reconstructed image have the same size is described only as an example. It should be understood that the size of the quantization parameter map and the size of the corresponding original image or reconstructed image may also have a certain scaling ratio.
[0388] The second method is to divide the reconstructed image into multiple basic units of equal size for constructing a quantization parameter map. In this case, the basic unit in the reconstructed image is also the second image block; wherein the size of the basic unit should be less than or equal to the size of the smallest coding unit; for any one of the multiple basic units, the quantization parameter value of the coding unit containing the basic unit is used as the quantization parameter value of the basic unit; since the multiple basic units in the reconstructed image correspond one-to-one to the multiple second elements in the quantization parameter map, the value of any second element in the quantization parameter map is the quantization parameter value of the corresponding basic unit in the reconstructed image, that is, the quantization parameter value of the coding unit containing the basic unit.
[0389] For better understanding, it can be simply understood that both the original image and the reconstructed image are divided into multiple basic units, the basic unit obtained by dividing the original image is called basic unit i, and the basic unit obtained by dividing the reconstructed image is called basic unit j; multiple basic units i correspond one-to-one to multiple basic units j, and multiple basic units j correspond one-to-one to multiple second elements, so multiple basic units i correspond one-to-one to multiple second elements, and the value of any second element in the multiple second elements is the quantization parameter value of the corresponding basic unit i, and the quantization parameter value of the basic unit i is the quantization parameter value of the coding unit including the basic unit i in the original image. Or another way of explaining is that the multiple coding units in the original image correspond one-to-one to the multiple reconstructed blocks in the reconstructed image, and the quantization parameter value of any reconstructed block in the multiple reconstructed blocks in the reconstructed image is the quantization parameter value of the corresponding coding unit; the reconstructed image is divided into multiple basic units of equal size for constructing a quantization parameter map, and the multiple basic units in the reconstructed image correspond one-to-one to the multiple second elements in the quantization parameter map, so the value of any second element in the quantization parameter map is the quantization parameter value of the corresponding basic unit in the reconstructed image, that is, the quantization parameter value of the reconstructed block containing the basic unit. In this way of expression, the position of any second element in the quantization parameter map in the quantization parameter map is the same as the position of the corresponding basic unit in the original image or the reconstructed image. For example, the size of the reconstructed image is W×H, and the coding units of the original image corresponding to the reconstructed image are divided as follows: Fig.14 As shown, with the size of coding unit 2 M×N as the size of the basic unit, the reconstructed image is divided into R×S basic units of size M×N; thus, the constructed quantization parameter map is a two-dimensional array of R rows and S columns, as shown Fig.15 As shown, there are a total of R×S second elements in the two-dimensional array, and the R×S second elements correspond to the R×S basic units one by one. The position of any second element in the two-dimensional array is the same as the position of the corresponding basic unit in the reconstructed image, and the value of any second element in the R×S second elements is the quantization parameter value of the corresponding basic unit; specifically, Fig.15 The value of the second element in is 22. Fig.14 The quantization parameter value of coding unit 1 in Fig.15 The value of the second element in is 24. Fig.14 The quantization parameter value of coding unit 2 in Fig.15 The value of the second element in is 26. Fig.14 The quantization parameter value of coding unit 3 in Fig.15 The value of the second element in is 18. Fig.14 The quantization parameter value of coding unit 4 in Fig.15 The value of the second element in is 20. Fig.14The quantization parameter value of coding unit 5 in Fig.15 The value of the second element in is 16. Fig.14 It should be understood that the size of the quantization parameter map can be the same as the size of the corresponding original image or reconstructed image, or can have a certain scaling ratio; wherein, Fig.15 The quantization parameter map shown is a scaled-down version of the original or reconstructed image.
[0390] It should be noted that the quantization parameter map can also be used to characterize the fidelity of a preset area of the reconstructed image. In this case, the value of the second element in the quantization parameter map is the quantization parameter value of some coding units in the original image, that is, only the quantization parameter values of some coding units are used when constructing the quantization parameter map. It should be understood that the quantization parameter map used to characterize the fidelity of the preset area of the reconstructed image can also adopt the representation method described above. In this case, the multiple basic units used to construct the quantization parameter map are obtained by dividing the preset area of the reconstructed image.
[0391] In an embodiment of the present application, the encoding end divides the original image into multiple coding units, and encodes the multiple coding units obtained by dividing the original image to obtain a first code stream; the decoding end decodes the first code stream to obtain a reconstructed image and target quantization parameter information of the original image, and the target quantization parameter information includes the quantization parameter values of all or part of the multiple coding units; according to the target quantization parameter information, a quantization parameter map of the reconstructed image can be constructed; and the quantization parameter map of the reconstructed image is a form of fidelity map, when the target quantization parameter information includes the quantization parameter values of all the coding units in the multiple coding units, the quantization parameter map of the reconstructed image is a fidelity map of the entire reconstructed image; when the target quantization parameter information includes the quantization parameter values of part of the coding units in the multiple coding units, the quantization parameter map of the reconstructed image is a fidelity map of a preset area of the reconstructed image; therefore, the quantization parameter map of the reconstructed image can be used to characterize the fidelity of the reconstructed image or to characterize the fidelity of a preset area of the reconstructed image; therefore, the embodiment of the present application can obtain distortion intensity information of the encoded image at the decoding end.
[0392] In an embodiment of the present application, the decoding end can obtain the quantization parameter values of each area (second image block) of the reconstructed image by decoding the first code stream obtained by encoding the original image; and a quantization parameter map of the reconstructed image can be constructed according to the quantization parameter values of all or part of the multiple second image blocks of the reconstructed image, and the quantization parameter map of the reconstructed image can be used to represent the distortion between at least a part of the area of the original image and at least a part of the area of the reconstructed image. Therefore, the embodiment of the present application can obtain the distortion intensity information of the encoded image at the decoding end.
[0393] In a possible design, the quantization parameter map of the reconstructed image includes a plurality of second elements, the plurality of second image blocks correspond one-to-one to the plurality of second elements, the value of any second element among the plurality of second elements is the quantization parameter value of the second image block corresponding to the any second element, the position of the any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of the any second element in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image. The second element may also be referred to as a pixel point of the quantization parameter map.
[0394] In one possible design, the second image block includes three color components, the quantization parameter map of the reconstructed image is a three-dimensional array including three dimensions of color component, width and height, the two-dimensional array under any color component A in the quantization parameter map of the reconstructed image includes multiple second elements, the value of any second element among the multiple second elements is the quantization parameter value of the color component A of the second image block corresponding to the any second element, the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in the reconstructed image, or the position of the two-dimensional array of any second element under any color component A in the quantization parameter map of the reconstructed image is determined according to the position of the second image block corresponding to the any second element in a preset area of the reconstructed image.
[0395] In one possible design, constructing the quantization parameter map of the reconstructed image according to the target quantization parameter information includes: when the target quantization parameter information includes quantization parameter values of some coding units among the multiple coding units, obtaining the quantization parameter value of the target coding unit according to the quantization parameter value of the some coding units and / or the reference quantization parameter map, wherein the reference quantization parameter map is the quantization parameter map of the reference image of the reconstructed image, and the target coding unit is the coding unit among the multiple coding units except the some coding units; obtaining the quantization parameter map of the reconstructed image according to the quantization parameter value of the some coding units and the quantization parameter value of the target coding unit.
[0396] Specifically, when the target quantization parameter information includes the quantization parameter values of all the coding units among the multiple coding units, the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of all the coding units; when the target quantization parameter information includes the quantization parameter values of some of the coding units among the multiple coding units, the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of the some of the coding units, or the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of the some of the coding units and a reference quantization parameter map, wherein the reference quantization parameter map is a quantization parameter map of a reference image of the reconstructed image.
[0397] It should be understood that the reference quantization parameter map is a quantization parameter map of a reference image of a reconstructed image, and the reference image of the reconstructed image is also a reference image of a coding picture. The current mainstream video image coding and decoding standards do not guarantee that a true and effective quantization parameter value can be derived for each coding unit at the decoding end. Taking the H.265 standard solution as an example, if the residual quantization of a coding unit is all 0, that is, no residual will be transmitted for it, the decoding end will skip the dequantization operation. In order to avoid transmitting useless quantization parameter information, the encoding end will not pass the quantization parameter value of the coding unit to the decoding end. At this time, the decoding end cannot obtain the quantization parameter value of the coding unit from the first bitstream at all. In addition, even if the quantization parameter value of a coding unit can be obtained through a decoding operation at the decoding end, the quantization parameter value cannot accurately represent the degree of distortion of the coding unit. Specifically, if the width of the residual of a coding unit is less than the quantization step size (the quantization step size is calculated based on the quantization parameter value), the encoding end will quantize the residual of the coding unit to all 0. This situation is more common in P and B frame encoding, and occurs when the reference image quality is high and the quantization step size of the coding unit is set large. At this time, although the decoder can obtain the quantization parameter value of the coding unit through the decoding operation, the quantization parameter value cannot reflect the actual distortion of the coding unit.
[0398] The scheme for constructing a quantization parameter map provided in the embodiment of the present application can be combined with any existing mainstream video image coding and decoding standard (H.264, H.265, H.266, AVS3, etc.), and a quantization parameter map is constructed according to the quantization parameter information obtained during the decoding process. In addition, the inaccurate quantization parameter value identified in the quantization parameter map can be corrected to obtain a corrected quantization parameter map, that is, a fidelity map, and finally the fidelity map is applied to the reconstructed image. Specifically, the quantization parameter value of the coding unit in the current original image can be used to correct the quantization parameter value of other coding units in the current original image, or the quantization parameter map of the reference image of the current original image (hereinafter referred to as the reference quantization parameter map) can be used to correct the inaccurate quantization parameter data. After the quantization parameter map of each reconstructed image is constructed and corrected, it will be saved and reserved for use as an input for constructing a quantization parameter map of a subsequent reconstructed image, that is, reserved for use as an input for constructing a quantization parameter map of a subsequent reconstructed image.
[0399] Among them, when the target quantization parameter information includes the quantization parameter values of all coding units in a plurality of coding units in a reconstructed image, a quantization parameter map of the reconstructed image can be constructed based on the quantization parameter values of all coding units, and in this case, the quantization parameter map of the reconstructed image is used to characterize the fidelity of the entire reconstructed image; or, the quantization parameter values of the coding units in a preset area of the reconstructed image can be selected from all coding units, and the quantization parameter map of the reconstructed image can be constructed based on the quantization parameter values of the coding units in the preset area of the reconstructed image, and in this case, the quantization parameter map of the reconstructed image is used to characterize the fidelity of the preset area of the reconstructed image.
[0400] Among them, when the target quantization parameter information includes quantization parameter values of some coding units among multiple coding units in the reconstructed image, and the quantization parameter values of some coding units completely include the quantization parameter values of coding units in a preset area of the reconstructed image, a quantization parameter map of the reconstructed image can be constructed according to the quantization parameter values of the coding units in the preset area of the reconstructed image; or, the quantization parameter values can be corrected according to the quantization parameter values of some coding units among the multiple coding units or the reference quantization parameter map to obtain the quantization parameter values of all coding units among the multiple coding units, and then, the quantization parameter map of the reconstructed image is constructed according to the quantization parameter values of all coding units.
[0401] Among them, when the target quantization parameter information includes quantization parameter values of some coding units among multiple coding units in the reconstructed image, and the quantization parameter values of some coding units do not completely include the quantization parameter values of coding units in a preset area of the reconstructed image, the quantization parameter values can be corrected according to the quantization parameter values of some coding units among the multiple coding units or the reference quantization parameter map to obtain the quantization parameter values of all coding units among the multiple coding units; then, the quantization parameter map of the reconstructed image can be constructed according to the quantization parameter values of all coding units; or, the quantization parameter values of coding units in the preset area of the reconstructed image can be selected from all coding units, and the quantization parameter map of the reconstructed image can be constructed according to the quantization parameter values of the coding units in the preset area of the reconstructed image.
[0402] In an embodiment of the present application, when the target quantization parameter information includes the quantization parameter values of all coding units among a plurality of coding units, a quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of all coding units, and the quantization parameter map of the reconstructed image obtained in this case can be used to characterize the fidelity of the entire reconstructed image; when the target quantization parameter information includes the quantization parameter values of some coding units among a plurality of coding units, the quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some coding units, and the quantization parameter map of the reconstructed image obtained in this case can be used to characterize the fidelity of a preset area of the reconstructed image; when the target quantization parameter information includes the quantization parameter values of some coding units among a plurality of coding units, the quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some coding units and the reference quantization parameter map A quantization parameter map of a reconstructed image. Since the reference quantization parameter map is a quantization parameter map of a reference image of the reconstructed image, the quantization parameter value of any one of the multiple coding units except the part of the coding units can be obtained according to the reference quantization parameter map. In this way, the quantization parameter of any one of the multiple coding units can be obtained. The quantization parameter map of the reconstructed image obtained in this case can be used to characterize the fidelity of the entire reconstructed image or to characterize the fidelity of a preset area of the reconstructed image. Therefore, in any case in which the target quantization parameter information obtained by decoding includes the quantization parameter values of all or part of the coding units in the multiple coding units, the embodiment of the present application can obtain a quantization parameter map of the reconstructed image for characterizing the fidelity of the reconstructed image or for characterizing the fidelity of the preset area of the reconstructed image.
[0403] In a possible design, obtaining the quantization parameter value of the target coding unit according to the quantization parameter values of the partial coding units includes: determining the quantization parameter value of the target coding unit according to the quantization parameter value of at least one coding unit in the partial coding units.
[0404] When the decoding end cannot obtain the quantization parameter value of a coding unit from the first bitstream, the quantization parameter value of the spatial neighborhood of the coding unit is used to fill the coding unit. Specifically, assuming that the decoding end cannot obtain the quantization parameter value of the target coding unit from the first bitstream, the quantization parameter value of at least one coding unit in the coding units of the quantization parameter values obtained by decoding can be used to determine the quantization parameter value of the target coding unit, including: taking any one of the quantization parameter values obtained by decoding as the quantization parameter value of the target coding unit, and calculating the average value of several quantization parameter values obtained by decoding, and taking the average value as the quantization parameter value of the target coding unit. For example, the quantization parameter values of the coding units at the left side, above, upper left, etc. of the target coding unit in the original image can be used as the quantization parameter value of the target coding unit.
[0405] In this way, the quantization parameter values of the coding units whose quantization parameter values cannot be obtained from the first code stream can be filled according to the quantization parameter values of some coding units among the multiple coding units, so as to obtain the quantization parameter values of all the coding units among the multiple coding units; and then, the quantization parameter map of the reconstructed image can be constructed according to the quantization parameter values of all the coding units; or, the quantization parameter values of the coding units in the preset area of the reconstructed image are selected from all the coding units, and the quantization parameter map of the reconstructed image is constructed according to the quantization parameter values of the coding units in the preset area of the reconstructed image.
[0406] In an embodiment of the present application, when the decoding end cannot obtain the quantization parameter value of a certain coding unit from the first bitstream, the quantization parameter value of the spatial neighborhood of the coding unit can be used for filling. Specifically, when the target quantization parameter information includes the quantization parameter values of some coding units among multiple coding units, the quantization parameter value of any one coding unit other than the part of the coding units can also be determined according to the quantization parameter value of at least one coding unit among the part of the coding units, so that the quantization parameter value of any one coding unit among the multiple coding units can be ensured; and the quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some or all coding units among the multiple coding units; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of all coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of the entire reconstructed image; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of some coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of the preset area of the reconstructed image.
[0407] In a possible design, the reference quantization parameter map includes multiple reference elements, and the value of any reference element among the multiple reference elements is the quantization parameter value of the coding unit in the reference image; the quantization parameter value of the target coding unit is obtained according to the reference quantization parameter map, including: taking the value of the target element as the quantization parameter value of any target coding unit among the target coding units, wherein the target element is a reference element in the reference quantization parameter map, and the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image, or the position of the target element in the reference quantization parameter map is determined according to the position of any target coding unit in the reconstructed image and the motion vector of any target coding unit. The reference element is another name for the second element.
[0408] When the decoder cannot obtain the quantization parameter value of a certain coding unit from the first bitstream, the quantization parameter value of the temporal neighborhood of the coding unit can be used for filling. Assuming that the decoder cannot obtain the quantization parameter value of the target coding unit from the first bitstream, the filling methods include:
[0409] Method 1: Use the value of the target element in the reference quantization parameter map of the reconstructed image as the quantization parameter value of the target coding unit, and the position of the target element in the reference quantization parameter map is determined according to the position of the target coding unit in the reconstructed image.
[0410] Specifically, a reference coding unit is determined based on the target coding unit on a reference image of the reconstructed image, wherein a position of the reference coding unit in the reference image is the same as a position of the target coding unit in the reconstructed image; and a value of a target element in a reference quantization parameter map is used as a quantization parameter value of the target coding unit, wherein the target element is a second element in the reference quantization parameter map corresponding to the reference coding unit.
[0411] Method 2: Use the value of the target element in the reference quantization parameter map of the reconstructed image as the quantization parameter value of the target coding unit. The position of the target element in the reference quantization parameter map is determined according to the position of the target coding unit in the original image and the motion vector of the target coding unit.
[0412] Specifically, the coordinates of the target coding unit on the reconstructed image are (x, y), and the coordinates of the target coding unit on the reconstructed image (x, y) and the motion vector (mvx, mvy) of the target coding unit are used for offset to obtain (x+mvx, y+mvy), and the reference coding unit is determined at the (x+mvx, y+mvy) position of the reference image of the reconstructed image; the value of the target element in the reference quantization parameter map is used as the quantization parameter value of the target coding unit, wherein the target element is the second element in the reference quantization parameter map corresponding to the reference coding unit.
[0413] It should be noted that in the various filling methods mentioned above, the representation method of the reference quantization parameter map and the representation method of the quantization parameter map of the reconstructed image are the same, that is, it can be assumed that the quantization parameter map of the reconstructed image and the reference quantization parameter map have the same size as the corresponding original image, or are scaled proportionally.
[0414] Furthermore, a plurality of quantization parameter values may be obtained by using the aforementioned various methods, and an arithmetic average operation may be performed on the plurality of quantization parameter values to determine the quantization parameter value of the target coding unit.
[0415] In this way, the quantization parameter values of the coding units whose quantization parameter values cannot be obtained from the first code stream can be filled according to the reference quantization parameter map of the reconstructed image, so as to obtain the quantization parameter values of all the coding units in the multiple coding units; and then, the quantization parameter map of the reconstructed image can be constructed according to the quantization parameter values of all the coding units; or, the quantization parameter values of the coding units in the preset area of the reconstructed image are selected from all the coding units, and the quantization parameter map of the reconstructed image is constructed according to the quantization parameter values of the coding units in the preset area of the reconstructed image.
[0416] In an embodiment of the present application, when the quantization parameter value of a certain coding unit cannot be obtained from the first bitstream, the quantization parameter of the temporal neighborhood of the coding unit can be used for filling. Specifically, for any coding unit whose quantization parameter value cannot be obtained from the first bitstream; the value of the target element in the reference quantization parameter map can be used as the quantization parameter value of the coding unit, and the position of the target element in the reference quantization parameter map is determined according to the position of the coding unit in the reconstructed image, or the position of the target element in the reference quantization parameter map is determined according to the position of the coding unit in the reconstructed image and the motion vector of the coding unit. In this way, the quantization parameter value of any coding unit among the multiple coding units can be ensured; and the quantization parameter map of the reconstructed image can be obtained according to the quantization parameter values of some or all coding units among the multiple coding units; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of all coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of the entire reconstructed image; when the quantization parameter map of the reconstructed image is obtained according to the quantization parameter values of some coding units among the multiple coding units, the obtained quantization parameter map of the reconstructed image can also be used to characterize the fidelity of the preset area of the reconstructed image.
[0417] In one possible design, the method further includes: storing the reconstructed image and a quantization parameter map of the reconstructed image in association with each other, so as to use the reconstructed image as a reference image and the quantization parameter map of the reconstructed image as a reference quantization parameter map.
[0418] After a quantization parameter map is constructed and modified, it will be saved and used as an input for constructing a quantization parameter map of a subsequent coded image; for example, a quantization parameter map buffer is designed to store the quantization parameter map.
[0419] The decoder of any mainstream video codec solution includes a decoded image buffer to store decoded images, and a reference image management mechanism to manage the addition and removal of decoded images. In the present application, each quantization parameter map corresponds to a decoded image and records the quantization parameter information of the decoded image. Therefore, its quantization parameter map can be managed in exactly the same way as managing a decoded image. In other words, a decoded image and its quantization parameter map are managed in a decoded image buffer and a reference quantization parameter map buffer, respectively, and the management operations of the two are completely synchronized.
[0420] In an embodiment of the present application, the quantization parameter map of the reconstructed image can be stored and used as a reference quantization parameter map for constructing a quantization parameter map of a subsequent decoded image, thereby facilitating the construction of a quantization parameter map of a subsequent decoded image.
[0421] In one possible design, the method also includes: processing the reconstructed image or a preset area of the reconstructed image according to a quantization parameter map of the reconstructed image to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determining whether to apply the reconstructed image according to the quantization parameter map of the reconstructed image.
[0422] Specifically, the decoding end determines the signal distortion of a reconstructed image or a preset area in the reconstructed image based on the reconstruction map of the fidelity map, and applies it to different business environments. For example, in a video surveillance scenario, the degree of distortion of a certain image area in the reconstructed image is determined based on the reconstruction map of the fidelity map. If it is greater than a preset threshold, the reconstructed image will not be applied. For another example, based on the degree of distortion of a certain image area in the reconstructed image, one of the image enhancement methods can be selected from a group of image enhancement methods according to preset rules and applied to the image area to improve the image quality.
[0423] Among them, when processing the reconstructed image or the preset area of the reconstructed image, in the learning-based post-processing enhancement algorithm, the image can be divided into B distortion ranges according to the degree of distortion and trained separately to obtain multiple image enhancement models, where B is an integer greater than 1. The decoding end can determine the degree of distortion of different areas of the reconstructed image based on the quantization parameter map of the reconstructed image, and select different models for image enhancement in different areas. Using a training model that better matches the distortion distribution can obtain a better image quality improvement effect. For another example, in the image enhancement algorithm of a single model, the quantization parameter map can be used as additional input information of the network during training and use. Under the guidance of the quantization parameter map, the network can output an image with a better enhancement effect.
[0424] In an embodiment of the present application, the decoding end can process the reconstructed image or a preset area of the reconstructed image according to the reconstruction map of the fidelity map to improve the image quality of the reconstructed image or the preset area of the reconstructed image; or determine whether to apply the reconstructed image according to the reconstruction map of the fidelity map; thereby facilitating the application of the reconstructed image.
[0425] Fig.16 A schematic block diagram of an encoding device provided in an embodiment of the present application; the encoding device includes a video encoder and a fidelity map encoder, wherein:
[0426] A video encoder, used for encoding the original image to obtain a first code stream;
[0427] A fidelity map encoder is used to encode a fidelity map to obtain a second code stream, wherein the fidelity map is used to represent the distortion between at least a partial area of the original image and at least a partial area of the reconstructed image, and the reconstructed image is obtained after decoding the first code stream.
[0428] in, Fig.16 The compressed code stream is a general term for the code stream transmitted from the encoding end to the decoding end, and the compressed code stream includes the first code stream and the second code stream.
[0429] In a possible design, the encoding device also includes a fidelity map calculator, which is used to: divide the original image into multiple first image blocks, and divide the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the original image is the same as the division strategy for dividing the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; or divide a preset area of the original image into multiple first image blocks, and divide the preset area of the reconstructed image into multiple second image blocks, wherein the division strategy for dividing the preset area of the original image is the same as the division strategy for dividing the preset area of the reconstructed image, and the multiple first image blocks correspond one-to-one to the multiple second image blocks; calculate the fidelity value of any second image block according to any second image block among the multiple second image blocks and the first image block corresponding to the any second image block, and the fidelity map includes the fidelity value of any second image block, and the fidelity value of any second image block is used to represent the distortion between any second image block and the first image block corresponding to the any second image block.
[0430] In one possible design, the fidelity map includes multiple first elements, the multiple second image blocks correspond one-to-one to the multiple first elements, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0431] In one possible design, the second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions: color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
[0432] In a possible design, the fidelity map encoder is specifically used to: perform entropy encoding on any of the first elements to obtain the second code stream, and the entropy encoding of the any of the first elements is independent of the entropy encoding of other first elements; or, determine the probability distribution of the value of any of the first elements or the predicted value of any of the first elements according to the value of at least one first element among the encoded first elements, and perform entropy encoding on any of t...
Claims
1. A coding method, characterized in that: include: Encoding the original image to obtain a first code stream; Encoding a fidelity map to obtain a second code stream, wherein the fidelity map is used to represent distortion between at least a portion of the original image and at least a portion of the reconstructed image, wherein the reconstructed image is obtained by decoding the first code stream; Among them, if the encoding is lossless encoding, the reconstructed image of the fidelity map is the same as the fidelity map, and the reconstructed image of the fidelity map is obtained by decoding the second bitstream; if the encoding is lossy encoding, the reconstructed image of the fidelity map includes the coding distortion generated by encoding the fidelity map.
2. The method according to claim 1, characterized in that The method further comprises: Dividing the original image into a plurality of first image blocks, and dividing the reconstructed image into a plurality of second image blocks, wherein a division strategy for dividing the original image is the same as a division strategy for dividing the reconstructed image, and the plurality of first image blocks correspond one to one to the plurality of second image blocks; or dividing the preset area of the original image into a plurality of first image blocks, and dividing the preset area of the reconstructed image into a plurality of second image blocks, wherein a division strategy for dividing the preset area of the original image is the same as a division strategy for dividing the preset area of the reconstructed image, and the plurality of first image blocks correspond to the plurality of second image blocks in a one-to-one manner; The fidelity value of any second image block is calculated based on any second image block among the multiple second image blocks and a first image block corresponding to the any second image block, and the fidelity map includes the fidelity value of any second image block, and the fidelity value of any second image block is used to represent the distortion between the any second image block and the first image block corresponding to the any second image block.
3. The method according to claim 2, characterized in that The fidelity map includes multiple first elements, the multiple second image blocks correspond to the multiple first elements one-to-one, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
4. The method according to claim 2, characterized in that: The second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions of color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
5. The method according to claim 3 or 4, characterized in that: The step of encoding the fidelity map to obtain a second bitstream includes: performing entropy coding on any of the first elements to obtain the second bitstream, wherein the entropy coding on any of the first elements is independent of the entropy coding on other first elements; or, determining a probability distribution of a value of any first element or a predicted value of any first element according to a value of at least one first element among the encoded first elements, and performing entropy encoding on any first element according to the probability distribution of the value of any first element or the predicted value of any first element to obtain the second bitstream; The second code stream includes code streams of the multiple first elements.
6. The method according to claim 3 or 4, characterized in that: The step of encoding the fidelity map to obtain a second bitstream includes: quantizing any of the first elements to obtain a quantized first element; Encoding the quantized first element to obtain the second code stream; The second code stream includes code streams of the multiple first elements.
7. A decoding method, characterized in that: include: Decoding the first code stream to obtain a reconstructed image of the original image; Decoding a second bitstream to obtain a reconstructed image of a fidelity map, wherein the second bitstream is obtained by encoding the fidelity map, and the reconstructed image of the fidelity map is used to represent distortion between at least a partial area of the original image and at least a partial area of the reconstructed image; If the encoding is lossless encoding, the reconstructed image of the fidelity map is the same as the fidelity map; if the encoding is lossy encoding, the reconstructed image of the fidelity map includes encoding distortion generated by encoding the fidelity map.
8. The method according to claim 7, characterized in that The fidelity map includes a fidelity value of any second image block among a plurality of second image blocks, and the fidelity value of any second image block is used to represent the distortion between the any second image block and an original image block corresponding to the any second image block.
9. The method according to claim 8, characterized in that The fidelity map includes multiple first elements, the multiple second image blocks correspond to the multiple first elements one-to-one, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
10. The method according to claim 8, characterized in that The second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions of color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
11. The method according to claim 9 or 10, characterized in that: The decoding of the second bitstream to obtain a reconstructed image of the fidelity image includes: Decoding the second code stream to obtain a reconstruction fidelity value of any one of the first elements; A reconstruction map of the fidelity map is obtained according to the reconstruction fidelity value of any one of the first elements.
12. The method according to claim 9 or 10, characterized in that: The second bitstream is obtained by encoding the quantized first element; and the decoding of the second bitstream to obtain a reconstructed image of the fidelity image includes: Decoding the second bitstream to obtain a reconstruction fidelity value of the quantized first element; Dequantizing the reconstruction fidelity value of the quantized first element to obtain a reconstruction fidelity value of any first element; A reconstruction map of the fidelity map is obtained according to the reconstruction fidelity value of any one of the first elements.
13. A coding device, characterized in that: include: A video encoder, used for encoding the original image to obtain a first code stream; a fidelity map encoder, configured to encode a fidelity map to obtain a second code stream, wherein the fidelity map is used to represent distortion between at least a portion of an area of the original image and at least a portion of an area of a reconstructed image, wherein the reconstructed image is obtained by decoding the first code stream; Among them, if the encoding is lossless encoding, the reconstructed image of the fidelity map is the same as the fidelity map, and the reconstructed image of the fidelity map is obtained by decoding the second bitstream; if the encoding is lossy encoding, the reconstructed image of the fidelity map includes the coding distortion generated by encoding the fidelity map.
14. The encoding device according to claim 13, characterized in that The encoding device further comprises a fidelity map calculator, the fidelity map calculator being configured to: Dividing the original image into a plurality of first image blocks, and dividing the reconstructed image into a plurality of second image blocks, wherein a division strategy for dividing the original image is the same as a division strategy for dividing the reconstructed image, and the plurality of first image blocks correspond one to one to the plurality of second image blocks; or dividing the preset area of the original image into a plurality of first image blocks, and dividing the preset area of the reconstructed image into a plurality of second image blocks, wherein a division strategy for dividing the preset area of the original image is the same as a division strategy for dividing the preset area of the reconstructed image, and the plurality of first image blocks correspond to the plurality of second image blocks in a one-to-one manner; The fidelity value of any second image block is calculated based on any second image block among the multiple second image blocks and a first image block corresponding to the any second image block, and the fidelity map includes the fidelity value of any second image block, and the fidelity value of any second image block is used to represent the distortion between the any second image block and the first image block corresponding to the any second image block.
15. The encoding device according to claim 14, characterized in that The fidelity map includes multiple first elements, the multiple second image blocks correspond to the multiple first elements one-to-one, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
16. The encoding device according to claim 14, characterized in that The second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions of color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
17. The encoding device according to claim 15 or 16, characterized in that The fidelity map encoder is specifically used for: performing entropy coding on any of the first elements to obtain the second bitstream, wherein the entropy coding on any of the first elements is independent of the entropy coding on other first elements; or, determining a probability distribution of a value of any first element or a predicted value of any first element according to a value of at least one first element among the encoded first elements, and performing entropy encoding on any first element according to the probability distribution of the value of any first element or the predicted value of any first element to obtain the second bitstream; The second code stream includes code streams of the multiple first elements.
18. The encoding device according to claim 15 or 16, characterized in that The fidelity map encoder is specifically used for: quantizing any of the first elements to obtain a quantized first element; Encoding the quantized first element to obtain the second code stream; The second code stream includes code streams of the multiple first elements.
19. A decoding device, characterized in that: include: A video decoder, configured to decode the first bitstream to obtain a reconstructed image of the original image; a fidelity map decoder, configured to decode a second bitstream to obtain a reconstructed map of a fidelity map, wherein the second bitstream is obtained by encoding the fidelity map, and the reconstructed map of the fidelity map is used to represent distortion between at least a partial area of the original image and at least a partial area of the reconstructed image; Among them, if the encoding is lossless encoding, the reconstructed image of the fidelity map is the same as the fidelity map, and the reconstructed image of the fidelity map is obtained by decoding the second bitstream; if the encoding is lossy encoding, the reconstructed image of the fidelity map includes the coding distortion generated by encoding the fidelity map.
20. The decoding device according to claim 19, characterized in that The fidelity map includes a fidelity value of any second image block among a plurality of second image blocks, and the fidelity value of any second image block is used to represent the distortion between the any second image block and an original image block corresponding to the any second image block.
21. The decoding device according to claim 20, characterized in that The fidelity map includes multiple first elements, the multiple second image blocks correspond to the multiple first elements one-to-one, the value of any first element among the multiple first elements is the fidelity value of the second image block corresponding to the any first element, the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of any first element in the fidelity map is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
22. The decoding device according to claim 20, characterized in that The second image block includes three color components, and the fidelity map is a three-dimensional array including three dimensions of color component, width and height. The two-dimensional array under any color component A in the fidelity map includes multiple first elements, and the value of any first element among the multiple first elements is the fidelity value of the color component A of the second image block corresponding to the any first element. The position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in the reconstructed image, or the position of the two-dimensional array under any color component A in the fidelity map of any first element is determined according to the position of the second image block corresponding to the any first element in a preset area of the reconstructed image.
23. The decoding device according to claim 21 or 22, characterized in that The fidelity map decoder is specifically used for: Decoding the second code stream to obtain a reconstruction fidelity value of any one of the first elements; A reconstruction map of the fidelity map is obtained according to the reconstruction fidelity value of any one of the first elements.
24. The decoding device according to claim 21 or 22, characterized in that The second code stream is obtained by encoding the quantized first element; the fidelity map decoder is specifically used to: Decoding the second bitstream to obtain a reconstruction fidelity value of the quantized first element; Dequantizing the reconstruction fidelity value of the quantized first element to obtain a reconstruction fidelity value of any first element; A reconstruction map of the fidelity map is obtained according to the reconstruction fidelity value of any one of the first elements.
25. An encoder (20), characterized in that The method comprises a processing circuit for executing the method of any one of claims 1 to 6.
26. A decoder (30), characterized in that Comprising a processing circuit for executing the method of any one of claims 7-12.
27. An encoder, characterized in that include: one or more processors; A non-transitory computer-readable storage medium, coupled to the processor and storing a program executed by the processor, wherein the program, when executed by the processor, causes the encoder to perform the method of any one of claims 1-6.
28. A decoder, characterized in that: include: one or more processors; A non-transitory computer-readable storage medium, coupled to the processor and storing a program executed by the processor, wherein the program, when executed by the processor, causes the decoder to perform the method of any one of claims 7-12.
29. A non-transitory computer-readable storage medium, characterized in that: The method comprises a program code which, when executed by a computer device, is used to perform the method according to any one of claims 1 to 6 or 7 to 12.
30. A non-transitory storage medium, characterized in that: A bit stream comprising a bit stream encoded according to the method according to any one of claims 1-6.
Citation Information
Patent Citations
Video coding method and device, video decoding method and device and electronic equipment
CN109120937A