Image decoding method, image encoding method, and corresponding device
By dynamically adjusting the quantization parameters according to the characteristics of each pixel point during image encoding and decoding, the quantization distortion problem in the prior art is solved, and higher image decoding authenticity and accuracy are achieved, while maintaining the constant of the compression ratio.
Patent Information
- Application Number
- JP2024543258
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-01-19
- Filing Date
- 2023-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-01-17
AI Technical Summary
In the existing video encoding technology, the same quantization parameters (QP) are used in the image encoding and decoding process to quantize all pixel points, resulting in large quantization distortion, affecting the authenticity and accuracy of the image.
During the image encoding and decoding process, the quantization parameters (QP) are dynamically adjusted according to the characteristics of each pixel point, that is, quantization and inverse quantization are performed at the pixel level, thereby maintaining a constant compression ratio, while reducing decoding distortion and improving the authenticity and accuracy of the image.
Through pixel-level quantization parameter adjustment, the decoding distortion of image frames can be effectively reduced, the authenticity and accuracy of image decoding can be improved, while maintaining the constant compression ratio.
Smart Images

Figure 0007678943000013 
Figure 0007678943000014 
Figure 0007678943000015
Abstract
Description
[Technical field]
[0001] FIELD OF THE DISCLOSURE Embodiments of the present invention relate to video codecs, and in particular to an image decoding method, an image encoding method, and a corresponding device. [Background technology]
[0002] In the field of video codecs, video compression (ie, video codec) techniques can be used to compress the amount of video data, thereby enabling efficient transmission or storage of the video.
[0003] To encode and decode a video is to encode and decode the image of each frame of a video. Taking an image of one frame as an example, in the encoding device, an image encoder encodes the image, obtains a bitstream corresponding to the image, and transmits it to a decoding device, and in the decoding device, an image decoder analyzes the bitstream to obtain an image. Currently, an image is divided into one or more coding units (CUs), and an image encoder predicts each CU, determines a residual value between the predicted value of the CU and the true value of the CU, and sequentially transforms, quantizes, and encodes the residual value to obtain a bitstream. Accordingly, an image decoder predicts each CU, sequentially dequantizes and inversely transforms the decoding result of the bitstream corresponding to the CU to obtain a residual value corresponding to the CU, and calculates the sum of the predicted value and the residual value of the CU to obtain a reconstruction value of the CU.
[0004] In the process of image encoding and decoding, quantization can realize many-to-one mapping of signal values, effectively reducing the space of signal values and achieving better compression effect. As can be understood, encoding devices and decoding devices perform quantization and dequantization processes based on a quantization parameter (QP). Currently, one QP is set for one CU, and the encoding device obtains the QP of each CU and quantizes the residual values and transform coefficients of the CU based on the QP, and correspondingly, the decoding device obtains the QP of the CU and dequantizes the quantized coefficients obtained by analyzing the bitstream based on the QP. However, if the same QP is used to quantize all pixel points in one CU, that is, if the same degree of quantization is performed on all pixel points in the CU, the quantization distortion (image distortion due to quantization) in the image encoding and decoding process will be large. Summary of the Invention
[0005] The embodiment of the present invention provides an image decoding method, an encoding method and a video encoding device, which can reduce the decoding distortion of an image frame while ensuring a certain compression rate, and improve the authenticity and accuracy of image decoding. To achieve the above objective, the embodiment of the present invention adopts the following technical solutions:
[0006] According to a first aspect, an embodiment of the present invention provides an image decoding method performed by a decoding device, the method comprising: For any one pixel point or any plurality of parallel inverse quantization pixel points in a current coding block, determining a quantization parameter QP value of the pixel point, wherein the QP values of at least two pixel points among the pixel points in the current coding block are different; and dequantizing the pixel point based on a QP value of the pixel point.
[0007] According to the decoding method provided in the embodiment of the present invention, a video decoder can determine the QP value of each pixel point for a coding block, and then inverse quantize each pixel point according to the QP value of each pixel point, that is, perform inverse quantization on a pixel-by-pixel basis, thereby reducing the decoding distortion of an image frame while ensuring a certain compression ratio, and improving the authenticity and accuracy of image decoding.
[0008] In one possible implementation, a method for decoding an image according to an embodiment of the present invention comprises the steps of: obtaining a QP value of the current coding block; determining that the predicted QP value of the pixel point is a QP value of the current coding block.
[0009] In one possible implementation, determining the QP value of the pixel point comprises: If the pixel point is any one of the target pixel points or any multiple parallel dequantized target pixel points in the current coding block, adjust a predicted QP value of the pixel point, and set the adjusted predicted QP value as the QP value of the pixel point. Tosu This includes:
[0010] In one possible implementation, adjusting the predicted QP value of the pixel point comprises: Obtaining information of reconstructed pixel points around the pixel point; and adjusting a predicted QP value of the pixel point based on information of reconstructed pixel points surrounding the pixel point.
[0011] In one possible implementation, adjusting the predicted QP value of the pixel point comprises: If the pixel point satisfies a first preset condition, set the preset QP value as the QP value of the pixel point; otherwise, setting the predicted QP value of the pixel point as the QP value of the pixel point; The first preset condition is: the pixel point is a luminance pixel point; the pixel point is a chromaticity pixel point; the bit depth of the pixel point is less than or equal to a bit depth threshold; the predicted QP value of the pixel point is equal to or less than an adjustable QP maximum value, and the adjustable QP maximum value is equal to or less than a QP maximum value; and information of reconstructed pixel points surrounding the pixel point is less than or equal to a first preset threshold.
[0012] In one possible implementation, the step of dequantizing the pixel point based on the QP value of the pixel point comprises: The method includes the step of inverse quantizing the pixel point based on the adjusted predicted QP value.
[0013] In one possible implementation, the target pixel point is any one or more pixel points in the current coding block.
[0014] In one possible implementation, the current coding block includes at least pixel points of a first portion and / or pixel points of a second portion, and the target pixel points are any one or more of the pixel points of the second portion.
[0015] In one possible implementation, the target pixel point is any one or more of the pixel points at the first location in the second portion of pixel points.
[0016] In one possible implementation, the target pixel point is any one or more of the pixel points at the second location in the second portion of pixel points.
[0017] In one possible implementation, the prediction mode of the current coding block is a pixel-wise prediction mode; the current coding block includes at least pixel points of a first portion and / or pixel points of a second portion; the pixel points of the second portion include pixel points at the first location and / or pixel points at the second location; The pixel point at the first location and the pixel point at the second location are determined based on a pixel-by-pixel prediction mode of the current coding block.
[0018] In one possible implementation, adjusting the predicted QP value of the pixel point based on information of reconstructed pixel points around the pixel point comprises: If the pixel point satisfies a second preset condition and a third preset condition, adjust a predicted QP value of the pixel point according to a first QP offset amount and a distortion reference QP value; If the pixel point does not satisfy the second preset condition, or if the pixel point satisfies the second preset condition but does not satisfy the third preset condition and the fourth preset condition, A predicted QP value of the pixel point is set as a QP value of the pixel point; The distortion reference QP value represents a QP value corresponding to a perceptible distortion, the second preset condition is that the predicted QP value of the pixel point is greater than the distortion reference QP value and is less than or equal to an adjustable QP maximum value, the third preset condition is that information of reconstructed pixel points around the pixel point is less than or equal to a first threshold, and the fourth preset condition is that information of reconstructed pixel points around the pixel point is greater than a second threshold and the first threshold is less than or equal to the second threshold.
[0019] In one possible implementation, if the pixel point satisfies the second preset condition and the third preset condition, the adjusted QP value satisfies: finalQP=max(initQP‐offset1,jndQP) Here, finalQP represents the predicted QP value after the adjustment, initQP represents the predicted QP value of the pixel point, offset1 represents the first QP offset amount, jndQP represents the distortion reference QP value, and max represents the maximum value.
[0020] In one possible implementation, the reconstructed pixel points around said pixel point are: A pixel point within a square region having the pixel point as a center and a side length of a first preset value, or The pixel points are included within a diamond-shaped region having the pixel point as a center and a diagonal length of a second preset value.
[0021] In one possible implementation, the information of the reconstructed pixel point includes at least one of the following information of the reconstructed pixel point: pixel value, reconstructed residual value, gradient value, flatness information or texture information or complexity information, background luminance, contrast or motion amount, where the reconstructed residual value includes a residual value after inverse quantization or a difference between a reconstructed value and a predicted value, and the gradient value includes a horizontal gradient, a vertical gradient or an average gradient.
[0022] In one possible implementation, the value of the information of the reconstructed pixel point comprises at least one of an original value, an absolute value, an average value, or a difference.
[0023] In one possible implementation, obtaining information of reconstructed pixel points around the pixel point comprises: obtaining information of a predicted pixel point of the pixel point; If the predicted pixel point is not a reconstructed pixel point in the current coding block, The method includes determining a difference between information of the predicted pixel point and information of reconstructed pixel points around the predicted pixel point, or an absolute value of the difference, as information of the reconstructed pixel points around the pixel point.
[0024] In one possible implementation, the prediction mode of the current coding block is a block prediction mode, and the image decoding method according to an embodiment of the present invention further comprises: obtaining region division information of the current coding block, the region division information including a number N of regions and position information of region boundary lines, where N is an integer equal to or greater than 2; Dividing the current coding block into N regions based on the region division information.
[0025] In one possible implementation, the step of obtaining region division information of the current coding block comprises: obtaining predefined region partition information of the current coding block; or The method includes parsing a bitstream to obtain region partition information of the current coding block.
[0026] According to a second aspect, an embodiment of the present invention provides an image coding method performed by a coding device, the method comprising: For any one pixel point or any plurality of parallel quantization pixel points in a current coding block, determining a quantization parameter QP value for the pixel point, wherein the QP values of at least two pixel points in the current coding block are different; quantizing the pixel point based on a QP value of the pixel point.
[0027] According to the encoding method provided in the embodiment of the present invention, a video encoder can determine a QP value of each pixel point for a coding block, and then quantize each pixel point according to the QP value of each pixel point, that is, perform pixel-by-pixel quantization, thereby reducing the decoding distortion of an image frame while ensuring a certain compression ratio, and improving the authenticity and accuracy of image decoding.
[0028] For various possible implementations of the image coding method, please refer to the description of various possible implementations of the image decoding method.
[0029] According to a third aspect, an embodiment of the present invention provides an image decoding device applied to a decoding device, the decoding device comprising respective modules for implementing the method according to the first aspect and one of its possible realization manners, such as a quantization parameter QP determination unit and an inverse quantization unit.
[0030] For technical solutions and beneficial effects of the third aspect, please refer to the description of the first aspect and any one of its possible implementations. The decoding device has a function of implementing the actions in the method examples of the first aspect and any one of the possible implementations of the first aspect. The function may be implemented by hardware, or may be implemented by executing corresponding software by hardware. The hardware or software includes one or more modules corresponding to the above function.
[0031] According to a fourth aspect, an embodiment of the present invention provides an image encoding device applied in an encoding device, the encoding device comprising respective modules for implementing the method according to the second aspect and one of its possible realization manners, such as a quantization parameter QP determination unit and a quantization unit.
[0032] For technical solutions and beneficial effects of the fourth aspect, please refer to the description of the second aspect and any one of its possible implementations. The encoding device has a function of implementing the actions in the method examples of the second aspect and any one of the possible implementations of the second aspect. The function may be implemented by hardware, or may be implemented by executing corresponding software by hardware. The hardware or software includes one or more modules corresponding to the above function.
[0033] According to a fifth aspect, an embodiment of the present invention provides an electronic device including a processor and a memory, the memory being configured to store computer instructions, and the processor being configured to retrieve and execute the computer instructions from the memory to perform the method according to any one of the first aspect, the second aspect, and possible implementations thereof. For example, the electronic device may be a video encoder or an encoding device including a video encoder. As another example, the electronic device may be a video decoder or a decoding device including a video decoder.
[0034] According to a sixth aspect, an embodiment of the present invention provides a computer readable storage medium having stored thereon a computer program or instructions which, when executed by a computing device or a storage system in which the computing device is arranged, performs a method according to the first aspect, the second aspect and any one of their possible implementations.
[0035] According to a seventh aspect, an embodiment of the present invention provides a computer program product comprising instructions which, when executed on a computing device or processor, cause the computing device or processor to execute the instructions to perform a method according to the first aspect, the second aspect and any one of its possible implementations.
[0036] According to an eighth aspect, an embodiment of the present invention provides a chip comprising a memory and a processor, the memory configured to store computer instructions, and the processor configured to retrieve and execute the computer instructions from the memory to perform a method according to the first aspect, the second aspect and any one of its possible implementations.
[0037] According to a ninth aspect, an embodiment of the present invention provides a video codec system comprising an encoding device and a decoding device, the decoding device configured to perform a method according to the first aspect and any one of its possible implementation manners, and the encoding device configured to perform a method according to the second aspect and any one of its possible implementation manners.
[0038] According to a tenth aspect, an embodiment of the present invention provides an image decoding method performed by a decoding device, the method comprising: For any one pixel point or any plurality of parallel inverse quantization pixel points in a current coding block, determining a quantization parameter QP value of the pixel point, wherein the QP values of at least two pixel points among the pixel points in the current coding block are different; determining a quantization step Qstep of the pixel point based on the QP value of the pixel point; For the selected quantizer combination, inverse quantizing the level value of the pixel point using the Qstep of the pixel point.
[0039] In one possible implementation, the combination of quantizers includes one or more quantizers, which may be uniform or non-uniform quantizers, and the level values of the pixel points are obtained by analyzing a bitstream.
[0040] In one possible implementation, the step of determining a quantization step Qstep for the pixel point based on the QP value of the pixel point comprises: Determining Qstep based on the QP value of the pixel point by at least one of equation derivation or lookup.
[0041] In one possible implementation, determining Qstep by at least one of formula derivation or lookup includes:
number
number
number
[0042] In one possible implementation, the inverse quantization formula for the uniform quantizer is:
number
number
[0043] In one possible implementation, f may be 0.5 or some other fixed value, or f may be adaptively determined based on the QP value, the prediction mode, and whether or not a transform is performed.
[0044] According to an eleventh aspect, an embodiment of the present invention provides an image encoding method performed by an encoding device, the method comprising: For any one pixel point or any plurality of parallel quantization pixel points in a current coding block, determining a quantization parameter QP value for the pixel point, wherein the QP values of at least two pixel points in the current coding block are different; determining a quantization step Qstep of the pixel point based on the QP value of the pixel point; quantizing the pixel point using the Qstep of the pixel point for the selected quantizer combination.
[0045] According to a twelfth aspect, an embodiment of the present invention provides an image decoding device, the decoding device comprising: A quantization parameter QP determination unit is used for determining a quantization parameter QP value of any one pixel point or any multiple parallel inverse quantization pixel points in a current coding block, and determining a quantization step Qstep of the pixel point according to the QP value of the pixel point; a dequantization unit, which is used to dequantize the level value of the pixel point using the Qstep of the pixel point for the selected quantizer combination; The QP values of at least two of the pixel points in the current coding block are different.
[0046] According to a thirteenth aspect, an embodiment of the present invention provides an image encoding apparatus, the encoding apparatus comprising: A quantization parameter QP determination unit is used for determining a quantization parameter QP value of any one pixel point or any multiple parallel quantization pixel points in a current coding block, and determining a quantization step Qstep of the pixel point according to the QP value of the pixel point; a quantization unit used to quantize the pixel point using a Qstep for the pixel point for the selected quantizer combination; The QP values of at least two of the pixel points in the current coding block are different.
[0047] According to a fourteenth aspect, an embodiment of the present invention provides a video codec system comprising an encoding device and a decoding device, the decoding device configured to perform the method according to the tenth aspect and any one of its possible implementation manners, and the encoding device configured to perform the method according to the eleventh aspect.
[0048] According to a fifteenth aspect, an embodiment of the present invention provides an electronic device including a processor and a memory, the memory configured to store computer instructions, and the processor configured to retrieve and execute the computer instructions from the memory to perform a method according to any one of the tenth aspect, the eleventh aspect and possible implementation manners thereof.
[0049] According to a sixteenth aspect, an embodiment of the present invention provides a computer readable storage medium having stored thereon a computer program or instructions which, when executed by a computing device or a storage system in which the computing device is arranged, performs a method according to the tenth aspect, the eleventh aspect and any one of their possible implementations.
[0050] According to a seventeenth aspect, an embodiment of the present invention provides an image decoding method performed by a decoding device, the method comprising: If a prediction mode of the current coding block is an intra block copy (IBC) prediction mode, dividing the current coding block into a number of transform blocks (TB) and a number of prediction blocks (PB); and a step of referring to reconstructed pixel values in a reconstructed TB to the left of the current TB when performing motion compensation on a PB in each TB from the second TB of the current coding block.
[0051] In one possible implementation, the size of the current coding block is 16x2, the size of the TB is 8x2 and the size of the PB is 2x2.
[0052] In one possible implementation, the splitting of the current coding block into several transform blocks (TB) and several prediction blocks (PB) comprises: Obtaining region partition information of the current coding block; Dividing the current coding block into a number of transform blocks (TB) and a number of prediction blocks (PB) based on region division information of the current coding block.
[0053] In one possible implementation, obtaining region partition information of the current coding block comprises: Obtaining predefined region partition information of the current coding block; or When the method is performed by a decoding device, it includes parsing a bitstream to obtain region partition information for the current coding block.
[0054] According to an eighteenth aspect, an embodiment of the present invention provides an image encoding method performed by an encoding device, the method comprising: If a prediction mode of the current coding block is an intra block copy prediction mode, dividing the current coding block into several transform blocks (TB) and several prediction blocks (PB); and a step of referring to reconstructed pixel values in a reconstructed TB to the left of the current TB when performing motion compensation on a PB in each TB from the second TB of the current coding block.
[0055] In one possible implementation, the coding scheme of the block vector (BV) or block vector difference (BVD) corresponding to the current block by the coding device is: When horizontal motion estimation is performed but vertical motion estimation is not performed, this includes transmitting horizontal BV or BVD in the bitstream and not transmitting vertical BV or BVD.
[0056] In one possible implementation, the encoding of the block vector (BV) or block vector difference (BVD) corresponding to the current block by the encoding device comprises a fixed length encoding.
[0057] In one possible implementation, the prediction block of the current coding block is Obtaining a matching block based on a block vector (BV) or a block vector difference (BVD), and performing processing on the matching block to generate a final predicted block; Performing processing on the matching block includes performing predictive filtering and / or light compensation processing on the matching block.
[0058] According to a nineteenth aspect, an embodiment of the present invention provides an electronic device including a processor and a memory, the memory configured to store computer instructions, and the processor configured to retrieve and execute the computer instructions from the memory to perform a method according to any one of the seventeenth aspect, the eighteenth aspect and possible implementation manners thereof.
[0059] According to a twentieth aspect, an embodiment of the present invention provides a computer readable storage medium having stored thereon a computer program or instructions which, when executed by a computing device or a storage system in which the computing device is arranged, performs a method according to the seventeenth aspect, the eighteenth aspect and any one of their possible implementations.
[0060] The embodiments of the present invention may be further combined to provide more implementation methods based on the implementation methods provided in each aspect above. [Brief description of the drawings]
[0061] [Figure 1] 1 is an exemplary block diagram of a video codec system according to one embodiment of the present invention; [Diagram 2] FIG. 2 is an exemplary block diagram of a video encoder according to one embodiment of the present invention; [Diagram 3] FIG. 2 is an illustrative block diagram of a video decoder according to one embodiment of the present invention; [Figure 4] 2 is a flowchart of video encoding / decoding according to an embodiment of the present invention. [Diagram 5] 2 is a flowchart of an image decoding method according to an embodiment of the present invention. [Figure 6] 2 is a flowchart of an image decoding method according to an embodiment of the present invention. [Figure 7] 2 is a flowchart of an image decoding method according to an embodiment of the present invention. [Figure 8A] FIG. 2 is a schematic diagram of a distribution of pixel points according to an embodiment of the present invention; [Figure 8B] FIG. 2 is a schematic diagram of a distribution of pixel points according to an embodiment of the present invention; [Figure 9A] FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 9B] FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 9C] FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 9D]FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 10A] FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 10B] FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 11A] FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 11B] FIG. 2 is a schematic diagram of pixel point division in pixel-wise prediction mode according to an embodiment of the present invention; [Figure 12] 2 is a flowchart of an image decoding method according to an embodiment of the present invention. [Figure 13A] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 13B] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 13C] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 13D] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 14A] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 14B] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 14C] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 15A] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 15B] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 15C] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 15D]FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 15E] FIG. 2 is a schematic diagram of a region division of a coding block according to an embodiment of the present invention; [Figure 16] 1 is a flowchart of an image encoding method according to an embodiment of the present invention. [Figure 17] FIG. 2 is a schematic diagram of residual group division according to an embodiment of the present invention; [Figure 18] FIG. 2 is a schematic diagram of TB / PB division according to one embodiment of the present invention. [Figure 19] FIG. 2 is a schematic structural diagram of a decoding device according to the present invention; [Figure 20] 1 is a schematic structural diagram of an encoding device according to the present invention; [Figure 21] 1 is a schematic structural diagram of an electronic device according to the present invention; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0062] The term "and / or" in this specification is merely a description of the relationship of related objects, and indicates that three types of relationships may exist. For example, A and / or B can indicate three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0063] The terms "first" and "second" in the description of the embodiments of the present invention and the claims are intended to distinguish different objects, and are not intended to describe a specific order of the objects. For example, the terms "first preset value" and "second preset value" are intended to distinguish different preset values, and are not intended to describe a specific order of the preset values.
[0064] In the embodiments of the present invention, terms such as "exemplary" or "for example" are used to denote an example, illustration, or explanation. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as preferred or advantageous over other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to specifically present the related concept.
[0065] In describing the embodiments of the present invention, unless otherwise specified, "plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units, and a plurality of systems refers to two or more systems.
[0066] The image decoding and encoding methods provided in the embodiments of the present invention can also be applied to video decoding and video encoding, although it should be understood that a video comprises a series of pictures, and decoding and encoding a video essentially involves decoding and encoding all of the pictures contained in the video.
[0067] As can be understood, quantization in the image coding process refers to mapping the continuous values (or a large number of discrete values) of a signal to a finite number of discrete values, and quantization can realize many-to-one mapping of signal values. In video coding, after the residual signal is transformed, the transform coefficients generally have a large dynamic range, so that the signal value space can be effectively reduced and better compression effect can be obtained by quantizing the transform coefficients. However, due to the many-to-one mapping mechanism, distortion is inevitably introduced in the quantization process, which is the fundamental cause of distortion in video coding.
[0068] Inverse quantization is the inverse process of quantization: mapping the quantized coefficients to a reconstructed signal in the input signal space, where the reconstructed signal is an approximation of the input signal.
[0069] Quantization includes scalar quantization (SQ) and vector quantization. Scalar quantization is the most basic quantization method, and the input of scalar quantization is a one-dimensional scalar signal. The scalar quantization process involves firstly dividing the input signal space into a series of non-intersecting intervals, selecting one representative signal from each interval, and then, for each input signal, scalar quantizing the input signal into the representative signal of the interval in which the input signal is located. Here, the interval length is called the quantization step (Qstep), the interval index is the level value (Level), i.e., the value after quantization, and the parameter representing the quantization step is the quantization parameter (QP).
[0070] The simplest scalar quantization method is uniform scalar quantization, which divides the input signal space into equally spaced intervals, and the representative signal for each interval is the interval midpoint.
[0071] The optimal scalar quantizer is the Lloyd-Max quantizer, which takes into account the distribution of the input signal, the interval division is uneven, the representative signal of each interval is the probability center of gravity of the interval, and the boundary point between two adjacent intervals is the midpoint of the representative signals of these two intervals.
[0072] The following describes the system architecture applied in the embodiment of the present application, and FIG. 1 is an exemplary block diagram of a video codec system according to the present invention. In this specification, the term "video encoder / decoder" generally refers to both a video encoder and a video decoder. In the present invention, the term "video codec" or "codec" generally refers to video encoding or video decoding. The video encoder 100 and the video decoder 200 in the video codec system 1 are configured to predict motion information, such as a motion vector of a currently encoded / decoded image block or its sub-block, according to various method examples described in any one of the multiple new inter-prediction modes provided in the present invention, so that the predicted motion vector is maximally close to the motion vector obtained using a motion estimation method, and there is no need to transmit the difference of the motion vector during encoding, which can further improve the encoding / decoding performance.
[0073] As shown in FIG. 1, video codec system 1 includes encoding device 10 and decoding device 20. Encoding device 10 generates encoded video data. Encoding device 10 may therefore be referred to as a video encoding device. Decoding device 20 may decode the encoded video data generated by encoding device 10. Decoding device 20 may therefore be referred to as a video decoding device. Various embodiments of encoding device 10, decoding device 20, or both may include one or more processors and memory coupled to the one or more processors. The memory may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or any other medium that may be used to store desired program code in the form of instructions or data structures accessible by a computer, as described herein.
[0074] Encoding device 10 and decoding device 20 may be a variety of devices, including a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video game console, an in-vehicle computer, or the like.
[0075] Decoding device 20 may receive the encoded video data from encoding device 10 via link 30. Link 30 may include one or more media or devices that may move encoded video data from encoding device 10 to decoding device 20. In one example, link 30 may include one or more communications media that enable encoding device 10 to transmit encoded video data directly to decoding device 20 in real time. In this example, encoding device 10 may modulate the encoded video data according to a communications standard (e.g., a wireless communication protocol) and transmit the modulated video data to decoding device 20. The one or more communications media may include wireless and / or wired communications media, such as a radio frequency (RF) spectrum or one or more physical transmission paths. The one or more communications media may form part of a packet-based network, such as a local area network (LAN), a wide area network (WAN), or a global network (e.g., the Internet). The one or more communications media may include routers, switches, base stations, or other devices that facilitate communication from encoding device 10 to decoding device 20.
[0076] In another example, the encoded data may be output from output interface 140 to storage device 40. Similarly, the encoded data may be accessed from storage device 40 via input interface 240. Storage device 40 may be a distributed or locally accessed data storage medium, such as a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0077] In another example, storage device 40 corresponds to a file server or another intermediate storage device that may hold the encoded video generated by encoding device 10. Decoding device 20 may access the stored video data from storage device 40 by streaming or downloading. The file server may be any type of server capable of storing and transmitting encoded video data to decoding device 20. Exemplary file servers include a network server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Decoding device 20 may access the encoded video data through any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a wireless-fidelity (Wi-Fi) connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination thereof suitable for accessing encoded video data stored in a file server. The transmission of the encoded video data from storage device 40 may be a streaming transmission, a download transmission, or a combination thereof.
[0078] The image decoding method provided by the present invention may be applied to video encoding and decoding to support various multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), encoding video data stored on a data storage medium, decoding video data stored on a data storage medium, or other applications. In some embodiments, the video codec system 1 may be configured to support unidirectional or bidirectional video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0079] The video codec system 1 shown in FIG. 1 is merely exemplary, and the techniques of the present invention are applicable to video codec settings (e.g., video encoding and video decoding) that do not necessarily include data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory and streamed over a network. A video encoding device may encode data and store it in memory, and / or a video decoding device may read data from memory and decode it. In many examples, encoding and decoding are performed by devices that do not communicate with each other, but rather encode data into memory and / or retrieve data from memory and decode it.
[0080] 1, encoding device 10 includes video source 120, video encoder 100, and output interface 140. In some examples, output interface 140 may include a modulator / demodulator (modem) and / or a transmitter. Video source 120 may include a video capture device (e.g., a camera), a video archive containing previously captured video data, a video feed-in interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources of video data.
[0081] Video encoder 100 may encode video data from video source 120. In some embodiments, encoding device 10 transmits the encoded video data directly to decoding device 20 via output interface 140. In other examples, the encoded video data may be stored on storage device 40 for access by decoding device 20 for decoding and / or playback.
[0082] In the example of Figure 1, decoding device 20 includes an input interface 240, a video decoder 200, and a display device 220. In some examples, input interface 240 includes a receiver and / or a modem. Input interface 240 can receive encoded video data via link 30 and / or from storage device 40. Display device 220 may be integrated with decoding device 20 or may be external to decoding device 20. In general, display device 220 displays the decoded video data. Display device 220 may include a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0083] 1, in some aspects, the video encoder 100 and the video decoder 200 may be integrated with an audio encoder and decoder, respectively, and may include appropriate multiplexer-demultiplexer units or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams. In some examples, where applicable, the demultiplexer (MUX-DEMUX) units may conform to the ITU H.223 multiplexer protocol or other protocols, such as the user datagram protocol (UDP).
[0084] Each of the video encoder 100 and the video decoder 200 may be implemented as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the present invention is implemented in part in software, the device may store instructions for the software in a suitable non-volatile computer-readable storage medium and execute the instructions in hardware using one or more processors to implement the techniques of the present invention. Any of the above (including hardware, software, a combination of hardware and software, etc.) may be considered as one or more processors. The video encoder 100 and the video decoder 200 may be included in one or more encoders or decoders, and any of the encoders or decoders may be integrated as part of a composite encoder / decoder (encoder / decoder) in the corresponding device.
[0085] In the present invention, the video encoder 100 may generally be referred to as a "signaling" or another device that "sends" some information, for example to the video decoder 200. The terms "signaling" or "sending" may generally refer to the transmission of syntax elements and / or other data for decoding the compressed video data. This transmission may occur in real time or near real time. Alternatively, this communication may occur after a period of time has elapsed, for example, when storing syntax elements in an encoded bitstream to a computer-readable storage medium at the time of encoding, and a decoding device may then retrieve the syntax elements any time after they are stored to the medium.
[0086] JCT-VC developed the H.265 (HEVC) standard. HEVC standardization is based on an evolution model of a video decoder called the HEVC Test Model (HEVC model, HM). The latest standard document of H.265 can be obtained from http: / / www.itu.int / rec / T-REC-H.265. The latest version of the standard document is H.265(12 / 16), which is incorporated herein by reference. HM assumes that the video decoder has some additional capabilities with respect to the existing algorithm of ITU-TH.264 / AVC. For example, H.264 provides nine types of intra-prediction coding modes, and HM can provide up to 35 types of intra-prediction coding modes.
[0087] JVET is working on the development of the H.266 standard. The H.266 standardization process is based on an evolution model of a video decoding device called the H.266 Test Model. The algorithm description of H.266 can be obtained from http: / phenix.int-evry.fr / jvet, and the latest algorithm description is included in JVET-F1001-v2, which is incorporated herein by reference. Also, the reference software of the JEM Test Model can be obtained from https: / / jvet.hhi.fraunhofer.de / svn / svn_HMJEMSoftware / , which is also incorporated herein by reference.
[0088] In general, the HM operation model describes that a video frame or image can be divided into a sequence of treeblocks or largest coding units (LCUs) that contain both luma and chroma samples, and LCUs are also called coding tree units (CTUs). Treeblocks have a similar purpose to macroblocks in the H.264 standard. A slice contains multiple treeblocks that are consecutive in decoding order. A video frame or image can be divided into one or more slices. Each treeblock can be divided into coding units (CUs) based on a quadtree. For example, a treeblock that is the root node of the quadtree may be divided into four child nodes, and each child node may be divided into four other child nodes as a parent node. The ultimately indivisible child nodes as leaf nodes of the quadtree include decoded nodes such as decoded video blocks. Syntax data associated with the decoded bitstream may define the maximum number of times a treeblock can be divided and may define the minimum size of a decoded node.
[0089] The size of a CU corresponds to the size of a decoding node and must be square in shape. The size of a CU may range from 8x8 pixels to 64x64 pixels, or even larger than the size of a treeblock.
[0090] A video sequence typically includes a series of video frames or pictures. A group of pictures (GOP), for example, includes a series of one or more video pictures. Syntax data describing the number of pictures included in a GOP may be included in the header information of the GOP, in the header information of one or more pictures, or elsewhere. Each slice of a picture may include slice syntax data describing the coding mode of the corresponding picture. Video encoder 100 typically operates on video blocks within individual video slices to encode video data. Video blocks may correspond to decoding nodes within a CU. Video blocks may have a fixed or variable size and may vary in size according to specified decoding criteria.
[0091] In the present invention, "NxN" and "N by N" may be used interchangeably to refer to the pixel size of a video block along the vertical and horizontal dimensions, such as 16x16 pixels or 16 by 16 pixels. Generally, a 16x16 block has 16 pixel points in the vertical direction (y=16) and 16 pixel points in the horizontal direction (x=16). Similarly, an NxN block usually has N pixel points in the vertical direction and N pixel points in the horizontal direction, where N represents a non-negative integer value. The pixels in a block may be arranged in rows and columns. Also, a block does not necessarily have to have the same number of pixel points in the horizontal and vertical directions. For example, a block may include NxM pixel points, where M may not be equal to N.
[0092] After intra / inter predictive decoding of a CU is used, video encoder 100 may calculate residual data of the CU. The CU may include pixel data in the spatial domain (also called the pixel domain) or may include coefficients in a transform domain after applying a transform (e.g., a discrete cosine transform (DCT), an integer transform, a discrete wavelet transform, or a conceptually similar transform) to the residual video data. The residual data may correspond to pixel differences between pixels of an uncoded image and a predicted value corresponding to the CU. Video encoder 100 may form a CU including the residual data and generate transform coefficients of the CU.
[0093] After generating the transform coefficients via any transform, video encoder 100 may perform quantization of the transform coefficients to possibly reduce the amount of data used to represent the coefficients to provide further compression. Quantization may reduce the bit depth for some or all of the coefficients. For example, during quantization, an n-bit value may be truncated to an m-bit value, where n is greater than m.
[0094] In some possible implementations, the video encoder 100 may scan the quantized transform coefficients in a predefined scan order to generate a serialized vector that can be entropy coded. In other possible implementations, the video encoder 100 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 100 may generate a context-adaptive variable length encoding (context-based adaptive variable-length code, CAVLC) encoding (context-based adaptive binary arithmetic coding, CABAC) encoding(syntax-based adaptive binary arithmetic coding, SBAC), probability interval partitioning entropy (PIPE) encoding , or other entropy encoding According to the method, we convert a 1-dimensional vector into entropy encoding Video encoder 100 may also entropy code syntax elements associated with the encoded video data such that video decoder 200 decodes the video data.
[0095] To perform CABAC, video encoder 100 may assign a context in a context model to a transmitted symbol. The context may relate to whether neighboring values of the symbol are non-zero. To perform CAVLC, video encoder 100 may select a variable length code for the transmitted symbol. encoding Codewords in variable-length code (VLC) may be constructed such that shorter codes correspond to more likely symbols and longer codes correspond to less likely symbols. The use of VLC, as opposed to using codewords of the same length for each symbol transmitted, can achieve the goal of saving bitrate. Probabilities in CABAC may be determined based on the context assigned to the symbol.
[0096] In an embodiment of the present invention, the video encoder can perform inter prediction to reduce temporal redundancy between images. In the present invention, the CU currently being decoded by the video decoder can be referred to as the current CU. In the present invention, the image currently being decoded by the video decoder can be referred to as the current image.
[0097] 2 is an exemplary block diagram of a video encoder according to the present invention. Video encoder 100 is configured to output video to post-processing entity 41. Post-processing entity 41 represents an example of a video entity that can process encoded video data from video encoder 100, such as a media aware network element (MANE) or a splicing / editing device. In some cases, post-processing entity 41 may be an example of a network entity. In some video encoding systems, post-processing entity 41 and video encoder 100 may be parts of separate devices, and in other embodiments, the functions described with respect to post-processing entity 41 may be performed by the same device that includes video encoder 100. In one example, post-processing entity 41 is an example of storage device 40 of FIG. 1.
[0098] In the example of FIG. 2, the video encoder 100 includes a prediction processing unit 108, a filter unit 106, a decoded picture buffer (DPB) 107, an adder 112, a transformer 101, a quantizer 102, and an entropy coder 103. The prediction processing unit 108 includes an inter predictor 110 and an intra predictor 109. For image block reconstruction, the video encoder 100 further includes an inverse quantizer 104, an inverse transformer 105, and an adder 111. The filter unit 106 represents one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although FIG. 2 illustrates the filter unit 106 as an in-loop filter, in other embodiments, the filter unit 106 may be implemented as a post-loop filter. In one example, the video encoder 100 may further include a video data memory and a splitting unit (not shown).
[0099] The video data memory may store video data to be encoded by the components of the video encoder 100. Video data may be obtained from the video source 120 and stored in the video data memory. The DPB 107 may be a reference picture memory that stores reference video data for the video encoder 100 to encode the video data in intra- and inter-codec modes. The video data memory and the DPB 107 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous dynamic random access memory (SDRAM), magnetic random access memory (MRAM), resistive random access memory (RRAM), or another type of memory device. The video data memory and the DPB 107 may be provided by the same memory device or by separate memory devices. In various embodiments, the video data memory may be on-chip with other components of the video encoder 100 or off-chip with respect to those components.
[0100] As shown in FIG. 2, the video encoder 100 receives video data and stores the video data in a video data memory. A division unit divides the video data into several image blocks, which may be further divided into smaller blocks, such as image block division based on a quadtree structure or a binary tree structure. This division may include dividing into slices, tiles, or other large units. The video encoder 100 generally represents a component for encoding image blocks in a video slice to be encoded. The slice may be divided into multiple image blocks (or into sets of image blocks called tiles). The prediction processing unit 108 may select one of multiple possible codec modes to be used for the current image block, for example one of multiple intra-codec modes or one of multiple inter-codec modes. The prediction processing unit 108 may supply the intra- and inter-codec blocks to an adder 112 to generate a residual block, and to an adder 111 to reconstruct a coding block to be used as a reference image.
[0101] An intra predictor 109 in prediction processing unit 108 may perform intra predictive coding of the current image block relative to one or more neighboring blocks in the same frame or slice as the current block to be encoded to remove spatial redundancy. An inter predictor 110 in prediction processing unit 108 may perform inter predictive coding of the current image block relative to one or more predictive blocks in one or more reference images to remove temporal redundancy.
[0102] Specifically, the inter predictor 110 may be configured to determine an inter prediction mode for encoding a current image block. For example, the inter predictor 110 may use a bit rate-distortion analysis to calculate bit rate-distortion values of various inter prediction modes in a set of candidate inter prediction modes, and select an inter prediction mode with optimal bit rate-distortion characteristics therefrom. The bit rate-distortion analysis typically determines the amount of distortion (or error) between a coding block and its original uncoded block, and the bit rate (i.e., the number of bits) for generating the coding block. For example, the inter predictor 110 may determine an inter prediction mode in a set of candidate inter prediction modes for encoding the current image block with the smallest bit rate-distortion cost as the inter prediction mode for performing inter prediction on the current image block.
[0103] The inter predictor 110 is configured to predict the motion information (e.g., motion vector) of one or more sub-blocks in the current image block based on the determined inter prediction mode, and obtain or generate a prediction block of the current image block using the motion information (e.g., motion vector) of one or more sub-blocks in the current image block. The inter predictor 110 may locate the prediction block pointed to by the motion vector in a reference image list. The inter predictor 110 may also generate syntax elements related to the image block and the video slice for use when the video decoder 200 decodes the image block of the video slice. Alternatively, in one example, the inter predictor 110 performs a motion compensation process using the motion information of each sub-block to generate a prediction block of each sub-block, and obtain a prediction block of the current image block. It should be understood that the inter predictor 110 here performs motion estimation and motion compensation processes.
[0104] Specifically, after selecting an inter prediction mode for the current image block, the inter predictor 110 may provide information indicating the selected inter prediction mode of the current image block to the entropy encoder 103, such that the entropy encoder 103 encodes the information indicating the selected inter prediction mode.
[0105] The intra predictor 109 may perform intra prediction on the current image block. Specifically, the intra predictor 109 may determine an intra prediction mode for encoding the current block. For example, the intra predictor 109 may use a bit rate-distortion analysis to calculate bit rate-distortion values for various test intra prediction modes, and select an intra prediction mode with optimal bit rate-distortion characteristics from among the test modes. In any case, after selecting an intra prediction mode for the image block, the intra predictor 109 may provide information indicative of the selected intra prediction mode of the current image block to the entropy encoder 103, such that the entropy encoder 103 encodes the information indicative of the selected intra prediction mode.
[0106] After prediction processing unit 108 generates a prediction block for a current image block by inter-prediction or intra-prediction, video encoder 100 subtracts the prediction block from the current image block to be encoded to form a residual image block. Adder 112 represents one or more components that perform this subtraction operation. The residual video data in the residual block may be included in one or more transform units (TUs) and applied to transformer 101. Transformer 101 converts the residual video data into residual transform coefficients, such as using a discrete cosine transform (DCT) or a conceptually similar transform. Transformer 101 may convert the residual video data from a pixel value domain to a transform domain, such as a frequency domain.
[0107] The transformer 101 may transmit the resulting transform coefficients to the quantizer 102, which quantizes the transform coefficients to further reduce the bit rate. In some examples, the quantizer 102 may then perform a scan on a matrix containing the quantized transform coefficients. Alternatively, the entropy coder 103 may perform the scan.
[0108] After quantization, the entropy coder 103 entropy codes the quantized transform coefficients. For example, the entropy coder 103 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. After being entropy coded by the entropy coder 103, the coded bitstream may be transmitted to the video decoder 200 or archived for later transmission or retrieval by the video decoder 200. The entropy coder 103 may also entropy code syntax elements of the current image block to be coded.
[0109] The inverse quantizer 104 and the inverse transformer 105 apply inverse quantization and inverse transformation, respectively, to reconstruct the residual block in the pixel domain, e.g., for later use as a reference block in a reference image. The adder 111 adds the reconstructed residual block to the prediction block generated by the inter predictor 110 or the intra predictor 109 to generate a reconstructed image block. A filter unit 106 may be applied to the reconstructed image block to reduce distortions, such as block artifacts. The reconstructed image block is then stored in the decoded picture buffer 107 as a reference block and may be used by the inter predictor 110 as a reference block for performing inter prediction on blocks in subsequent video frames or images.
[0110] It should be understood that other structural changes of the video encoder 100 may also be used to encode the video bitstream. For example, for some image blocks or image frames, the video encoder 100 may directly quantize the residual signal without needing to be processed by the transformer 101 and, accordingly, by the inverse transformer 105. Alternatively, for some image blocks or image frames, the video encoder 100 may not generate residual data and, accordingly, may not need to be processed by the transformer 101, the quantizer 102, the inverse quantizer 104, and the inverse transformer 105. Alternatively, the video encoder 100 may directly store the reconstructed image block as a reference block without needing to be processed by the filter unit 106. Alternatively, the quantizer 102 and the inverse quantizer 104 in the video encoder 100 may be integrated together.
[0111] 3 is an exemplary block diagram of a video decoder 200 according to the present invention. In the example of FIG. 3, the video decoder 200 includes an entropy decoder 203, a prediction processing unit 208, an inverse quantizer 204, an inverse transformer 205, an adder 211, a filter unit 206, and a DPB 207. The prediction processing unit 208 may include an inter predictor 210 and an intra predictor 209. In some examples, the video decoder 200 may perform a decoding process that is substantially inverse to the encoding process described for the video encoder 100 of FIG. 2.
[0112] In the decoding process, the video decoder 200 receives from the video encoder 100 an encoded video bitstream representing image blocks and associated syntax elements of encoded video slices. The video decoder 200 may receive video data from a network entity 42 and may optionally store the video data in a video data memory (not shown). The video data memory may store video data to be decoded by components of the video decoder 200, such as an encoded video bitstream. The video data stored in the video data memory may be obtained, for example, from the storage device 40, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium. The video data memory may function as a decoded picture buffer (DPB) for storing encoded video data from the encoded video bitstream. Thus, although the video data memory is not shown in FIG. 3, the video data memory and the DPB 207 may be the same memory or may be separately provided memories. Video data memory and DPB 207 may be formed by any of a variety of memory devices, such as DRAM, including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. In various embodiments, the video data memory may be integrated on-chip with other components of video decoder 200 or may be off-chip relative to those components.
[0113] Network entity 42 may be a server, a MANE, a video editor / splicing device, or other device for implementing one or more of the techniques described above. Network entity 42 may or may not include a video encoder, such as video encoder 100. Network entity 42 may implement some of the techniques described in this invention before transmitting the encoded video bitstream to video decoder 200. In some video decoding systems, network entity 42 and video decoder 200 may be parts of separate devices, and in other cases, the functions described with respect to network entity 42 may be performed by the same device that includes video decoder 200. In some cases, network entity 42 may be an example of storage device 40 of FIG. 1.
[0114] The entropy decoder 203 of the video decoder 200 entropy decodes the bitstream to generate quantized coefficients and some syntax elements. The entropy decoder 203 forwards the syntax elements to a prediction processing unit 208. The video decoder 200 may receive the syntax elements at the video slice level and / or the image block level.
[0115] If the video slice is decoded as an intra-decoded (I) slice, intra predictor 209 of prediction processing unit 208 may generate a prediction block of an image block of the current video slice based on the intra prediction mode signaled by the signaling and data from a previously decoded block of the current frame or image. If the video slice is decoded as an inter-decoded (i.e., B or P) slice, inter predictor 210 of prediction processing unit 208 may determine an inter prediction mode for decoding a current image block of the current video slice based on a syntax element received from entropy decoder 203, and decode (e.g., perform inter prediction) the current image block based on the determined inter prediction mode. Specifically, inter predictor 210 may determine whether to predict a current image block of the current video slice using a new inter prediction mode, and if the syntax element indicates to predict the current image block using the new inter prediction mode, predict motion information of the current image block or a sub-block of the current image block of the current video slice based on the new inter prediction mode (e.g., the new inter prediction mode specified by the syntax element or the default new inter prediction mode). Thus, through a motion compensation process, a prediction block of the current image block or a sub-block of the current image block may be obtained or generated using motion information of the predicted current image block or a sub-block of the current image block. Here, the motion information may include reference image information and a motion vector, and the reference image information may include, but is not limited to, unidirectional / bidirectional prediction information, a reference image list number, and a reference image index corresponding to the reference image list. In the case of inter prediction, the prediction block may be generated from one reference image in the reference image list. The video decoder 200 may construct the reference image lists, i.e., list 0 and list 1, based on the reference images stored in the DPB 207. The reference frame index of the current image may be included in the reference frame list 0 and / or list 1.In some examples, the video encoder 100 may transmit a signal indicating whether to use a new inter prediction mode to decode a particular syntax element of a particular block, or may transmit a signal indicating whether and which new inter prediction mode to use to decode a particular syntax element of a particular block. Note that the inter predictor 210 here performs motion compensation processing.
[0116] The inverse quantizer 204 inverse quantizes the quantized transform coefficients provided to the bitstream and decoded by the entropy decoder 203. Inverse quantization may include using a quantization parameter calculated by the video encoder 100 for each image block in a video slice to determine the degree of quantization to apply and determining the degree of inverse quantization to apply. The inverse transformer 205 applies an inverse transform to the transform coefficients, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to produce residual blocks in the pixel domain.
[0117] After the inter predictor 210 generates a prediction block for the current image block or a sub-block of the current image block, the video decoder 200 obtains a reconstructed block, i.e., a decoded image block, by adding a residual block from the inverse transformer 205 and the corresponding prediction block generated by the inter predictor 210. The adder 211 represents a component that performs this summing operation. If necessary, a loop filter (in the decoding loop or after the decoding loop) may be used to smooth pixel transitions or otherwise improve video quality. The filter unit 206 may represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although FIG. 3 illustrates the filter unit 206 as an in-loop filter, in other embodiments, the filter unit 206 may be implemented as a post-loop filter. In one example, the filter unit 206 is configured to reconstruct blocks to reduce blockiness and output the result as a decoded video bitstream. Decoded image blocks in a given frame or image may also be stored in DPB 207, which stores reference images for subsequent motion compensation. DPB 207 may be part of a memory that stores decoded video for later display on a display device (e.g., display device 220 of FIG. 1), or may be separate from such memory.
[0118] It should be understood that other structural variations of the video decoder 200 may be used to decode the encoded video bitstream. For example, the video decoder 200 may generate an output video bitstream without being processed by the filter unit 206, or for some image blocks or frames, the entropy decoder 203 of the video decoder 200 does not decode quantized coefficients, which accordingly do not need to be processed by the inverse quantizer 204 and the inverse transformer 205.
[0119] The techniques of the present invention may be performed by any of the video encoders or video decoders described herein, such as the video encoder 100 and video decoder 200 shown and described in Figures 1-3. That is, in a possible embodiment, the video encoder 100 described in Figure 2 may perform certain techniques described below when performing inter prediction during encoding of a block of video data. In another possible embodiment, the video decoder 200 described in Figure 3 may perform certain techniques described below when performing inter prediction during decoding of a block of video data. Thus, references to a general "video encoder" or "video decoder" may include the video encoder 100, the video decoder 200, or another video encoding or decoding unit.
[0120] 1-3 are merely examples according to embodiments of the present invention, and in some examples, video encoder 100, video decoder 200, and video codec system may include more or fewer components or units without limiting the present invention.
[0121] In the following, based on the video codec system shown in Figures 1 to 3, an embodiment of the present invention provides a possible video encoding / decoding embodiment. Figure 4 is a flowchart of a video encoding / decoding process according to the present invention. As shown in Figure 4, the video encoding / decoding embodiment includes process (1) to process (5), which are performed by the above-mentioned encoding device 10, Video Encoder 100, decoding device 20, or video decoder 200.
[0122] Process (1): Divide an image of one frame into one or more parallel coding units that do not overlap each other. There is no dependency between the one or more parallel coding units, and they can be coded and decoded completely in parallel / independently, as shown in FIG. 4, for parallel coding unit 1 and parallel coding unit 2.
[0123] Process (2): For each parallel coding unit, it may be further divided into one or more independent coding units that do not overlap with each other, and each independent coding unit may not depend on each other, but may share some parallel coding unit header information.
[0124] For example, the width of an independent coding unit is w_lcu and the height is h_lcu. If a parallel coding unit is split into one independent coding unit, the size of the independent coding unit is exactly the same as the size of the parallel coding unit, otherwise the width of the independent coding unit must be greater than the height (unless it is an edge region).
[0125] Typically, the independent coding unit may be a fixed w_lcu×h_lcu, where w_lcu and h_lcu are both the Nth power of 2 (N≧0). For example, the size of the independent coding unit may be 128×4, 64×4, 32×4, 16×4, 8×4, 32×2, 16×2, or 8×2, etc.
[0126] As a possible example, the independent coding units may be fixed at 128x4. If the size of the parallel coding units is 256x8, the parallel coding units may be divided evenly into four independent coding units, if the size of the parallel coding units is 288x10, the parallel coding units may divide the first / second row into two 128x4 independent coding units and one 32x4 independent coding unit, and the third row into two 128x2 independent coding units and one 32x2 independent coding unit.
[0127] It should be noted that the independent coding unit may include three components of luminance Y, chrominance Cb, and chrominance Cr, or three components of red (R), green (G), and blue (B), or may include only one of the components. When the independent coding unit includes three components, the sizes of these three components may be completely the same or different, which is specifically related to the input format of the image.
[0128] Process (3): For each independent coding unit, it may be further divided into one or more non-overlapping coding units, and each coding unit within the independent coding unit may be dependent on each other, for example, multiple coding units may be cross-referenced for precoding and decoding.
[0129] If the size of the coding unit and the independent coding unit is the same (i.e., the independent coding unit is divided into only one coding unit), the size may be any size described in process (2).
[0130] When an independent coding unit is divided into multiple non-overlapping coding units, possible examples of division include horizontal equal division (where the height of the coding unit is the same as the independent coding unit but the width is different, which may be 1 / 2, 1 / 4, 1 / 8, 1 / 16, etc.), vertical equal division (where the width of the coding unit is the same as the independent coding unit but the height is different, which may be 1 / 2, 1 / 4, 1 / 8, 1 / 16, etc.), horizontal and vertical equal division (quadtree division), etc., of which horizontal equal division is preferred.
[0131] If the coding unit has a width w_cu and a height h_cu, the width must be greater than the height (unless it is an edge region). In general, the coding unit may be a fixed w_cu×h_cu, where w_cu and h_cu are both 2 to the power of N (N is 0 or greater), such as 16×4, 8×4, 16×2, 8×2, 8×1, 4×1, etc.
[0132] As one possible example, the coding unit may be a fixed 16x4. If the size of the independent coding unit is 64x4, the independent coding unit may be divided evenly into four coding units, and if the size of the independent coding unit is 72x4, the coding unit is divided into four 16x4 and one 8x4.
[0133] Note that the coding unit may include three components of luminance Y, chrominance Cb, and chrominance Cr (or three components of red R, green G, and blue B), or may include only one of the three components. When the coding unit includes three components, the sizes of the three components may be completely the same or different, which is specifically related to the image input format.
[0134] Note that process (3) is an optional step in the video coding method, and the video encoder / decoder may encode / decode the residual coefficients (or residual values) of the independent coding units obtained in process (2).
[0135] Process (4): The coding unit may be further divided into one or more non-overlapping prediction groups (PGs). The PGs may be abbreviated as Groups, and each PG is encoded and decoded according to a selected prediction mode to obtain a predicted value of each PG, and the predicted value of each PG constitutes a predicted value of the entire coding unit, and a residual value of the coding unit can be obtained based on the predicted value of the coding unit and the original value of the coding unit.
[0136] Process (5): Group the coding units according to the residual values of the coding units, obtain one or more non-overlapping residual blocks (RBs), and encode and decode the residual coefficients of each RB according to the selected mode to form a residual coefficient stream. Specifically, the residual coefficients can be transformed or not transformed.
[0137] The selection mode of the residual coefficient encoding and decoding method in process (5) includes, but is not limited to, a semi-fixed length encoding method, an exponential Golomb encoding method, a Golomb-Rice encoding method, a truncated unary (TU) encoding method, a run length encoding (RLE) method, a direct encoding method of the original residual value, etc.
[0138] For example, a video encoder may directly encode the coefficients in the RBs.
[0139] Also, for example, the video encoder can perform a transform such as DCT, DST, or Hadamard transform on the residual block and encode the transformed coefficients.
[0140] As a possible example, when the RB is small, the video encoder may directly quantize each coefficient in the RB in a uniform manner and perform binarization encoding. When the RB is large, the video encoder may further divide the RB into multiple coefficient groups (CGs), uniformly quantize each CG, and then perform binarization encoding. In some embodiments of the present invention, the CGs and the quantization groups may be the same.
[0141] Hereinafter, the residual coefficient coding part in the semi-fixed length coding method will be described as an example. First, the maximum value of the absolute value of the residual in one RB block is set as a modified maximum. Next, the number of coding bits of the residual coefficient in the RB block is determined, and the number of coding bits of the residual coefficient in the same RB block is the same. For example, if the critical limit (CL) of the current RB block is 2 and the current residual coefficient is 1, 2 bits are required to code the residual coefficient 1, which is represented as 01. If the CL of the current RB block is 7, it indicates that 8 bits of the residual coefficient and 1 sign bit are coded. The determination of CL is to find the minimum M value that satisfies that all the residuals of the current sub-block are within the range of [-2^(M-1), 2^(M-1)]. If the two boundary values of -2^(M-1) and 2^(M-1) exist simultaneously, M needs to be increased by 1, that is, M+1 bits are required to code all the residuals of the current RB block. If only one of the two boundary values, -2^(M-1) and 2^(M-1), exists, one Trailing bit needs to be coded to determine whether the boundary value is -2^(M-1) or 2^(M-1). If neither -2^(M-1) nor 2^(M-1) exists, there is no need to code the Trailing bit.
[0142] In some special cases, the video encoder may directly encode the original values of the image instead of the residual values.
[0143] A coding block in an embodiment of the present invention corresponds to one image block in an image, and the coding block may be a coding unit divided in the above process (3), or a prediction group into which the coding unit is divided.
[0144] An image decoding method and an image encoding method according to an embodiment of the present invention will be described in detail below with reference to schematic configuration diagrams of a video codec system shown in FIG. 1, a video encoder shown in FIG. 2, and a video decoder shown in FIG.
[0145] Fig. 5 is a flowchart of an image decoding method according to the present invention, which may be applied to the video codec system shown in Fig. 1, and which may be executed by a decoding device 20. Specifically, the decoding method may be executed by a video decoder 200 included in the decoding device 20. As shown in Fig. 5, the image decoding method according to an embodiment of the present invention includes the following steps:
[0146] In S501, for any one pixel point or any multiple parallel inverse quantization pixel points in the current coding block, the QP value of the pixel point is determined.
[0147] Here, the QP values of at least two pixel points among the pixel points in the current coding block are different. Parallel dequantization pixel points means that for some pixel points, these pixel points can be dequantized in parallel.
[0148] Any one pixel point or any multiple parallel dequantized pixel points in the current coding block is the pixel point currently being processed by the video decoder, and for ease of explanation, may also be referred to as the current pixel point in the following embodiments.
[0149] The video stream to be decoded can be decoded to obtain one or more image frames contained in the video stream. An image frame includes one or more image blocks. In an embodiment of the present invention, a current coding block corresponds to an image block of an image to be processed (which may be any image frame in the video stream), and a coding block may be a coding unit.
[0150] In the prior art, the same QP value is used for all pixel points in one coding block (e.g., the above current coding block). That is, in the process of dequantizing the current coding block, the QP value is the QP value of the granularity of the coding block, so the resulting image distortion is relatively large. In contrast, in the embodiment of the present invention, for one coding block (e.g., the above current coding block), the video decoder determines the QP value of the current pixel point of the coding block, and assigns different QP values to at least two of the pixel points of the coding block. That is, in the process of dequantizing the current coding block, the QP value is the QP value of the pixel granularity, and the difference between different pixel points in the same coding block is fully taken into account. By assigning an appropriate QP value to each pixel point, the decoding distortion of the image frame can be reduced and the dequantization effect of the current coding block can be improved (the dequantization effect can be evaluated by the quality of the image obtained by decoding).
[0151] In S502, the current pixel point is inverse quantized based on the QP value of the current pixel point.
[0152] Dequantizing the current pixel point specifically means dequantizing the level value of the current pixel point, where the level value of the current pixel point is obtained by the video decoder analyzing the bitstream.
[0153] In the encoding process, the video encoder predicts a current coding block to obtain a residual value of the current coding block. The residual value is a difference between a true value of the current coding block and a predicted value of the current coding block, and may be called a residual coefficient of the current coding block. Then, the residual coefficient of the current coding block is transformed and then quantized to obtain a quantized coefficient of the current coding block. Alternatively, the video encoder does not transform the residual coefficient of the current coding block, but directly quantizes the residual coefficient to obtain a quantized coefficient of the current coding block. In this case, the quantized coefficient may be called a level value or a quantized residual coefficient. In the embodiment of the present invention, for ease of description, the quantized value is collectively called a level value.
[0154] In an embodiment of the present invention, the step of inverse quantizing the level value of the current pixel point based on the QP value of the current pixel point may include a step of determining a quantization step Qstep of the current pixel point based on the QP value of the current pixel point, and a step of inverse quantizing the level value of the current pixel point using the Qstep of the current pixel point for the selected quantizer combination.
[0155] Optionally, the quantizers are uniform or non-uniform, with the quantizer combination being determined by mark information contained in the bitstream.
[0156] The video decoder can determine Qstep based on the QP value by at least one of formula derivation or lookup, and three possible implementation methods are provided below.
[0157] Method 1:
number
[0158] Method 2:
number
[0159] Method 3:
number
[0160] Optionally, quantization and dequantization may be achieved using conventional scalar quantization methods in H.265 described below. Quantization:
number
number
[0161] It can be seen that the larger the QP value and the larger the Qstep value, the coarser the quantization, the greater the image distortion due to quantization, and the smaller the bit rate of coefficient coding.
[0162] [0, 1-f) represents the quantization dead zone, and the parameter f is related to the length of the quantization dead zone. The smaller f is, the longer the quantization dead zone is, and the level value after quantization is closer to the zero point. When f=0.5, the above quantization and inverse quantization formulas correspond to rounding off, and the quantization distortion is minimized. When f<0.5, the smaller f is, the larger the quantization distortion is, and the smaller the bit rate of coefficient coding is. In H.265, f=1 / 3 is selected for I frames, and f=1 / 6 is selected for B / P frames.
[0163] By way of example, the quantization or inverse quantization formula of the uniform quantizer can refer to the above quantization and inverse quantization formula, and the parameter f can be taken in the following manner:
[0164] Method 1: f is set to 0.5 or some other fixed value.
[0165] Method 2: f may be adaptively determined based on the QP value, the prediction mode, and whether or not to perform a transform.
[0166] As described above, according to the decoding method provided in the embodiment of the present invention, the video decoder determines the QP value of each pixel point for the coding block, and then inverse quantizes each pixel point according to the QP value of each pixel point, that is, performs inverse quantization on a pixel-by-pixel basis, thereby reducing the decoding distortion of the image frame while ensuring a certain compression rate, and improving the authenticity and accuracy of image decoding.
[0167] Optionally, referring to FIG. 5, as shown in FIG. 6, before determining the QP value of the current pixel point, the decoding method provided in the embodiment of the present invention further includes S503-S504.
[0168] In S503, the QP value of the current coding block is obtained.
[0169] In one embodiment, the QP value of the current coding block may be obtained by analyzing the bitstream. Because the probability of a small QP occurring is higher than a large QP in near-lossless compression techniques, the video encoder can directly encode the QP value of a coding block using Truncated Unary (TU), Truncated Rice (TR), or Exponential Golomb, so that the video decoder can obtain the QP value of the coding block by analyzing the bitstream.
[0170] In another embodiment, the QP value of the current coding block may be obtained based on the predicted QP value of the current coding block and the QP offset amount. For example, the QP value of the current coding block may be obtained by the following method 1 or method 2.
[0171] Here, the process of the former embodiment includes S1 to S3.
[0172] In S1, the predicted QP value of the current coding block is obtained.
[0173] Optionally, the predicted QP value of the current coding block may be calculated based on the QP values of surrounding blocks of the current coding block.
[0174] For example, the predicted QP value of the current coding block may be determined based on the QP value of the reconstructed block to the left of the current coding block and the QP value of the reconstructed block above, as follows:
number
[0175] At S2, the bitstream is parsed to obtain the QP offset of the current coding block.
[0176] In the video encoding process, the video encoder determines a predicted QP value of the current encoding block, determines the difference between the true QP value of the current encoding block and the predicted QP value, obtains a QP offset amount (which can be written as deltAQP) of the current encoding block, and then encodes the QP offset amount using variable length coding, and transmits the QP offset amount to the video decoder through a bitstream. Thus, after obtaining the bitstream, the video decoder can analyze the bitstream to obtain the QP offset amount of the current encoding block.
[0177] In S3, the sum of the predicted QP value of the current coding block and the QP offset amount is set as the QP value of the current coding block. QP = predQP + deltaQP. Here, QP represents the QP value currently being coded, predQP represents the predicted QP value of the current coding block, and deltaQP represents the QP offset amount of the current coding block.
[0178] The process of the latter embodiment includes steps S10 to S30.
[0179] At S10, a predicted QP value of the current coding block is obtained.
[0180] For the explanation of S10, please refer to the relevant explanation of S1, which will not be repeated here.
[0181] In S20, a QP offset value of the current coding block is determined based on the derivation information of the current coding block.
[0182] Here, the derived information includes at least one of flatness information of the current coding block, remaining space in the bitstream buffer, or distortion constraint information.
[0183] In the video encoding process, the video encoder uses a code control algorithm to derive the QP offset amount of the current encoding block based on the derivation information of the current encoding block, but the video encoder does not transmit the QP offset amount in the bitstream. Thus, in the video decoding process, the video decoder uses the same method as the video encoder to derive the QP offset amount of the current encoding block. The above S20 can use any method for deriving the QP offset amount known to those skilled in the art, and will not be described in detail in this specification.
[0184] In S30, the sum of the predicted QP value of the current coding block and the QP offset amount is set as the QP value of the current coding block.
[0185] In an embodiment of the present invention, the predicted QP value of the current coding block can be obtained based on more information, for example, the QP offset amount of the current coding block is derived based on the QP of the coding block previous to the current coding block, the number of coding bits of the previous coding block (prev Block Rate), the target bit rate (target Rate), flatness information of the current coding block, and the fullness of the current bitstream buffer (rc Fullness).
[0186] In S504, the predicted QP value of the current pixel point is determined to be the QP value of the current coding block.
[0187] The obtained QP value of the current coding block is taken as the initial QP value (i.e., predicted QP value) of each pixel point of the current coding block, and the QP value of each pixel is obtained by adjusting or not adjusting the predicted QP value.
[0188] Based on S503-S504, as shown in FIG. 7, the determining the QP value of the current pixel point (ie, S501) specifically includes S5011-S5012.
[0189] In S5011, if the current pixel point is a target pixel point in the current coding block, the predicted QP value of the current pixel point is adjusted, and the adjusted QP value is set as the QP value of the current pixel point.
[0190] In the embodiment of the present invention, the target pixel point is one or more designated pixel points in the current coding block, and these designated pixel points can be understood as pixel points or candidate pixel points whose QP values should be adjusted. A QP adjustment policy is executed for the candidate pixel points. The target pixel point will be described later.
[0191] In S5012, if the current pixel point is a pixel point other than the target pixel point in the current coding block, the predicted QP value of the current pixel point is set as the QP value of the current pixel point.
[0192] In the embodiment of the present invention, pixel points other than the target pixel point in the current coding block are not adjusted, there is no need to execute the QP adjustment policy, and the QP values of these pixel points are the QP value of the current coding block.
[0193] As shown in FIG. 7, in one embodiment, S5011 in the above embodiment may be realized by S601 to S602.
[0194] In S601, information on reconstructed pixel points around the current pixel point is obtained.
[0195] The reconstructed pixel points around the current pixel point may be understood as the reconstructed pixel points adjacent to the current pixel point, including pixel points within a square region centered on the current pixel point and having a side length of a first predetermined value, or pixel points within a diamond region centered on the current pixel point and having a diagonal length of a second predetermined value.
[0196] The first preset value and the second preset value may be set according to actual requirements, and the first preset value and the second preset value may be equal or different, for example, the first preset value and the second preset value may be 3 or 5.
[0197] 8A is a schematic diagram illustrating the division of a square region centered on a current pixel point, showing two possible cases. In case 1, the reconstructed pixel point refers to a pixel point in a square region centered on the current pixel point and having a side length of 3, such as surrounding pixel point 1 shown in FIG. 8A. In case 2, the reconstructed pixel point refers to a pixel point in a square region centered on the current pixel point and having a side length of 5, such as surrounding pixel point 2 shown in FIG. 8A.
[0198] 8B is a schematic diagram illustrating the division of a diamond region centered on a current pixel point, showing two possible cases. In case 1, the reconstructed pixel point refers to a pixel point in a diamond region centered on the current pixel point and having a diagonal length of 3, such as the surrounding pixel point 1 shown in 8B. In case 2, the reconstructed pixel point refers to a pixel point in a diamond region centered on the current pixel point and having a diagonal length of 5, such as the surrounding pixel point 2 shown in 8B.
[0199] 8A and 8B are merely examples of embodiments of the present invention for describing the reconstructed pixel points around a current pixel point, and should not be understood as limiting the present invention. In some other possible examples, the reconstructed pixel points around a current pixel point may be one or two pixel points adjacent to the current pixel point above and below, or adjacent to the current pixel point to the left and right.
[0200] The information of the reconstructed pixel points around the current pixel point may include at least one of the following information of the reconstructed pixel point: pixel value, reconstructed residual value, gradient value, flatness information, texture information, complexity information, background luminance, contrast, or motion amount. Here, the reconstructed residual value includes a residual value after inverse quantization, or a difference between a reconstructed value and a predicted value. The gradient value includes a horizontal gradient, a vertical gradient, or an average gradient. The motion amount can be represented by a motion vector.
[0201] Furthermore, the value of the information of the reconstructed pixel point mentioned above may include at least one of an original value, an absolute value, an average value, or a difference.
[0202] In one embodiment, the above method for obtaining information of reconstructed pixel points surrounding the current pixel point may include steps 1 to 2.
[0203] In step 1, the information of the predicted pixel point of the current pixel point is obtained.
[0204] The predicted pixel point of the current pixel point is a reconstructed pixel point. Optionally, the reconstructed pixel point may be a reconstructed pixel point in the current coding block, or may be a reconstructed pixel point other than the current coding block. For example, when the prediction mode is an intra prediction mode, the reconstructed pixel point is a pixel point around the coding block in the current image frame, and when the prediction mode is an inter prediction mode, the reconstructed pixel point may be a reconstructed block on a reference frame of the current image frame.
[0205] In step 2, if the predicted pixel point is a reconstructed pixel point in the current encoding block, the information of the predicted pixel point is taken as the information of the reconstructed pixel points surrounding the pixel point; otherwise, the difference between the information of the predicted pixel point and the information of the reconstructed pixel points surrounding the predicted pixel point or the absolute value of the difference is taken as the information of the reconstructed pixel points surrounding the current pixel point.
[0206] In S602, the predicted QP value of the current pixel point is adjusted based on the information of the reconstructed pixel points surrounding the current pixel point.
[0207] The adjusted predicted QP value becomes the final QP value of the pixel point.
[0208] In an embodiment of the present invention, a QP value adjustment parameter table may be set for a current pixel point of a current coding block. For example, the following Table 1 shows some parameters required to adjust the QP value: [Table 1]
[0209] Specifically, S602 includes the following S6021 to S6023.
[0210] In S6021, if the current pixel point satisfies the second preset condition and the third preset condition, adjust the predicted QP value of the current pixel point according to the first QP offset amount and the distortion reference QP value.
[0211] Here, the second preset condition is that the predicted QP value of the current pixel point is greater than the distortion reference QP value, and the predicted QP value of the current pixel point is less than or equal to the adjustable QP maximum value (i.e., maxAdjustQP in Table 1), and the adjustable QP maximum value is less than or equal to the QP maximum value. The third preset condition is that the information of the reconstructed pixel points around the current pixel point is less than or equal to a first threshold value (i.e., thres1 in Table 1).
[0212] Optionally, referring to Table 1, if the current pixel point satisfies the second preset condition and the third preset condition, the adjusted QP value of the current pixel point satisfies: finalQP=max(initQP‐offset1,jndQP) Here, finalQP represents the predicted QP value after adjustment, initQP represents the predicted QP value of the current pixel point, offset1 represents the first QP offset amount, jndQP represents the distortion reference QP value, and max represents taking the maximum value.
[0213] In one case, the distortion-based QP value is obtained by analyzing a bitstream, for example, the bitstream conveys a distortion-based QP value, for example, 20.
[0214] In another case, the distortion-based QP value is derived based on flatness or texture information, background luminance, and contrast information of surrounding reconstructed coding blocks.
[0215] In yet another case, the distortion reference QP value may be a value, for example 15, that is preset by the video encoder or video decoder.
[0216] That is, the distortion-based QP value may be transmitted in a bitstream, may be derived by a video encoder or a video decoder during a video encoding / decoding process, or may be a preset value. In the embodiment of the present invention, by introducing the distortion-based QP value into the process of determining the QP value of a pixel point, each pixel point can satisfy the judgment information corresponding to the perceptible distortion, reduce the image distortion, and improve the subjective quality of the image.
[0217] In S6022, if the current pixel point satisfies the second preset condition and the fourth preset condition, adjust the predicted QP value of the current pixel point according to the second QP offset amount and the QP maximum value.
[0218] Here, the fourth preset condition is that the information of the reconstructed pixel points around the current pixel point is greater than the second threshold (i.e., thres2 in Table 1), and the first threshold is less than or equal to the second threshold.
[0219] In the above S6021 and S6022, it is known that the predicted QP value of the current pixel point needs to be adjusted.
[0220] In S6023, in cases other than S6021 and S6022, the predicted QP value of the current pixel point is set as the QP value of the current pixel point, that is, there is no need to adjust the predicted QP value of the current pixel point.
[0221] As can be understood, the above S6023 includes Case 1 to Case 3 as follows.
[0222] In case 1, the predicted QP value of the current pixel point is less than or equal to the distortion reference QP value.
[0223] In case 2, the current pixel point satisfies a second preset condition (i.e., the predicted QP value of the current pixel point is greater than the distortion reference QP value and the predicted QP value of the current pixel point is less than or equal to the adjustable QP maximum value), and the information of the reconstructed pixel points around the current pixel point is greater than the first threshold and less than or equal to the second threshold.
[0224] In case 3, the predicted QP value of the current pixel point is greater than the adjustable QP maximum value.
[0225] As shown in FIG. 7, in another embodiment, S5011 in the above embodiment may be realized by S603.
[0226] In S603, if the current pixel point satisfies a first preset condition, the preset QP value is the QP value of the current pixel point; otherwise, the predicted QP value of the current pixel point is the QP value of the current pixel point.
[0227] The first preset condition includes at least one of: the current pixel point is a luma pixel point; the current pixel point is a chroma pixel point; the bit depth of the current pixel point is less than or equal to a bit depth threshold (bdThres in Table 1); the predicted QP value of the current pixel point is less than or equal to the adjustable QP maximum value (maxAdjustQP in Table 1); and information of reconstructed pixel points around the current pixel point is less than or equal to a first preset threshold (threS1 in Table 1).
[0228] For example, the first preset condition may be that the current pixel point is a luma pixel point and the bit depth of the current pixel point is equal to or less than a bit depth threshold, or that the current pixel point is a chroma pixel point and the predicted QP value of the current pixel point is equal to or less than an adjustable QP maximum value. Specifically, one or more of the above conditions may be combined to form the first preset condition according to actual needs.
[0229] In one possible implementation, the target pixel point is any one or more pixel points in the current coding block, i.e., the QP value of each pixel point in the current coding block needs to be adjusted, and for each pixel point, its QP value can be determined using the above S601-S602 or S603.
[0230] In yet another possible implementation, the target pixel point is any one or more of the pixel points of the second part of the current coding block. In an embodiment of the present invention, the current coding block includes at least pixel points of the first part and / or pixel points of the second part, the pixel points of the first part are set to pixel points whose QP value does not need to be adjusted, i.e., the QP value of each pixel point of the pixel points of the first part is its predicted QP value, and the pixel points of the second part are set to pixel points whose QP value needs to be adjusted, i.e., the QP value of each pixel point of the pixel points of the second part is determined using the above S601-S602 or S603.
[0231] In another possible implementation, the current coding block includes at least pixel points of a first portion and / or pixel points of a second portion, the pixel points of the first portion are set to pixel points whose QP values do not need to be adjusted, and the pixel points of the second portion are set to pixel points whose QP values should be adjusted.
[0232] In one possible case, when the prediction mode of the current coding block is a pixel-wise prediction mode, the pixel point of the second part may include a pixel point of a first position and / or a pixel point of a second position. The pixel point of the first position and the pixel point of the second position are determined based on the pixel-wise prediction mode of the current coding block. Typically, the pixel point of the second position may include a pixel point corresponding to a starting point of pixel-wise prediction in the horizontal direction.
[0233] If the target pixel point is one or more of the pixel points at the first position, and the current pixel point is one or more of the pixel points at the first position, the QP value of the current pixel point may be determined using S601 to S602 above, and if the current pixel point is one or more of the pixel points at the second position, the QP value of the current pixel point may be determined using S603 above.
[0234] In an embodiment of the present invention, which pixel points of the second part of the current coding block are set as pixel points of the first position and which pixel points are set as pixel points of the second position are related to the pixel-wise prediction mode of the current coding block.
[0235] As an example, referring to FIG. 9A to FIG. 9D, a current coding block is taken as an example of a 16×2 (width w is 16, height h is 2) coding block, and the prediction mode of the coding block is a pixel-by-pixel prediction mode, which includes four modes: pixel-by-pixel prediction mode 1, pixel-by-pixel prediction mode 2, pixel-by-pixel prediction mode 3, and pixel-by-pixel prediction mode 4. Here, ≡ indicates that the predicted value of the current pixel point is the average of the reconstructed values of the pixels on either side of the current pixel point, ||| means that the predicted value of the current pixel point is the average of the reconstructed values of the pixels above and below the current pixel point, > indicates that the predicted value of the current pixel point is the reconstructed value of the pixel point to the left of the current pixel point; V indicates that the predicted value of the current pixel point is the reconstructed value of the pixel point above it.
[0236] In an embodiment of the present invention, for ease of explanation, a pixel point predicted based on the pixel points on either side thereof is referred to as a first type pixel point, a pixel point predicted based on the pixel points on either side thereof is referred to as a second type pixel point, a pixel point predicted based on the pixel point on the left side thereof is referred to as a third type pixel point, and a pixel point predicted based on the pixel point on the upper side thereof is referred to as a fourth type pixel point.
[0237] As shown in Figure 9A, the prediction mode of the current coding block is pixel-wise prediction mode 1, so the current coding block includes pixel points of a first portion and pixel points of a second portion, and the pixel points of the second portion are all defined as pixel points of a first position, that is, the pixel points of the second portion do not include pixel points of a second position, where the pixel points of the first position include pixel points of a first type and pixel points of a fourth type.
[0238] As shown in Figure 9B, the prediction mode of the current coding block is pixel-wise prediction mode 2, so the current coding block includes pixel points of a first portion and pixel points of a second portion, some of the pixel points of the second portion are defined as pixel points of a first position, and another part is defined as pixel points of a second position, where the pixel points of the first position include pixel points of a second type and pixel points of a third type, and the pixel points of the second position include pixel points of a third type and pixel points of a fourth type.
[0239] As shown in Figure 9C, the prediction mode of the current coding block is pixel-wise prediction mode 3, so all pixel points of the current coding block are pixel points of the second part, some of the pixel points of the second part are defined as pixel points of the first position, and another part is defined as pixel points of the second position, where the pixel points of the first position include pixel points of a third type, and the pixel points of the second position include pixel points of the third type and pixel points of a fourth type.
[0240] As shown in Figure 9D, the prediction mode of the current coding block is pixel-wise prediction mode 4, so the current coding block includes pixel points of a first portion and pixel points of a second portion, and the pixel points of the second portion are all defined as pixel points of a first position, that is, the pixel points of the second portion do not include pixel points of a second position, where the pixel points of the first position include pixel points of a fourth type.
[0241] As an example, referring to Figures 10A and 10B, taking an example of a current coding block being an 8x2 coding block (width w is 8, height h is 2), the prediction mode of the coding block is a pixel-by-pixel prediction mode, and the pixel-by-pixel prediction mode of the coding block includes two modes: pixel-by-pixel prediction mode 1 and pixel-by-pixel prediction mode 2.
[0242] As shown in Figure 10A, the prediction mode of the current coding block is pixel-wise prediction mode 1, so the current coding block includes pixel points of a first portion and pixel points of a second portion, and the pixel points of the second portion are all defined as pixel points of a first position, that is, the pixel points of the second portion do not include pixel points of a second position, where the pixel points of the first position include pixel points of a fourth type.
[0243] As shown in Figure 10B, the prediction mode of the current coding block is pixel-wise prediction mode 2, so that all pixel points of the current coding block are pixel points of the second part, some of the pixels of the second part are defined as pixel points of the first position, and another part is defined as pixel points of the second position, where the pixel points of the first position include pixel points of a third type, and the pixel points of the second position include pixel points of the third type and pixel points of a fourth type.
[0244] As an example, referring to Figures 11A and 11B, taking the current coding block as an 8x1 coding block (width w is 8 and height h is 1), the prediction mode of the coding block is a pixel-by-pixel prediction mode, and the pixel-by-pixel prediction mode of the coding block includes two modes: pixel-by-pixel prediction mode 1 and pixel-by-pixel prediction mode 2.
[0245] As shown in FIG. 11A, the prediction mode of the current coding block is pixel-wise prediction mode 1, therefore, all pixel points of the current coding block are pixel points of the first part, and the pixel points of the first part are pixel points of the fourth type.
[0246] 11B, the prediction mode of the current coding block is pixel-wise prediction mode 2, so that all pixel points of the current coding block are pixel points of the second part, some of the pixel points of the second part are defined as pixel points of the first position, and another part is defined as pixel points of the second position, where the pixel points of the first position include pixel points of a third type, and the pixel points of the second position include pixel points of the third type and pixel points of a fourth type.
[0247] In combination with FIG. 5, as shown in FIG. 12, in a possible case, when the prediction mode of the current coding block is a block prediction mode, before determining the QP value of the current pixel, the decoding method provided in the embodiment of the present invention may further include S505 to S506.
[0248] In S505, region division information of the current coding block is obtained, and the region division information includes the number N of regions and position information of region boundary lines, where N is an integer of 2 or more.
[0249] The region division information may be called a division template.
[0250] Optionally, the method for obtaining region partition information of the current coding block comprises a step of obtaining predefined region partition information of the current coding block, or a step of analyzing a bitstream to obtain the region partition information of the current coding block, or a step of being derived by a video decoder.
[0251] In S506, the current coding block is divided into N regions based on the region division information.
[0252] Optionally, the block prediction mode may include a block-based inter prediction mode, a block-based intra prediction mode, or an intra block copy (IBC) prediction mode.
[0253] Obtain a prediction block of a current coding block according to a block prediction mode, and obtain a residual block thereof. Optionally, in terms of whether to perform a transform on the residual block, the block prediction mode may include a block prediction mode that does not perform a transform and a locked prediction mode that performs a transform. The block prediction mode that does not perform a transform refers to not performing a transform on the residual block determined according to the block prediction mode, and the block prediction mode that performs a transform refers to performing a transform on the residual block determined according to the block prediction mode.
[0254] For block prediction modes that do not perform transformation, pixel points in the current coding block may be reconstructed sequentially in a "pixel-wise" or "region-wise" manner, and the QP values of later reconstructed pixel points may be adjusted using information of previously reconstructed pixel points.
[0255] The pixel-by-pixel reconstruction of pixel points in the current coding block is similar to the above-mentioned pixel-by-pixel prediction mode, and therefore the method of determining the QP value of a pixel point of the current coding block in the case of "pixel-by-pixel" reconstruction is similar to the method of determining the QP value in the above-mentioned pixel-by-pixel prediction mode, i.e., the above S601-S602 or S603 can be used to determine the QP value of a pixel point of the current coding block.
[0256] The "region-by-region" reconstruction of pixel points in the current coding block allows pixels of the same region to be reconstructed in parallel, and the idea is to divide the current coding block into N regions (N≧2) and then reconstruct them region-by-region sequentially.
[0257] Specifically, for the "region unit" reconstruction method in the block prediction mode without transformation, the current coding block can be divided into N regions according to the number N of regions and the position information of the region boundary in the region division information. It should be noted that the QP value of the pixel point of at least one region among the N regions is determined based on the information of the reconstructed pixel point of at least one other region. Here, the other region is a region other than at least one region among the N regions, or a region other than the current coding block. That is, the N regions have a reconstruction order, that is, the reconstruction process between the sub-regions in the N regions has a dependency relationship, for example, one region needs to be reconstructed first (the corresponding other region is a region other than the current coding block), and then another region needs to be reconstructed based on the reconstruction result of the region (that is, the region is a region other than the other region).
[0258] Optionally, the number N of regions and the position information of the region borders may be derived based on information of the current coding block or information of reference pixels of the current coding block.
[0259] As an example, suppose N=2 (i.e., divide the current coding block into two regions), the current coding block includes a first region and a second region, and the pixel points in the first region include at least one of a pixel point in a horizontal slice at any position, a pixel point in a vertical slice at any position, or a pixel point in a diagonal slice at any position of the current coding block, where the width of the slice is less than or equal to 2, and if the slice is located at the boundary of the current coding block, the width of the slice is equal to 1. The pixel points in the second region are pixel points other than the first region in the current coding block. The reconstruction order of the first region and the second region is to first reconstruct the second region and then reconstruct the first region.
[0260] By way of example, it can be understood that for the current coding block, the pixel points of the upper boundary are pixel points of the upper slice, the pixel points of the lower boundary are pixel points of the lower slice, the pixel points of the left boundary are pixel points of the left slice, and the pixel points of the right boundary are pixel points of the right slice.
[0261] 13A to 13D are schematic diagrams of some exemplary division results of a current coding block. In FIG. 13A: The pixel points in the first region are the pixel points on the upper boundary of the current coding block, the pixel points on the lower boundary of the current coding block, In FIG. 13B, the pixel points in the first region include pixel points on the bottom boundary and pixel points on the right boundary of the current coding block; in FIG. 13C, the pixel points in the first region include pixel points on the right boundary of the current coding block; and in FIG. 13D, the pixel points in the first region include pixel points on the bottom boundary of the current coding block.
[0262] For the block prediction mode in which the transformation is performed, all pixel points in the current coding block need to be reconstructed in parallel, so the current coding block can be divided into N regions (N≧2) and the pixels in the same region can be reconstructed in parallel.
[0263] Specifically, for a block prediction mode in which transformation is performed, the current coding block may be divided into N regions based on the number N of regions and the position information of region borders in the region division information.
[0264] The number of regions N and the position information of the region boundary line may be derived based on the information of the current coding block or the information of the reference pixel of the current coding block. The region division information of the current coding block may be determined based on the pixel points of the upper adjacent row and / or the pixel points of the left adjacent column of the current coding block (i.e., the reference pixel points of the current coding block). Specifically, based on the pixel points of the upper adjacent row and / or the pixel points of the left adjacent column of the current coding block, an object edge in the current coding block is predicted, and the current coding block is divided into several regions based on the object edge. For example, based on the pixel points of the upper adjacent row and / or the pixel points of the left adjacent column of the current coding block, a gradient algorithm is used to predict pixel points where pixel values of the rows and / or columns of the current coding block suddenly change, and the pixel points where the pixel values suddenly change are set as the positions of the region boundary lines, thereby determining the number of regions N.
[0265] According to the region division information of the current coding block determined by the above method, the region division method of the current coding block can include at least one of horizontal division, vertical division, and diagonal division. For example, if there are one or more pixel points where a sudden change occurs in the row pixel value of the current coding block and there is no pixel point where a sudden change occurs in the column pixel value of the current coding block, the region division method of the current coding block is vertical division, if there are one or more pixel points where a sudden change occurs in the column pixel value of the current coding block and there is no pixel point where a sudden change occurs in the row pixel value of the current coding block, the region division method of the current coding block is horizontal division, and if there are one or more pixel points where a sudden change occurs in the row pixel value of the current coding block and there is also one or more pixel points where a sudden change occurs in the column pixel value of the current coding block, the region division method of the current coding block is diagonal division.
[0266] Referring to Figures 14A to 14C, one division method is to divide the current coding block into two regions, where Figure 14A shows division into two regions vertically, with point A1 being the abrupt-changing pixel point in the row, Figure 14B shows division into two regions horizontally, with point B1 being the abrupt-changing pixel point in the column, and Figure 14C shows division into two regions diagonally, with point C1 being the abrupt-changing pixel point in the row, and point D1 being the abrupt-changing pixel point in the column.
[0267] Referring to Figure 15A-Figure 15E, one division method is to divide the current coding block into three regions, Figure 15A shows division into three regions vertically, and points A2 and A3 are row abrupt-changing pixel points, Figure 15B shows division into three regions horizontally, and points B2 and B3 are column abrupt-changing pixel points, Figure 15C-Figure 15E shows division into three regions diagonally. In Figure 15C, points C2 and C3 are row abrupt-changing pixel points, and points D2 and D3 are column abrupt-changing pixel points, in Figure 15D, points C4 and C5 are row abrupt-changing pixel points, and point D4 is column abrupt-changing pixel points, in Figure 15E, point C6 is row abrupt-changing pixel point, and points D5 and D6 are column abrupt-changing pixel points.
[0268] 14A to 14C and 15A to 15E are merely examples of some division results of the current coding block, and do not limit the embodiments of the present invention. The division method of the current coding block may be a combination of multiple division methods.
[0269] Based on the above S505 to S506, in one embodiment, for a block prediction mode that does not perform transformation, in an example where the current coding block is divided into two (i.e., N=2) regions, i.e., the current coding block includes a first region and a second region (see Figures 10A and 10B), as shown in Figure 12, determining the QP value of the current pixel point (i.e., S501) specifically includes S5013 to S5014.
[0270] In S5013, if the current pixel point is any one of the pixel points in the first region, the predicted QP value of the current pixel point is adjusted, and the adjusted QP value is set as the QP value of the current pixel point.
[0271] Specifically, obtain information of reconstructed pixel points around the current pixel point, and adjust the predicted QP value of the pixel point according to the information of reconstructed pixel points around the current pixel point. For the specific process, refer to the relevant description of S601 to S602 (wherein S602 includes S6021 to S6023) in the above embodiment, and the description is omitted here.
[0272] In S5014, if the current pixel point is any one of the pixel points in the second region, the predicted QP value of the current pixel point is set as the QP value of the current pixel point. The pixel point in the second region needs to be reconstructed, and at this time, there may be no reconstructed pixel point around it, so the predicted QP value of the pixel point in the second region is not adjusted, that is, the predicted QP value of the pixel point in the second region is set as the QP value of the pixel point.
[0273] In another embodiment, for a block prediction mode that performs the above transformation, for a current pixel point in any one of the N regions, determining a QP value for the current pixel point (i.e., S501), as shown in FIG. 12, specifically includes S5015 to S5016.
[0274] In S5015, the QP offset amount of the current pixel point is obtained.
[0275] Optionally, parse the bitstream to obtain the offset of the current pixel point. It can be understood that in the process of the video encoder encoding the image, after the video encoder predicts the QP offset amount of each pixel point, the QP offset amount of each pixel point can be embedded into the bitstream and transmitted to the video decoder.
[0276] Alternatively, optionally, a QP offset amount for the current pixel point is determined based on derived information, the derived information including index information of a region in which the current pixel is located, and / or a distance from the current pixel to a region boundary of the region in which the current pixel is located, where the distance includes any of a horizontal distance, a vertical distance, or a Euclidean distance.
[0277] Therefore, the derived QP offset amount of the current pixel point is any one of the third QP offset amount, the fourth QP offset amount, and the sum of the third QP offset amount and the fourth QP offset amount.
[0278] Here, the third QP offset amount is derived according to the index information of the region where the current pixel is located, and the third QP offset amount can be regarded as a region-level QP offset amount. It should be understood that the third QP offset amounts of pixel points in the same region are the same, and the third QP offset amounts of pixel points in different regions are different.
[0279] The fourth QP offset amount is derived based on the distance from the current pixel point to the region boundary of the region where the current pixel point is located. The fourth QP offset amount can be regarded as a pixel-level QP offset amount. When the distance corresponding to the pixel point is different, the QP offset amount of the current pixel element may be different.
[0280] According to the configuration of the video encoding device, one of the third QP offset amount, the fourth QP offset amount, or the sum of the third QP offset amount and the fourth QP offset amount may be selected as the QP offset amount for the pixel point.
[0281] In S5016, the predicted QP value of the current pixel point is adjusted based on the QP offset amount of the current pixel point, and the adjusted QP value is set as the QP value of the current pixel point.
[0282] As described above, the video decoder may determine the QP value of pixel point granularity for the pixel points in the coding block, and inverse quantize each pixel point on a pixel basis according to the QP value of each pixel point. This can reduce the decoding distortion of the image frame while maintaining a certain compression ratio, and improve the authenticity and accuracy of image decoding.
[0283] Accordingly, in the image encoding method, the video encoder first obtains the QP, Qstep and residual coefficient of the pixel point, adaptively selects a quantizer, quantizes the residual coefficient, and finally adjusts the quantization coefficient to obtain a final level value, thereby realizing the encoding of the image frame.
[0284] Based on the video encoder 100 shown in Fig. 2, the present invention further provides an image encoding method. Fig. 16 is a flowchart of the image encoding method according to the present invention. The image encoding method may be performed by the video encoder 100, or may be performed by an encoding device supporting the function of the video encoder 100 (encoding device 10 shown in Fig. 1). Here, an example is described in which the video encoder 100 implements the encoding method. The image encoding method includes the following steps S1601 to S1602.
[0285] In S1601, the QP value of any one pixel point or any multiple parallel quantization pixel points in the current coding block is determined.
[0286] Here, the QP values of at least two of the pixel points of the current coding block are different.
[0287] In S1602, the pixel point is quantized based on the QP value of the pixel point.
[0288] Quantization is the inverse process of inverse quantization. For the quantization of QP values in the encoding method, the corresponding processes in the decoding method of FIGS. 5 to 15A to 15E can be referred to, and a description thereof will be omitted here.
[0289] As described above, the encoding method provided in the embodiment of the present invention allows a video encoder to determine the QP value of each pixel point for a coding block, and then quantize each pixel point according to the QP value of each pixel point, that is, perform pixel-by-pixel quantization, thereby ensuring a certain compression rate while reducing the decoding distortion of an image frame, and improving the authenticity and accuracy of image decoding.
[0290] As can be understood, for block prediction mode, the current coding block is divided into N (N≧2) regions according to the method described in the embodiment, pixel-by-pixel quantization or parallel quantization of multiple pixel points is performed on pixel points in each region, and level values (i.e., quantized residual coefficients or quantized parameter coefficients) are obtained, and then the parameter coefficients are coded.
[0291] When the coding block is divided into N regions, the adjustment methods of the QP values of different regions may be different, so that the distributions of the residual coefficients after quantization are also different, and therefore a region-based residual group coding method can be designed.
[0292] Specifically, the residual coefficients of each region can be divided into several residual groups, and it should be noted that each residual group cannot cross regions. Then, the code length parameter of the residual group is coded, and the coding method can be fixed-length coding or variable-length coding. Then, each residual coefficient in the residual group is coded using fixed-length coding, and the code length of the fixed-length coding is determined by the code length parameter of the residual group, and the code length parameters of different residual groups can be different.
[0293] For example, referring to FIG. 17, assuming that the current coding block is a 16×2 coding block and the current coding block is divided into two regions, region 1 and region 2, the residual coefficients corresponding to region 1 may be divided into n (n≧1) residual groups, and the residual coefficients corresponding to region 2 may be divided into m (m≧1) residual groups, and each residual group does not span regions. Note that the residual coefficients correspond one-to-one to pixel points, and grouping the residual coefficients corresponding to a region means grouping the pixel points included in the region. As shown in FIG. 17, region 1 includes 15 pixel points, and exemplarily, the region 1 may be divided into one residual group, that is, the 15 pixel points may be divided into one residual group. The region 1 may also be divided into two residual groups, for example, the first 8 pixel points of region 1 may be divided into one residual group 1, and the last 7 pixel points of region 1 may be divided into another residual group 2. Alternatively, the region 1 may be divided into three residual groups, for example, one residual group for every five adjacent pixel points, to obtain three residual groups, such as residual group 1, residual group 2 and residual group 3.
[0294] Optionally, the prediction modes for the video encoder to predict the current coding block may include pixel-wise prediction modes and block prediction modes, where the block prediction modes may be inter prediction modes of a block, intra prediction modes of a block, or intra block copy (IBC) prediction modes. The IBC prediction modes are briefly described below.
[0295] The IBC technique is to search for a matching block of a current coding block from a reconstruction region of a current frame, aiming at removing spatial non-local redundancy. The prediction process in the IBC prediction mode can be divided into two processes: motion estimation and motion compensation. Motion estimation is a process in which an encoding device searches for a matching block of a current coding block, estimates the relative displacement between the current coding block and its matching block, i.e., a block vector (BV) or block vector difference (BVD) corresponding to the current coding block, and transmits the BV or BVD in a bitstream. Motion compensation is a process in which a prediction block is generated based on a matching block, and includes, for example, operations such as weighting and predictive filtering on the matching block.
[0296] Optionally, the method for the video encoder to obtain the predictive block of the current coding block may include the following method.
[0297] Method 1: If a pixel point in the prediction block is unavailable, allow padding by the pixel point above or to the left, or allow padding to a default value.
[0298] Method 2: Obtain a matching block based on BV or BVD, and perform some processing (e.g., predictive filtering, light compensation, etc.) on the matching block to generate a final predicted block.
[0299] Optionally, in the IBC prediction mode, the video encoder divides the current coding block into several transform blocks (TB) and several prediction blocks (PB), and one TB may include one or more PBs. Exemplarily, referring to FIG. 18, taking the current coding block as an example of a 16×2 coding block, the current coding block is divided into two TBs, TB1 and TB2, and the sizes of TB1 and TB2 are both 8×2. Each TB includes four PBs, and the size of each PB is 2×2, and each TB is reconstructed sequentially. One pixel reconstruction method may refer to the reconstructed pixel values in the reconstruction TB on the left side of the current TB when performing motion compensation on the PBs in each TB from the second TB of the current coding block.
[0300] Optionally, the encoding method of the BV or BVD by the video encoder is the following method 1 and / or method 2.
[0301] Method 1: When only horizontal motion estimation is performed, only horizontal BV or BVD is transmitted in the bitstream, and vertical BV or BVD does not need to be transmitted.
[0302] Method 2: The encoding method of BV or BVD may be fixed-length encoding or variable-length encoding.
[0303] In addition, the code length of fixed-length coding or the binarization method of variable-length coding is obtained based on one or more of the position information, size information (including width, height, or area) of the current coding block, division mode information or TB / PB division method, position information or size information (including width, height, or area) of the current TB, and position information or size information (including width, height, or area) of the current PB.
[0304] Optionally, the step of obtaining the current coding block may be to first obtain a prediction BV (block vector prediction, BVP), and then obtain BVD, where BV=BVP+BVD.
[0305] Here, the BVP may be obtained based on one or more of the BV or BVD, position information, size information (including width or height or area), partition mode information or TB / PB partition mode, BV or BVD, position information or size information (including width or height or area) of the coding block.
[0306] It should be understood that, to realize the functions in the above embodiments, the video encoder / video decoder includes corresponding hardware structures and / or software modules for performing the respective functions. Those skilled in the art should easily understand that the present invention can be implemented in the form of hardware, or in the form of a combination of hardware and computer software, with reference to the example units and method steps described with reference to the embodiments disclosed in the present invention. Whether a function is implemented by hardware or driven by computer software depends on the specific application scenario and design constraints of the technical solution.
[0307] FIG. 19 is a schematic structural diagram of a decoding device according to an embodiment of the present invention, where the decoding device 1900 includes a QP determination unit 1901 and an inverse quantization unit 1902. The decoding device 1900 is configured to realize the functions of the video decoder or the decoding device in the above embodiment of the decoding method, and thus can also realize the beneficial effects of the above embodiment of the decoding method. In the embodiment of the present invention, the decoding device 1900 may be the decoding device 20 or the video decoder 200 shown in FIG. 1, may be the video decoder 200 shown in FIG. 3, or may be a module (e.g., a chip) applied to the decoding device 20 or the video decoder 200.
[0308] The QP determination unit 1901 and the inverse quantization unit 1902 are used to implement the decoding method according to any of the embodiments of Figures 5 to 15A to 15E. For detailed description of the QP determination unit 1901 and the inverse quantization unit 1902, reference can be made directly to the relevant description in the embodiment of the method shown in Figures 5 to 15A to 15E, and the description will be omitted here.
[0309] Fig. 20 is a schematic structural diagram of an encoding device according to the present invention, where the encoding device 2000 includes a QP determination unit 2001 and a quantization unit 2002. The encoding device 2000 is configured to realize the functions of the video encoder or encoding device in the above encoding method embodiment, and thus can also realize the beneficial effects of the above encoding method embodiment. In the embodiment of the present invention, the encoding device 2000 may be the encoding device 10 or the video encoder 100 shown in Fig. 1, may be the video encoder 100 shown in Fig. 2, or may be a module (e.g., a chip) applied to the encoding device 10 or the video encoder 100.
[0310] The QP determination unit 2001 and the quantization unit 2002 are used to implement the encoding method provided in Figures 16 to 18. For a more detailed description of the above QP determination unit 2001 and the quantization unit 2002, reference can be made directly to the relevant description in the embodiment of the method shown in Figures 4 to 18, and the description will be omitted here.
[0311] The present invention further provides an electronic device, and FIG. 21 is a schematic structural diagram of the electronic device according to the present invention. As shown in FIG. 21, the electronic device 2100 includes a processor 2101 and a communication interface 2102. The processor 2101 and the communication interface 2102 are coupled to each other. It can be understood that the communication interface 2101 may be a transceiver or an input / output interface. Optionally, the electronic device 2100 may further include a memory 2103 configured to store instructions to be executed by the processor 2101, to store input data required for the processor 2101 to execute the instructions, or to store data generated after the processor 2101 executes the instructions.
[0312] When the electronic device 2100 is used to implement the methods shown in Figures 5 to 15A to 15E, the processor 2101 and the interface circuit 2102 are used to perform the functions of the above-mentioned QP determination unit 1901 and the inverse quantization unit 1902.
[0313] When the electronic device 2100 is used to implement the methods shown in FIGS. 16 to 18, the processor 2101 and the interface circuit 2102 are used to perform the functions of the QP determination unit 2001 and the quantization unit 2002 mentioned above.
[0314] In the embodiment of the present invention, the specific connection medium between the communication interface 2102, the processor 2101, and the memory 2103 is not limited. In the embodiment of the present invention, in FIG. 21, the communication interface 2102, the processor 2101, and the memory 2103 are connected via a bus 2104, and the bus is shown by a thick line in FIG. 21. The connection manner between other components is only an exemplary explanation and is not limited. The bus may be classified into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is shown in FIG. 21, but it does not indicate that there is only one bus or only one type of bus.
[0315] The memory 2103 may be configured to store software programs and modules, such as program instructions / modules corresponding to the decoding or encoding methods provided in the embodiments of the present invention. The processor 2101 executes various functional applications and data processing by executing the software programs and modules stored in the memory 2103. The communication interface 2102 may also be configured to communicate signaling and data with another device. In the present invention, the electronic device 2100 may have multiple communication interfaces 2102.
[0316] It may be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), neural processing unit (NPU), or graphic processing unit (GPU), other general-purpose processor, digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0317] The steps of the method in the embodiment of the present invention may be implemented by hardware or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, which may be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable hard disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may be a component of the processor. The processor and the storage medium may be located in an ASIC. The ASIC may also be located in a network device or a terminal device. Of course, the processor and the storage medium may exist as separate components in the network device or the terminal device.
[0318] In the above embodiment, all or part may be realized by software, hardware, firmware, or any combination thereof. When realized by using a software program, all or part may be realized in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the procedure or function according to the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or may be transmitted from one computer-readable storage medium to another, for example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) manner or a wireless (e.g., infrared, radio, microwave, etc.) manner. The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device such as a server, data center, etc. in which one or more available media are integrated. The available medium may be a magnetic medium (e.g., a floppy disk, magnetic disk, or magnetic tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0319] From the description of the above embodiments, it is obvious to those skilled in the art that for convenience and simplicity of description, only the division of each of the above functional modules is described as an example, and in actual application, the above functions can be assigned to different functional modules to be completed as necessary, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. For the specific operation processes of the above systems, devices, and units, reference can be made to the corresponding processes in the above method embodiments, and will not be repeated in this specification.
[0320] In some embodiments provided by the present invention, it should be understood that the disclosed system, device, and method can be realized in other ways. For example, the above-mentioned device embodiments are merely examples, and the division of modules or units is merely a division of logical functions, and may be divided in other ways when actually implemented. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Meanwhile, the coupling or direct coupling or communication connection between each other shown or discussed may be an indirect coupling or communication connection via some interfaces, devices or units, and may be electrical, mechanical, or other forms.
[0321] The units described as separate components may or may not be physically separated. Also, the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the objective of the solution of this embodiment.
[0322] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The integrated unit may be realized in the form of hardware or in the form of a software functional unit.
[0323] The integrated unit may be realized in the form of a software functional unit and stored in a computer-readable storage medium when sold or used as an independent product. Based on such understanding, the technical solution of the present invention may be essentially or in part contributing to the prior art, or all or part of the technical solution may be embodied in the form of a software product. The computer software product is stored in one storage medium and includes some instructions for causing a computer device (which may be a personal computer, a server, a network device, etc.) or a processor to execute all or part of the steps of the method according to each embodiment of the present invention. The above storage medium includes various media capable of storing program code, such as a flash memory, a removable hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0324] The above is merely a specific embodiment of the present invention, and the protection scope of the present invention is not limited thereto. Any changes or replacements within the technical scope disclosed in the present invention shall be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims. [Explanation of symbols]
[0325] 1. Video Codec System 10 Coding equipment 20 Decoding Device 30 Links 35 max 40 Storage device 41 Post-processing entities 42 Network Entities 100 Video Encoder 101 Converter 102 Quantizer 103 Entropy Encoder 104 Inverse quantizer 105 Inverse converter 106 Filter unit 107 Decoded Picture Buffer 108 Prediction Processing Unit 109 Intra Predictor 110 Inter Predictor 111 Adder 112 Adder 120 Video Sources 140 Output Interface 200 Video Decoder 203 Entropy Decoder 204 Inverse quantizer 205 Inverse converter 206 Filter unit 208 Prediction Processing Unit 209 Intra Predictor 210 Inter Predictor 211 Adder 220 Display device 240 Input Interface 1900 Decoding device 1901 QP Decision Unit 1902 Inverse Quantization Unit 2000 encoding device 2001 QP Decision Unit 2002 Quantization Unit 2100 Electronic equipment 2101 Processor 2102 Communication Interface 2103 Memory 2104 Bus
Claims
1. 1. An image decoding method performed by a decoding device, comprising: For any one pixel point or any plurality of parallel inverse quantization pixel points in a current coding block, determining a quantization parameter QP value of the pixel point, wherein the QP values of at least two pixel points in the current coding block are different; and dequantizing the pixel point based on a QP value of the pixel point; obtaining a QP value of the current coding block; and setting the QP value of the current coding block as a predicted QP value of the pixel point, Determining the QP value for the pixel point comprises: If the pixel point is any one of the target pixel points or any multiple parallel dequantized target pixel points in the current coding block, adjusting a predicted QP value of the pixel point, and setting the adjusted predicted QP value as a QP value of the pixel point; Adjusting the predicted QP value of the pixel point includes: Obtaining information of reconstructed pixel points around the pixel point; and adjusting a predicted QP value of the pixel point based on information of reconstructed pixel points surrounding the pixel point.
2. An image decoding method comprising:
2. Adjusting the predicted QP value of the pixel point includes: If the pixel point satisfies a first preset condition, the preset QP value is the QP value of the pixel point; otherwise, setting the predicted QP value of the pixel point as the QP value of the pixel point; The first preset condition is: the pixel point is a luminance pixel point; the pixel point is a chromaticity pixel point; the bit depth of the pixel point is less than or equal to a bit depth threshold; the predicted QP value of the pixel point is equal to or less than an adjustable QP maximum value, and the adjustable QP maximum value is equal to or less than a QP maximum value; and information of the reconstructed pixel points surrounding the pixel point is less than or equal to a first preset threshold.
2. The method of claim 1 .
3. The step of inverse quantizing the pixel point based on the QP value of the pixel point includes: and dequantizing the pixel points based on the adjusted predicted QP values.
2. The method of claim 1 .
4. the target pixel point is any one or more pixel points in the current coding block; 2. The method of claim 1 .
5. the current coding block includes at least pixel points of a first portion and / or pixel points of a second portion, and the target pixel points are any one or more pixel points of the pixel points of the second portion; 2. The method of claim 1 .
6. The target pixel point is any one or more pixel points among the pixel points at the first position in the pixel points of the second portion.
6. The method of claim 5 .
7. The target pixel point is any one or more pixel points among pixel points at a second position in the pixel points of the second portion.
6. The method of claim 5 .
8. the prediction mode of the current coding block is a pixel-wise prediction mode; the current coding block comprises at least pixel points of a first portion and / or pixel points of a second portion, the pixel points of the second portion include pixel points at the first location and / or pixel points at the second location; the pixel point at the first location and the pixel point at the second location are determined based on a pixel-wise prediction mode of the current coding block.
2. The method of claim 1 .
9. The reconstructed pixel points around the pixel point are A pixel point within a square region having the pixel point as a center and a side length of a first preset value, or a pixel point within a diamond-shaped region having the pixel point as a center and a diagonal line length of a second preset value; 2. The method of claim 1 .
10. The information of the reconstructed pixel point includes at least one of a pixel value, a reconstruction residual value, a gradient value, flatness information, texture information, or complexity information, background luminance, contrast, or motion amount of the reconstructed pixel point.
2. The method of claim 1 .
11. The value of the information of the reconstructed pixel point includes at least one of an original value, an absolute value, an average value, or a difference.
11. The method of claim 10.
12. The prediction mode of the current coding block is a block prediction mode, and the method further comprises: obtaining region division information of the current coding block, the region division information including a number N of regions and position information of region boundary lines, where N is an integer equal to or greater than 2; Dividing the current coding block into N regions based on the region division information.
2. The method of claim 1 .
13. The step of obtaining region division information of the current coding block includes: obtaining predefined region partition information of the current coding block; or Parsing a bitstream to obtain region partition information of the current coding block; 13. The method of claim 12.
14. 1. An image coding method performed by a coding device, comprising: For any one pixel point or any plurality of parallel quantization pixel points in a current coding block, determining a quantization parameter QP value for the pixel point, wherein the QP values of at least two pixel points in the current coding block are different; quantizing the pixel point based on a QP value of the pixel point; obtaining a QP value of the current coding block; and setting the QP value of the current coding block as a predicted QP value of the pixel point, Determining the QP value for the pixel point comprises: If the pixel point is any one of the target pixel points or any multiple parallel quantized target pixel points in the current coding block, adjusting a predicted QP value of the pixel point, and setting the adjusted predicted QP value as a QP value of the pixel point; Adjusting the predicted QP value of the pixel point includes: Obtaining information of reconstructed pixel points around the pixel point; and adjusting a predicted QP value of the pixel point based on information of reconstructed pixel points surrounding the pixel point.
13. An image coding method comprising:
15. A quantization parameter QP determination unit is used for determining a quantization parameter QP value of any one pixel point or any multiple parallel dequantization pixel points in a current coding block; an inverse quantization unit used for inverse quantizing the pixel point based on a QP value of the pixel point; the QP values of at least two pixel points among the pixel points in the current coding block are different; The QP determination unit is further used for obtaining a QP value of the current coding block, and setting the QP value of the current coding block as a predicted QP value of the pixel point; When determining the QP value of the pixel point, the QP determination unit specifically: If the pixel point is any one of the target pixel points or any multiple parallel dequantized target pixel points in the current coding block, a predicted QP value of the pixel point is adjusted, and the adjusted predicted QP value is used as the QP value of the pixel point; When adjusting the predicted QP value of the pixel point, the QP determination unit specifically comprises: and obtaining information of reconstructed pixel points around the pixel point and adjusting a predicted QP value of the pixel point based on the information of the reconstructed pixel points around the pixel point.
2. An image decoding device comprising:
16. A quantization parameter QP determination unit is used for determining a quantization parameter QP value of any one pixel point or any multiple parallel quantization pixel points in a current coding block; a quantization unit used to quantize the pixel point based on a QP value of the pixel point; the QP values of at least two pixel points among the pixel points in the current coding block are different; The QP determination unit is further used for obtaining a QP value of the current coding block, and setting the QP value of the current coding block as a predicted QP value of the pixel point; When determining the QP value of the pixel point, the QP determination unit specifically: If the pixel point is any one of the target pixel points or any multiple parallel quantized target pixel points in the current coding block, a predicted QP value of the pixel point is adjusted, and the adjusted predicted QP value is used as the QP value of the pixel point; When adjusting the predicted QP value of the pixel point, the QP determination unit specifically comprises: and obtaining information of reconstructed pixel points around the pixel point and adjusting a predicted QP value of the pixel point based on the information of the reconstructed pixel points around the pixel point.
1. An image encoding device comprising:
17. A video codec system comprising an encoding device and a decoding device, said encoding device communicatively connected to said decoding device, said decoding device configured to implement a method according to any one of claims 1 to 13, and said encoding device configured to implement a method according to claim 14.
1. A video codec system comprising:
18. A memory configured to store computer instructions, and a processor configured to retrieve and execute the computer instructions from the memory to perform the method of any one of claims 1 to 14.
1. An electronic device comprising:
19. A computer readable storage medium having stored thereon a computer program or instructions, which, when executed by an electronic device, performs the method according to any one of claims 1 to 14. A computer-readable storage medium comprising:
Citation Information
Patent Citations
Image encoding apparatus, image processing apparatus, and image encoding method
JP2019114868A
Image processing apparatus, image processing method, and program
JP2019201288A
DPCM with Adaptive Range and PCM Escape Mode
US20080226183A1
Multidimensional quantization techniques for video coding / decoding systems
US20180063544A1
Adaptive video signal processing apparatus
US5594679A