Scalar quantizer decision scheme for dependent quantization dependencies
The proposed scalar quantizer decision scheme in video encoding and decoding separates normal and bypass coding bins using SIG-based state transitions, enhancing throughput and coding efficiency, addressing the inefficiencies in existing technologies.
Patent Information
- Application Number
- JP2025147719
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-20
- Filing Date
- 2025-09-05
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2039-09-13
AI Technical Summary
Existing video encoding and decoding technologies face challenges in achieving high throughput and coding efficiency due to the interleaving of normal and bypass coding bins in dependent scalar quantization, leading to reduced performance in hardware implementations.
A scalar quantizer decision scheme that separates normal and bypass coding bins and uses SIG-based state transitions to determine the scalar quantizer, ensuring high throughput and coding efficiency by grouping normal coding bins first, followed by bypass bins, and using context modeling based on the selected state.
This approach maintains high throughput similar to HEVC and VTM-1 designs while improving coding efficiency, reducing bit rates by approximately 3-4% compared to VTM-1.0 anchors.
Smart Images

Figure 2026000962000001_ABST
Abstract
Description
[Technical Field]
[0001] The present embodiments generally relate to methods and apparatus for video encoding or decoding. [Background technology]
[0002] To achieve high compression efficiency, image and video coding schemes typically use prediction and transformation to exploit spatial and temporal redundancy in video content. Generally, intra- or inter-prediction is used to exploit intra- or inter-picture correlation, and then the difference between the original block and the predicted block, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to entropy coding, quantization, transformation, and prediction. Summary of the Invention [Problem to be solved by the invention]
[0003] According to an embodiment, a method of video decoding comprises: accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order; accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the decoding of the first transform coefficient; and accessing a second parameter set associated with a first transform coefficient in a block of the picture, the first transform coefficient preceding a second transform coefficient in decoding order, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the decoding of the first transform coefficient; A method is provided, comprising: entropy decoding the first and second parameter sets, wherein the first scan path is performed before other scan paths of the entropy decoded transform coefficients of the block, and each of the first and second parameter sets for the first and second transform coefficients includes at least one of a gt1 flag and a gt2 flag, wherein the gt1 flag indicates whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicates whether the absolute value of the corresponding transform coefficient is greater than 2; and reconstructing the block in response to the decoded transform coefficients.
[0004] According to an embodiment, a method of video coding comprises accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, and accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modelling of at least parameters of the second parameter set for the second transform coefficient depending on the coding of the first transform coefficient; and entropy coding the first and second parameter sets in a first scan path of a block, the first scan path being performed before other scan paths of the entropy coded transform coefficients of the block, each of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 2.
[0005] According to another embodiment, there is provided an apparatus for video decoding, comprising one or more processors, wherein the one or more processors access a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order, and access a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in a normal mode, context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the decoding of the first transform coefficient, and a first skip of the block of the picture. an apparatus configured to entropy decode the first and second parameter sets in a scan pass, the first scan pass being performed before other scan passes of entropy decoded transform coefficients of the block, each of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2, and reconstructing the block in response to the decoded transform coefficients. The apparatus may further include one or more memories coupled to the one or more processors.
[0006] According to another embodiment, there is provided an apparatus for video coding comprising one or more processors, the one or more processors accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, and accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on coding of the first transform coefficient. and entropy coding the first and second parameter sets in a first scan path of the block, the first scan path being performed before other scan paths of the entropy coded transform coefficients of the block, and each of the first and second parameter sets for the first and second transform coefficients includes at least one of a gt1 flag and a gt2 flag, the gt1 flag being configured to indicate whether an absolute value of a corresponding transform coefficient is greater than 1 and the gt2 flag being configured to indicate whether an absolute value of the corresponding transform coefficient is greater than 2.
[0007] According to another embodiment, an apparatus for video decoding comprises: means for accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order; means for accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on decoding of the first transform coefficient; and means for accessing a first parameter set associated with a first transform coefficient in a block of the picture in decoding order; An apparatus is provided, comprising: means for entropy decoding the first and second parameter sets, wherein the first scan path is performed before other scan paths of the entropy decoded transform coefficients of the block, and each set of the first and second parameter sets for the first and second transform coefficients includes at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 2; and means for reconstructing the block in response to the decoded transform coefficients.
[0008] According to another embodiment, an apparatus for video coding comprises means for accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, means for accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode and context modelling of at least parameters of the second parameter set for the second transform coefficient depending on the coding of the first transform coefficient, and means for accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode and and means for entropy coding the first and second parameter sets in a first scan path of a block, the first scan path being performed before other scan paths of entropy coding transform coefficients of the block, each set of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2.
[0009] According to another embodiment, a signal comprising coded video includes accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, and accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the coding of the first transform coefficient, and accessing the block. a signal is formed by entropy coding the first and second parameter sets in a first scan path of a block, the first scan path being performed before other scan paths of the entropy coded transform coefficients of the block, each set of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 2. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 shows a block diagram of a system in which aspects of the present embodiment can be implemented. [Figure 2] FIG. 2 shows a block diagram of an embodiment of a video encoder. [Figure 3] FIG. 3 shows a block diagram of an embodiment of a video decoder. [Figure 4] FIG. 4 is an example diagram showing two scalar quantizers used in the dependent quantization proposed in JVET-J0014. [Figure 5] FIG. 5 is an example diagram showing state transitions and quantizer selection for dependent quantization proposed in JVET-J0014. [Figure 6]FIG. 6 is an example of a diagram showing the order of coefficient bins in CG, as proposed in JVET-J0014. [Figure 7] Figure 7 is an example diagram showing SIG-based state transitions proposed in JVET-K0319. [Figure 8] FIG. 8 is an example of a diagram showing the order of coefficient bins in CG, as proposed in JVET-K0319. [Figure 9] FIG. 9 is an example diagram illustrating the order of coefficient bins in a CG according to one embodiment. [Figure 10] FIG. 10 is an example diagram illustrating a state transition based on SUM(SIG, gt1, gt2) according to one embodiment. [Figure 11] FIG. 11 is an example diagram illustrating XOR(SIG, gt1, gt2) based state transitions according to one embodiment. [Figure 12] FIG. 12 is an example diagram illustrating gt1-based state transitions according to one embodiment. [Figure 13] FIG. 13 illustrates a process for encoding a current block according to one embodiment. [Figure 14] FIG. 14 illustrates a process for decoding a current block according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] FIG. 1 shows a block diagram of an example system in which various aspects and embodiments can be implemented. System 100 can be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, connected home appliances, and servers. The elements of system 100, either singly or in combination, can be embodied in a single integrated circuit, multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or separate components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described herein.
[0012] System 100 includes at least one processor 110 configured to execute instructions loaded therein, for example, to implement various aspects described herein. Processor 110 may include embedded memory, input / output interfaces, and various other circuits as known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes storage 140, which may include non-volatile and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. Storage 140 may include, by way of non-limiting example, internal storage, removable storage, and / or network-accessible storage.
[0013] System 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 130 represents a module or modules that may be included in a device that performs encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Furthermore, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software known to those skilled in the art.
[0014] Program code to be loaded into processor 110 or encoder / decoder 130 to perform aspects described herein may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during execution of processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, expressions, operations, and computational logic.
[0015] In some embodiments, memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing unit (e.g., the processing unit can be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140, and can be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as MPEG-2, HEVC, or VVC.
[0016] Input to the elements of system 100 can be provided via various input devices as shown in block 105. Such input devices include, but are not limited to, (i) an RF section that receives RF signals transmitted over the air by, for example, a broadcaster, (ii) a composite input, (iii) a USB input, and / or (iv) an HDMI input.
[0017] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF section can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to as a channel (for example), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section can include, for example, a tuner that performs various of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and filtering again to the desired frequency band an RF signal transmitted over a wired (e.g., cable) medium. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0018] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, as desired, within a separate interface IC or within processor 110. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operates in combination with memory and storage elements to process the data stream as desired for display on an output device.
[0019] The various elements of system 100 may be provided within an integrated housing in which the various elements are interconnected and may transmit data therebetween using a suitable connection arrangement 115, e.g., an internal bus as known in the art, including an I2C bus, wiring, and a printed circuit board.
[0020] System 100 includes a communication interface 150 that enables communication with other devices over a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented in a wired and / or wireless medium, for example.
[0021] In various embodiments, data is streamed to system 100 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received via communication channel 190 and communication interface 150, which are adapted for Wi-Fi communication. Communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 100 using a set-top box that delivers data via an HDMI connection in input block 105. Still other embodiments provide streamed data to system 100 using an RF connection in input block 105.
[0022] System 100 can provide output signals to various output devices, including display 165, speakers 175, and other peripherals 185. In various example embodiments, other peripherals 185 include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, control signals are transmitted between system 100 and display 165, speakers 175, or other peripherals 185 using signaling such as AV Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 using communication channel 190 via communication interface 150. Display 165 and speakers 175 can be integrated in a single unit with other components of system 100 within an electronic device, such as a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
[0023] Display 165 and speakers 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which display 165 and speakers 175 are external components, output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a composite (COMP) output.
[0024] Figure 2 shows an example of a video encoder 200, such as a High Efficiency Video Coding (HEVC) encoder. Figure 2 may also show an encoder that improves on the HEVC standard or employs technology similar to HEVC, such as the Versatile Video Coding (VVC) encoder under development by the Joint Video Exploration Team (JVET).
[0025] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "encoded" or "coded" can be used interchangeably, and the terms "image," "picture," and "frame" can be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side, while "decoded" is used on the decoder side.
[0026] Before being encoded, the video sequence may undergo a pre-encoding process (201), such as applying a color transformation to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and may be attached to the bitstream.
[0027] To encode a video sequence with one or more pictures, a picture is divided into one or more slices, where each slice can contain one or more slice segments (202). In HEVC, slice segments are organized into coding units, prediction units, and transform units. The HEVC specification distinguishes between "blocks" and "units," where a "block" addresses a specific region of a sample array (e.g., luma, Y), and a "unit" includes an ordered block of all coded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data (e.g., motion vectors) associated with the block.
[0028] In HEVC coding, a picture is divided into square coding tree blocks (CTBs) with a configurable size (typically 64x64, 128x128, or 256x256 pixels), and contiguous sets of coding tree blocks are grouped into slices. A coding tree unit (CTU) contains the CTB of a coded color component. The CTB is the root of a quadtree division into coding blocks (CBs), which can be divided into one or more prediction blocks (PBs), forming the root of a quadtree division into transform blocks (TBs). Transform blocks (TBs) larger than 4x4 are divided into 4x4 subblocks of quantized coefficients called coefficient groups (CGs). Corresponding to the coding blocks, prediction blocks, and transform blocks, a coding unit (CU) includes a prediction unit (PU) and a tree-structured set of transform units (TUs), where a PU contains prediction information for all color components and a TU contains the residual coding syntax structure for each color component. The sizes of the CB, PB, and TB of the luma components apply to the corresponding CU, PU, and TU. In this application, the term "block" may be used to refer to, for example, any of CTU, CU, PU, TU, CG, CB, PB, and TB. Furthermore, "block" may also be used to refer to macroblocks and partitions defined in H.264 / AVC or other video coding standards, or more generally to refer to arrays of data of various sizes.
[0029] In encoder 200, pictures are coded by encoder elements as described below. A picture to be coded is processed, for example, in units of CUs. Each coding unit is coded using either intra mode or inter mode. When a coding unit is coded in intra mode, intra prediction is performed (260). In inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines (205) whether intra mode or inter mode is used to code the coding unit and indicates the intra / inter decision with a prediction mode flag. A prediction residual is calculated (210) by subtracting the prediction block from the original image block.
[0030] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as the motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. As a non-limiting example, context-based adaptive binary arithmetic coding (CABAC) can be used to code the syntax elements into the bitstream.
[0031] To encode with CABAC, the values of non-binary syntax elements are mapped to a binary sequence called a bin string via a binarization process. For a bin, a context model is selected. A "context model" is a probability model for one or more bins, selected from a selection of available models depending on the statistics of recently coded symbols. The context model for each bin is identified by a context model index (also called a "context index"), with different context indexes corresponding to different context models. A context model stores the probability that each bin is "1" or "0" and can be adaptive or static. A static model triggers the coding engine with equal probability for bins "0" and "1." In the adaptive coding engine, the context model is updated based on the actual coded value of the bin. The operating modes corresponding to the adaptive and static models are called normal and bypass modes, respectively. Based on the context, the binary arithmetic coding engine encodes or decodes the bin according to the corresponding probability model.
[0032] A scan pattern transforms a two-dimensional block into a one-dimensional array and defines the processing order of the samples or coefficients. A scan path is an iteration over the transform coefficients in a block (according to a selected scan pattern) to code a particular syntax element.
[0033] In HEVC, the scan path across a TB then consists of processing each CG in turn according to the scan pattern (diagonal, horizontal, vertical), and the 16 coefficients within each CG are scanned according to the same considered scan order. The scan starts from the last significant coefficient of the TB and processes all coefficients up to the DC coefficient. The CGs are scanned sequentially. Up to five scan paths are applied to a CG. Each scan path encodes the syntax elements of the coefficients within the CG as follows: · Significant Coefficient Flag (SIG, significant_coeff_flag): Significance of the coefficient (zero / non-zero). · Coefficient absolute level flag greater than 1 (gt1, coeff_abs_level_greater1_flag): Indicates whether the absolute value of the coefficient level is greater than 1. · Coefficient absolute level greater than 2 flag (gt2, coeff_abs_level_greater2_flag): Indicates whether the absolute value of the coefficient level is greater than 2. · Coefficient sign flag (coeff_sign_flag): Sign of significant coefficient (0: positive, 1: negative). Coefficient remaining absolute level (coeff_abs_level_remaining): The remaining value of the coefficient level absolute value (if the value is greater than the value coded in the previous pass). For a particular CG, up to 8 coeff_abs_level_greater1_flag can be coded, and coeff_abs_level_greater2_flag is coded only for the first coefficient of the CG whose magnitude is greater than 1.
[0034] In each scan path, syntax is coded only if necessary, as determined by the previous scan path. For example, if a coefficient is not significant, the remaining scan paths are not needed for that coefficient. The bins of the first three scan paths are coded in normal mode, and the context model index depends on the position of the particular coefficient in the TB and the value of nearby previously coded coefficients covered by the local template. The bins of scan paths 4 and 5 are coded in bypass mode, so that all bypass bins in the CG are grouped together.
[0035] The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder can also bypass both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process. In direct PCM coding, no prediction is applied and coded unit samples are coded directly into the bitstream.
[0036] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (255) to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture, for example, to perform deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).
[0037] Figure 3 shows a block diagram of an example video decoder 300, such as an HEVC decoder. In the decoder 300, a bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs a decoding pass that is the inverse of the encoding pass as described in Figure 2, which performs video decoding as part of encoding the video data. Figure 3 may also show a decoder that improves on the HEVC standard, such as a VVC decoder, or a decoder that employs HEVC-like technology.
[0038] In particular, the decoder input includes a video bitstream, such as may be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, picture partitioning information, and other coded information. If CABAC is used for entropy coding, a context model is initialized in the same manner as the encoder context model, and syntax elements are decoded from the bitstream based on the context model.
[0039] The picture partitioning information indicates how a picture is divided, e.g., the size of CTUs and how the CTUs are possibly divided into CUs and, if applicable, PUs. Thus, the decoder can, for example, divide a picture into CTUs (335) and divide each CTU into CUs according to the decoded picture partitioning information. The transform coefficients are inverse quantized (340) and inverse transformed (350) to decode the prediction residual.
[0040] The decoded prediction residual is combined with the prediction block (355) to reconstruct an image block. The prediction block can be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0041] The decoded picture may further undergo a post-decoding process (385), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (201). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0042] Dependent scalar quantization, proposed in an article titled "Description of SDR, HDR and 360° video coding technology proposed by Fraunhofer HHI," Document JVET-J0014, 10th Meeting: San Diego, USA, April 10-20, 2018 (hereinafter "JVET-J0014"), involves switching between two scalar quantizers with different reconstruction levels for quantization. Compared to traditional independent scalar quantization (used in HEVC and VTM-1), the set of allowable reconstruction values for a transform coefficient depends on the value of the transform coefficient level preceding the current transform coefficient level in the reconstruction order.
[0043] The dependent scalar quantization approach is realized by (a) defining two scalar quantizers with different reconstruction levels, and (b) defining a process for switching between the two scalar quantizers.
[0044] The two scalar quantizers used, denoted by Q0 and Q1, are shown in Figure 4. The position of the available reconstruction levels is uniquely specified by the quantization step size Δ. Ignoring the fact that the actual reconstruction of the transform coefficients uses integer arithmetic, the two scalar quantizers Q0 and Q1 can be characterized as follows: Q0: The reconstruction levels of the first quantizer Q0 are given by even integer multiples of the quantization step size Δ. When this quantizer is used, the reconstructed transform coefficients t′ are calculated according to: t'=2·k·Δ, where k denotes the associated transform coefficient level. Note that the term "transform coefficient level" (k) refers to the quantized transform coefficient value, e.g., it corresponds to TransCoeffLevel as described in the residual_coding syntax structure below. The term "reconstructed transform coefficient" (t') refers to the dequantized transform coefficient value. Q1: The reconstruction levels of the second quantizer Q1 are given by odd integer multiples of the quantization step size Δ, and furthermore, the reconstruction levels are equal to zero. The mapping of transform coefficient levels k to reconstructed transform coefficients t′ is specified by: t'=(2·k-sgn(k))·Δ, Here, sgn(·) denotes the sign function, and sgn(x)=(k==0?0:(k<0?-1:1)).
[0045] The scalar quantizer used (Q0 or Q1) is not explicitly signaled in the bitstream. Instead, the quantizer used for the current transform coefficient is determined by the parity of the transform coefficient level that precedes the current transform coefficient in the coding / reconstruction order.
[0046] As shown in FIG. 5, switching between the two scalar quantizers (Q0 and Q1) is performed via a state machine with four states. The state can take on four different values: 0, 1, 2, and 3. The state is uniquely determined by the parity of the transform coefficient level preceding the current transform coefficient in the coding / reconstruction order. At the start of dequantization of a transform block, the state is set to 0. The transform coefficients are reconstructed in scan order (i.e., in the same order as they are entropy coded / decoded). After the current transform coefficient is reconstructed, the state is updated as shown in Table 1, where k denotes the value of the transform coefficient level. Note that the next state depends only on the current state and the parity of the current transform coefficient level k (k & 1). If k represents the value of the current transform coefficient level, the state update can be described as follows: state = stateTransTable[state][k&1], where stateTransTable represents the table shown in FIG. 5 and Table 1, and the operator & specifies the bitwise "and" operator in two's complement arithmetic.
[0047] The state uniquely specifies the scalar quantizer to be used. If the state for the current transform coefficient is equal to 0 or 1, scalar quantizer Q0 is used. Otherwise (state is equal to 2 or 3), scalar quantizer Q1 is used. [Table 1]
[0048] More generally, the quantizer for the transform coefficients can be selected from three or more scalar quantizers, the state machine can have five or more states, or quantizer switching can be handled via other possible mechanisms.
[0049] A coefficient coding scheme combined with dependent scalar quantization was also proposed in JVET-J0014, whereby the context modeling of a quantized coefficient depends on the quantizer used. Specifically, the significance flag (SIG) and the greater-than-one flag (gt1) each have two sets of context models, and the set selected for a particular SIG or gt1 depends on the quantizer used for the associated coefficient. Therefore, the coefficient coding proposed in JVET-J0014 requires a complete reconstruction of the absolute level (absLevel) of a quantized coefficient before moving to the next scan position in order to know the parity used to determine the quantizer and, therefore, the context set for the next coefficient. That is, entropy decoding of coefficient n (SIG, gt1, . . . , gt4, sign flag, absolute residual level) must be completed to obtain the context model for entropy decoding coefficient (n+1). As a result, some normal coding bins of coefficient (n+1) need to wait for the decoding of some bypass coding bins of coefficient n, and thus the bypass coding bins of different coefficients are interleaved with the normal coding bins as shown in FIG. 6.
[0050] As shown in Figure 6, the coefficient coding design in JVET-J0014 may reduce throughput compared to the HEVC or VTM-1 designs, as explained below: 1. More regular coding bins. Regular coding bins are slower than bypassed coding bins due to context selection and interval subdivision calculations. In JVET-J0014 coefficient coding, the number of regular coding bins for a coefficient group (CG) is up to 80 compared to 25 in VTM-1 (16 for SIG, 8 for GT1, and 1 for GT2). 2. Bypass coded bins are not grouped. Grouping bypass bins into longer chains increases the number of bins processed per cycle, thereby reducing the number of cycles required to process a single bypass bin. However, in JVET-J0014 coefficient coding, bypass coded bins for CGs are not grouped; instead, they are interleaved with the normal coded bins for each coefficient.
[0051] This application is directed to a scalar quantizer decision scheme that achieves approximately the same level of throughput as the coefficient coding designs of HEVC and VTM-1, while retaining most of the gains provided by dependent scalar quantization.
[0052] In contribution JVET-K0319 ("CE7-Related:TCQ with High Throughput Coefficient Coding," Document JVET-K0319, JVET 11th Meeting: Ljubljana, Slovenia, July 10-18, 2018, hereafter referred to as "JVET-K0319"), the parity-based state transitions proposed in JVET-J0014 are replaced with SIG-based state transitions as shown in Figure 7 and Table 2, while the other dependent scalar quantization designs in JVET-J0014 remain unchanged. By doing this, the scalar quantizer used to quantize the current transform coefficient is determined by the SIG of the quantized coefficient preceding the current transform coefficient in scan order. [Table 2]
[0053] The coefficient coding proposed in JVET-K0319 is based on HEVC and VTM-1 coefficient coding. The difference is that SIG and gt1 each have two sets of context models, and the entropy coder selects a specific SIG or gt1 context set according to the quantizer used by the associated coefficient. Therefore, changing the scalar quantizer of the dependent scalar quantization from parity-based to SIG-based enables a high-throughput design similar to HEVC and VTM-1 coefficient coding. The proposed order of coefficient bins in CG is shown in Figure 8. That is, in the coefficient coding proposed in JVET-K0319, the normal coding bins are up to 25 per CG, which remains the same as HEVC and VTM-1 coefficient coding, and all bypass coding bins within a CG are grouped together.
[0054] The dependent quantization approach in JVET-J0014 was tested with CE7's Test 7.2.1 software, and the simulation results showed a 4.99% AI (all intra), 3.40% RA (random access), and 2.70% LDB (low latency B) BD rate reduction compared to the VTM-1.0 anchor. However, the simulation results for JVET-K0319 showed a 3.98% AI, 2.58% RA, and 1.80% LDB BD rate reduction compared to the VTM-1.0 anchor. That is, the scalar quantizer used to quantize transform coefficients in JVET-K0319 (switching based only on SIG) may reduce coding efficiency compared to the one proposed in JVET-J0014 (switching based on full absLevel of quantized coefficients).
[0055] This application proposes several alternative determination schemes for the scalar quantizer used for dependent scalar quantization to achieve a suitable trade-off between high throughput and coding efficiency. Instead of using absolute levels or parity of SIG values, state transitions and context model selection based on regular coding bins are proposed. Below, several embodiments for determining the scalar quantizer used for dependent scalar quantization are described.
[0056] In the case of the dependent scalar quantization proposed in JVET-J0014, the absolute level (absLevel) of a quantized coefficient needs to be fully reconstructed before moving to the next scan position in order to know the parity used to determine the quantizer for the next coefficient. Therefore, bypass coding bins within a CG are not grouped; they are interleaved with the normal coding bins for each coefficient. Furthermore, as shown in the syntax table below, in contrast to HEVC and VTM-1, the maximum number of normal coding bins per transform coefficient level is increased (in the approach proposed in JVET-J0014, up to five normal coding bins per transform coefficient level can occur). Changes relative to VTM-1 are italicized. The entropy coder selects a particular SIG or gt1 context set according to the "state" used to determine the quantizer, depending on the transform coefficient level information. The bin coding order is shown by the following syntax, where the function getSigCtxId(xC, yC, state) is used to derive the context of the syntax sig_coeff_flag based on the current coefficient scan position (xC, yC) and state, decodeSigCoeffFlag(sigCtxId) is for decoding the syntax sig_coeff_flag with the associated context sigCtxId, getGreater1CtxId(xC, yC, state) is used to derive the context of the syntax abs_level_gt1_flag based on the current coefficient scan position (xC, yC) and state, and decodeAbsLevelGt1Flag(greater1CtxId) is for decoding the syntax abs_level_gt1_flag with the associated context greater1CtxId. [Table 3] JPEG2026000962000005.jpg244170
[0057] As previously explained, these syntax changes pose potential problems for high-throughput hardware implementation. In our proposal, we propose an alternative approach to achieve nearly the same level of throughput as HEVC while supporting dependent scalar quantization. Below, we use HEVC as an example to illustrate the proposed changes.
[0058] In one embodiment, the maximum number of normal coding bins per transform coefficient level is kept at 3 instead of 5 (SIG, gt1, and gt2 are normal coded). For each CG, the normal coding bins and bypass coding bins are separated in coding order, with all normal coding bins of the CG transmitted first, followed by the bypass coding bins. The proposed ordering of coefficient bins in a CG is shown in Figure 9. The bins of a CG are coded in multiple passes across the scan positions of the CG. Pass 1: Coding of significance (SIG, sig_coeff_flag), flags greater than 1 (gt1, abs_level_gt1_flag), and flags greater than 2 (gt2, abs_level_gt2_flag), in coding order. Flags greater than 1 are only present if sig_coeff_flag is equal to 1. Coding of flags greater than 2 (abs_level_gt2_flag) is only performed for scan positions where abs_level_gt1_flag is equal to 1. The values of gt1 and gt2 are inferred to be 0 if they are not present in the bitstream. The SIG, gt1, and gt2 flags are coded in normal mode, and the choice of context modeling for SIG depends on which state is selected for the associated coefficient. Pass 2: Coding of the syntax element abs_level_remaining for all scan positions where abs_level_gt2_flag is equal to 1. Non-binary syntax elements are binarized and the resulting bins are coded in the bypass mode of the arithmetic coding engine. Pass 3: Encoding the signs (coeff_sign_flag) of all scan positions where sig_coeff_flag is equal to 1. The signs are encoded in bypass mode.
[0059] The above embodiments show proposed modifications compared to HEVC. The modifications can also be based on other solutions. For example, if JVET-J0014 is used as a base, the context modeling of both SIG and gt1 depends on the quantizer selection, and flags greater than x (gtx, where x=3 and 4) can be coded in pass 1 or pass 2. If other normal coding bins exist, such as gt5, gt6, and gt7 flags, they can be coded in pass 1 or pass 2. Furthermore, the sign of pass 3 (coeff_sign_flag) can also be coded in normal mode.
[0060] At the March 2019 meeting, JVET adopted a new residual coding process for transform skip residual blocks. When transform skip (TS) is enabled, the transform of the prediction residual is skipped. The residual levels of a coefficient group (CG) are coded as follows during three passes through the scan position: Pass 1: The following flags are signaled: osig_coeff_flag ocoeff_sign_flag Flags greater than o1 (abs_level_gtx_flag[0]) o Parity (par_level_flag) flag Pass 2: The following flags are signaled: Flags greater than o3 (abs_level_gtx_flag[1]) Flags greater than o5 (abs_level_gtx_flag[2]) Flags greater than o7 (abs_level_gtx_flag[3]) Flags greater than o9 (abs_level_gtx_flag[4]) Pass 3: Use Golomb-Rice coding, bypassing coding of the remaining absolute levels (abs_remainder).
[0061] The above proposed embodiment can also be applied to this newly adopted TS residual coding, for example, flag positions greater than 3 are moved to the first pass as shown below: Pass 1: The following flags are signaled: osig_coeff_flag ocoeff_sign_flag Flags greater than o1 (abs_level_gtx_flag[0]) o Parity (par_level_flag) flag Flags greater than o3 (abs_level_gtx_flag[1]) Pass 2: The following flags are signaled: Flags greater than o5 (abs_level_gtx_flag[2]) Flags greater than o7 (abs_level_gtx_flag[3]) Flags greater than o9 (abs_level_gtx_flag[4]) Pass 3: Use Golomb-Rice coding, bypassing coding of the remaining absolute levels (abs_remainder).
[0062] Embodiment 1 - Scalar quantizer decision scheme based on function SUM(SIG, gt1, gt2)
[0063] To solve the problem of high-throughput hardware implementation, a complete reconstruction of the absolute level (absLevel) of the quantized coefficients is not performed to determine the state, and the switching between the two scalar quantizers does not depend on the parity of the absolute level of the complete transform coefficients. As mentioned above, a scalar quantizer determined solely by SIG, as in JVET-K0319, may reduce coding efficiency. In one embodiment, we propose to determine the scalar quantizer based on the function SUM(SIG, gt1, gt2), which takes into account both the SIG, gt1, and gt2 values of the current transform coefficient. [Table 4]
[0064] The possible combinations of the conversion coefficients SIG, gt1, and gt2 values are shown in Table 3. There is a one-to-one correspondence mapping from these four different combinations to four possible marking level values, where m denotes the marking value for these four possible cases. The function to derive the marking value m from the values of SIG, gt1, and gt2 can be written as follows: m=SUM(SIG,gt1,gt2)=SIG+gt1+gt2.
[0065] As shown in Figure 10, the switch between the two scalar quantizers is uniquely determined by the parity of the marking value m of the transform coefficient level preceding the current transform coefficient in the coding / reconstruction order. At the start of the inverse quantization of a transform block, the state is set to 0. After the normal coding bin of the current transform coefficient is reconstructed, the state is updated as shown in Figure 10 and Table 4. Note that the next state depends only on the current state and the parity (m & 1) of the marking value m of the current transform coefficient level. The state update can be described as follows: state = stateTransTable[state][m&1], where stateTransTable represents the table shown in FIG. 10 and Table 4, and the operator & specifies the bitwise "and" operator in two's complement arithmetic. [Table 5] [Table 6] JPEG2026000962000009.jpg95170
[0066] After transformation and quantization, the magnitude of most transform coefficients is usually very low. As shown in Tables 3 and 4, when the absolute level of the transform coefficients is less than 3, the proposed method can achieve almost the same results as JVET-J0014. Meanwhile, the proposed method does not require complete reconstruction of the absolute level (absLevel) of the quantized coefficients, which solves the problem of high-throughput hardware implementation.
[0067] The coding order and presence of bins, as well as the details of the reconstruction of transform coefficient levels from the transmitted data, are shown in the syntax table above. For ease of explanation, the different paths of scan positions are commented out in the syntax table. Changes relevant to HEVC and VTM-1 are shown in italics. The entropy coder selects the context set for a particular SIG according to the "state" that depends on the transform coefficient level information and is used to determine the quantizer.
[0068] Embodiment 2 - Scalar quantizer decision scheme based on the function XOR(SIG, gt1, gt2)
[0069] In another embodiment, the SIG, gt1, and gt2 values of the current transform coefficient are considered and a scalar quantizer is selected based on the function XOR(SIG, gt1, gt2). The function to derive the exclusive-or value x from the SIG, gt1, and gt2 values can be written as follows: x=XOR(SIG,gt1,gt2)=SIG^gt1^gt2. The exclusive OR values x corresponding to the possible combinations of values of the transform coefficients SIG, gt1, and gt2 are presented in Table 5. [Table 7]
[0070] Compared to the first embodiment, the switching between the two scalar quantizers is uniquely determined by the exclusive OR value x of the SIG, gt1 and gt2 flags. The state update can be written as follows: state = stateTransTable[state][x], Here, stateTransTable represents the table shown in Figure 11 and Table 6. The rest of the state machine remains similar to the previous approach. [Table 8]
[0071] As shown in Tables 5 and 6, when the absolute level of the transform coefficients is less than 3, the proposed method can achieve almost the same results as JVET-J0014. Meanwhile, the proposed method does not require a complete reconstruction of the absolute levels (absLevel) of the quantized coefficients, which solves the problem of high-throughput hardware implementation.
[0072] Embodiment 3 - Scalar quantizer decision scheme based on one of the regular coding bins
[0073] According to the above embodiment, all normal coding bins of the current transform coefficient (SIG, gt1, and gt2) are considered to determine the scalar quantizer. In another embodiment, the switching between two scalar quantizers can be based on one of the normal coding bins, for example, the gt1 flag. In this embodiment, as shown in FIG. 12, the previous state transition can be replaced by the gt1-based state transition, while the other design remains unchanged. Thereby, the scalar quantizer used to quantize the current transform coefficient is determined by the gt1 flag of the quantized coefficient preceding the current transform coefficient in scan order. The state update can be described as follows: state = stateTransTable[state][gt1], Here, stateTransTable represents the table shown in FIG.
[0074] In this embodiment, the CG bins are coded with three scan paths over the following scan positions in the CG: a first pass of sig, gt1, and gt2, a second pass of the remaining absolute levels, and a third pass of the sign information. In a variant, the CG bins are coded with four scan paths over the following scan positions in the CG: a first pass of sig and gt1, a second pass of gt2, a third pass of the remaining absolute levels, and a fourth pass of the sign information. This variant can further reduce the dependency between bins compared to the three scan paths proposed in the previous embodiment. [Table 9]
[0075] Alternatively, the scalar quantizer used to quantize the current transform coefficient is determined by the gt2 flag of the quantized coefficient preceding the current transform coefficient in scan order. More generally, the scalar quantizer used to quantize the current transform coefficient is determined by one regular coding bin (e.g., the gtx flag) of the quantized coefficient preceding the current transform coefficient in scan order.
[0076] The above examples show some embodiments based on HEVC that use three normal coding bins (SIG, gt1, gt2) for coefficients. If the normal coding bins of the coefficients are different from those of HEVC, the proposed embodiments can be implemented by taking a different number (more or less than 3) of normal coding bins per transform coefficient.
[0077] In the above, sum and exclusive-or functions are considered, the proposed embodiment can also be implemented using different state update derived (1 / 0) functions from the normal coding bins for each transform coefficient level.
[0078] Above, the description is mainly about inverse quantization. Note that the quantization is adjusted accordingly. The scalar quantizer used (Q0 or Q1) is not explicitly signaled in the bitstream. For example, the quantization module on the encoder side selects the quantizer to use for the current transform coefficient based on its state. If the state of the current transform coefficient is equal to 0 or 1, the scalar quantizer Q0 is used. Otherwise (the state is equal to 2 or 3), the scalar quantizer Q1 is used. The state is uniquely determined by the transform coefficient level information using the methods described in Figures 10 to 12, and the decoder selects the same quantizer to properly decode the bitstream.
[0079] 13 illustrates a method (1300) for encoding a current coding unit according to an embodiment. In step 1305, an initial state is set to zero. To encode a coding unit, the coefficients within the coding unit are scanned. The scan path of the coding unit processes each CG of the coding unit in turn according to a scan pattern (diagonal, horizontal, vertical), and the coefficients within each CG are also scanned according to a similarly considered scan order. The scan starts (1315) with the last CG with a significant coefficient of the coding unit and processes all coefficients up to the first CG with a DC coefficient.
[0080] If the CG does not contain a last significant or DC coefficient (1320), a flag (coded_sub_block_flag) indicating whether the CG contains non-zero coefficients is coded (1325). For CGs that contain a last non-zero level or DC coefficient, coded_sub_block_flag is inferred to be equal to 1 and is not displayed in the bitstream.
[0081] If coded_sub_block_flag is true (1330), three scan paths are applied to the CG. In the first pass (1335-1360), for the coefficient, the SIG flag (sig_coeff_flag) is coded (1335). To code the SIG flag, a context mode index is determined using a state, e.g., sigCtxId=getSigCtxId(state). If the SIG flag is true (1340), the gt1 flag (abs_level_gt1_flag) is coded (1345). If the gt1 flag is true (1350), the gt2 flag (abs_level_gt2_flag) is coded (1355). Based on one or more of the SIG, gt1, and gt2 flags, the state is updated (1360), e.g., using the methods described in Figures 10-12.
[0082] In the second scan path (1365, 1370), the encoder checks whether the gt2 flag is true (1365). If true, the remaining absolute level (abs_level_remaining) is coded (1370). In the third scan path (1375, 1380), the encoder checks whether the SIG flag is true (1375). If true, the sign flag (coeff_sign_flag) is coded (1380). In step 1385, the encoder checks whether there are more CGs to be processed. If so, it moves to the next CG to be processed (1390).
[0083] 14 illustrates a method (1400) for decoding a current coding unit according to an embodiment. In step 1405, an initial state is set to zero. Similar to the encoder side, to decode a coding unit, the coefficient positions within the coding unit are scanned. The scan starts with the last significant coefficient of the coding unit (1415) and processes all coefficients up to the first significant coefficient of the coding unit.
[0084] If the CG does not contain a last significant coefficient or a DC coefficient (1420), a flag (coded_sub_block_flag) indicating whether the CG contains any non-zero coefficients is decoded (1425). For a CG that contains a last non-zero level or a DC coefficient, coded_sub_block_flag is inferred to be equal to 1.
[0085] If coded_sub_block_flag is true (1430), three scan paths are applied to the CG. In the first pass (1435-1460), the SIG flag (sig_coeff_flag) is decoded for the coefficient (1435). To decode the SIG flag, a context mode index is determined using the state, e.g., sigCtxId=getSigCtxId(state). If the SIG flag is true (1440), the gt1 flag (abs_level_gt1_flag) is decoded (1445). If the gt1 flag is true (1450), the gt2 flag (abs_level_gt2_flag) is decoded (1455). Based on one or more of the SIG, gt1, and gt2 flags, the state is updated (1460), e.g., using the methods described in Figures 10-12.
[0086] In the second scan path (1470, 1475), the decoder checks whether the gt2 flag is true (1470). If true, the remaining absolute level (abs_level_remaining) is decoded (1475). In the third scan path (1480, 1485), the decoder checks whether the SIG flag is true (1480). If true, the sign flag (coeff_sign_flag) is decoded (1485). In step 1487, the decoder calculates transform coefficients based on the available SIG, gt1, gt2, sign flag, and the remaining absolute value.
[0087] The decoder checks (1490) whether there are any more CGs to be processed. If so, it moves to the next CG to be processed (1495). If all coefficients are entropy decoded, the transform coefficients are dequantized (1497) using dependent scalar quantization. The scalar quantizer (Q0 or Q1) used for the transform coefficients is determined by the state, which, together with information about the decoded transform coefficient levels, is derived using the methods described in Figures 10-12.
[0088] Various methods are described herein, each of which comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be varied or combined. Furthermore, terms such as “first,” “second,” etc. may be used in various embodiments to vary elements, components, steps, actions, etc., e.g., “first decode” and “second decode.” The use of such terms does not imply a varied order of actions, unless specifically required. Thus, in this example, the first decode need not be performed before the second decode, but could occur, for example, before, during, or during an overlapping period with the second decode.
[0089] Various methods and other aspects described herein can be used to modify modules, such as the entropy encoding and decoding modules (245, 330) of video encoder 200 and decoder 300, as shown in Figures 2 and 3. Furthermore, the aspects are not limited to VVC or HEVC, but can be applied to, for example, other standards and recommendations, and any such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described herein can be used individually or in combination.
[0090] Various numerical values are used in this application. The particular values are for illustrative purposes and the described aspects are not limited to these particular values.
[0091] According to an embodiment, a method of video decoding comprises accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order, and accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the decoding of the first transform coefficient, and a first parameter set associated with a first transform coefficient in a first scan path of the block. A method is provided, comprising: entropy decoding the first and second parameter sets, wherein the first scan path is performed before other scan paths of the entropy decoded transform coefficients of the block, and each of the first and second parameter sets for the first and second transform coefficients includes at least one of a gt1 flag and a gt2 flag, wherein the gt1 flag indicates whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicates whether the absolute value of the corresponding transform coefficient is greater than 2; and reconstructing the block in response to the decoded transform coefficients.
[0092] According to an embodiment, a method of video coding comprises accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, and accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modelling of at least parameters of the second parameter set for the second transform coefficient depending on the coding of the first transform coefficient; and entropy coding the first and second parameter sets in a first scan path of a block, the first scan path being performed before other scan paths of the entropy coded transform coefficients of the block, each of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 2.
[0093] According to another embodiment, there is provided an apparatus for video decoding, comprising one or more processors, wherein the one or more processors access a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order, and access a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in a normal mode, context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the decoding of the first transform coefficient, and a first skip of the block of the picture. an apparatus configured to entropy decode the first and second parameter sets in a scan pass, the first scan pass being performed before other scan passes of entropy decoded transform coefficients of the block, each of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2, and reconstructing the block in response to the decoded transform coefficients. The apparatus may further include one or more memories coupled to the one or more processors.
[0094] According to another embodiment, there is provided an apparatus for video coding comprising one or more processors, the one or more processors accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, and accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on coding of the first transform coefficient. and entropy coding the first and second parameter sets in a first scan path of the block, the first scan path being performed before other scan paths of the entropy coded transform coefficients of the block, and each of the first and second parameter sets for the first and second transform coefficients includes at least one of a gt1 flag and a gt2 flag, the gt1 flag being configured to indicate whether an absolute value of a corresponding transform coefficient is greater than 1 and the gt2 flag being configured to indicate whether an absolute value of the corresponding transform coefficient is greater than 2.
[0095] According to another embodiment, an apparatus for video decoding comprises: means for accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order; means for accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on decoding of the first transform coefficient; and means for accessing a first parameter set associated with a first transform coefficient in a block of the picture in decoding order; An apparatus is provided, comprising: means for entropy decoding the first and second parameter sets, wherein the first scan path is performed before other scan paths of the entropy decoded transform coefficients of the block, and each set of the first and second parameter sets for the first and second transform coefficients includes at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 2; and means for reconstructing the block in response to the decoded transform coefficients.
[0096] According to another embodiment, an apparatus for video coding comprises means for accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, means for accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode and context modelling of at least parameters of the second parameter set for the second transform coefficient depending on the coding of the first transform coefficient, and means for accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode and and means for entropy coding the first and second parameter sets in a first scan path of a block, the first scan path being performed before other scan paths of entropy coding transform coefficients of the block, each set of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2.
[0097] According to another embodiment, a signal comprising coded video includes accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order, and accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the coding of the first transform coefficient, and accessing the block. a signal is formed by entropy coding the first and second parameter sets in a first scan path of a block, the first scan path being performed before other scan paths of the entropy coded transform coefficients of the block, each set of the first and second parameter sets for the first and second transform coefficients including at least one of a gt1 flag and a gt2 flag, the gt1 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 1 and the gt2 flag indicating whether the absolute value of the corresponding transform coefficient is greater than 2.
[0098] According to an embodiment, the context modeling of at least parameters of the second parameter set of the second transform coefficients depends on the decoding of the first parameter set of the first transform coefficients and is independent of parameters (1) used to represent the first transform coefficients and (2) entropy coded in bypass mode.
[0099] According to an embodiment, a SIG flag is also coded or decoded in the first scan path, said SIG flag indicating whether the corresponding transform coefficient is zero.
[0100] According to an embodiment, the gt1 flag is encoded or decoded in a first scan path, and the gt2 flag is encoded or decoded in a second scan path.
[0101] According to an embodiment, an inverse quantizer for inverse quantizing the second transform coefficients is selected between two or more quantizers based on the first transform coefficients.
[0102] According to an embodiment, the inverse quantizer is selected based on the first parameter set for the first transform coefficients.
[0103] According to an embodiment, the context modeling of at least the parameters in the second parameter set for the second transform coefficients depends on the decoding of the first and second transform coefficients.
[0104] According to an embodiment, the inverse quantizer is selected based on the sum of the SIG, gt1, and gt2 flags.
[0105] According to an embodiment, the inverse quantizer is selected based on an XOR function of the SIG, gt1, and gt2 flags.
[0106] According to an embodiment, the inverse quantizer is selected based on the gt1 flag, the gt2 flag, or a gtx flag indicating whether the absolute value of the corresponding transform coefficient is greater than x.
[0107] According to an embodiment, (1) parameters used to represent transform coefficients in the block and (2) parameters coded in bypass mode are entropy coded or decoded in one or more scan paths after the parameters (1) used to represent transform coefficients in the block and (2) coded in normal mode.
[0108] According to an embodiment, the context modeling of the SIG, gt1, gt2 or gtx flags is based on the quantizer or the state used in the selection of the quantizer.
[0109] Embodiments provide computer programs comprising instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding or decoding method according to any of the above-described embodiments. One or more of the embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the above-described methods. One or more embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the above-described methods. One or more embodiments also provide methods and apparatus for transmitting or receiving a bitstream generated according to the above-described methods.
[0110] Various implementations involve decoding. As used herein, "decoding" can encompass all or some of the processes performed on a received encoded sequence to generate a final output suitable for display, for example. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally will be clear based on the context of a particular description and will be well understood by those skilled in the art.
[0111] Various implementations involve encoding. Similar to the above discussion of "decoding," "encoding" as used in this application can encompass, for example, all or part of the processes performed on an input video sequence to generate an encoded bitstream.
[0112] It should be noted that the syntax elements used herein, e.g., sig_coeff_flag, abs_level_gt1_flag, etc., are descriptive terms and therefore do not preclude the use of other syntax element names.
[0113] The implementations and aspects described herein may be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even when discussed in the context of a single type of implementation (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. The method may be implemented, for example, in an apparatus, such as a processor, which generally refers to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the transfer of information between end users.
[0114] References to "one embodiment" or "embodiment," or "one implementation" or "implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation," as well as any other variations, in various places throughout this document are not necessarily all referring to the same embodiment.
[0115] Additionally, the application may refer to "determining" various portions of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0116] Additionally, the application may refer to "accessing" various portions of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0117] Additionally, the application may refer to "receiving" various portions of information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is typically included in some manner, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0118] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded as many times as there are items listed, as would be apparent to one of ordinary skill in the art.
[0119] As will be apparent to those skilled in the art, implementations can generate a wide variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a wide variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Claims
1. 1. A method comprising: accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order; accessing a second parameter set associated with the second transform coefficient, wherein the first and second parameter sets are entropy coded in a normal mode and context modeling of at least parameters of the second parameter set for the second transform coefficient depends on decoding of the first transform coefficient; entropy decoding the first and second parameter sets in a first scan path of the block, the first scan path being performed before one or more other scan paths of entropy decoded transform coefficients of the block, each set of the first and second parameter sets for the first and second transform coefficients including at least one of a first flag and a second flag, the first flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1, and the second flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2; reconstructing the block in response to the decoded transform coefficients; A method comprising:
2. 1. A method comprising: accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order; accessing a second parameter set associated with the second transform coefficient, wherein the first and second parameter sets are entropy coded in a normal mode and context modeling of at least parameters of the second parameter set for the second transform coefficient depends on coding of the first transform coefficient; entropy coding the first and second parameter sets in the first scan path of the block, the first scan path being performed before one or more other scan paths of entropy coded transform coefficients of the block, each set of the first and second parameter sets for the first and second transform coefficients including at least one of a first flag and a second flag, the first flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1, and the second flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2; A method comprising:
3. 1. An apparatus comprising one or more processors, the one or more processors comprising: accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in decoding order; accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in a normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on decoding of the first transform coefficient; entropy decoding the first and second parameter sets in the first scan path of the block, the first scan path being performed before one or more other scan paths of entropy decoded transform coefficients of the block, each of the first and second parameter sets for the first and second transform coefficients including at least one of a first flag and a second flag, the first flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1, and the second flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2; reconstructing the block in response to the decoded transform coefficients. The apparatus is configured to:
4. 1. An apparatus comprising one or more processors, the one or more processors comprising: accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order; accessing a second parameter set associated with the second transform coefficient, the first and second parameter sets being entropy coded in a normal mode, and context modeling of at least parameters of the second parameter set for the second transform coefficient depending on the coding of the first transform coefficient; Entropy coding the first and second parameter sets in the first scan path of the block, the first scan path being performed before one or more other scan paths of the entropy coded transform coefficients of the block, each of the first and second parameter sets for the first and second transform coefficients including at least one of a first flag and a second flag, the first flag indicating whether the absolute value of the corresponding transform coefficient is greater than 1, and the second flag indicating whether the absolute value of the corresponding transform coefficient is greater than 2. The apparatus is configured to:
5. 5. The method of claim 1 or 2, or the apparatus of claim 3 or 4, wherein the context modeling of at least parameters in the second parameter set of the second transform coefficients depends on decoding of the first parameter set of the first transform coefficients and is independent of parameters (1) used to represent the first transform coefficients and (2) entropy coded in bypass mode.
6. 6. The method of claim 1, 2 or 5, or the apparatus of claim 3, 4 or 5, wherein a SIG flag is also coded or decoded in the first scan path, the SIG flag indicating whether the corresponding transform coefficient is zero.
7. The method of any one of claims 1, 2, 5 and 6, or the apparatus of any one of claims 3 to 6, wherein the first flag is encoded or decoded in the first scan path, and the second flag is encoded or decoded in the second scan path.
8. The method of any one of claims 1, 2 and 5 to 7, or the apparatus of any one of claims 3 to 7, wherein an inverse quantizer that inverse quantizes the second transform coefficients is selected from two or more quantizers based on the first transform coefficients.
9. The method of any one of claims 1, 2 and 5 to 8 or the apparatus of any one of claims 3 to 8, wherein the inverse quantizer is selected based on the first parameter set for the first transform coefficients.
10. The method or apparatus according to any one of claims 8 to 9, wherein the inverse quantizer is selected based on the sum of the SIG, first and second flags.
11. The method or apparatus of any one of claims 8 to 9, wherein the inverse quantizer is selected based on an XOR function of the SIG, first and second flags.
12. 10. The method according to claim 8, wherein the inverse quantizer is selected based on the first flag, the second flag, or another flag, the other flag indicating whether the absolute value of the corresponding transform coefficient is greater than x.
13. 13. The method of claim 1, 2, or 5-12, or the apparatus of claim 3, wherein parameters (1) used to represent transform coefficients in the block and (2) coded in bypass mode are entropy coded or decoded in one or more scan paths after parameters (1) used to represent transform coefficients in the block and (2) coded in normal mode.
14. A method according to any one of claims 8 to 12 or an apparatus according to any one of claims 8 to 12, wherein context modelling of the SIG, the first flag, the second flag or another flag is based on the quantizer or a state used in selecting the quantizer.
15. 15. The method of claim 1, 2 or 5 to 14, or any of the apparatuses of claims 1, 2 or 5 to 14, wherein each set of the first and second parameters for the first and second transform coefficients further includes a third flag, the third flag indicating whether the absolute value of the corresponding transform coefficient is greater than 3.
16. 1. A signal containing encoded video, comprising: accessing a first parameter set associated with a first transform coefficient in a block of a picture, the first transform coefficient preceding a second transform coefficient in the block of the picture in coding order; accessing a second parameter set associated with the second transform coefficient, wherein the first and second parameter sets are entropy coded in a normal mode and context modeling of at least parameters of the second parameter set for the second transform coefficient depends on the coding of the first transform coefficient; entropy coding the first and second parameter sets in the first scan path of the block, the first scan path being performed before one or more other scan paths of entropy coded transform coefficients of the block, each set of the first and second parameter sets for the first and second transform coefficients including at least one of a first flag and a second flag, the first flag indicating whether an absolute value of a corresponding transform coefficient is greater than 1, and the second flag indicating whether an absolute value of the corresponding transform coefficient is greater than 2; The signal is formed by executing
Citation Information
Patent Citations
advanced arithmetic coder
JP2018521556A
Coding sign information of video data
US20170142448A1