Flexible allocation of regular bins in residual coding for video coding
By allocating regular bins based on a budget among coding groups using CABAC encoding and decoding, the video encoding process addresses suboptimal compression efficiency in conventional methods, enhancing overall coding performance.
Patent Information
- Application Number
- JP2025077190
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-23
- Filing Date
- 2025-05-07
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2040-03-06
AI Technical Summary
Conventional video encoding methods face challenges in optimizing the use of regular bins in residual coding, leading to suboptimal compression efficiency due to high-level constraints on CABAC coding.
Implementing a video encoding and decoding process that allocates a limited number of regular bins based on a budget among coding groups, using CABAC encoding and decoding, to respect high-level constraints and improve compression efficiency.
Enhances compression efficiency by optimizing the use of regular bins in residual coding, adhering to high-level constraints and improving the overall coding performance.
Smart Images

Figure 0007793094000001 
Figure 0007793094000002 
Figure 0007793094000003
Abstract
Description
[Technical Field]
[0001] Technical Field At least one of the present embodiments generally relates to conventional bin assignment in residual coding for video encoding or decoding. [Background technology]
[0002] background To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to exploit spatial and temporal redundancy in video content. Generally, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, followed by transforming, quantizing, and entropy coding the difference between the original block and the predicted block (often expressed as a prediction error or prediction residual). To reconstruct the video, the compressed data is decoded by the inverse process of entropy coding, quantizing, transforming, and predicting. Summary of the Invention
[0003] overview One or more of the present embodiments handle high-level constraints on the maximum usage of conventional CABAC coding of bins in the residual coding process of blocks and their coefficient groups so that the high-level constraints are respected and compression efficiency is improved compared to current approaches.
[0004] According to a first aspect of at least one embodiment, a video encoding method includes a residual encoding process that encodes a syntax element representing a picture area including a coding group using a limited number of regular bins, the coding being CABAC encoding, and the number of regular bins being determined based on a budget allocated among a number of coding groups.
[0005] According to a second aspect of at least one embodiment, a video decoding method includes a residual decoding process that analyzes a bitstream representing a picture area including coded groups using regular bins, the decoding being performed using CABAC decoding, and the number of regular bins is determined based on a budget allocated among a number of coded groups.
[0006] According to a third aspect of at least one embodiment, an apparatus includes an encoder for encoding picture data for at least one block of a picture or video, the encoder configured to perform a residual encoding process that encodes syntax elements representing a picture area including coded groups using a limited number of regular bins, the coding being CABAC encoding, and the number of regular bins being determined based on a budget allocated among a number of coded groups.
[0007] According to a fourth aspect of at least one embodiment, an apparatus includes a decoder for decoding picture data for at least one block of a picture or video, the decoder configured to perform a residual decoding process that analyzes a bitstream representing a picture area including coded groups using regular bins, the decoding being performed using CABAC decoding, and the number of regular bins being determined based on a budget allocated among a number of coded groups.
[0008] According to a fifth aspect of at least one embodiment, there is provided a computer program comprising program code instructions executable by a processor, the computer program performing at least the steps of a method according to the first or second aspect.
[0009] According to a sixth aspect of at least one embodiment, there is provided a computer program product stored on a non-transitory computer readable medium and comprising program code instructions executable by a processor, the computer program product performing at least the steps of a method according to the first or second aspect. [Brief explanation of the drawings]
[0010] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] 1 shows a block diagram of an example video encoder 100, such as a High Efficiency Video Coding (HEVC) encoder. [Figure 2] 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. [Figure 3] 1 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. [Figure 4A] 1 shows an example of a coding tree unit and a coding tree in the compressed domain. [Figure 4B] 1 shows an example of division of a CTU into coding units, prediction units, and transform units. [Figure 5] Illustrates the use of two scalar quantizers in dependent scalar quantization. [Figure 6A] 1 illustrates an example mechanism for switching between scalar quantizers. [Figure 6B] 1 illustrates an example mechanism for switching between scalar quantizers. [Figure 6C] 1 shows an example of inter-CG and inter-coefficient scan order as used in VVC for an 8x8 transform block. [Figure 7A] 1 shows examples of syntax elements for coding / parsing for a transform block. [Figure 7B] An example of CG level residual coding syntax is shown below. [Figure 8A] Here is an example of CU-level syntax: [Figure 8B] Here is an example of the transform_tree syntax structure: [Figure 8C] This shows the syntax arrangement at the translation unit level. [Figure 9A] Illustrates the CABAC decoding process. [Figure 9B] 1 shows the CABAC encoding process. [Figure 10] 1 illustrates a CABAC encoding process including a CABAC optimizer, in accordance with an embodiment of the present principles; [Figure 11] 1 shows an example of a flowchart for a CABAC optimizer used in an encoding process in accordance with an embodiment of the present principles; [Figure 12] 10 shows an example of a modified process for encoding / decoding coding groups using the CABAC optimizer. [Figure 13A] 1 illustrates an example embodiment in which the budget for a regular bin is determined at the transform block level. [Figure 13B] 1 illustrates an example of an embodiment in which the budget of a regular bin is determined at the transform block level according to the position of the last significant coefficient. [Figure 14A] 10 illustrates an example embodiment in which the budget for a regular bin is determined at the transform unit level. [Figure 14B] An example of an embodiment is shown in which the budget of normal bins is determined at the transform unit level and allocated between transform blocks as a function of the relative surfaces between different transform blocks and as a function of the normal bins used in the already coded / analyzed transform blocks of the considered transform unit. [Figure 14C] 10 illustrates an example embodiment of a process for TB-level residual coding / analysis. [Figure 15A] 10 illustrates an example embodiment in which the regular bin budget is allocated at the coding unit level. [Figure 15B] 10 illustrates an example embodiment of encoding / parsing a transform tree associated with a current CU based on a regular bin budget allocated at the CU level. [Figure 15C] 10 illustrates an example of an alternative embodiment of the coding tree coding / parsing process when the budget is fixed at the CU level. [Figure 15D] 10 illustrates the transform unit coding / parsing process adapted for the present embodiment when the regular bin budget is fixed at the CU level. [Figure 15E] 10 illustrates the transform unit coding / parsing process adapted for the present embodiment when the regular bin budget is fixed at the CU level. [Figure 16A]1 shows an example of an embodiment in which the budget for normal bins is allocated at a level higher than the CG level, and is for using a first type of CG coding / analysis process. [Figure 16B] 10 shows an example of an embodiment in which the budget for regular bins is allocated at a level higher than the CG level, and is for use with a second type of CG coding / analysis process. DETAILED DESCRIPTION OF THE INVENTION
[0011] Detailed Description Various embodiments relate to entropy coding of quantized transform coefficients. This stage of a video codec is also called the residual coding step. At least one embodiment aims to optimize the coding efficiency of a video codec under the constraint of a maximum number of regular coding bins per single area.
[0012] Various methods and other aspects described in this application can be used to modify at least the entropy coding and / or decoding modules (145, 230) of video encoder 100 and decoder 200 as shown in Figures 1 and 2. Moreover, although the aspects describe principles related to particular drafts of the VVC (Versatile Video Coding) or HEVC (High Efficiency Video Coding) specifications, they are not limited to VVC or HEVC and may be applied, for example, to other standards and recommendations, and extensions of such standards and recommendations (including VVC and HEVC), whether pre-existing or later developed. Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.
[0013] Figure 1 shows a block diagram of an example video encoder 100, such as an HEVC encoder. Figure 1 also shows an encoder that incorporates improvements to the HEVC standard or employs HEVC-like technology, such as the Joint Exploration Model (JEM) or VVC Test Model (VTM) encoders under development by the Joint Video Exploration Team (JVET) in VVC.
[0014] Before being encoded, a video sequence may undergo pre-encoding processing (101). For example, this processing may be performed by applying a color transform to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0), or by performing a remapping of the input picture components to obtain a signal distribution that is more compression-resilient (e.g., using histogram equalization of one of the color components). Metadata associated with the pre-processing may be carried in the bitstream.
[0015] In HEVC, to encode a video sequence having one or more pictures, a picture is partitioned (102) into one or more slices, and each slice may include one or more slice segments. Slice segments are organized into coding units, prediction units, and transform units. The HEVC specification distinguishes between "blocks" and "units," where a "block" addresses a specific area of a sample array (e.g., luma, Y), and a "unit" includes all coded color components (Y, Cb, Cr, or monochrome) associated with a block (e.g., motion vectors), syntax elements, and an array of prediction data.
[0016] For HEVC coding, a picture is partitioned into square coding tree blocks (CTBs) with configurable sizes, and a contiguous set of coding tree blocks is grouped into slices. A coding tree unit (CTU) contains the CTB of a coded color component. The CTB is the root of a quadtree partitioned into coding blocks (CBs), which can be partitioned into one or more prediction blocks (PBs), forming the root of a quadtree partitioned into transform blocks (TBs). Corresponding to the coding blocks, prediction blocks, and transform blocks, a coding unit (CU) contains a tree-structured set of prediction units (PUs) and transform units (TUs), where a PU contains prediction information for all color components and a TU contains a residual coding syntax structure for each color component. The sizes of the CBs, PBs, and TBs of the luma component apply to the corresponding CUs, PUs, and TUs. In this application, the term "block" can be used to refer to, for example, any of the CTUs, CUs, PUs, TUs, CBs, PBs, and TBs. In addition, "block" can also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, or more generally to refer to arrays of data of various sizes.
[0017] In the example encoder 100, a picture is encoded by the encoder elements as described below. The picture to be encoded is processed in units of CUs. Each CU is encoded using intra or inter mode. When a CU is encoded in intra mode, intra prediction (160) is performed. In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) whether to use intra or inter mode for encoding the CU and indicates the intra / inter decision with a prediction mode flag. A prediction residual is calculated by subtracting (110) the prediction block from the original image block.
[0018] Intra modes, CUs are predicted from neighboring reconstructed samples within the same slice. HEVC provides a set of 35 intra prediction modes, including one DC mode, one planar mode, and 33 angular prediction modes. The intra prediction reference is reconstructed from rows and columns adjacent to the current block. The reference spans more than twice the block size horizontally and vertically, using samples available from previously reconstructed blocks. When an angular prediction mode is used for intra prediction, the reference samples can be copied along the direction indicated by the angular prediction mode.
[0019] The luma intra-prediction modes applicable to the current block can be coded using two different options: If the applicable mode is included in a constructed list of three most probable modes (MPM), the mode is signaled by an index in the MPM list; otherwise, the mode is signaled by a fixed-length binarization of the mode index. The three most probable modes are derived from the intra-prediction modes of the upper and left neighboring blocks.
[0020] For an inter CU, the corresponding coded block is further partitioned into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information about how the inter prediction is performed. Motion information (e.g., motion vectors and reference picture indexes) can be signaled in two ways: "merge mode" and "advanced motion vector prediction (AMVP)."
[0021] In merge mode, a video encoder or decoder assembles a candidate list based on already coded blocks, and the video encoder signals an index to one of the candidates in the candidate list. At the decoder side, motion vectors (MVs) and reference picture indices are reconstructed based on the signaled candidates.
[0022] In AMVP, a video encoder or decoder assembles a candidate list based on motion vectors determined from previously coded blocks. The video encoder then signals an index into the candidate list to identify a motion vector predictor (MVP) and a motion vector difference (MVD). At the decoder side, the motion vector (MV) is reconstructed as MVP+MVD. Also, the applicable reference picture index is explicitly coded in the PU syntax for AMVP.
[0023] The prediction residual is then transformed (125) and quantized (130), including at least one embodiment for adapting chroma quantization parameters, described below. The transformation is typically based on a separable transform. For example, a DCT transform is applied first horizontally and then vertically. While in earlier codecs the variety of 2D transforms for a given block size is typically limited, in more recent codecs such as JEM, the transforms used in both directions can be different (e.g., DCT in one direction and DST in the other), thereby allowing a wide variety of 2D transforms.
[0024] The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder can also avoid both the transform and quantization (i.e., the residual is coded directly without applying either the transform or quantization process). In direct PCM coding, no prediction is applied, and coded unit samples are coded directly and embedded in the bitstream.
[0025] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. Combining the decoded residual with the prediction block (155) reconstructs an image block. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).
[0026] Figure 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. In the example decoder 200, a bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding path that is the inverse of the encoding path as described in Figure 1, which performs video decoding as part of encoding the video data. Figure 2 also shows a decoder that has improvements to the HEVC standard or that employs HEVC-like technology, such as a JEM or VVC decoder.
[0027] Specifically, the decoder's input includes a video bitstream, which may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, picture partitioning information, and other coding information. The picture partitioning information indicates the size of the CTUs and how the CTUs are divided into CUs, and possibly PUs, if applicable. The decoder can then divide (235) the picture into CTUs and each CTU into CUs according to the decoded picture partitioning information. The transform coefficients are then inverse quantized (240) and inverse transformed (250), including at least one embodiment for adapting chroma quantization parameters, as described below, to decode the prediction residual.
[0028] An image block is reconstructed by combining the decoded prediction residual with the prediction block (255). The prediction block can result from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (270). As described above, AMVP and merge mode techniques can be used to derive motion vectors for motion compensation, which can use interpolation filters to calculate interpolated values for sub-integer samples of the reference block. An in-loop filter (265) is applied to the reconstructed picture. The filtered image is stored in a reference picture buffer (280).
[0029] The decoded picture may undergo further post-decoding processing (285), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (101). The post-decoding processing may use metadata that is derived in the pre-encoding processing and signaled in the bitstream.
[0030] FIG. 3 shows a block diagram of an example system in which various aspects and embodiments may be implemented. System 300 may be embodied as a device including various components described below and configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, encoders, transcoders, and servers. The elements of system 300, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 300 are distributed across multiple ICs and / or discrete components. In various embodiments, the elements of system 300 are communicatively coupled through an internal bus 310. In various embodiments, system 300 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 300 is configured to implement one or more of the aspects described in this document, such as video encoder 100 and video decoder 200, modified as described above and below.
[0031] System 300 includes at least one processor 301 configured to execute instructions loaded therein to implement various aspects described herein, for example. Processor 301 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 300 includes at least one memory 302 (e.g., a volatile memory device and / or a non-volatile memory device). System 300 includes storage 304, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. Storage 304 may include, by way of non-limiting example, internal storage, attached storage, and / or network-accessible storage.
[0032] System 300 includes an encoder / decoder module 303 configured to process data to provide, for example, encoded video or decoded video, which may include its own processor and memory. Encoder / decoder module 303 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, encoder / decoder module 303 may be implemented as a separate element of system 300 or may be incorporated within processor 301 as a combination of hardware and software, as is known to those skilled in the art.
[0033] Program code to be loaded into the processor 301 or the encoder / decoder 303 to perform various aspects described in this document may be stored in the storage device 304 and then loaded into the memory 302 for execution by the processor 301. According to various embodiments, one or more of the processor 301, the memory 302, the storage device 304, and the encoder / decoder module 303 may store one or more of various items during execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and arithmetic logic.
[0034] In some embodiments, memory internal to the processor 301 and / or encoder / decoder module 303 is used to store instructions and provide working memory for processing necessary during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 301 or the encoder / decoder module 303) is used for one or more of these functions. The external memory may be memory 302 and / or storage 304, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, external high-speed dynamic volatile memory, such as RAM, is used as working memory for video coding and decoding operations, such as those for MPEG-2, HEVC, or VVC.
[0035] Input to the elements of system 300 may be provided through various input devices as shown in block 309. Such input devices include, but are not limited to, (i) an RF section that receives RF signals transmitted over the air, for example by a broadcast station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0036] In various embodiments, the input devices of block 309 have respective associated input processing elements as known in the art. For example, the RF section may be associated with elements necessary to (i) select a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a band of frequencies), (ii) downconvert the selected signal, (iii) bandlimit again to a narrower band of frequencies (for example) to select a signal frequency band (which in some embodiments may also be referred to as a channel), (iv) demodulate the downconverted and bandlimited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, including, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs these various functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and then perform frequency selection by filtering, downconverting, and filtering again to obtain the desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements (e.g., inserting amplifiers and analog-to-digital converters, etc.). In various embodiments, the RF section includes an antenna.
[0037] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 300 to other electronic devices through USB and / or HDMI connections. It should be understood that various aspects of the input processing (e.g., Reed-Solomon error correction) can be implemented, for example, within a separate input processing IC or within processor 301, as desired. Similarly, aspects of the USB or HDMI interface processing can be implemented, as desired, within a separate interface IC or within processor 301. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements (e.g., including processor 301 and encoder / decoder 303 in conjunction with memory and storage elements) to process the data stream for presentation on an output device, as desired.
[0038] The various elements of system 300 may be provided within an integrated housing in which the various elements may be interconnected and data may be transmitted between them using any suitable connection arrangement, such as, for example, an internal bus (including an I2C bus, wires, and printed circuit boards) known in the art.
[0039] System 300 includes a communication interface 305 that enables communication with other devices over a communication channel 320. Communication interface 305 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 320. Communication interface 305 may include, but is not limited to, a modem or a network card, and communication channel 320 may be implemented within a wired and / or wireless medium, for example.
[0040] Data is streamed to system 300 using a Wi-Fi network, such as IEEE 802.11, in various embodiments. The Wi-Fi signal in these embodiments is received over communication channel 320 and communication interface 305 adapted for Wi-Fi communication. Communication channel 320 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 300 using a set-top box that transmits data over the HDMI connection of input block 309. Still other embodiments provide streamed data to system 300 using the RF connection of input block 309.
[0041] System 300 can provide output signals to various output devices, including display 330, speakers 340, and other peripheral devices 350. Other peripheral devices 350, in various example embodiments, include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 300. In various embodiments, control signals are communicated between system 300 and display 330, speakers 340, or other peripheral devices 350 using signaling such as AV.Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. Output devices can be communicatively coupled to system 300 through respective interfaces 306, 307, and 308 via dedicated connections. Alternatively, output devices can be connected to system 300 via communication interface 305 using communication channel 320. Display 330 and speakers 340 can be integrated into a single unit along with other components of system 300, such as an electronic device (e.g., a television). In various embodiments, the display interface 306 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0042] For example, if the RF portion of input 309 is part of a separate set-top box, display 330 and speakers 340 can be selectively isolated from one or more of the other components. In various embodiments in which display 330 and speakers 340 are external components, output signals can be provided via dedicated output connections (including, for example, an HDMI port, a USB port, or a COMP output). Implementations described herein can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even when discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., an apparatus or a program). An apparatus can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in an apparatus (e.g., a processor), which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0043] 4A shows an example of a coding tree unit and a coding tree in the compressed domain. In the HEVC video compression standard, motion-compensated temporal prediction is employed to exploit the redundancy that exists between consecutive pictures of a video. To achieve this, a picture is partitioned into so-called coding tree units (CTUs), whose sizes are typically 64x64, 128x128, or 256x256 pixels. Each CTU is represented by a coding tree in the compressed domain (e.g., a quadtree decomposition of the CTU). Each leaf is called a coding unit (CU).
[0044] 4B shows an example of division of a CTU into coding units, prediction units, and transform units. Each CU is then given several intra- or inter-prediction parameters (prediction information). To achieve this, each CU is spatially partitioned into one or more prediction units (PUs), and each PU is assigned several prediction information such as motion vectors. An intra- or inter-coding mode is assigned at the CU level.
[0045] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "encoded" and "coded" can be used interchangeably, and the terms "image," "picture," and "frame" can be used interchangeably. Typically, and although not universally applicable, the term "reconstructed" is used on the encoder side, and the term "decoded" is used on the decoder side. The term "block" or "picture block" can be used to refer to any one of CTU, CU, PU, TU, CB, PB, and TB. In addition, the term "block" or "picture block" can be used to refer to macroblocks, partitions, and sub-blocks as specified in H.264 / AVC or other video coding standards, and more generally to refer to arrays of samples of many sizes.
[0046] FIG. 5 illustrates the use of two scalar quantizers in dependent scalar quantization. Dependent scalar quantization, as proposed in JVET (contribution JVET-J0014), uses two scalar quantizers with different reconstruction levels for quantization. Compared to traditional scalar quantization (e.g., as used in HEVC and VTM-1), the main advantage of this approach is that the set of allowable reconstruction values of a transform coefficient depends on the value of the transform coefficient level preceding the current transform coefficient level in the reconstruction order. The dependent scalar quantization approach is realized by (a) defining two scalar quantizers with different reconstruction levels and (b) defining a process for switching between the two scalar quantizers. The two scalar quantizers used are shown in FIG. 5 (represented by Q0 and Q1). The location of the available reconstruction levels is uniquely specified by the quantization step size Δ. Ignoring the fact that the actual reconstruction of the transform coefficients uses integer arithmetic, the two scalar quantizers Q0 and Q1 can be characterized as follows: Q0: The reconstruction levels of the first quantizer Q0 are given by even integer multiples of the quantization step size Δ. When this quantizer is used, the dequantized transform coefficients t′ are given by t'=2·k·Δ where k represents the associated quantized coefficient (transmitted quantization index). Q1: The reconstruction levels of the second quantizer Q1 are given by odd integer multiples of the quantization step size Δ in addition to the reconstruction levels equal to zero. The dequantized transform coefficients t′ are given by t'=(2·k-sgn(k))·Δ where sgn(·) is calculated as a function of the quantized coefficient k as follows: sgn(x)=(k==0?0:(k<0?-1:1)) is the sign function defined as
[0047] The scalar quantizer used (Q0 or Q1) is not explicitly signaled in the bitstream. Instead, the quantizer used for the current transform coefficient is determined by the parity of the quantized coefficient preceding the current transform coefficient in the coding / reconstruction order and the state of a finite state machine introduced below.
[0048] Figures 6A and 6B show an example of a mechanism for switching between scalar quantizers in VVC. Switching between two scalar quantizers (Q0 and Q1) is realized via a finite state machine with four states (labeled as 0, 1, 2, or 3, respectively), as shown in Figure 6A. The state of the finite state machine considered for a given quantized coefficient is uniquely determined by the parity of the quantized coefficient k preceding the current quantized coefficient in the coding / reconstruction order and the state of the finite state machine considered when processing this preceding coefficient. At the start of dequantization for a transform block, the state is set equal to 0. The transform coefficients are reconstructed in scan order (i.e., the same order in which they are entropy decoded). After the current transform coefficient is reconstructed, the state is updated, where k is the quantized coefficient. The next state is state=stateTransTable[current_state][k&1] where stateTransTable represents the state transition table shown in FIG. 6B and the operator & specifies the bitwise “and” operator in two's complement arithmetic.
[0049] Moreover, the quantized coefficients contained in so-called transform blocks (TB) can be entropy coded and decoded as explained below.
[0050] First, the transform block is divided into 4x4 sub-blocks of quantized coefficients called coding groups (sometimes also named coefficient groups, abbreviated as CG). The entropy coding / decoding is made up of several scan passes, which traverse the TB according to the diagonal scan order shown by Figure 6C.
[0051] Transform coefficient coding in VVC involves five major steps: scanning, final significant coefficient coding, significance map coding, coefficient level residual coding, absolute level and sign data coding.
[0052] 6C shows an example of inter-CG and inter-coefficient scan orders as used in VVC for an 8x8 transform block. A scan pass over a TB involves processing each CG sequentially according to a diagonal scan order, and the 16 coefficients within each CG are similarly scanned according to the scan order considered. The scan pass starts with the last significant coefficient of the TB and processes all coefficients down to the DC coefficient.
[0053] The entropy coding of the transform coefficients involves up to seven syntax elements in the list below. - coded_sub_block_flag: coefficient group (CG) significance - sig_flag: Coefficient significance (zero / non-zero) - gt1_flag: Indicates whether the absolute value of the coefficient level is greater than 1 - par_flag: indicates the parity of the coefficient greater than 1 - gt3_flag: Indicates whether the absolute value of the coefficient level is greater than 3 - remainder: remainder for the absolute value of the coefficient level (if it is greater than the one coded in the previous pass) - abs_level: Value of absolute value of coefficient level (if no CABAC bins are signaled for the current coefficient for the maximum number of bin budget problem) sign_data: signs of all significant coefficients included in the considered CG. It contains a series of bins, each signaling the sign (0: positive, 1: negative) of each non-zero transform coefficient.
[0054] Once the absolute value of a quantized coefficient is known by decoding a subset of the above elements (apart from its sign), no further syntax elements for that coefficient relating to its absolute value are coded. Similarly, the sign flag is only signaled for non-zero coefficients.
[0055] All scan passes required for a given CG are coded until all quantized coefficients of that CG can be reconstructed before proceeding to the next CG.
[0056] The overall decrypted TB analysis process consists of the following major steps: 1. Decode the last significant coefficient coordinates, which includes the syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix and last_sig_coeff_y_suffix, which provides the decoder with the spatial location (x and y coordinates) of the last non-zero coefficient in the entire TB.
[0057] Then, for each successive CG from the CG containing the last significant coefficient of the TB to the top left CG of the TB, the following steps are applied: 2. Decode the CG significance flag (called coded_sub_block_flag in the VVC specification) 3. Decode the significant coefficient flag for each coefficient of the CG considered. This corresponds to the syntax element sig_flag. It indicates which coefficients of the CG are non-zero.
[0058] The next analysis step concerns the coefficient level for the known non-zero coefficients of the considered CG. This involves the following syntax elements: 4. gt1_flag: This flag indicates whether the absolute value of the current coefficient is greater than 1. If it is not greater than 1, the absolute value is equal to 1. 5. par_flag: This flag indicates whether the current quantized coefficient is even or not. If the gt1_flag of the current quantized coefficient is true, it is coded. If par_flag is zero, the quantized coefficient is even, otherwise the quantized coefficient is odd. After par_flag is parsed at the decoder side, the partially decoded quantized coefficient is set equal to (1 + gt1_flag + par_flag). 6. gt3_flag: This flag indicates whether the absolute value of the current coefficient is greater than 3. If it is not greater than 3, the absolute value is equal to 1 + gt1_flag + par_flag. If 1 + gt1_flag + par_flag is 2 or greater, gt3_flag is coded. When gt3_flag is analyzed, the quantized coefficient value is calculated on the decoder side as 1 + gt1_flag + par_flag + (2 * gt3_flag). 7. remainder: This encodes the absolute value of the coefficient. This is the case if the partially decoded absolute value is equal to but greater than 4. Note that in the VVC Draft 3 example, for each coding group, the maximum number of regular coding bins in the budget is fixed. Thus, for some coefficients, only the elements sig_flag, gt1_flag, and par_flag can be signaled, while for other coefficients, gt3_flag can also be signaled. The coded and analyzed remainder values are therefore calculated on the flags already decoded for the coefficient considered, and therefore as a function of the partially decoded quantized coefficients. 8. abs_level: This indicates the absolute value of the coefficients that are not coded with any flags (among sig_flag, gt1_flag, papr_flag, or gt3_flag) of the considered CG for the general problem of the maximum number of coding bins. This syntax element is Rice-Golomb binarized and bypass coded, just like the syntax element "remainder." 9. sign_flag: This indicates the sign of the non-zero coefficient, which is bypass coded.
[0059] Figure 7A shows syntax elements for coding / parsing for transform blocks according to VVC Draft 4. The coding / pairing involves a four-pass process. As can be seen, for a given transform block, the position of the last significant coefficient is coded first. Then, for each CG from the first CG and the CG containing the last significant coefficient (except this CG), the significance of the CG is signaled (coded_sub_block_flag), and in case the CG is significant, residual coding for that CG is performed as shown by Figure 7B.
[0060] Figure 7B shows the CG level residual coding syntax as specified in VVC Draft 4. The syntax signals the previously introduced syntax elements sig_flag, gt1_flag, par_flag, gt3_flag, remainder, abs_level, and sign_data of the considered CG. The figure shows the signaling and parsing of the syntax elements sig_flag, gt1_flag, par_flag, gt3_flag, remainder, and abs_level according to VVC Draft 4. EP stands for "equi-probable," meaning that the corresponding bin is not arithmetically coded, but is coded in bypass mode. Bypass mode involves direct writing / parsing of bits, which are generally equal to the binary syntax element (bin) we wish to code or parse.
[0061] Moreover, VVC Draft 4 specifies a hierarchical syntactic arrangement from the CU level down to the residual subblock (CG) level.
[0062] 8A shows the CU level syntax as specified in the VVC Draft 4 example. The syntax includes signaling the coding mode of the CU being considered. cu_skip_flag signals if the CU is in merge-skip mode. If the CU is not in merge-skip mode, cu_pred_mode_flag indicates the variable value of cuPredMode and therefore if the prediction mode of the current CU is coded via intra-prediction (cuPredMode=MODE_INTRA) or inter-prediction (cuPredMode=MODE_INTER).
[0063] In the case of non-skip and non-intra modes, the next flag indicates the use of intra block copy (IBC) mode for the current CU. The rest of the syntax contains intra- or inter-predicted data for the current CU. The end of the syntax is dedicated to signaling the transform tree associated with the current CU. This transform tree signaling begins with a syntax element called cu_cbf, which indicates that some non-null residual data is coded for the current CU. In the case where this flag is equal to true, the transform_tree associated with the current CU is signaled according to the syntax sequence shown in FIG. 8B.
[0064] FIG. 8B shows a syntax structure called transform_tree for an example of VVC Draft 4. The syntax basically includes signaling whether a CU contains one or several transform units (TUs). First, if the CU size is larger than the maximum allowed TU size, the CU is divided into four sub-transform trees in a quadtree manner. Otherwise, if ISP (intra sub-partitioning) or SBT (sub-block transform) is not used for the current CU, the current transform tree is not partitioned and contains exactly one TU, which is signaled through the transform unit syntax table of FIG. 8C. Otherwise, if the CU is in intra mode and ISP (intra sub-partitioning) mode is used, the CU is divided into several (actually, two) TUs. Each of the two TUs is signaled one after the other through the transform unit syntax table of FIG. 8C. Otherwise, if the CU is in inter mode and SBT (sub-block transform) mode is used, the CU is divided into two TUs according to the SBT-related decoded syntax at the CU level. Each of the two TUs is signaled in succession through the translation unit syntax table of FIG. 8C.
[0065] Figure 8C shows the transform unit level syntax array for the VVC Draft 4 example. The syntax is made up of the following: First, the tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr flags indicate, respectively, that non-null residual data corresponding to each component is contained in each transform block of the current TU. For each component, if the corresponding cbf flag is true, the transform block level residual_tb syntax table is used to code the corresponding residual transform block. The transform_unit level syntax also includes some coded, transformed, and quantized block-related syntax (i.e., delta QP information, if any, and the type of transform used to code the considered TU).
[0066] Figures 9A and 9B illustrate the CABAC decoding and encoding processes, respectively. CABAC stands for Context-Adaptive Binary Arithmetic Coding (CABAC) and is a form of entropy coding used, for example, in HEVC or VVC, to provide lossless compression with good compression efficiency. The input to the process in Figure 9A includes a coded bitstream, typically conforming to the HEVC specification or a further evolution thereof. At any point in the decoding process, the decoder knows which syntax element to decode next. This is fully specified in the standardized bitstream syntax and decoding process. Furthermore, the decoder also knows how to binarize the current syntax element to be decoded (i.e., represented as a sequence of binary symbols, called bins, each equal to "1" or "0") and how each bin in the bin string is coded.
[0067] Thus, the first stage of the CABAC decoding process (left side of FIG. 9A) decodes a series of bins. For each bin, the decoder knows whether it is coded according to bypass mode or normal mode. In bypass mode, a bit is simply read from the bitstream and the resulting value is assigned to the current bin. This mode has the advantage of being straightforward and therefore fast, and not requiring intensive resource utilization. It is efficient and therefore typically used for bins with a uniform statistical distribution (i.e., equal probability of being equal to "1" and equal probability of being equal to "0").
[0068] Conversely, if the current bin is not coded in bypass mode, it means that it is coded in so-called normal coding (i.e., through context-based arithmetic coding), which is much more resource intensive.
[0069] In that case, the decoding of the considered bin proceeds as follows: First, a context for the decoding of the current bin is obtained. The context is given by the context modeler of FIG. 9A. The goal of the context is to obtain the conditional probability that the current bin has a value of "0" given some context prior or information X. The prior X here is the value of some already decoded syntax element that is synchronously available at both the encoder and decoder sides when the current bin is being decoded.
[0070] Typically, the prior X used for decoding a bin is specified in the standard and is chosen because it is statistically correlated with the current bin to be decoded. What is interesting about using this context information is that it reduces the rate cost of encoding the bin. This is due to the fact that the conditional entropy of a bin is lower given X because of the correlation between the bin and X. The following relationship is well known in information theory: H(bin│X) <H(bin)
[0071] This means that if the bin and X are statistically correlated, the conditional entropy of the bin given X is lower than the entropy of the bin. Therefore, the context information X is used to derive the probability that the bin is "0" or "1." Given these conditional probabilities, the conventional decoding engine of FIG. 14 performs arithmetic decoding of the binary-valued bins. The bin values are then used to update the conditional probability values associated with the current bin given the current context information X. This is called the context model update step of FIG. 9A. Updating the context model for each bin as the bin is decoded (or coded) allows for progressive refinement of the context modeling for each binary element. Thus, the CABAC decoder progressively learns the statistical behavior of each of the conventional coding bins.
[0072] The context modeller and context model update steps are exactly the same operations on the encoder and decoder sides.
[0073] Regular arithmetic decoding of the current bin or its bypass decoding will result in a sequence of decoded bins depending on how they were coded.
[0074] The second stage of CABAC decoding, shown on the right side of Figure 9A, involves converting this sequence of binary symbols into higher-level syntax elements. The syntax elements can take the form of flags, in which case they directly take the value of the current decoding bin. On the other hand, if the binarization of the current syntax element corresponds to a set of several bins according to the considered standard, a conversion step called "Binary Codeword to Syntax Element" in Figure 9A is performed.
[0075] This step reverses the binarization step performed by the encoder as shown in Figure 9B. Thus, the inverse transformation performed here involves obtaining the values of these syntax elements based on their respective decoded binary versions.
[0076] The encoder 100 of FIG. 1, the decoder 200 of FIG. 2 and the system 1000 of FIG. 3 are adapted to implement at least one of the embodiments described below.
[0077] In current video coding systems (e.g., VVC Draft 4), some hard constraints are imposed on the coefficient group coding process, and a hard-coded maximum number of regular CABAC bins can be adopted for the syntax elements sig_flag, gt1_flag, and parity_flag on the one hand, and for the syntax element gt3_flag on the other hand. More precisely, the 4x4 coding group (or coefficient group or CG) level constraint ensures that the maximum number of regular bins per unit area must be analyzed by the CABAC decoder engine. However, the restriction on the use of regular bins in CABAC coding may lead to rate distortion and non-optimal video compression. In fact, if only a limited number of regular bins is allowed for a picture, the utilization of this total budget may be significantly lower in cases where some of the pictures considered for coding contain significant areas coded in skip mode (thus without residual), while other significant areas of the picture adopt residual coding. In such cases, the low-level (ie, 4x4 CG level) constraint on the maximum number of regular bins may cause the total number of regular bins in the picture to be far below the allowable picture-level threshold.
[0078] At least one embodiment relates to handling high-level constraints on the maximum usage of regular CABAC coding bins in the residual coding process of blocks and their coefficient groups so that the high-level constraints are respected and compression efficiency is improved compared to current approaches. In other words, the budget of regular coding bins is allocated over a picture area larger than the CG, thus covering a large number of CGs, the number being determined from the average allowable number of regular bins per unit area. For example, an average of 1.75 regular coding bins per sample can be allowed. This budget can then be distributed more efficiently over smaller units. In different embodiments, the higher-level constraints are set at the transform block, transform unit, coding unit, coding tree unit, or picture level. These embodiments can be implemented by a CABAC optimizer such as that shown in FIG. 10.
[0079] In at least one embodiment, a higher level constraint is set at the transform unit level. In such an embodiment, the number of regular bins allowed for the entire transform unit for encoding / decoding is determined from the average allowed number of regular bins per unit area. From the regular bin budget obtained for the TU, the number of regular bins allowed for each transform block (TB) of the TU is derived. A transform block is a set of transform coefficients belonging to the same TU and the same color component. Then, taking into account the number of regular bins allowed in a transform block, residual coding or decoding is applied under the constraint of this number of regular bins allowed for all TBs. Therefore, a modified residual coding and decoding process is proposed herein, in which the regular bin budget at the TB level is considered instead of the regular bin budget at the CG level. For this purpose, several embodiments are proposed.
[0080] 10 illustrates a CABAC encoding process including a CABAC optimizer, according to an embodiment of the present principles. In at least one embodiment, the CABAC optimizer 190 operates on a budget representing the maximum number of bins coded (or to be coded) using conventional coding for a set of coding groups. This budget is determined, for example, by multiplying the per-sample budget of a conventional coding bin by the surface of the data unit being considered (i.e., the number of samples contained in that surface). For example, in the case of a fixed conventional coding bin budget at the TB level, the allowed number of conventional coding bins per sample is multiplied by the number of samples contained in the considered transform block. The output of the CABAC optimizer 190 controls coding in conventional or bypass coding mode and therefore significantly influences coding efficiency.
[0081] FIG. 11 shows an example of a flowchart of a CABAC optimizer used in an encoding process according to an embodiment of the present principles. This flowchart is executed for each new set of considered coded groups. Thus, according to different embodiments, this flowchart may occur for each picture, each CTU, each CU, each TU, or each TB. In step 191, processor 301 assigns a budget for normal coding determined as described above. In step 192, processor 301 loops through the set of coded groups and checks in step 193 whether the last coded group has been reached. In step 194, for each coded group, the processor processes the coded group with the input budget of the normal coding bin. The budget of the normal coding bin decreases during the processing of the considered coded group (i.e., the budget of the normal coding bin is decremented each time a bin is coded in normal mode). The budget of the normal coding bin is returned as an output parameter of the coded group coding or decoding process.
[0082] FIG. 12 shows an example of a modified process for encoding / decoding a coded group using a CABAC optimizer. As introduced above, the allocation of the number of regular CABAC bins dedicated to residual coding is performed at a higher level than the 4×4 coded group. To achieve this, the process for encoding / decoding a coded group is modified compared to the conventional function by adding an input parameter representing the current budget of regular bins when the CG is being coded or decoded. This budget is therefore modified by coding or decoding the current CG according to the number of regular bins used to code / decode the current CG. The input / output parameter representing the budget of regular bins is called numRegBins_in_out in FIG. 12. The process for coding the residual data itself for a given CG can be the same as the conventional scheme, except that the allowed number of regular CABAC bins is provided by external means. As in the example of FIG. 12, each time a regular bin is coded / decoded, this budget is decremented by one.
[0083] Therefore, at least one embodiment of the present disclosure includes treating the common bin budget as an input / output parameter of the coefficient group coding / analysis function. Thus, the common bin budget can be determined at a level higher than the CG coding / analysis level. Different embodiments propose determining the budget at different levels (i.e., transform block, transform unit, coding unit, coding tree unit, or picture level).
[0084] 13A shows an example of an embodiment in which the budget of regular bins is determined at the transform block level. In such an embodiment, the budget is determined as a function of two main parameters: the size of the current transform block and a base budget of regular bins fixed for a given unit of picture area. In this specification, the area unit considered is the sample. Since the current VVC Draft 2 allows 32 regular bins for a 4x4 CG, that means that an average of two regular bins are allowed per component sample. The proposed transform block level regular bin allocation is based on this average rate of regular bins per single area.
[0085] In at least one embodiment, the budget allocated to the considered transform block is calculated as the product of the transform block surface and the normal allocated number per sample. The budget is then passed to the CG residual coding / analysis process for each significant coefficient group in the considered transform block.
[0086] 13B shows an example of an embodiment in which the budget of regular bins is determined at the transform block level according to the position of the last significant coefficient. According to this variant, the number of regular bins allowed for the current transform block is calculated as a function of the position of the last significant coefficient in the considered transform block. The advantage of such an approach is that the allocation of regular bins better adapts to the energy contained in the considered transform block. Typically, a high-energy transform block can be allocated more bins, and a low-energy transform block can be allocated fewer regular bins.
[0087] Figure 14A shows an example embodiment in which the budget for regular bins is determined at the transform unit level. As shown, unitary_budget (i.e., the average rate of regular bins allowed per sample) is still equal to 2. The TU-level budget for regular bins is determined by multiplying this rate by the total number of samples contained in the considered transform unit (i.e., in the example of a 4:2:0 color format, width * height * 3 / 2). This total budget is then used for coding all transform blocks contained in the considered transform unit. The budget is passed as an input / output parameter to the residual transform block coding / analysis process for each color component. In other words, each time a transform block of the considered transform unit is coded / analyzed, the budget is reduced by the number of bins that are normal to the coding / analysis of successively coded / analyzed transform blocks in that transform unit.
[0088] FIG. 14B shows an example embodiment in which the budget for regular bins is determined at the transform unit level and allocated among transform blocks as a function of the relative surface area between different transform blocks and as a function of the regular bins used in already coded / analyzed transform blocks of the considered transform unit. In this embodiment, the regular bins allocated to the luma TB are 2 / 3 of the total budget at the TU level. Next, the Cb TB budget is half of the remaining budget after the luma TB is coded / analyzed. Finally, the Cr TB budget is set to the remaining TU-level budget after the first two TBs are coded / analyzed. Furthermore, note that in a further embodiment, the budget for an entire transform unit can be shared among the transform blocks of that TU based on the knowledge that one or more TBs of that TU are coded with a null residual. This is known, for example, through analysis of the tu_cbf_luma, tu_cbf_cb, or tu_cbf_cr flags. Therefore, the residual coding / analysis process at the TB level is adapted to this variant, as shown by the modified process in Figure 14C. Such an embodiment results in better coding efficiency when the budget is allocated at a lower level of the hierarchy. In fact, if some CGs or TBs are of very low energy, a reduced number of bins can be employed in those CGs or TBs compared to others. Therefore, a larger number of regular bins can be used in CGs or TBs where a larger number of bins need to be coded / analyzed.
[0089] FIG. 15A shows an example of an embodiment in which the common bin budget is allocated at the coding unit level. The budget is calculated similarly to the embodiment in which the budget is fixed at the previous TU level, but is calculated based on the common bin sample rate and CU size. The calculated common bin budget is then passed to the transform tree coding / analysis procedure of FIG. 15B. FIG. 15B shows adapted coding / analysis of the transform tree associated with the current CU based on the common bin budget assigned at the CU level. Here, the CU level is essentially passed to the coding of each transform unit contained in the considered transform tree in succession. The budget is also reduced by the number of common bins employed in the coding of each TU. FIG. 15C shows an alternative embodiment of the coding tree coding / analysis process when the budget is fixed at the CU level. Typically, this process updates the total CU-level budget as a function of the common bins used in the TUs already coded / analyzed when coding the current TU of the considered transform tree. For example, in the case of the SBT (sub-block transform) mode of inter-CU, VVC Draft 4 describes that a CU is divided into two TUs, one of which has a null residual. In this case, this embodiment proposes that the total budget of the normal bin at the CU level is dedicated to the TU with a non-null residual, thereby achieving better coding efficiency for such inter-CUs in SBT mode. Furthermore, in the case of ISP (intra sub-partitioning), this VVC partitioning mode of intra-CU divides a CU into several TUs. In this case, the coding / parsing of a TU attempts to reuse the unused normal bin budget allocated to the preceding TU within the same CU. Figures 15D and 15E show the transform unit coding / parsing process adapted for this embodiment when the normal bin budget is fixed at the CU level. They are adaptations of the processes in Figures 14A and 14B, respectively, for this embodiment.
[0090] In at least one embodiment, the budget of the regular CABAC bins used in residual coding is fixed at the CTU level. This allows for better utilization of the total allowed regular bin budget, and therefore improves coding efficiency compared to previous embodiments. Typically, bins initially dedicated to skip CUs can be advantageously used for other non-skip CUs.
[0091] In at least one embodiment, the budget of the regular CABAC bins used in residual coding is fixed at the picture level. This allows for better utilization of the total allowed regular bin budget, and therefore improves coding efficiency compared to previous embodiments. Typically, bins initially dedicated to skipping a CU can be advantageously used for other non-skip CUs.
[0092] In at least one embodiment, the average rate of the allowed normal bins for a given picture is fixed according to the temporal layer / depth to which the considered picture belongs. For example, in a random access coding structure, the temporal structure of a group of pictures conforms to a hierarchical B-picture arrangement. In this structure, pictures are organized in scalable temporal layers. Pictures from higher layers depend on reference pictures from lower temporal layers. In contrast, pictures from lower layers do not depend on any pictures from higher layers. Higher temporal layers are typically coded with higher quantization parameters and therefore lower quality than pictures from lower layers. Moreover, pictures from lower layers have a significant impact on the overall coding efficiency across the entire sequence. Therefore, it is interesting to encode these pictures with optimal coding efficiency. According to this embodiment, it is proposed to allocate a higher sample-based rate of the normal CABAC bins to pictures from lower layers than to pictures from higher temporal layers.
[0093] 16A and 16B show an example of an embodiment in which the budget for regular bins is allocated at a level higher than the CG level, and two types of CG coding / analysis processes are used. The first type, shown in FIG. 16A, is called "all_bypass" and involves coding all bins associated with the CG in bypass mode. This implies that the magnitude of the transform coefficients of the CG is coded only through the syntax element abs_level. The second type, shown in FIG. 16B, is called "all_regular" and involves coding all bins corresponding to the syntax elements sig_flag, gt1_flag, par_flag, and gt3_flag in regular mode. In this embodiment, considering the budget for regular bins fixed at a level higher than the CG level, the TB coding process can switch between the "all_regular" CG coding mode and the "all_bypass" CG coding mode according to whether the budget of the currently considered regular bin is fully used.
[0094] Various implementations involve decoding. "Decoding," as used herein, may encompass all or some of the processes performed on a received encoded sequence, e.g., to generate a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders of various implementations described herein, such as, for example, the embodiments presented in the diagrams of FIGS. 10-16.
[0095] As a further example, in one embodiment, "decoding" refers to entropy decoding only, in another embodiment, "decoding" refers to differential decoding only, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether a reference to a "decoding process" is intended to refer specifically to a subset of operations or to a broader decoding process in general will be clear based on the context of the particular description and is believed to be well understood by one of ordinary skill in the art.
[0096] Various implementations involve encoding. In a manner similar to the above discussion of "decoding," "encoding," as used herein, may encompass all or some of the processes performed on an input video sequence, e.g., to generate an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as, e.g., partitioning, differential encoding, transforming, quantization, and entropy coding. In various embodiments, such processes also or alternatively include processes performed by the encoders of various implementations described herein, such as, e.g., the embodiments of the figures in Figures 10-16.
[0097] As a further example, in one embodiment, "encoding" refers to entropy encoding only, in another embodiment, "encoding" refers to differential encoding only, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether a reference to an "encoding process" is intended to refer specifically to a subset of operations or to a broader encoding process in general will be clear based on the context of the particular description and is believed to be well understood by one of ordinary skill in the art.
[0098] It should be noted that the syntax elements used herein are descriptive terms, and therefore do not preclude the use of other syntax element names.
[0099] This application describes a variety of aspects, including tools, features, embodiments, models, techniques, and the like. Many of these aspects are described with specificity, often in a manner that may be considered limiting, at least to illustrate their individual characteristics. However, this is for clarity of description and does not limit the application or scope of the aspects. Indeed, all different aspects can be combined or interchanged to provide additional aspects. Moreover, aspects can also be combined or interchanged with aspects described in prior applications. The aspects described and contemplated in this application can be implemented in many different forms. While the diagrams in Figures 1, 2, and 3 above provide some embodiments, other embodiments are contemplated, and the discussion of the figures does not limit the breadth of implementations.
[0100] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image," "picture," and "frame" can be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side and the term "decoded" is used on the decoder side.
[0101] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.
[0102] In this application, various numerical values are used, for example, with respect to block sizes. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0103] References to "one embodiment," "an embodiment," "one implementation," or "implementation," and other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of "one embodiment," "in an embodiment," "in one implementation," or "in an implementation," and any other variations thereof, appearing in various places throughout this specification are not necessarily all referring to the same embodiment.
[0104] Additionally, this application or its claims may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0105] Additionally, this application or its claims may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, predicting information, or inferring information.
[0106] Additionally, this application or its claims may refer to "receiving" various pieces of information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory or optical media storage). Furthermore, "receiving" typically involves various methods during an operation, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0107] It should be understood that the use of any of the following terms " / ," "and / or," and "at least one of" is intended to encompass, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," the selection of only the first listed option (A), the selection of only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B and / or C" and "at least one of A, B, and C," such a statement is intended to encompass the selection of only the first listed option (A), the selection of only the second listed option (B), the selection of only the third listed option (C), the selection of only the first and second listed options (A and B), the selection of only the first and third listed options (A and C), the selection of only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be expanded as many times as the number of items listed, as would be readily apparent to one of ordinary skill in this and related arts.
[0108] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using a high frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Claims
1. 1. A video coding method comprising: a residual coding process for coding a syntax element representing a picture area including coded groups representing sub-blocks of quantized coefficients of a transform block of a transform unit using a limited number of regular bins, wherein the coding is CABAC coding, and the number of regular bins is determined based on a budget allocated to the transform unit based on the number of samples of the transform unit.
2. The method of claim 1 , wherein the transform unit comprises at least two transform blocks.
3. 2. The method of claim 1, wherein the budget allocated among multiple coding groups for a block being coded is decremented each time a binary element is coded using a conventional CABAC coding mode.
4. 1. A video decoding method comprising: a residual decoding process that decodes a bitstream comprising syntax elements representing picture areas including coded groups representing sub-blocks of quantized coefficients of a transform block of a transform unit using regular bins, wherein the decoding is CABAC decoding, and the number of regular bins is determined based on a budget allocated to the transform unit based on the number of samples of the transform unit.
5. The method of claim 4 , wherein the transform unit comprises at least two transform blocks.
6. 5. The method of claim 4, wherein the budget allocated among multiple coding groups for a block being decoded is decremented each time a binary element is decoded using a normal CABAC decoding mode.
7. 1. A video encoding device comprising: a processor configured to encode video using a residual encoding process that encodes syntax elements representing picture areas including coded groups representing sub-blocks of quantized coefficients of a transform block of a transform unit using a limited number of regular bins, wherein the encoding is CABAC encoding, and the number of regular bins is determined based on a budget allocated to the transform unit based on the number of samples of the transform unit.
8. A video decoding device comprising: a processor configured to decode video using a residual decoding process that decodes a bitstream including syntax elements representing picture areas including coded groups representing sub-blocks of quantized coefficients of a transform block of a transform unit using regular bins, wherein the decoding is CABAC decoding, and the number of regular bins is determined based on a budget allocated to the transform unit based on the number of samples of the transform unit.
9. A computer program comprising program code instructions for carrying out the steps of the method according to at least one of claims 1 to 6 when the computer program is executed by a processor.
10. A non-transitory computer-readable storage medium storing instructions for performing the steps of the method according to at least one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Method and Apparatus for Providing Rate Control for Panel-Based Real Time Video Encoder
US20080151998A1
Frame buffer compression for video processing devices
US20100220783A1
Throughput improvement for cabac coefficient level coding
US20130182757A1