Flexible allocation of regular bins in residual coding for video coding
By allocating the budget of normal CABAC bins over larger picture areas and using a CABAC optimizer to manage encoding, the challenges of suboptimal video compression and low bin budget utilization in current techniques are addressed, resulting in improved compression efficiency and video coding performance.
Patent Information
- Application Number
- JP2024111154
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-23
- Filing Date
- 2024-07-10
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2040-03-06
AI Technical Summary
Current video encoding techniques face challenges in achieving high compression efficiency due to constraints on the maximum usage of normal CABAC coding bins in residual coding, leading to suboptimal video compression and low utilization of the allocated bin budget.
The proposed solution involves allocating the budget of normal CABAC bins over a larger picture area, determining the number of normal bins based on a budget allocated among encoding or coded groups, and using a CABAC optimizer to manage the encoding process, allowing for more efficient distribution of bins across smaller units such as transform blocks or coding units.
This approach improves compression efficiency by allowing for better utilization of the allocated bin budget, adapting to the energy distribution in video content, and optimizing the use of normal CABAC bins across different coding units, thereby enhancing overall video coding performance.
Smart Images

Figure 0007679532000001 
Figure 0007679532000002 
Figure 0007679532000003
Abstract
Description
Technical Field
[0001] Technical Field At least one of the present embodiments generally relates to the assignment of normal bins in residual coding for video encoding or decoding.
Background Art
[0002] Background To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to exploit spatial and temporal redundancies in video content. Generally, intra or inter prediction is used to utilize intra or inter-frame dependencies, and then the difference between the original block and the predicted block (often represented as a prediction error or prediction residue) is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by the reverse processes for entropy coding, quantization, transformation, and prediction.
Summary of the Invention
[0003] Summary One or more of the present embodiments address high-level constraints on the maximum usage of normal CABAC coding of bins in the residual coding process of blocks and their coefficient groups such that high-level constraints are respected and compression efficiency is improved compared to current techniques.
[0004] According to a first aspect of at least one embodiment, a video encoding method includes a residual encoding process that encodes a syntax element representing a picture area including encoding groups using a limited number of normal bins, the encoding is CABAC encoding, and the number of normal bins is determined based on a budget allocated among a number of encoding groups.
[0005] According to a second aspect of at least one embodiment, the video decoding method includes a residual decoding process of analyzing a bitstream representing a picture area including coded groups using normal bins, the decoding is performed using CABAC decoding, and the number of normal bins is determined based on a budget allocated among a number of coded groups.
[0006] According to a third aspect of at least one embodiment, the apparatus includes an encoder for encoding picture data for at least one block of a picture or video, the encoder is configured to execute a residual encoding process of encoding a syntax element representing a picture area including coded groups using a limited number of normal bins, the encoding is CABAC encoding, and the number of normal bins is determined based on a budget allocated among a number of coded groups.
[0007] According to a fourth aspect of at least one embodiment, the apparatus includes a decoder for decoding picture data for at least one block of a picture or video, the decoder is configured to execute a residual decoding process of analyzing a bitstream representing a picture area including coded groups using normal bins, the decoding is performed using CABAC decoding, and the number of normal bins is determined based on a budget allocated among a number of coded groups.
[0008] According to a fifth aspect of at least one embodiment, there is presented a computer program including program code instructions executable by a processor, the computer program implementing the steps of the method according to at least the first or second aspect.
[0009] According to a sixth aspect of at least one embodiment, there is presented a computer program product stored on a non-transitory computer-readable medium and including program code instructions executable by a processor, the computer program product implementing the steps of the method according to at least the first or second aspect.
Brief Description of the Drawings
[0010] Brief Description of the Drawings
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Figure 8C
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13A
Figure 13B
Figure 14A
Figure 14B
Figure 14C
Figure 15A
Figure 15B
Figure 15C
Figure 15D
Figure 15E
Figure 16A
Figure 16B
[0011] **Detailed Description** Various embodiments relate to entropy coding of quantized transform coefficients. This stage of the video codec is also referred to as the residual coding step. At least one embodiment aims to optimize the coding efficiency of the video codec under the constraint of the maximum number of normal coding bins per single area.
[0012] The various methods and other aspects described in this application can be used to modify at least the entropy coding and / or decoding modules (145, 230) of the video encoder 100 and decoder 200 as shown in FIGS. 1 and 2. Moreover, while this aspect describes principles related to specific drafts of the VVC (Versatile Video Coding) or HEVC (High Efficiency Video Coding) specifications, it is not limited to VVC or HEVC, and can be applied to other standards and recommendations, as well as extended versions of such standards and recommendations (including VVC and HEVC), whether existing or developed in the future. Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.
[0013] FIG. 1 shows a block diagram of an example of a video encoder 100 such as an HEVC encoder. Further, FIG. 1 also shows an encoder in which improvements have been made to the HEVC standard, or an encoder that employs the same technology as HEVC, such as a Joint Exploration Model (JEM) or VVC Test Model (VTM) encoder developed under the Joint Video Exploration Team (JVET) for VVC.
[0014] Before being encoded, the video sequence can undergo pre-encoding processing (101). For example, this processing can be performed by applying a color conversion to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or by performing remapping of the input picture components to obtain a signal distribution with faster resilience to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the pre-processing and placed in the bitstream.
[0015] In HEVC, to encode a video sequence having one or more pictures, the pictures are partitioned into one or more slices (102), and each slice can include one or more slice segments. The slice segments are organized into coding units, prediction units, and transform units. The HEVC specification distinguishes between "blocks" and "units", where a "block" addresses a specific area of the sample array (e.g., luma, Y), and a "unit" includes an array block of all encoded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data associated with the block (e.g., motion vectors).
[0016] For encoding with HEVC, a picture is partitioned into square coding tree blocks (CTBs) having configurable sizes, and a consecutive set of coding tree blocks is grouped into slices. A coding tree unit (CTU) encloses the CTBs of the encoded color components. A CTB is the root of a quadtree partitioned into coding blocks (CBs), and a coding block can be partitioned into one or more prediction blocks (PBs) and forms the root of a quadtree partitioned into transform blocks (TBs). Corresponding to the coding block, prediction block, and transform block, a coding unit (CU) includes a tree structure set of prediction units (PUs) and transform units (TUs), a PU includes prediction information for all color components, and a TU includes a residual coding syntax structure for each color component. The sizes of the CB, PB, and TB of the luma component apply to the corresponding CU, PU, and TU. In the present application, the term "block" can be used to refer to any one of, for example, a CTU, CU, PU, TU, CB, PB, and TB. In addition, "block" can also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, and more generally, can also be used to refer to an array of data of various sizes.
[0017] In the example of encoder 100, a picture is encoded by encoder elements as described below. A picture to be encoded is processed in units of CUs. Each CU is encoded using either an intra or inter mode. When a CU is encoded in the intra mode, intra prediction (160) is performed. In the inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder determines (105) whether to use the intra mode or the inter mode for encoding the CU, and the prediction mode flag indicates the intra / inter decision. The prediction residual is calculated by subtracting the prediction block from the original image block (110).
[0018] In intra mode, the CU is predicted from the reconstructed neighboring samples within the same slice. In HEVC, a set of 35 intra prediction modes is available, including one DC mode, one planar mode, and 33 angular prediction modes. Intra prediction references are reconstructed from the rows and columns adjacent to the current block. The references use available samples from previously reconstructed blocks and extend horizontally and vertically by more than twice the block size. When an angular prediction mode is used for intra prediction, the reference samples can be copied along the direction indicated by the angular prediction mode.
[0019] The applicable luma intra prediction modes for the current block can be coded using two different options. If the applicable modes are included in the construction list of the three most probable modes (MPM), the mode is signaled by the index of the MPM list. Otherwise, the mode is signaled by the fixed-length binary of the mode index. The three most probable modes are derived from the intra prediction modes of the upper and left neighboring blocks.
[0020] In the case of an inter CU, the corresponding coded block is further partitioned into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU encapsulates information about how the inter prediction is performed. Motion information (e.g., motion vectors and reference picture indices) can be signaled in two ways, namely, "merge mode" and "advanced motion vector prediction (AMVP)".
[0021] In merge mode, the video encoder or decoder assembles a candidate list based on already coded blocks, and the video encoder signals an index for one of the candidates within the candidate list. On the decoder side, the motion vector (MV) and reference picture index are reconstructed based on the signaled candidate.
[0022] In AMVP, the video encoder or decoder assembles a candidate list based on motion vectors determined from already coded blocks. Then, the video encoder signals the index of the candidate list to identify the motion vector predictor (MVP) and signals the motion vector difference (MVD). On the decoder side, the motion vector (MV) is reconstructed as MVP + MVD. Also, applicable reference picture indices are explicitly coded in the PU syntax for AMVP.
[0023] Then, transform (125) and quantization (130) of the prediction residual are performed, including at least one embodiment for adapting the chroma quantization parameter described below. The transform is generally based on a separable transform. For example, the DCT transform is first applied horizontally and then vertically. In previous codecs, the various 2D transforms for a given block size are usually limited, but in recent codecs such as JEM, the transforms used in both directions can be different (e.g., DCT in one direction and DST in the other), thereby enabling a wide variety of 2D transforms.
[0024] The quantized transform coefficients along with the motion vectors and other syntax elements are entropy coded (145) to output a bitstream. Also, the encoder can skip the transform and directly apply quantization to the non-transformed residual signal on a 4×4 TU basis. The encoder can also avoid both the transform and quantization (i.e., the residual is directly coded without applying either the transform process or the quantization process). In direct PCM coding, prediction is not applied, and the coding unit samples are directly coded and embedded in the bitstream.
[0025] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. The image block is reconstructed by combining the decoded residual and the prediction block (155). The in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (180).
[0026] FIG. 2 shows a block diagram of an example of a video decoder 200 such as an HEVC decoder. In the example of decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding path that is the reverse of the encoding path as described in FIG. 1 that performs video decoding as part of the encoding of video data. Also, FIG. 2 shows a decoder in which improvements have been made to the HEVC standard, or a decoder that employs a technology similar to HEVC such as a JEM or VVC decoder.
[0027] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, picture partitioning information, and other coded information. The picture partitioning information indicates the size of the CTU and the way the CTU is divided into CUs, and, if applicable, probably the way it is divided into PUs. Thus, the decoder can divide the picture into CTUs and each CTU into CUs according to the decoded picture partitioning information (235). The transform coefficients are inverse quantized (240) and inverse transformed (250), including at least one embodiment for adapting chroma quantization parameters as described below, to decode the prediction residual.
[0028] The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. The prediction block can be obtained from intra prediction (260) or motion compensated prediction (i.e., inter prediction) (275) (270). As described above, the AMVP and merge mode techniques can be used to derive motion vectors for motion compensation, and in motion compensation, an interpolation filter can be used to calculate interpolation values for sub-integer samples of the reference block. The in-loop filter (265) is applied to the reconstructed picture. The filtered image is stored in the reference picture buffer (280).
[0029] The decoded picture can further undergo post-decoding processing (285), for example, inverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping process executed in the pre-encoding processing (101). The post-decoding processing can use metadata, which is derived in the pre-encoding processing and signaled in the bitstream.
[0030] Figure 3 shows a block diagram of an example of a system in which various aspects and embodiments are implemented. System 300 can be embodied as a device that includes various components described below and is configured to execute one or more of the aspects described in this application. Examples of such devices include, but are not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, encoders, transcoders, and servers, and various other electronic devices. The elements of system 300 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 300 are distributed across multiple ICs and / or discrete components. In various embodiments, the elements of system 300 are communicatively coupled through internal bus 310. In various embodiments, system 300 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 300 is configured to implement one or more of the aspects described in this document, such as video encoder 100 and video decoder 200, as described above and modified as described below.
[0031] System 300 includes at least one processor 301 configured to execute instructions loaded therein to implement various aspects described, for example, in this document. The processor 301 may include an embedded memory, an input / output interface, and various other circuits as known in the art. System 300 includes at least one memory 302 (e.g., a volatile memory device and / or a non-volatile memory device). System 300 includes a storage device 304 that may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 304 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0032] System 300 includes an encoder / decoder module 303 configured to process data to provide encoded video or decoded video, for example. The encoder / decoder module 303 may include its own processor and memory. The encoder / decoder module 303 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. In addition, the encoder / decoder module 303 can be implemented as a separate element of the system 300 or incorporated within the processor 301 as a combination of hardware and software, as is known to those skilled in the art.
[0033] The program code scheduled to be loaded into the processor 301 or the encoder / decoder 303 to execute the various aspects described in this document can be stored in the storage device 304 and then loaded into the memory 302 for execution by the processor 301. According to various embodiments, one or more of the processor 301, the memory 302, the storage device 304, and the encoder / decoder module 303 can store one or more of the various items during the execution of the processes described in this document. Such stored items can include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstream, matrix, variable, and intermediate or final results from equations, formulas, operations, and operation logic processing.
[0034] In some embodiments, the internal memory of the processor 301 and / or the encoder / decoder module 303 is used to provide a working memory for storing instructions and processing what is necessary during encoding or decoding. However, in other embodiments, for one or more of these functions, an external memory to the processing device (e.g., the processing device can be the processor 301 or the encoder / decoder module 303) is used. The external memory can be the memory 302 and / or the storage device 304 and can be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, an external high-speed dynamic volatile memory such as RAM is used as a working memory for video encoding and decoding operations such as those for MPEG-2, HEVC, or VVC.
[0035] Inputs to the elements of system 300 can be provided through various input devices as shown in block 309. Such input devices include, but are not limited to, (i) an RF portion that receives RF signals wirelessly transmitted, for example, by a broadcast station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0036] In various embodiments, the input device of block 309 has respective associated input processing elements as known in the art. For example, the RF portion can be associated with elements necessary for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) again bandlimiting to a narrower frequency band (e.g., to select a signal frequency band, which in some embodiments can also be called a channel), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various embodiments includes one or more elements for performing these functions, such as, for example, a frequency selector, a signal selector, a bandlimiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF portion can include a tuner, which performs these various functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one embodiment of a set-top box, the RF portion and its associated input processing elements receive an RF signal transmitted over a wired (e.g., cable) medium and then perform frequency selection by filtering, downconverting, and filtering again to obtain a desired frequency band. Various embodiments perform reordering of the (and other) elements described above, removing some of these elements, and / or adding other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements (e.g., inserting an amplifier and an analog-to-digital converter). In various embodiments, the RF portion includes an antenna.
[0037] In addition, the USB and / or HDMI terminals may each include an interface processor for connecting the system 300 to other electronic devices through a USB and / or HDMI connection. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented, for example, in a separate input processing IC or within the processor 301 as needed. Similarly, aspects of USB or HDMI interface processing can be implemented, as needed, in a separate interface IC or within the processor 301. The demodulated, error-corrected, and de-multiplexed streams are provided to various processing elements (including, for example, the processor 301 and the encoder / decoder 303 in conjunction with memory and storage elements) to process a data stream for presentation on an output device as needed.
[0038] The various elements of the system 300 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected, and data can be transmitted between the various elements using a suitable connection arrangement, such as an internal bus (including I2C buses, wiring, and printed circuit boards) known in the art.
[0039] The system 300 includes a communication interface 305 that enables communication with other devices via a communication channel 320. The communication interface 305 can include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 320. The communication interface 305 can include, but is not limited to, a modem or a network card, and the communication channel 320 can be implemented, for example, within a wired and / or wireless medium.
[0040] In various embodiments, data is streamed to system 300 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals in these embodiments are received on communication channel 320 and communication interface 305 adapted for Wi-Fi communication. The communication channel 320 in these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments use a set-top box that transfers data on the HDMI connection of input block 309 to provide the streamed data to system 300. Still other embodiments use the RF connection of input block 309 to provide the streamed data to system 300.
[0041] System 300 can provide output signals to various output devices, including display 330, speaker 340, and other peripheral devices 350. Other peripheral devices 350 include, in various examples of embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of system 300. In various embodiments, control signals are transmitted between system 300 and display 330, speaker 340, or other peripheral devices 350 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 300 through their respective interfaces 306, 307, and 308 via dedicated connections. Alternatively, the output devices can be connected to system 300 via communication channel 320 through communication interface 305. Display 330 and speaker 340 can be integrated into a single unit with other components of system 300 of an electronic device (e.g., a television, etc.). In various embodiments, display interface 306 includes a display driver, such as a timing controller (T Con) chip, for example.
[0042] For example, if the RF portion of input 309 is part of a separate set-top box, display 330 and speaker 340 can be selectively separated from one or more of the other components. In various embodiments where display 330 and speaker 340 are external components, the output signal can be provided via a dedicated output connection (including, for example, an HDMI port, a USB port, or a COMP output). The implementations described herein can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., only discussed as a method), the implementation of the features discussed can also be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. The method can be implemented, for example, with an apparatus (e.g., a processor, etc.) that generally refers to a processing device, including a computer, a microprocessor, an integrated circuit, or a programmable logic device. Also, the processor includes, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant (``PDA''), and other devices that facilitate the communication of information among end users.
[0043] FIG. 4A shows an example of a coded tree unit and a coded tree in the compression region. In the HEVC video compression standard, motion-compensated temporal prediction is adopted to utilize the redundancy existing between consecutive pictures of a video. To achieve this, a picture is partitioned into so-called coded tree units (CTUs), the size of which is typically 64×64, 128×128, or 256×256 pixels. Each CTU is represented by a coded tree in the compression region (e.g., a quadtree partition of the CTU). Each leaf is called a coding unit (CU).
[0044] FIG. 4B shows an example of the division of a CTU into coding units, prediction units, and transform units. Each CU is then given some intra or inter prediction parameter prediction information). To achieve this, each CU is spatially partitioned into one or more prediction units (PUs), and some prediction information such as motion vectors is assigned to each PU. The intra or inter coding mode is assigned at the CU level.
[0045] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "encoded" and "coded" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Generally speaking, although it cannot be said categorically, the term "reconstructed" is used on the encoder side, and the term "decoded" is used on the decoder side. The term "block" or "picture block" can be used to refer to any one of CTU, CU, PU, TU, CB, PB, and TB. In addition, the term "block" or "picture block" can be used to refer to macroblocks, partitions, and subblocks as specified in H.264 / AVC or other video coding standards, and more generally, can be used to refer to an array of samples of many sizes.
[0046] Figure 5 shows the use of two scalar quantizers in dependency scalar quantization. Dependency scalar quantization uses two scalar quantizers with different reconstruction levels for quantization, as proposed in JVET (submission JVET-J0014). Compared to conventional scalar quantization (such as that used in HEVC and VTM-1), the main effect of this approach is that the set of allowable reconstruction values for the transform coefficients depends on the values of the transform coefficient levels that precede the current transform coefficient level in the reconstruction order. The dependency scalar quantization technique is realized by (a) defining two scalar quantizers with different reconstruction levels and (b) defining a process for switching between the two scalar quantizers. The two scalar quantizers used are shown in Figure 5 (represented by Q0 and Q1). The location of the available reconstruction levels is uniquely specified by the quantization step size Δ. Ignoring the fact that the actual reconstruction of the transform coefficients uses integer arithmetic, the two scalar quantizers Q0 and Q1 are characterized as follows. Q0: The reconstruction level of the first quantizer Q0 is given by an even integer multiple of the quantization step size Δ. When this quantizer is used, the inverse quantized transform coefficient t’ is t’ = 2·k·Δ calculated according to, where k represents the associated quantized coefficient (the transmitted quantization index). Q1: The reconstruction level of the second quantizer Q1 is given by an odd integer multiple of the quantization step size Δ in addition to a reconstruction level equal to zero. The inverse quantized transform coefficient t’ is t’=(2·k - sgn(k))·Δ computed as a function of the quantized coefficient k, where sgn(·) is sgn(x)=(k == 0? 0 : (k < 0? -1 : 1)) the sign function defined as.
[0047] The scalar quantizer (Q0 or Q1) used is not explicitly signaled in the bitstream. Instead, the quantizer used for the current transform coefficient is determined by the parity of the quantized coefficients preceding the current transform coefficient in the coding / reconstruction order and the state of the finite state machine introduced below.
[0048] Figures 6A and 6B show an example of a mechanism for switching between scalar quantizers in VVC. The switching between the two scalar quantizers (Q0 and Q1) is realized via a finite state machine having four states (labeled as 0, 1, 2, or 3 respectively) as shown in Figure 6A. The state of the finite state machine considered for a given quantized coefficient is uniquely determined by the parity of the quantized coefficient k preceding the current quantized coefficient in the coding / reconstruction order and the state of the finite state machine considered when processing this preceding coefficient. At the start of inverse quantization for a transform block, the state is set equal to 0. The transform coefficients are reconstructed in scan order (i.e., the same order as the order in which they are entropy decoded). After the current transform coefficient is reconstructed, the state is updated. k is a quantized coefficient. The next state is state=stateTransTable[current_state][k&1] depending on the current state and the parity of the current quantized coefficient k (k&1) as shown, where stateTransTable represents the state transition table shown in Figure 6B and the operator & specifies the bit 'product' operator in two's complement arithmetic.
[0049] Moreover, the quantized coefficients included in a so-called transform block (TB) can be entropy encoded and decoded as described below.
[0050] First, the transform block is divided into 4×4 sub-blocks of quantized coefficients called coded groups (sometimes also named coefficient groups and abbreviated as CGs). Entropy coding / decoding is made up of several scan paths, and the scan paths scan the TB according to the diagonal scan order shown in FIG. 6C.
[0051] The transform coefficient coding in VVC involves five main steps, namely, scanning, final significant coefficient coding, significance map coding, coefficient level residue coding, and absolute level and sign data coding.
[0052] FIG. 6C shows an example of the scan order between CGs and between coefficients as used in the VVC of an 8×8 transform block. The scan path across the TB includes processing each CG sequentially according to the diagonal scan order, and the 16 coefficients within each CG are also scanned according to the considered scan order. The scan path starts from the last significant coefficient of the TB and processes all coefficients up to the DC coefficient.
[0053] The entropy coding of transform coefficients includes up to seven syntax elements in the following list. - coded_sub_block_flag: Significance of the coefficient group (CG) - sig_flag: Significance of the coefficient (zero / non-zero) - gt1_flag: Indicates whether the absolute value of the coefficient level is greater than 1 - par_flag: Indicates the parity of the coefficient greater than 1 - gt3_flag: Indicates whether the absolute value of the coefficient level is greater than 3 - remainder: Residual value for the absolute value of the coefficient level (if the value is greater than that coded in the previous path) - abs_level: Value of the absolute value of the coefficient level (if the CABAC bin is not signaled for the current coefficient for the maximum number of bin-budget problem) - sign_data: The signs of all significant coefficients included in the considered CG. It contains a series of bins, each of which signals the sign (0: positive, 1: negative) of each non-zero transform coefficient.
[0054] When the absolute value of the quantized coefficient is known by decoding a subset of the above elements (apart from the signs), no further syntax element coding for that coefficient with respect to its absolute value is done. Similarly, the sign flag is signaled only for non-zero coefficients.
[0055] All the scan paths required for a given CG are coded until all the quantized coefficients of that CG can be reconstructed before proceeding to the next CG.
[0056] The overall decoding TB analysis process is made up of the following main steps. 1. Decode the final significant coefficient coordinates. This includes syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix and last_sig_coeff_y_suffix. This provides the decoder with the spatial positions (x and y coordinates) of the final non-zero coefficients in the entire TB.
[0057] Then, for each consecutive CG from the CG containing the final significant coefficient of the TB to the top-left CG of the TB, the following steps are applied. 2. Decode the CG significance flag (referred to as coded_sub_block_flag in the VVC specification) 3. Decode the significant coefficient flag for each coefficient of the considered CG. This corresponds to the syntax element sig_flag. This indicates which coefficients of the CG are non-zero.
[0058] The next analysis stage is related to the coefficient level for the coefficients known as non-zero in the considered CG. This involves the following syntax elements. 4.gt1_flag: This flag indicates whether the absolute value of the current coefficient is greater than 1. If it is not greater than 1, the absolute value is equal to 1. 5.par_flag: This flag indicates whether the current quantized coefficient is even. If the gt1_flag of the current quantized coefficient is true, it is coded. When the par_flag is zero, the quantized coefficient is even; otherwise, the quantized coefficient is odd. After the par_flag is analyzed on the decoder side, the partially decoded quantized coefficient is set equal to (1 + gt1_flag + par_flag). 6.gt3_flag: This flag indicates whether the absolute value of the current coefficient is greater than 3. If it is not greater than 3, the absolute value is equal to 1 + gt1_flag + par_flag. When 1 + gt1_flag + par_flag is 2 or more, the gt3_flag is coded. When the gt3_flag is analyzed, the quantized coefficient value becomes 1 + gt1_flag + par_flag + (2 * gt3_flag) on the decoder side. 7.remainder: This codes the absolute value of the coefficient. This applies when the partially decoded absolute value is equal to or greater than 4. Note that in the example of VVC draft 3, for each coding group, the maximum number of normal coding bins' budget is fixed. Therefore, for some coefficients, only elements such as sig_flag, gt1_flag, and par_flag can be signaled, while for other coefficients, gt3_flag can also be signaled. Therefore, the coded and analyzed values of the remainder are calculated for the flags that have already been decoded for the coefficients under consideration, and thus are calculated as a function of the partially decoded quantized coefficient. 8.abs_level: This indicates the absolute value of the coefficients for which none of the considered CG flags (among sig_flag, gt1_flag, papr_flag, or gt3_flag) are coded, with respect to the problem of the maximum number of normal coding bins. This syntax element, like the remainder syntax element, is subject to Rice-Golomb binary evolution and bypass coding. 9.sign_flag: This indicates the sign of the non-zero coefficients. This is bypass-coded.
[0059] Figure 7A shows the syntax elements for coding / parsing for a transform block according to VVC draft 4. The coding / pairing involves a 4-pass process. As can be seen from the figure, for a given transform block, the position of the last significant coefficient is coded first. Then, for each CG (excluding this CG) containing the last significant coefficient and each CG from the first CG, the significance of the CG is signaled (coded_sub_block_flag), and in the case where the CG is significant, residual coding for that CG is performed as shown in Figure 7B.
[0060] Figure 7B shows the CG-level residual coding syntax as specified in VVC draft 4. The syntax signals the syntax elements of the considered CG, namely sig_flag, gt1_flag, par_flag, gt3_flag, remainder, abs_level, and sign_data, as introduced previously. The figure shows the signaling and parsing of the syntax elements sig_flag, gt1_flag, par_flag, gt3_flag, remainder, and abs_level according to VVC draft 4. EP means "equi-probable", meaning that the corresponding bin is not arithmetically coded but is coded in bypass mode. The bypass mode performs a direct write / parse of the bits, which is generally equal to the binary syntax element (bin) for which encoding or parsing is desired.
[0061] Moreover, VVC draft 4 specifies a hierarchical syntax array from the CU level to the residual sub-block (CG) level.
[0062] FIG. 8A shows the CU level syntax as specified in the example of VVC draft 4. The syntax includes signaling the coding mode of the CU under consideration. The cu_skip_flag is signaled when the CU is in the merge skip mode. When the CU is not in the merge skip mode, the cu_pred_mode_flag indicates the variable value of cuPredMode, and thus, when the prediction mode of the current CU is coded through intra prediction (cuPredMode = MODE_INTRA) or inter prediction (cuPredMode = MODE_INTER).
[0063] In the case of non - skip and non - intra modes, the following flags indicate the use of the intra block copy (IBC) mode for the current CU. The remainder of the syntax includes the intra or inter prediction data of the current CU. The end of the syntax is dedicated to signaling the transform tree associated with the current CU. This signaling of the transform tree starts from a syntax element called cu_cbf, which indicates that some non - null residual data is coded for the current CU. In the case where this flag is equal to true, the transform_tree associated with the current CU is signaled according to the syntax array shown in FIG. 8B.
[0064] Figure 8B shows the syntax structure of transform_tree for an example of VVC draft 4. The syntax basically includes signaling when a CU encloses one or several transform units (TUs). First, if the CU size is larger than the maximum allowable TU size, the CU is split into four sub-transform trees in a quadtree manner. Otherwise, and when ISP (intra-subdivision) or SBT (sub-block transform) is not used for the current CU, the current transform tree is not subdivided, encloses exactly one TU, and is signaled through the transform unit syntax table in Figure 8C. Otherwise, if the CU is in an intra mode and the ISP (intra-subdivision) mode is used, the CU is divided into several (actually, two) TUs. Each of the two TUs is successively signaled through the transform unit syntax table in Figure 8C. Otherwise, if the CU is in an inter mode and the SBT (sub-block transform) mode is used, the CU is divided into two TUs according to the decoded syntax related to SBT at the CU level. Each of the two TUs is successively signaled through the transform unit syntax table in Figure 8C.
[0065] Figure 8C shows the transform unit level syntax array in an example of VVC draft 4. The syntax is composed of the following. First, the tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr flags respectively indicate that non-null residual data corresponding to each component is enclosed in each transform block of the current TU. For each component, if the corresponding cbf flag is true, the residual_tb syntax table at the transform block level is used to code the corresponding residual transform block. Also, the syntax at the transform_unit level includes some coded, transformed, and quantized block-related syntax (i.e., delta QP information (if present) and the type of transform used to code the considered TU).
[0066] Figures 9A and 9B each show the CABAC decoding and encoding processes. CABAC represents Context-Adaptive Binary Arithmetic Coding, which is a form of entropy coding used, for example, in HEVC or VVC to provide lossless compression with excellent compression efficiency. The input to the process in Figure 9A includes the encoded bitstream, typically conforming to the HEVC specification or a further evolved version thereof. At any point in the decoding process, the decoder knows which syntax elements are to be decoded next. This is well-specified in the standardized bitstream syntax and the decoding process. Moreover, the decoder also knows how to binary evolve the current syntax element to be decoded (i.e., represented as a sequence of binary symbols called bins, each equal to either '1' or '0') and how each bin of the bin string is encoded.
[0067] Therefore, the first stage of the CABAC decoding process (the left side of Figure 9A) is to decode a series of bins. For each bin, the decoder knows whether each bin is encoded according to the bypass mode or the normal mode. In the bypass mode, the bit is simply read from the bitstream, and the resulting value is assigned to the current bin. This mode is straightforward and thus fast and has the advantage of not requiring intensive resource utilization. It is typically efficient and thus used for bins with a uniform statistical distribution (i.e., equal probability of being equal to '1' and equal probability of being equal to '0').
[0068] Conversely, if the current bin is not encoded in the bypass mode, it means that it is encoded in the so-called normal coding (i.e., through context-based arithmetic coding). This mode is much more resource-intensive.
[0069] In that case, the decoding of the bins to be considered proceeds as follows. First, a context for the decoding of the current bin is obtained. The context is given by the context model of FIG. 9A. The goal of the context is to obtain the conditional probability that the current bin has the value "0" considering some context prior or information X. Here, the prior X is some already decoded syntax element value that is synchronously available on both the encoder side and the decoder side when the current bin is being decoded.
[0070] Typically, the prior X used for bin decoding is specified in the standard and is chosen because it is statistically correlated with the current bin to be decoded. What is interesting in the use of this context information is that the rate cost of bin coding is reduced. This is based on the fact that if X is given because there is a correlation between the bin and X, the conditional entropy of the bin becomes lower. The following relationship is well-known in information theory. H(bin│X)<H(bin)
[0071] This means that the conditional entropy of the bin when X is known when the bin and X are statistically correlated is lower than the entropy of the bin. Therefore, the context information X is used to obtain the probability that the bin is "0" or the probability that it is "1". Considering these conditional probabilities, the normal decoding engine of FIG. 14 performs arithmetic decoding of the binary bin. Then, the value of the bin is used to update the value of the conditional probability associated with the current bin for which the current context information X is known. This is called the context model update step of FIG. 9A. As long as the bin is being decoded (or encoded), by updating the context model for each bin, progressive refinement of the context modeling for each binary element becomes possible. Therefore, the CABAC decoder progressively learns the statistical behavior of each of the normal coded bins.
[0072] The context model and the context model update step are exactly the same operations on the encoder side and the decoder side.
[0073] A series of decoded bins are obtained according to how they are coded, by normal arithmetic decoding or bypass decoding of the current bin.
[0074] The second stage of CABAC decoding shown on the right side of FIG. 9A involves converting this series of binary symbols into higher-level syntax elements. The syntax elements can take the form of flags, in which case the value of the current decoded bin is directly incorporated. On the other hand, when the binary evolution of the current syntax element corresponds to a set of several bins according to the standard specification under consideration, a conversion step called "binary codeword for syntax element" in FIG. 9A is performed.
[0075] This step proceeds in reverse to the binary evolution step performed by the encoder as shown in FIG. 9B. Therefore, the inverse transformation performed here involves obtaining the values of these syntax elements based on the decoded binary evolution versions of these syntax elements respectively.
[0076] The encoder 100 of FIG. 1, the decoder 200 of FIG. 2, and the system 1000 of FIG. 3 are adapted to implement at least one of the embodiments described below.
[0077] In the current video coding system (e.g., VVC draft 4), some hard constraints are imposed on the coefficient group coding process, and the maximum number of normal CABAC bins to be hard-coded can be adopted for syntax elements such as sig_flag, gt1_flag, and parity_flag on one hand, and for the syntax element gt3_flag on the other hand. More precisely, the 4×4 coding group (or coefficient group or CG) level constraint ensures that the CABAC decoder engine has to parse the maximum number of normal bins per unit area. However, due to the limitation of the use of normal bin CABAC coding, rate distortion may lead to suboptimal video compression. In fact, when only a limited number of normal bins are allowed for a picture, in a case where a significant area of the picture to be coded is included in a skip mode (therefore, no residual) while another significant area of the picture adopts residual coding, the utilization rate of this total budget can be significantly low. In such a case, due to the low-level (i.e., 4×4 CG level) constraint on the maximum number of normal bins, the total number of normal bins of the picture can be far below the acceptable picture level threshold.
[0078] At least one embodiment relates to handling high-level constraints on the maximum usage of normal CABAC coding of bins in the residual coding process of blocks and their coefficient groups such that high-level constraints are respected and compression efficiency is improved compared to current techniques. In other words, the budget of normal coding bins is allocated over a picture area larger than the CG, and thus a number of CGs are covered, the number being determined from the average allowable number of normal bins per unit area. For example, it is possible to allow an average of 1.75 normal coding bins per sample. This budget can then be distributed more efficiently over smaller units. In different embodiments, higher-level constraints are set at the transform block, transform unit, coding unit, coding tree unit, or picture level. These embodiments can be implemented by a CABAC optimizer as shown in FIG. 10.
[0079] In at least one embodiment, higher-level constraints are set at the transform unit level. In such an embodiment, the number of normal bins allowed for all transform units for encoding / decoding is determined from the average allowable number of normal bins per unit area. From the budget of normal bins obtained for the TU, the number of normal bins allowed for each transform block (TB) of the TU is derived. A transform block is a set of transform coefficients belonging to the same TU and the same color component. Next, considering the number of normal bins allowed in the transform block, residual coding or decoding is applied under this constraint on the number of normal bins allowed for all TBs. Thus, in this specification, a modified residual coding and decoding process is proposed, in which the budget of normal bins at the TB level is considered instead of the budget of normal bins at the CG level. For this purpose, several embodiments are proposed.
[0080] FIG. 10 shows a CABAC encoding process including a CABAC optimizer according to an embodiment of the present principle. In at least one embodiment, the CABAC optimizer 190 handles a budget representing the maximum number of bins encoded (or to be encoded) using normal encoding for a set of coding groups. This budget is determined, for example, by multiplying the surface of the data unit considered in the budget per sample of the normal coding bin (i.e., the number of samples included in that surface). For example, in the case of a budget of normal coding bins fixed at the TB level, the allowable number of normal coding bins per sample is multiplied by the number of samples included in the transform block considered. The output of the CABAC optimizer 190 controls the encoding in the normal coding mode or the bypass coding mode, and thus greatly affects the efficiency of the encoding.
[0081] FIG. 11 shows an example of a flowchart of a CABAC optimizer used in an encoding process according to an embodiment of the present principle. This flowchart is executed for each new set of coding groups considered. Thus, according to different embodiments, this flowchart can occur for each picture, each CTU, each CU, each TU, or each TB. In step 191, the processor 301 assigns a budget for the normal encoding determined as described above. In step 192, the processor 301 loops through the set of coding groups, and in step 193, checks whether the last coding group has been reached. In step 194, for each coding group, the processor processes the coding group with the input budget of the normal coding bin. The budget of the normal coding bin decreases during the processing of the coding groups considered (i.e., the budget of the normal coding bin is decremented each time a bin is encoded in the normal mode). The budget of the normal coding bin is returned as an output parameter of the coding group coding or decoding process.
[0082] FIG. 12 shows an example of a modified process for encoding / decoding coded groups using a CABAC optimizer. As introduced above, the assignment of the number of normal CABAC bins dedicated to residual coding is performed at a level higher than the 4×4 coded group. To do this, the process for encoding / decoding coded groups is modified by adding an input parameter that represents the current budget of normal bins when the CG is being encoded or decoded, compared to the conventional functionality. Thus, this budget is modified by encoding or decoding the current CG according to the number of normal bins used to encode / decrypt the current CG. The input / output parameter representing the budget of normal bins is called numRegBins_in_out in FIG. 12. The process for encoding the residual data itself of a given CG can be the same as the conventional method, except that the allowable number of normal CABAC bins is given by external means. Similar to the case of FIG. 12, this budget is decremented by 1 each time a normal bin is encoded / decoded.
[0083] Accordingly, at least one embodiment of the present disclosure includes processing the budget of normal bins as an input / output parameter of the coefficient group encoding / analysis function. Thus, the budget of normal bins can be determined at a level higher than the CG encoding / analysis level. Different embodiments propose to determine the budget at different levels (i.e., transform block, transform unit, coding unit, coding tree unit, or picture level).
[0084] FIG. 13A shows an example of an embodiment in which the budget of normal bins is determined at the transform block level. In such an embodiment, the budget is determined as a function of two main parameters: the size of the current transform block and the basic budget of a normal bin fixed for a given unit of the picture area. In this specification, the area unit considered is a sample. In the current VVC draft 2, 32 normal bins are allowed for 4×4 CG, which means that on average two normal bins are allowed per component sample. The proposed assignment of normal bins at the transform block level is based on this average rate of normal bins per single area.
[0085] In at least one embodiment, the budget assigned to the transform block under consideration is calculated as the product of the transform block surface and the normally assigned number per sample. Next, the budget is passed to the CG residual coding / analysis process for each significant coefficient group in the transform block under consideration.
[0086] FIG. 13B shows an example of an embodiment in which the budget of normal bins is determined at the transform block level according to the position of the final significant coefficients. According to this variant embodiment, the number of normal bins allowed for the current transform block is calculated as a function of the position of the final significant coefficients in the transform block under consideration. The advantage of such an approach is that it better adapts to the energy contained in the transform block under consideration when assigning normal bins. Typically, more bins can be assigned to high-energy transform blocks and fewer normal bins can be assigned to low-energy transform blocks.
[0087] FIG. 14A shows an example of an embodiment in which the budget of a normal bin is determined at the conversion unit level. As shown, the unitary_budget (i.e., the average rate of normal bins allowed per sample) is still equal to 2. The budget at the TU level of the normal bin is determined by multiplying this rate by the total number of samples encapsulated in the conversion unit considered (i.e., in the example of a 4:2:0 color format, width * height * 3 / 2). This total budget is then used for the coding of all the transform blocks encapsulated in the conversion unit considered. That budget is passed as input / output parameters to the residual transform block coding / analysis process for each color component. In other words, each time a transform block of the conversion unit considered is coded / analyzed, the budget decreases by the number of normal bins in the coding / analysis of the transform blocks that are successively coded / analyzed in that conversion unit.
[0088] Figure 14B shows an example of an embodiment in which the normal bin budget for a picture is determined at the conversion unit level and assigned between conversion blocks as a function of the relative surfaces between different conversion blocks and as a function of the normal bins used in the already coded / parsed conversion blocks of the conversion unit. In this embodiment, the normal bin assigned to the luma TB is 2 / 3 of the total budget at the TU level. Next, the Cb TB budget is half of the remaining budget after the luma TB has been coded / parsed. Finally, the Cr TB budget is set to the remaining TU level budget after the first two TBs have been coded / parsed. Moreover, in a further embodiment, it should be noted that the budget for the entire conversion unit can be shared between the conversion blocks of that TU based on the knowledge that one or more TBs of that TU are coded with a null residual. This is known, for example, through the parsing of the tu_cbf_luma, tu_cbf_cb, or tu_cbf_cr flags. Thus, the TB level residual coding / parsing process is adapted to this variant embodiment as shown by the modified process shown in Figure 14C. Such an embodiment provides better coding efficiency when the budget is assigned at a lower level of the hierarchy. In fact, if some CUs or TBs are of very low energy, for those CUs or TBs, a reduced number of bins can be employed compared to others. Thus, a larger number of normal bins can be used in CUs or TBs where a larger number of bin codings / parsings are required.
[0089] FIG. 15A shows an example of an embodiment in which the budget of a normal bin is assigned at the coding unit level. The budget is calculated in the same way as for the embodiment fixed at the previous TU level, but is calculated based on the sample rate and CU size of the normal bin. Then, the calculated budget of the normal bin is passed to the transform tree coding / analysis procedure of FIG. 15B. FIG. 15B shows the adapted coding / analysis of the transform tree associated with the current CU based on the budget of the normal bin assigned at the CU level. Here, the CU level is basically passed sequentially for the coding of each transform unit included in the transform tree under consideration. Also, the budget decreases by the number of normal bins employed in the coding of each TU. FIG. 15C shows a modified embodiment of the coding tree coding / analysis process when the budget is fixed at the CU level. Typically, this process updates the total budget at the CU level as a function of the normal bins used in the already coded / analyzed TUs when coding the current TU of the transform tree under consideration. For example, in the case of the SBT (sub-block transform) mode between CUs, VVC draft 4 states that the CU is divided into two TUs, one of which has a null residual. In that case, this embodiment proposes that the total budget of normal bins at the CU level is dedicated to the TU with non-null residuals, thereby making the coding efficiency for such inter-CUs in the SBT mode better. Moreover, in the case of ISP (intra-sub-partitioning), this VVC partitioning mode of the intra-CU divides the CU into several TUs. In that case, the coding / analysis of the TU tries to reuse the budget of the unused normal bins assigned to the preceding TU within the same CU. FIGS. 15D and 15E show the transform unit coding / analysis process adapted for this embodiment when the budget of the normal bin is fixed at the CU level. They are the adaptations of the processes of FIGS. 14A and 14B to this embodiment respectively.
[0090] In at least one embodiment, the budget of normal CABAC bins used in residual coding is fixed at the CTU level. This enables better utilization of the total budget of allowable normal bins, and thus improves the coding efficiency compared to previous embodiments. Typically, the bins initially dedicated to skip CUs can advantageously be used for other non-skip CUs.
[0091] In at least one embodiment, the budget of normal CABAC bins used in residual coding is fixed at the picture level. This enables even better utilization of the total budget of allowable normal bins, and thus improves the coding efficiency compared to previous embodiments. Typically, the bins initially dedicated to CU skip can advantageously be used for other non-skip CUs.
[0092] In at least one embodiment, the average rate of normal bins allowed for a given picture is fixed according to the temporal layer / depth to which the picture under consideration belongs. For example, in a random access coding structure, the temporal structure of a group of pictures conforms to a hierarchical B picture arrangement. In this configuration, pictures are organized in scalable temporal layers. Pictures from upper layers depend on reference pictures from lower temporal layers. In contrast, pictures from lower layers do not depend on any upper layer pictures. Upper temporal layers are typically coded with higher quantization parameters than lower layer pictures, and thus are coded with lower quality. Moreover, pictures from lower layers have a significant impact on the overall coding efficiency across the entire sequence. Thus, it is interesting to encode these pictures with optimal coding efficiency. According to this embodiment, it is proposed to assign a higher sample-based rate of normal CABAC bins to pictures from lower layers than to pictures from upper temporal layers.
[0093] Figures 16A and 16B show an example of an embodiment in which the budget of a normal bin is assigned at a level higher than the CG level and are for using two types of CG coding / analysis processes. The first type shown by Figure 16A is called "all_bypass" and includes coding all bins associated with CG in bypass mode. This implies that the magnitude of the CG conversion factor is coded only through the syntax element abs_level. The second type shown by Figure 16B is called "all_regular" and codes all bins corresponding to the syntax elements sig_flag, gt1_flag, par_flag, and gt3_flag in normal mode. In this embodiment, considering the budget of the normal bins fixed at a level higher than the CG level, the TB coding process can switch between the "all_regular" CG coding mode and the "all_bypass" CG coding mode according to whether the considered budget of the currently considered normal bins is fully utilized.
[0094] Various implementations involve decoding. "Decoding", as used in this application, may include all or part of the process performed on the received encoded sequence, for example, to generate a final output suitable for display. In various embodiments, such a process includes one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such a process may also or alternatively include the processes performed by the decoders of the various implementations described in this application, such as the embodiments presented in the figures of Figures 10 to 16.
[0095] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the description of "decoding process" is intended to refer to a specific subset of operations or generally to a broader decoding process will become apparent based on the context of the particular description and is considered to be well understood by those skilled in the art.
[0096] Various implementations involve encoding. In a manner similar to the above considerations regarding "decoding", "encoding", as used in this application, may include all or part of the processes performed on an input video sequence, for example, to generate an encoded bitstream. In various embodiments, such processes may include one or more of the processes typically performed by an encoder, such as segmentation, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes may also or alternatively include the processes performed by the encoders of the various implementations described in this application, such as the embodiments of the figures in FIGS. 10-16.
[0097] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the description of "encoding process" is intended to refer to a specific subset of operations or generally to a broader encoding process will become apparent based on the context of the particular description and is considered to be well understood by those skilled in the art.
[0098] Note that the syntactic elements used in this specification are descriptive terms. Thus, they do not exclude the use of other syntactic element names.
[0099] This application describes various aspects, including tools, features, embodiments, models, techniques, etc. Many of these aspects are described with particularity and are often described in a way that appears to be intended to impose limitations, at least to indicate individual characteristics. However, this is for the purpose of clarifying the description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined or exchanged to provide further aspects. Moreover, aspects can also be combined or exchanged with aspects described in previous applications. The aspects described and contemplated in this application can be implemented in many different forms. The figures of FIGS. 1, 2, and 3 above provide some embodiments, but other embodiments are also contemplated, and the discussion of the figures does not limit the breadth of the implementation forms.
[0100] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Generally speaking, although not always the case, the term "reconstructed" is used on the encoder side, and the term "decoded" is used on the decoder side.
[0101] Various methods are described herein, and each method includes one or more steps or actions for achieving the described method. The order and / or use of specific steps and / or actions can be modified or combined as long as a specific order of steps or actions is not required for the correct operation of the method.
[0102] In this application, for example, various numerical values are used with respect to the block size. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0103] References to "one embodiment", "an embodiment", "one implementation", or "an implementation" and other variations thereof mean that the particular features, structures, characteristics, etc. described in relation to the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment", "in an embodiment", "in one implementation", or "in an implementation" and other arbitrary variations thereof that appear in various places throughout this specification do not necessarily all refer to the same embodiment.
[0104] In addition, this application or its claims may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0105] Furthermore, this application or its claims may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, predicting information, or estimating information.
[0106] In addition, this application or its claims may refer to "receiving" various pieces of information. Receiving is intended to be a broad term, similar to "accessing". Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory or an optical media storage device). Furthermore, "receiving" typically involves in various ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0107] It should be understood that any use of the following, " / ", "and / or", "at least one of ~", for example, in the cases of "A / B", "A and / or B", "at least one of A and B", is intended to include the selection of only the first list option (A), the selection of only the second list option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B and / or C" and "at least one of A, B and C", such a description is intended to include the selection of only the first list option (A), the selection of only the second list option (B), the selection of only the third list option (C), the selection of only the first and second list options (A and B), the selection of only the first and third list options (A and C), the selection of only the second and third list options (B and C), or the selection of all three options (A, B and C). This can be extended for as many items as are listed, as will be readily apparent to those skilled in this and related arts.
[0108] As will be apparent to those skilled in the art, the implementation form can generate various signals formatted to convey information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described implementation forms. For example, the signal can be formatted to convey a bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the high-frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave having the encoded data stream. The information conveyed by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Claims
1. 1. A video encoding method comprising: a residual encoding process for encoding a syntax element representing a picture area including coding groups representing sub-blocks of quantized coefficients of a transform block using a limited number of regular bins, the coding being CABAC coding, and the number of regular bins being determined based on a budget allocated to a number of coding groups based on a number of samples in the plurality of coding groups.
2. 2. The method of claim 1, wherein the budget allocated among multiple coding groups for a block being coded is decremented each time a binary element is coded using a conventional CABAC coding mode.
3. 1. A video decoding method comprising: a residual decoding process that analyzes a bitstream representing a picture area including coded groups representing sub-blocks of quantized coefficients of a transform block using regular bins, the residual decoding process being performed using CABAC decoding, and a number of the regular bins being determined based on a budget allocated to a number of coded groups based on a number of samples in the plurality of coded groups.
4. 4. The method of claim 3, wherein for a block being decoded, the budget allocated among multiple coding groups is decremented each time a binary element is decoded using a normal CABAC decoding mode.
5. 1. A video encoding apparatus comprising: a residual encoding process for encoding a syntax element representing a picture area including coding groups representing sub-blocks of quantized coefficients of a transform block using a limited number of regular bins, wherein the encoding is CABAC encoding, and the number of regular bins is determined based on a budget allocated to a plurality of coding groups based on the number of samples in the plurality of coding groups.
6. 6. The video encoding apparatus of claim 5, wherein the budget allocated among multiple coding groups for a block being coded is decremented each time a binary element is coded using a conventional CABAC coding mode.
7. 1. A video decoding apparatus comprising: a residual decoding process for analyzing a bitstream representing a picture area including coded groups representing sub-blocks of quantized coefficients of a transform block using regular bins, wherein the residual decoding process is performed using CABAC decoding, and the number of regular bins is determined based on a budget allocated to a number of coded groups based on the number of samples in the number of coded groups.
8. 8. The video decoding apparatus of claim 7, wherein the budget allocated among multiple coding groups for a block being decoded is decremented each time a binary element is decoded using a normal CABAC decoding mode.
9. A computer program comprising program code instructions for carrying out the steps of the method according to at least one of claims 1 to 4, when the computer program is executed by a processor.
10. A non-transitory computer-readable storage medium having stored thereon instructions for performing the steps of the method according to at least one of claims 1 to 4 when executed by a processor.
Citation Information
Patent Citations
Method and Apparatus for Providing Rate Control for Panel-Based Real Time Video Encoder
US20080151998A1
Frame buffer compression for video processing devices
US20100220783A1
Throughput improvement for cabac coefficient level coding
US20130182757A1