Flexible bin assignment in residual coding for video coding.

JP2026062716A5Pending Publication Date: 2026-05-21INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2025-12-16
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face constraints in maximizing the use of normal bins in residual coding, limiting compression efficiency.

Method used

Implementing a method that allocates a limited number of normal bins based on a budget among encoding or decoding groups using Context-Adaptive Binary Arithmetic Coding (CABAC) to optimize residual coding processes.

Benefits of technology

Improves compression efficiency by respecting high-level constraints and optimizing the use of normal bins in residual coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for encoding / decoding video. [Solution] The encoding method is based on the CABAC encoding of bins, and a high level of constraint is imposed on the maximum normal CABAC encoding usage of bins. In other words, the budget for normal coded bins is allocated over a picture area larger than the coded group, and thus a large number of coded groups are covered, the number of which is determined from the average allowable number of normal bins per unit area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field At least one of the present embodiments generally relates to the assignment of normal bins in residual coding for video encoding or decoding.

Background Art

[0002] Background To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to exploit spatial and temporal redundancies in video content. In general, intra or inter prediction is used to utilize intra or inter-frame dependencies, and then transformation, quantization, and entropy coding of the difference between the original block and the predicted block (often represented as the prediction error or prediction residue) are performed. To reconstruct the video, the compressed data is decoded by a reverse process for entropy coding, quantization, transformation, and prediction.

Summary of the Invention

[0003] Summary One or more of the present embodiments address high-level constraints on the maximum use of normal CABAC coding of bins in the residual coding process of blocks and their coefficient groups such that high-level constraints are respected and compression efficiency is improved compared to current techniques.

[0004] According to a first aspect of at least one embodiment, a video encoding method includes a residual encoding process that encodes a syntax element representing a picture area including encoding groups using a limited number of normal bins, the encoding is CABAC encoding, and the number of normal bins is determined based on a budget allocated among a number of encoding groups.

[0005] According to a second aspect of at least one embodiment, the video decoding method includes a residual decoding process that analyzes a bitstream representing a picture area containing coded groups using normal bins, the decoding is performed using CABAC decoding, and the number of normal bins is determined based on a budget allocated among a number of coded groups.

[0006] According to a third aspect of at least one embodiment, the apparatus includes an encoder for encoding picture data for at least one block of a picture or video, the encoder being configured to perform a residual coding process that codes syntactic elements representing picture areas containing coding groups using a limited number of normal bins, the coding being CABAC coding, and the number of normal bins being determined based on a budget allocated among a number of coding groups.

[0007] According to a fourth aspect of at least one embodiment, the apparatus includes a decoder for decoding picture data for at least one block of a picture or video, the decoder is configured to perform a residual decoding process that analyzes a bitstream representing a picture area containing coded groups using normal bins, the decoding is performed using CABAC decoding, the number of normal bins is determined based on a budget allocated among a number of coded groups.

[0008] According to a fifth aspect of at least one embodiment, a computer program is presented which includes program code instructions executable by a processor, and the computer program performs steps of the method according to at least one aspect of the first or second embodiment.

[0009] According to a sixth aspect of at least one embodiment, a computer program product is presented which includes program code instructions stored on a non-temporary computer-readable medium and executable by a processor, and the computer program product performs steps of at least the first or second aspect of the method. [Brief explanation of the drawing]

[0010] Brief explanation of the drawing [Figure 1] A block diagram of 100 example video encoders, such as a High Efficiency Video Coding (HEVC) encoder, is shown. [Figure 2] A block diagram of an example video decoder 200, such as an HEVC decoder, is shown. [Figure 3] A block diagram of an example system in which various aspects and embodiments are implemented is shown. [Figure 4A] Examples of coded tree units and coded trees in a compressed domain are shown. [Figure 4B] An example of dividing CTU into coding units, prediction units, and transformation units is shown. [Figure 5] This demonstrates the use of two scalar quantizers in dependent scalar quantization. [Figure 6A] Here is an example of a mechanism for switching between scalar quantizers. [Figure 6B] Here is an example of a mechanism for switching between scalar quantizers. [Figure 6C] This shows an example of a scan sequence between computer graphics (CGs) and coefficients, as used in a VVC (Variable Computation Block) of an 8x8 transformation block. [Figure 7A] Examples of syntactic elements for coding / parsing transformation blocks are shown. [Figure 7B] An example of a CG-level residual coding syntax is shown. [Figure 8A] An example of CU-level syntax is shown. [Figure 8B] An example of the `transform_tree` syntax structure is shown. [Figure 8C] This shows the syntax array at the conversion unit level. [Figure 9A] This shows the CABACcrypt process. [Figure 9B] This shows the CABAC coding process. [Figure 10] This document shows a CABAC encoding process including a CABAC optimizer according to an embodiment of this principle. [Figure 11] An example of a flowchart of a CABAC optimizer used in an encoding process according to an embodiment of the present principle is shown. [Figure 12] An example of a modified process for encoding / decoding a coding group using a CABAC optimizer is shown. [Figure 13A] An example of an embodiment where the budget of a normal bin is determined at the transform block level is shown. [Figure 13B] An example of an embodiment where the budget of a normal bin is determined at the transform block level according to the position of the last significant coefficient is shown. [Figure 14A] An example of an embodiment where the budget of a normal bin is determined at the transform unit level is shown. [Figure 14B] An example of an embodiment where the budget of a normal bin is determined at the transform unit level and is allocated between transform blocks as a function of the relative surface between different transform blocks and as a function of the normal bins used in the already encoded / parsed transform blocks of the transform unit being considered. [Figure 14C] An example of an embodiment of a process for residual encoding / parsing at the TB level is shown. [[ID=二十一]] [[ID=二十二]] [Figure 15A] An example of an embodiment where the budget of a normal bin is allocated at the coding unit level is shown. [Figure 15B] An example of an embodiment of the encoding / parsing of a transform tree associated with a current CU based on the budget of a normal bin allocated at the CU level is shown. [Figure 15C] An example of a modified embodiment of the coding tree encoding / parsing process when the budget is fixed at the CU level is shown. [Figure 15D] The transform unit encoding / parsing process adapted for this embodiment when the budget of a normal bin is fixed at the CU level is shown. [Figure 15E] The transform unit encoding / parsing process adapted for this embodiment when the budget of a normal bin is fixed at the CU level is shown. [Figure 16A]This example illustrates an embodiment in which the budget for a normal bin is allocated at a higher level than the CG level, and is intended for use with a first type of CG coding / analysis process. [Figure 16B] This example illustrates an embodiment in which the budget for a normal bin is allocated at a higher level than the CG level, and is intended for use with a second type of CG coding / analysis process. [Modes for carrying out the invention]

[0011] Detailed explanation Various embodiments relate to the entropy coding of quantized transformation coefficients. This stage of the video codec is also called the residual coding step. At least one embodiment aims to optimize the coding efficiency of the video codec under the constraint of the maximum number of normal coding bins per single area.

[0012] Various methods and other embodiments described in this application can be used to modify at least the entropy coding and / or decoding modules (145, 230) of a video encoder 100 and decoder 200 as shown in Figures 1 and 2. Furthermore, although these embodiments describe principles relating to specific drafts of the VVC (Multipurpose Video Coding) or HEVC (High Efficiency Video Coding) specifications, they are not limited to VVC or HEVC and can be applied to other standards and recommendations, whether existing or to be developed in the future, as well as extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the embodiments described in this application can be used individually or in combination.

[0013] Figure 1 shows a block diagram of an example of a video encoder 100, such as an HEVC encoder. Figure 1 also shows encoders that employ technologies similar to HEVC, such as encoders that have been improved over the HEVC standard, or encoders such as the Joint Exploration Model (JEM) or VVC Test Model (VTM) encoder developed under the Joint Video Exploration Team (JVET) for VVC.

[0014] Before encoding, a video sequence can undergo pre-encoding processing (101). For example, this processing can be performed by applying a color conversion to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or by remapping the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the pre-processing and included in the bitstream.

[0015] In HEVC, to encode a video sequence having one or more pictures, the picture is divided into one or more slices (102), and each slice may contain one or more slice segments. The slice segments are organized into coded units, predicted units and transformed units. The HEVC specification distinguishes between “blocks” and “units,” where a “block” deals with a specific area of ​​a sample array (e.g., luma, Y), and a “unit” includes a block (e.g., a motion vector) and an array block of all encoded color components (Y, Cb, Cr, or monochrome), syntactic elements, and predicted data associated with it.

[0016] For encoding in HEVC, a picture is divided into square coding tree blocks (CTBs) of a configurable size, and a contiguous set of coding tree blocks is grouped into slices. A coding tree unit (CTU) contains a CTB of encoded color components. A CTB is the root of a quadtree divided into coding blocks (CBs), which can be divided into one or more prediction blocks (PBs) that form the root of a quadtree divided into transformation blocks (TBs). Corresponding to coding blocks, prediction blocks, and transformation blocks, a coding unit (CU) contains a tree structure set of prediction units (PUs) and transformation units (TUs), where the PUs contain prediction information for all color components, and the TUs contain residual coding syntactic structures for each color component. The sizes of the CBs, PBs, and TBs of the color components correspond to the corresponding CUs, PUs, and TUs. In this application, the term “block” can be used, for example, to refer to any of the CTUs, CUs, PUs, TUs, CBs, PBs, and TBs. In addition, the term "block" can also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, and more generally, to refer to arrays of data of various sizes.

[0017] In the example of encoder 100, the picture is encoded by encoder elements as described below. The picture to be encoded is processed in units of CUs. Each CU is encoded using intra or intermode. When a CU is encoded in intramode, intra-prediction (160) is performed. In intermode, motion estimation (175) and motion compensation (170) are performed. The encoder decides whether to use intramode or intermode for encoding the CU (105), and the prediction mode flag indicates the intra / inter decision. The prediction residual is calculated by subtracting the prediction block from the original image block (110).

[0018] In intra-mode, the CU is predicted from reconstructed neighboring samples within the same slice. HEVC offers a set of 35 intra-prediction modes, including one DC mode, one planar mode, and 33 angular prediction modes. The intra-prediction reference is reconstructed from rows and columns adjacent to the current block. The reference extends horizontally and vertically to more than twice the block size, using samples available from previously reconstructed blocks. When an angular prediction mode is used for intra-prediction, the reference sample can be copied along the direction indicated by the angular prediction mode.

[0019] The intra-predictive modes applicable to the current block can be coded using two different options. If the applicable mode is included in the construction list of three most probable modes (MPMs), the mode is signaled by its index in the MPM list. Otherwise, the mode is signaled by a fixed-length binary representation of the mode index. The three most probable modes are derived from the intra-predictive modes of the adjacent blocks above and to the left.

[0020] In the case of interCU, the corresponding coded block is further subdivided into one or more prediction blocks. Interpretation is performed at the PB level, and the corresponding PU contains information about how interpretation is performed. Motion information (e.g., motion vectors and reference picture indices) can be signaled in two ways: "Merge Mode" and "Advanced Motion Vector Prediction (AMVP)".

[0021] In merge mode, the video encoder or decoder assembles a candidate list based on already coded blocks, and the video encoder signals an index for one of the candidates in the candidate list. On the decoder side, the motion vector (MV) and reference picture index are reconstructed based on the signaled candidate.

[0022] In AMVP, the video encoder or decoder assembles a candidate list based on motion vectors determined from already coded blocks. The video encoder then signals an index from the candidate list to identify the Motion Vector Predictor (MVP) and the Motion Vector Difference (MVD). On the decoder side, the motion vector (MV) is reconstructed as MVP + MVD. Additionally, applicable reference picture indices are explicitly coded in the PU syntax for AMVP.

[0023] Next, the predicted residuals are transformed (125) and quantized (130), including at least one embodiment for adapting the chroma quantization parameters described below. The transformations are generally based on separable transformations. For example, a DCT transformation is first applied horizontally and then vertically. In earlier codecs, the variety of 2D transformations for a given block size is usually limited, but in recent codecs such as JEM, the transformations used in each direction can be different (e.g., DCT in one direction and DST in the other), thereby enabling a wide variety of 2D transformations.

[0024] The quantized conversion coefficients, motion vectors, and other syntactic elements are entropy coded (145) to output a bitstream. Alternatively, the encoder can skip the conversion and directly apply quantization to the unconverted residual signal on a 4x4TU basis. The encoder can also avoid both conversion and quantization (i.e., the residual is directly coded without applying either the conversion or quantization process). In direct PCM coding, no prediction is applied, and the coded unit sample is directly coded and embedded in the bitstream.

[0025] The encoder decodes the encoded blocks to provide a reference for further predictions. The quantized transformation coefficients are inversely quantized (140) and inversely transformed (150) to decode the prediction residuals. The image blocks are reconstructed by combining the decoded residuals with the prediction blocks (155). An in-loop filter (165) is applied to the reconstructed picture to perform deblocking / SAO (sample adaptive offset) filtering, for example, to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (180).

[0026] Figure 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. In the example decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding path that is the reverse of the encoding path described in Figure 1, in which video decoding is performed as part of the encoding of the video data. Figure 2 also shows decoders that are improvements over the HEVC standard, or decoders that employ HEVC-like techniques, such as JEM or VVC decoders.

[0027] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 100. The bitstream is first entropy-decoded to obtain transformation coefficients, motion vectors, picture segmentation information, and other coded information (230). The picture segmentation information indicates the size of the CTU and how the CTU is divided into CUs, and possibly into PUs, where applicable. Thus, the decoder can divide the picture into CTUs and each CTU into CUs according to the decoded picture segmentation information (235). The transformation coefficients are then dequantized (240) and inverse-transformed (250), including at least one embodiment for adapting the chroma quantization parameters described below, in order to decode the prediction residuals.

[0028] Image blocks are reconstructed by combining the decoded predicted residuals and predicted blocks (255). Predicted blocks can be obtained from intra-prediction (260) or motion-compensated prediction (i.e., inter-prediction) (275) (270). As described above, AMVP and merge-mode techniques can be used to derive motion vectors for motion compensation, and in motion compensation, interpolation filters can be used to calculate interpolated values ​​for sub-integer samples of the reference block. An in-loop filter (265) is applied to the reconstructed picture. The filtered image is stored in the reference picture buffer (280).

[0029] The decoded picture can undergo further post-decoded processing (285), such as reverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or reverse remapping, which is the reverse of the remapping process performed in the pre-encoded processing (101). The post-decoded processing may use metadata, which is derived in the pre-encoded processing and signaled in the bitstream.

[0030] Figure 3 shows a block diagram of an example of a system in which various aspects and embodiments are implemented. System 300 can be embodied as a device comprising various components described below and configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, encoders, transcoders, and servers. The elements of System 300 can be embodied individually or in combination in a single integrated circuit, a plurality of ICs and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 300 are distributed across a plurality of ICs and / or discrete components. In various embodiments, the elements of System 300 are communicably coupled through an internal bus 310. In various embodiments, System 300 is communicably coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, the system 300 is configured to implement one or more of the embodiments described in this document, such as the video encoder 100 and video decoder 200 as described above and as modified below.

[0031] System 300 includes at least one processor 301 configured to execute instructions loaded therein to implement, for example, various embodiments described in this document. The processor 301 may include embedded memory, input / output interfaces and various other circuits as known in the art. System 300 includes at least one memory 302 (e.g., a volatile memory device and / or a non-volatile memory device). System 300 includes, but is not limited to, a storage device 304 which may include non-volatile memory and / or volatile memory, including EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives and / or optical disk drives. Storage devices 304 may, in non-limiting examples, include internal storage devices, mounted storage devices and / or network-accessible storage devices.

[0032] System 300 includes, for example, an encoder / decoder module 303 configured to process data to provide encoded or decoded video, the encoder / decoder module 303 may include its own processor and memory. The encoder / decoder module 303 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 303 may be implemented as a separate element of System 300, as is known to those skilled in the art, or it may be incorporated into the processor 301 as a combination of hardware and software.

[0033] Program code intended to be loaded into the processor 301 or encoder / decoder 303 to perform the various embodiments described in this document may be stored in the storage device 304 and then loaded into memory 302 for execution by the processor 301. According to various embodiments, one or more of the processor 301, memory 302, storage device 304, and encoder / decoder module 303 may store one or more of various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or a portion of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0034] In some embodiments, internal memory of the processor 301 and / or encoder / decoder module 303 is used to store instructions and to provide working memory for processing what is needed during encoding or decoding. However, in other embodiments, memory outside the processing device (for example, the processing device may be the processor 301 or the encoder / decoder module 303) is used for one or more of these functions. The external memory may be memory 302 and / or storage device 304, and may be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, external high-speed dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as those for MPEG-2, HEVC, or VVC.

[0035] Inputs to the elements of system 300 can be provided through various input devices as shown in block 309. Such input devices include, but are not limited to, (i) an RF section for receiving RF signals transmitted wirelessly, for example by a broadcasting station, (ii) a composite input terminal, (iii) a USB input terminal and / or (iv) an HDMI input terminal.

[0036] In various embodiments, the input device of block 309 has respective associated input processing elements as known in the Art. For example, the RF portion may be associated with elements necessary for (i) selecting a desired frequency (also called selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band (for example) in order to select a signal frequency band (which in some embodiments may also be called a channel), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various embodiments includes one or more elements for performing these functions, for example, frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner, which performs these various functions, for example, down-converting the received signal to a lower frequency (e.g., an intermediate frequency or a frequency near the baseband) or to the baseband. In one embodiment of the set-top box, the RF section and its associated input processing elements receive an RF signal transmitted over a wired (e.g., cable) medium and then perform frequency selection by filtering, down-converting, and filtering again to obtain a desired frequency bandwidth. Various embodiments involve rearranging the order of the elements described above (and others), removing some of these elements, and / or adding other elements that perform similar or different functions. Adding elements may involve inserting elements between existing elements (e.g., inserting amplifiers and analog / digital converters). In various embodiments, the RF section includes an antenna.

[0037] In addition, USB and / or HDMI terminals may include their respective interface processors for connecting the system 300 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented, for example, in a separate input processing IC or within processor 301, as needed. Similarly, aspects of USB or HDMI interface processing can be implemented, for example, in a separate interface IC or within processor 301, as needed. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements (e.g., processor 301 and encoder / decoder 303 in conjunction with memory and storage elements) for processing the data stream for presentation on an output device, as needed.

[0038] Various elements of System 300 can be provided within an integrated enclosure. Within the integrated enclosure, various elements can be interconnected, and data can be transmitted between them using appropriate connection arrangements, such as internal buses (including I2C buses, wiring, and printed circuit boards) known in the art.

[0039] System 300 includes a communication interface 305 that enables communication with other devices via a communication channel 320. The communication interface 305 may include, but is not limited to, transceivers configured to transmit and receive data on the communication channel 320. The communication interface 305 may also include, but is not limited to, a modem or a network card, and the communication channel 320 may be implemented, for example, in a wired and / or wireless medium.

[0040] In various embodiments, the data is streamed to system 300 using a Wi-Fi network such as IEEE 802.11. In these embodiments, the Wi-Fi signal is received on a communication channel 320 and communication interface 305 adapted for Wi-Fi communication. In these embodiments, communication channel 320 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, the streamed data is provided to system 300 using a set-top box that transmits data over the HDMI connection of input block 309. In yet another embodiment, the streamed data is provided to system 300 using the RF connection of input block 309.

[0041] System 300 can provide output signals to various output devices, including a display 330, a speaker 340, and other peripheral devices 350. In various embodiments, the other peripheral devices 350 include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of System 300. In various embodiments, control signals are transmitted between System 300 and the display 330, speaker 340, or other peripheral devices 350 using signaling such as AV.Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. The output devices can be communicatively coupled to System 300 via dedicated connections through their respective interfaces 306, 307, and 308. Alternatively, the output devices can be connected to System 300 via a communication channel 320 through a communication interface 305. The display 330 and speaker 340 can be integrated into a single unit along with other components of System 300, such as an electronic device (e.g., a television). In various embodiments, the display interface 306 includes a display driver, such as a timing controller (TCon) chip.

[0042] For example, if the RF portion of input 309 is part of a separate set-top box, the display 330 and speaker 340 can be selectively isolated from one or more of the other components. In various embodiments where the display 330 and speaker 340 are external components, the output signals can be provided via dedicated output connections (e.g., including HDMI ports, USB ports, or COMP outputs). The implementations described herein can be implemented, for example, as methods or processes, apparatus, software programs, data streams, or signals. Even if an implementation is discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the feature discussed can also be implemented in other forms (e.g., apparatus or programs). Apparatus can be implemented, for example, with appropriate hardware, software, and firmware. Methods can be implemented in apparatus (e.g., a processor) that generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Furthermore, the processor also includes communication devices such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.

[0043] Figure 4A shows an example of coded tree units and coded trees in the compressed region. In the HEVC video compression standard, motion-compensated time prediction is employed to utilize the redundancy present between consecutive pictures of video. To achieve this, pictures are divided into so-called coded tree units (CTUs), typically with sizes of 64x64, 128x128, or 256x256 pixels. Each CTU is represented by a coded tree in the compressed region (e.g., a quadtree partition of the CTUs). Each leaf is called a coded unit (CU).

[0044] Figure 4B shows an example of dividing a CTU into coding units, prediction units, and transformation units. Each CU is then given several intra or intercoding parameters (prediction information). To achieve this, each CU is spatially divided into one or more prediction units (PUs), and each PU is assigned several pieces of prediction information, such as motion vectors. The intra or intercoding mode is assigned at the CU level.

[0045] In this application, the terms “reconstructed” and “decoded” are interchangeable, the terms “encoded” and “coded” are interchangeable, and the terms “image,” “picture,” and “frame” are interchangeable. Generally, though not always, the term “reconstructed” is used on the encoder side, and the term “decoded” is used on the decoder side. The term “block” or “picture block” can be used to refer to any one of CTU, CU, PU, ​​TU, CB, PB, and TB. In addition, the term “block” or “picture block” can be used to refer to macroblocks, partitions, and subblocks as specified in H.264 / AVC or other video coding standards, and more generally, to refer to arrays of many sizes of samples.

[0046] Figure 5 illustrates the use of two scalar quantizers in dependent scalar quantization. Dependent scalar quantization uses two scalar quantizers with different reconstruction levels for quantization, as proposed in JVET (contributed JVET-J0014). Compared to conventional scalar quantization (e.g., as used in HEVC and VTM-1), the main effect of this method is that the set of acceptable reconstruction values ​​for the transformation coefficients depends on the values ​​of the transformation coefficient levels preceding the current transformation coefficient level in the reconstruction order. The dependent scalar quantization method is realized by (a) defining two scalar quantizers with different reconstruction levels and (b) defining a process for switching between the two scalar quantizers. The two scalar quantizers used are shown in Figure 5 (represented by Q0 and Q1). The location of the available reconstruction levels is uniquely specified by the quantization step size Δ. Ignoring the fact that the actual reconstruction of the transformation coefficients uses integer arithmetic, the two scalar quantizers Q0 and Q1 can be characterized as follows: Q0: The reconstruction level of the first quantizer Q0 is given by an even integer multiple of the quantization step size Δ. When this quantizer is used, the inverse quantized transformation coefficient t' is, t' = 2·k·Δ It is calculated according to the formula, where k represents the relevant quantized coefficient (the transmitted quantization index). Q1: The reconstruction level of the second quantizer Q1 is given by an odd integer multiple of the quantization step size Δ, in addition to a reconstruction level equal to zero. The inverse quantized transformation coefficient t' is: t'=(2·k-sgn(k))·Δ As shown, it is calculated as a function of the quantized coefficient k, and in the formula, sgn(·) is, sgn(x)=(k==0?0:(k<0?-1:1)) This is the sign function defined as follows.

[0047] The scalar quantizer used (Q0 or Q1) is not explicitly signaled in the bitstream. Instead, the quantizer used for the current transformation coefficient is determined by the parity of the quantized coefficients preceding the current transformation coefficient in the coding / reconstruction order and the state of the finite state machine described below.

[0048] Figures 6A and 6B illustrate an example of a mechanism for switching between scalar quantizers in VVC. Switching between two scalar quantizers (Q0 and Q1) is achieved via a finite state machine with four states (labeled 0, 1, 2, or 3, respectively), as shown in Figure 6A. The state of the finite state machine considered for a given quantized coefficient is uniquely determined by the parity of the quantized coefficient k preceding the current quantized coefficient in the coding / reconstruction order and the state of the finite state machine considered when processing this preceding coefficient. At the start of inverse quantization of a transformation block, the state is set to equal to 0. The transformation coefficients are reconstructed in scan order (i.e., the same order in which they are entropy-decoded). After the current transformation coefficient has been reconstructed, the state is updated, where k is a quantized coefficient. The next state is: state=stateTransTable[current_state][k&1] As shown, it depends on the current state and the parity (k&1) of the current quantized coefficient k, where stateTransTable represents the state transition table shown in Figure 6B, and the operator & specifies the bitwise "multiplication" operator in two's complement arithmetic.

[0049] Furthermore, the quantized coefficients included in the so-called transformation block (TB) can be entropy coded and decoded as described below.

[0050] First, the transformation block is divided into 4x4 subblocks of quantized coefficients called coding groups (sometimes also called coefficient groups, abbreviated as CG). Entropy coding / decoding consists of several scan passes, which scan the TB in the diagonal scan order shown in Figure 6C.

[0051] The encoding of transformation coefficients in VVC involves five main steps: scanning, encoding of final significance coefficients, encoding of significance maps, encoding of coefficient-level residuals, and encoding of absolute-level and coded data.

[0052] Figure 6C shows an example of a scan order between CGs and coefficients, as used in a VVC of an 8x8 transformation block. The scan path across the TB involves sequentially processing each CG according to the diagonal scan order, and similarly, the 16 coefficients within each CG are scanned according to the scan order considered. The scan path processes all coefficients, starting from the last significance coefficient of the TB and ending with the DC coefficient.

[0053] The entropy coding of the conversion coefficients includes up to seven syntactic elements from the list below. - coding_sub_block_flag: Significance of the coefficient group (CG) - sig_flag: Significance of the coefficient (zero / non-zero) - gt1_flag: Indicates whether the absolute value of the coefficient level is greater than 1. - par_flag: Indicates parity of coefficients greater than 1 - gt3_flag: Indicates whether the absolute value of the coefficient level is greater than 3. - remainder: The remaining value relative to the absolute value of the coefficient level (if the value is greater than the one coded in the previous pass). - abs_level: The absolute value of the coefficient level (if the CABAC bin is not signaled for the current coefficients in the maximum number of bin budget problem) - sign_data: The sign of all significance coefficients included in the CG considered. It consists of a series of bins, each signaling the sign of each non-zero transformation coefficient (0: positive, 1: negative).

[0054] Once the absolute value of the quantized coefficient is known by decoding a subset of the above elements (apart from the sign), no further coding of syntactic elements for that coefficient with respect to that absolute value is performed. Similarly, the sign flag is signaled only for non-zero coefficients.

[0055] All scan passes required for a given CG are coded until all quantized coefficients of that CG can be reconstructed before proceeding to the next CG.

[0056] The overall decryption and TB analysis process consists of the following key steps: 1. Decode the coordinates of the final significance coefficients. This involves the syntactic elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. This provides the decoder with the spatial location (x and y coordinates) of the final non-zero coefficients in the entire TB.

[0057] Next, for each consecutive CG from the CG containing the final significance coefficient of TB to the CG in the upper left of TB, the following steps are applied. 2. Decode the CG significance flag (called coded_sub_block_flag in the VVC specification). 3. Decode the significance coefficient flag for each coefficient of the CG being considered. This corresponds to the syntax element sig_flag, which indicates which coefficients of the CG are non-zero.

[0058] The next analysis step relates to the coefficient levels of the coefficients of the CG under consideration, for coefficients known to be non-zero. This involves the following syntactic elements. 4. gt1_flag: This flag indicates whether the absolute value of the current coefficient is greater than 1. If it is not greater than 1, the absolute value is equal to 1. 5. par_flag: This flag indicates whether the current quantized coefficient is even or not. If gt1_flag of the current quantized coefficient is true, it is coded. If par_flag is zero, the quantized coefficient is even; otherwise, the quantized coefficient is odd. After par_flag is analyzed on the decoder side, the partially decoded quantized coefficient is set to equal to (1 + gt1_flag + par_flag). 6. gt3_flag: This flag indicates whether the absolute value of the current coefficient is greater than 3. If it is not greater than 3, it is equal to 1 + gt1_flag + par_flag, and if so, the absolute value. If 1 + gt1_flag + par_flag is 2 or greater, gt3_flag is coded. Once gt3_flag is analyzed, the quantized coefficient value is encoded by the decoder as 1 + gt1_flag + par_flag + (2 * It will become gt3_flag). 7. remainder: This encodes the absolute value of the coefficient. This is true when the partially decoded absolute value is greater than 4. Note that in the example in VVC Draft 3, the maximum number of budgets for the normal coding bins is fixed for each coding group. Thus, for some coefficients, only the elements sig_flag, gt1_flag, and par_flag can be signaled, while for other coefficients, gt3_flag can also be signaled. Thus, the coded and analyzed remainder values ​​are calculated on the flags that have already been decoded for the coefficients under consideration, and are therefore calculated as a function of the partially decoded quantized coefficients. 8.abs_level: This indicates the absolute value of the coefficient for which none of the flags of the CGs considered (sig_flag, gt1_flag, papr_flag, or gt3_flag) are coded, for the usual maximum number of coded bins problem. This syntactic element, like the syntactic element remainder, undergoes Rice-Golomb binary evolution and bypass coding. 9. sign_flag: This indicates the sign of non-zero coefficients. This is bypass-coded.

[0059] Figure 7A shows the syntactic elements for coding / analysis of transformation blocks by VVC Draft 4. The coding / pairing involves a four-pass process. As can be seen from the figure, for a given transformation block, the position of the final significance coefficient is coded first. Then, for each CG containing the final significance coefficient (excluding this CG) and from the first CG, the significance of the CG is signaled (coded_sub_block_flag), and in cases where the CG is significant, residual coding is performed for that CG, as shown in Figure 7B.

[0060] Figure 7B shows the CG level residual coding syntax as specified in VVC Draft 4. This syntax signals the syntactic elements sig_flag, gt1_flag, par_flag, gt3_flag, remainder, abs_level, and sign_data of the CGs considered, as previously introduced. The figure shows the signaling and parsing of the syntactic elements sig_flag, gt1_flag, par_flag, gt3_flag, remainder, and abs_level according to VVC Draft 4. EP stands for "equi-probable," meaning the bin is not arithmetically coded but is coded in bypass mode. Bypass mode involves direct bit writing / parsing, which generally corresponds to the binary syntactic element (bin) that you wish to encode or parse.

[0061] Furthermore, VVC draft 4 specifies a hierarchical syntax array from the CU level down to the residual subblock (CG) level.

[0062] Figure 8A shows the CU level syntax as specified in the example of VVC draft 4. The syntax includes signaling the coding mode of the CU under consideration. cu_skip_flag signals if the CU is in merge skip mode. If the CU is not in merge skip mode, cu_pred_mode_flag indicates the variable value of cuPredMode, and therefore if the prediction mode of the current CU is coded through intra prediction (cuPredMode=MODE_INTRA) or inter prediction (cuPredMode=MODE_INTER).

[0063] In non-skip and non-intra-mode cases, the following flag indicates the use of intra-block copy (IBC) mode for the current CU. The rest of the syntax contains intra- or inter-predicted data for the current CU. The end of the syntax is dedicated to signaling the transformation tree associated with the current CU. This transformation tree signaling begins with a syntactic element called cu_cbf, which indicates that some non-null residual data will be coded for the current CU. In cases where this flag is equal to true, the transform_tree associated with the current CU is signaled according to the syntactic array shown in Figure 8B.

[0064] Figure 8B shows the syntactic structure of transform_tree, an example from VVC draft 4. Its syntax essentially involves signaling when a CU contains one or more transformation units (TUs). First, if the CU size is greater than the maximum allowed TU size, the CU is divided into four subtransformation trees in a quadtree manner. Otherwise, and if ISP (Intra-Subpartitioning) or SBT (Subblock Transformation) is not used for the current CU, the current transformation tree is not partitioned and contains exactly one TU, which is signaled through the transformation unit syntax table in Figure 8C. Otherwise, if the CU is in intra-mode and ISP (Intra-Subpartitioning) mode is used, the CU is divided into several (actually two) TUs. Each of the two TUs is signaled successively through the transformation unit syntax table in Figure 8C. Otherwise, if the CU is in inter-mode and SBT (Subblock Transformation) mode is used, the CU is divided into two TUs according to the decoded syntax of the SBT-related syntax at the CU level. Each of the two TUs is signaled successively through the conversion unit syntax table in Figure 8C.

[0065] Figure 8C shows the transformation unit level syntax array in the example of VVC draft 4. The syntax consists of the following: First, the tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr flags indicate that the non-null residual data corresponding to each component is contained within each transformation block of the current TU, respectively. For each component, if the corresponding cbf flag is true, the transformation block-level residual_tb syntax table is used to encode the corresponding residual transformation block. The transform_unit level syntax also includes several encoded, transformed, and quantized block-related syntaxes (i.e., delta QP information (if any) and the type of transformation used to encode the TU being considered).

[0066] Figures 9A and 9B illustrate the CABAC decoding and encoding processes, respectively. CABAC stands for Context-Adaptive Binary Arithmetic Coding (CABAC), a form of entropy coding used in HEVC or VVC, for example, to provide lossless compression with excellent compression efficiency. The input to the process in Figure 9A contains a coded bitstream, typically conforming to the HEVC specification or a further evolution thereof. At any point in the decoding process, the decoder knows which syntactic element to decode next. This is well specified in the standardized bitstream syntax and decoding process. Furthermore, the decoder also knows how to binary-code the current syntactic element to be decoded (i.e., represented as a sequence of binary symbols called bins, each equal to "1" or "0") and how each bin in the bin string is encoded.

[0067] Therefore, the first stage of the CABAC decoding process (left side of Figure 9A) decodes a series of bins. For each bin, the decoder knows whether each bin is encoded according to bypass mode or normal mode. In bypass mode, bits are simply read from the bitstream, and the resulting values ​​are assigned to the current bin. This mode has the advantages of being straightforward, and therefore fast, and requiring no intensive resource utilization. It is typically efficient and therefore used for bins with a uniform statistical distribution (i.e., the probability of being equal to "1" is equal to the probability of being equal to "0").

[0068] Conversely, if the current bin is not coded in bypass mode, it means it is coded in so-called normal coding (i.e., through context-based arithmetic coding). This mode is far more resource-intensive.

[0069] In this example, the decoding of the bin under consideration proceeds as follows: First, a context is obtained for decoding the current bin. The context is given by the context modeler in Figure 9A. The goal of the context is to obtain a conditional probability that the current bin has a value of "0", taking into account some context prior or information X. Here, the prior X is the value of some already decoded syntactic element that is synchronously available on both the encoder and decoder sides when the current bin is being decoded.

[0070] Typically, the pryor X used for decoding a bin is chosen because it is specified in the standard and is statistically correlated with the current bin to be decoded. An interesting aspect of using this contextual information is the reduced rate cost of bin coding. This is based on the fact that given X, the conditional entropy of the bin is lower because there is a correlation between the bin and X. The following relationship is well known in information theory. H(bin│X) <H(bin)

[0071] This means that if the bin and X are statistically correlated, the conditional entropy of a bin for which X is known is lower than the entropy of the bin. Thus, the context information X is used to obtain the probability that the bin is "0" or "1". Taking these conditional probabilities into account, the normal decoding engine in Figure 14 performs arithmetic decoding of the binary value bins. The bin values ​​are then used to update the value of the conditional probability associated with the current bin for which the current context information X is known. This is called the context model update step in Figure 9A. As long as the bins are decoded (or coded), updating the context model for each bin allows for the gradual refinement of the context modeling for each binary element. Thus, the CABAC decoder gradually learns the statistical behavior of each of the normal coded bins.

[0072] The context modeler and context model update steps operate identically on both the encoder and decoder sides.

[0073] A set of decoded bins is obtained by either the normal arithmetic decoding or bypass decoding of the current bins, depending on how they were encoded.

[0074] The second stage of CABAC decoding, shown on the right side of Figure 9A, involves converting this set of binary symbols into higher-level syntactic elements. Syntactic elements can take the form of flags, in which case they directly incorporate the values ​​of the current decoded bins. On the other hand, if the binary code of the current syntactic element corresponds to a set of bins according to the standard specification being considered, a conversion step called "binary codeword for syntactic element" in Figure 9A is performed.

[0075] This step is the reverse of the binary transformation step performed by the encoder as shown in Figure 9B. Thus, the inverse transformation performed here involves obtaining the values ​​of those syntactic elements based on their respective decoded binary versions.

[0076] The encoder 100 in Figure 1, the decoder 200 in Figure 2, and the system 1000 in Figure 3 are adapted to implement at least one of the embodiments described below.

[0077] Current video coding systems (e.g., VVC Draft 4) impose some hard constraints on the coefficient group coding process, so that the maximum number of normal CABAC bins hardcoded can be, on one side, adopted for the syntactic elements sig_flag, gt1_flag, and parity_flag, and on the other side, adopted for the syntactic element gt3_flag. More precisely, the 4x4 coding group (or coefficient group or CG) level constraint ensures that the CABAC decoder engine must analyze the maximum number of normal bins per unit area. However, the limitation on the use of CABAC coding for normal bins can lead to rate distortion, resulting in suboptimal video compression. In fact, if only a limited number of normal bins are allowed for a picture, the utilization of this total budget can be significantly lower in cases where some of the pictures under consideration for coding contain significant areas coded in skip mode (and therefore without residuals), while other significant areas of the picture employ residual coding. In such cases, the low-level constraint (i.e., 4x4 CG level) on the maximum number of normal bins can result in the total number of normal bins for the picture falling far below the acceptable picture level threshold.

[0078] At least one embodiment relates to handling high-level constraints on the maximum usage of normal CABAC coding in bins in the residual coding process of blocks and their coefficient groups, such that high-level constraints are respected and compression efficiency is improved compared to current methods. In other words, the budget for normal coding bins is allocated over picture areas larger than CGs, and thus a large number of CGs are covered, the number of which is determined from the average allowed number of normal bins per unit area. For example, an average of 1.75 normal coding bins per sample may be allowed. This budget can then be distributed more efficiently over smaller units. In different embodiments, the higher-level constraints are set at the transform block, transform unit, coding unit, coding tree unit, or picture level. These embodiments can be implemented by a CABAC optimizer as shown in Figure 10.

[0079] In at least one embodiment, a higher level of constraint is set at the transformation unit level. In such an embodiment, the number of normal bins allowed for the entire transformation unit for encoding / decoding is determined from the average number of normal bins allowed per unit area. From the normal bin budget obtained for the TU, the number of normal bins allowed for each transformation block (TB) of the TU is derived. A transformation block is a set of transformation coefficients belonging to the same TU and the same color component. Residual encoding or decoding is then applied under this constraint of the number of normal bins allowed for the entire TB, taking into account the number of normal bins allowed in the transformation block. Accordingly, a modified residual encoding and decoding process is proposed herein, in which the normal bin budget at the TB level is considered instead of the normal bin budget at the CG level. For this purpose, several embodiments are proposed.

[0080] Figure 10 shows a CABAC coding process including a CABAC optimizer according to an embodiment of the present principle. In at least one embodiment, the CABAC optimizer 190 handles a budget representing the maximum number of bins encoded (or to be encoded) using normal coding for a set of coding groups. This budget is determined, for example, by multiplying the budget per sample of a normal coding bin by the surface of the data unit considered (i.e., the number of samples contained on that surface). For example, in the case of a budget for normal coding bins fixed at the TB level, the allowable number of normal coding bins per sample is multiplied by the number of samples contained in the transformation block considered. The output of the CABAC optimizer 190 controls coding in normal coding mode or bypass coding mode and therefore greatly affects the coding efficiency.

[0081] Figure 11 shows an example flowchart of the CABAC optimizer used in the encoding process according to an embodiment of the present principle. This flowchart is executed for each of the new sets of coding groups to be considered. Thus, according to different embodiments, this flowchart may occur for each picture, each CTU, each CU, each TU, or each TB. In step 191, the processor 301 allocates a budget for the normal coding determined as described above. In step 192, the processor 301 loops through the set of coding groups, and in step 193, checks if the last coding group has been reached. In step 194, for each coding group, the processor processes the coding group with the input budget of the normal coding bin. The budget of the normal coding bin is decremented during the processing of the coded groups to be considered (i.e., the budget of the normal coding bin is decremented each time the bin is coded in normal mode). The budget of the normal coding bin is returned as an output parameter of the coding group coding or decoding process.

[0082] Figure 12 shows an example of a modified process for encoding / decoding a coding group using a CABAC optimizer. As mentioned above, the allocation of the number of regular CABAC bins dedicated to residual coding is performed at a higher level than for 4x4 coding groups. To do this, the process for encoding / decoding a coding group is modified compared to the conventional function by adding an input parameter that represents the current budget of the regular bins when the CG is being coded or decoded. Thus, this budget is modified by coding or decoding the current CG according to the number of regular bins used to code / decode the current CG. The input / output parameter representing the budget of the regular bins is called numRegBins_in_out in Figure 12. The process for coding the residual data itself for a given CG can be the same as the conventional method, except that the allowed number of regular CABAC bins is given by an external means. As in the example in Figure 12, this budget is decremented by 1 each time a regular bin is coded / decoded.

[0083] Accordingly, at least one embodiment of the present disclosure includes processing the budget of a normal bin as an input / output parameter of a coefficient group coding / analysis function. Thus, the budget of a normal bin can be determined at a higher level than the CG coding / analysis level. Different embodiments propose determining the budget at different levels (i.e., at the transformation block, transformation unit, coding unit, coding tree unit, or picture level).

[0084] Figure 13A shows an example of an embodiment in which the budget for normal bins is determined at the transformation block level. In such an embodiment, the budget is determined as a function of two main parameters: the size of the current transformation block and the base budget for normal bins fixed for a given unit of picture area. In this specification, the area unit considered is a sample. In the current VVC draft 2, 32 normal bins are allowed for a 4x4 CG, which means that on average, two normal bins are allowed per component sample. The proposed allocation of normal bins at the transformation block level is based on this average rate of normal bins per single area.

[0085] In at least one embodiment, the budget allocated to the transformation block under consideration is calculated as the product of the transformation block surface and the usual allocated number per sample. The budget is then passed to the CG residual coding / analysis process for each significance coefficient group in the transformation block under consideration.

[0086] Figure 13B shows an example of an embodiment in which the budget for normal bins is determined at the transformation block level according to the position of the final significance coefficient. According to this modified embodiment, the number of normal bins allowed for the current transformation block is calculated as a function of the position of the final significance coefficient in the transformation block under consideration. The advantage of such a method is that when allocating normal bins, it is better suited to the energy contained in the transformation block under consideration. Typically, more bins can be allocated to high-energy transformation blocks, and fewer normal bins can be allocated to low-energy transformation blocks.

[0087] Figure 14A shows an example of an embodiment in which the budget for a normal bin is determined at the conversion unit level. As shown, unitary_budget (i.e., the average rate of a normal bin allowed per sample) is still equal to 2. The budget for a normal bin at the TU level is determined by multiplying this rate by the total number of samples contained in the conversion unit being considered (i.e., in the example of a 4:2:0 color format, width * height * 3 / 2). This total budget is then used for coding all transformation blocks contained within the transformation unit under consideration. The budget is passed as an input / output parameter to the residual transformation block coding / analysis process for each color component. In other words, each time a transformation block of the transformation unit under consideration is coded / analyzed, the budget decreases by the number of bins in the coding / analysis of the transformation blocks subsequently coded / analyzed in that transformation unit.

[0088] Figure 14B shows an example of an embodiment in which the budget of a normal bin is determined at the transformation unit level and allocated between transformation blocks as a function of the relative surface between different transformation blocks and as a function of the normal bin used in the already coded / analyzed transformation blocks of the transformation unit being considered. In this embodiment, the normal bin allocated to the luma TB is 2 / 3 of the total budget at the TU level. Next, the Cb TB budget is half of the remaining budget after the luma TB has been coded / analyzed. Finally, the Cr TB budget is set to the remaining budget at the TU level after the two initial TBs have been coded / analyzed. Furthermore, it should be noted that in further embodiments, the budget for the entire transformation unit can be shared between the transformation blocks of that TU, based on the knowledge that one or more TBs of that TU are coded with null residuals. This is known, for example, through the analysis of the tu_cbf_luma, tu_cbf_cb, or tu_cbf_cr flags. Accordingly, the TB-level residual coding / analysis process is adapted to this modified embodiment, as shown by the modified process in Figure 14C. Such embodiments result in better coding efficiency when the budget is allocated at a lower level of the hierarchy. In fact, if some CGs or TBs are of very low energy, a reduced number of bins may be employed for those CGs or TBs compared to others. Thus, a larger number of normal bins can be used for CGs or TBs where coding / analysis of a larger number of bins is required.

[0089] Figure 15A shows an example of an embodiment in which the budget for the normal bins is assigned at the coding unit level. This budget is calculated similarly to that for the previous embodiment where it is fixed at the TU level, but based on the sample rate and CU size of the normal bins. The calculated budget for the normal bins is then passed to the transformation tree coding / analysis procedure in Figure 15B. Figure 15B shows the adapted coding / analysis of the transformation tree associated with the current CU based on the budget for the normal bins assigned at the CU level. Here, the CU level is essentially passed successively to the coding of each transformation unit contained within the transformation tree under consideration. The budget is also reduced by the number of normal bins used in the coding of each TU. Figure 15C shows a modified embodiment of the coding tree coding / analysis process when the budget is fixed at the CU level. Typically, this process updates the total budget at the CU level as a function of the normal bins used in the already coded / analyzed TUs when coding the current TU of the transformation tree under consideration. For example, in the case of the SBT (Subblock Transformation) mode for an interCU, VVC draft 4 describes a case where the CU is divided into two TUs, one of which has a null residual. In that case, this embodiment proposes that the total budget of normal bins at the CU level is dedicated to the TU with the non-null residual, thereby improving the coding efficiency for such interCUs in SBT mode. Furthermore, in the case of ISP (Intra-Sub-Partitioning), this VVC partitioning mode for an intraCU divides the CU into several TUs. In that case, the coding / analysis of the TUs attempts to reuse the unused normal bin budgets allocated to preceding TUs within the same CU. Figures 15D and 15E show the transformation unit coding / analysis processes adapted for this embodiment when the normal bin budgets are fixed at the CU level. They are adaptations of the processes in Figures 14A and 14B, respectively, for this embodiment.

[0090] In at least one embodiment, the budget of the standard CABAC bins used in residual coding is fixed at the CTU level. This allows for better utilization of the total budget of the allowed standard bins, and thus improves coding efficiency compared to previous embodiments. Typically, bins initially designated for skipped CUs can, advantageously, be used for other non-skipped CUs.

[0091] In at least one embodiment, the budget of the standard CABAC bins used in residual coding is fixed at the picture level. This allows for better utilization of the total budget of the allowed standard bins, and thus improves coding efficiency compared to previous embodiments. Typically, bins initially designated for skipping CUs can, advantageously, be used for other non-skipping CUs.

[0092] In at least one embodiment, the average rate of the allowable normal bins for a given picture is fixed according to the time layer / depth to which the picture under consideration belongs. For example, in a random access coding structure, the temporal structure of a group of pictures conforms to a hierarchical B-picture sequence. In this configuration, the pictures are organized in scalable time layers. Pictures from higher layers depend on reference pictures from lower time layers. Conversely, pictures from lower layers do not depend on any pictures from higher layers. Higher time layers are typically coded with higher quantization parameters than lower layer pictures and are therefore coded with lower quality. Moreover, pictures from lower layers have a significant impact on the overall coding efficiency over the entire sequence. Therefore, it is interesting to encode these pictures with optimal coding efficiency. According to this embodiment, it is proposed to assign a higher sample-based rate of the normal CABAC bins to pictures from lower layers than to pictures from higher time layers.

[0093] Figures 16A and 16B illustrate an example of an embodiment in which the budget for normal bins is allocated at a higher level than the CG level, and for using two types of CG coding / analysis processes. The first type shown in Figure 16A is called "all_bypass" and involves coding all bins associated with CG in bypass mode. This implies that the magnitude of the CG conversion coefficients is coded only through the syntactic element abs_level. The second type shown in Figure 16B is called "all_regular" and codes all bins corresponding to the syntactic elements sig_flag, gt1_flag, par_flag, and gt3_flag in normal mode. In this embodiment, considering the budget for normal bins fixed at a higher level than the CG level, the TB coding process can switch between the "all_regular" CG coding mode and the "all_bypass" CG coding mode depending on whether the budget considered for the currently considered normal bins is sufficiently used.

[0094] Various implementations involve decoding. “Decoding,” as used in this application, may encompass all or part of the processes performed on a received encoded sequence to produce, for example, a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include, or alternatively, processes performed by the decoders of various implementations described in this application, such as the embodiments presented in Figures 10 to 16.

[0095] As further examples, in one embodiment, “decoding” refers only to entropy decoding; in another embodiment, “decoding” refers only to differential decoding; and in yet another embodiment, “decoding” refers to a combination of entropy decoding and differential decoding. Whether the term “decoding process” is intended to refer to a specific subset of operations or to a broader decoding process in general will become clear from the context of the particular description and is considered to be well understood by those skilled in the art.

[0096] Various implementations involve encoding. In a manner similar to the above discussion of "decoding," "encoding," as used in this application, may encompass all or part of the processes performed on an input video sequence to generate an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as segmentation, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also include, or alternatively, processes performed by the encoders of various implementations described in this application, such as the embodiments shown in Figures 10 to 16.

[0097] As further examples, in one embodiment, “encoding” refers only to entropy coding; in another embodiment, “encoding” refers only to differential coding; and in yet another embodiment, “encoding” refers to a combination of differential coding and entropy coding. Whether the term “encoding process” is intended to refer to a specific subset of operations or to a broader encoding process in general will become clear from the context of the particular description and is considered to be well understood by those skilled in the art.

[0098] It should be noted that the syntactic elements used in this specification are descriptive terms. Therefore, these syntactic elements do not preclude the use of other syntactic element names.

[0099] This application describes a variety of embodiments, including tools, features, embodiments, models, and methods. Many of these embodiments are described in a specific manner, often in a way that seems to be intended to be restrictive, at least in order to illustrate their individual characteristics. However, this is for the purpose of clarifying the description and does not limit the application or scope of those embodiments. In fact, all different embodiments can be combined or interchangeable to provide further embodiments. Moreover, embodiments can be combined or interchangeable with embodiments described in prior applications. The embodiments described and envisioned in this application can be implemented in many different forms. Figures 1, 2 and 3 above provide some embodiments, but other embodiments are envisioned, and the discussion of the figures does not limit the breadth of implementation forms.

[0100] In this application, the terms “reconstructed” and “decoded” are interchangeable, the terms “pixel” and “sample” are interchangeable, and the terms “image,” “picture,” and “frame” are interchangeable. Generally, although not always the case, the term “reconstructed” is used on the encoder side, and the term “decoded” is used on the decoder side.

[0101] This specification describes various methods, each of which includes one or more steps or actions to achieve the method described. Unless a particular order of steps or actions is required for the correct operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined.

[0102] In this application, various numerical values ​​are used, for example, with respect to block size. The specific values ​​are for illustrative purposes only, and the embodiments described are not limited to these specific values.

[0103] References to “one embodiment,” “embodiment,” “one implementation,” or “implementation,” and other variations, mean that certain features, structures, characteristics, etc., described in relation to an embodiment are included in at least one embodiment. Therefore, the appearance of phrases such as “in one embodiment,” “in one embodiment,” “in one implementation,” or “in one implementation,” and any other variations, appearing in various places throughout this specification, do not necessarily all refer to the same embodiment.

[0104] In addition, this application or its claims may refer to “determining” various pieces of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.

[0105] Furthermore, this application or its claims may refer to “accessing” various pieces of information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, computing information, predicting information, or estimating information.

[0106] In addition, this application or its claims may refer to “receiving” various pieces of information. Receiving is intended to be a broad term, as is “accessing.” Receiving information may include, for example, one or more of the following: accessing information or retrieving information (e.g., from memory or optical media storage). Furthermore, “receiving” typically involves various ways of being involved in an operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0107] The use of any of the following " / ", "and / or", or "at least one of ~" should be understood to include, for example, the selection of only the first list option (A), only the second list option (B), or both options (A and B) in the cases of "A / B", "A and / or B", and "at least one of A and B". As further examples, in the cases of "A, B and / or C" and "at least one of A, B and C", such descriptions are intended to include the selection of only the first list option (A), only the second list option (B), only the third list option (C), only the first and second list options (A and B), only the first and third list options (A and C), only the second and third list options (B and C), or all three options (A, B and C). This can be extended to as many items as there are listed, as will be readily apparent to those skilled in the art in this and related fields.

[0108] As will be apparent to those skilled in the art, the implementation can generate a variety of signals formatted to convey information that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the implementations described. For example, a signal can be formatted to convey a bitstream of an embodiment described. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the high-frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave having the encoded data stream. The information conveyed by the signal may be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

Claims

1. A video decoding method, Regarding the coding unit (CU), this involves determining the budget for a normal CABAC (Context-Adaptive Binary Arithmetic Coding) bin, The aforementioned CU is divided into multiple conversion units (TU), The process of decoding the plurality of TUs such that the budget of the normal CABAC bins available to decode the current TU among the plurality of TUs is determined based on the number of normal CABAC bins used to decode one or more previously processed TUs of the same CU, or based on whether the one or more previously processed TUs contain residual data. Methods that include...

2. The method according to claim 1, wherein the partitioning of the CU is performed according to a subblock transform (SBT) mode, and the entire budget of the normal CABAC bin determined for the CU is dedicated to decoding a particular TU among the plurality of TUs that is signaled to contain non-null residual data.

3. The method according to claim 1, wherein the partitioning of the CU is performed according to an intra-sub-partitioning (ISP) mode, and the decoding of the plurality of TUs includes sequentially carrying over and reusing unused portions of the budget from preceding TUs for decoding subsequent TUs within the same CU.

4. A video decoding device, Memory and At least one processor and The at least one processor is Regarding coded units (CUs), the budget for a standard CABAC bin is determined, The aforementioned CU is divided into multiple conversion units (TU), A video decoding device configured to decode a plurality of TUs such that the budget of the normal CABAC bins available for decoding the current TU among the plurality of TUs is determined based on the number of normal CABAC bins used to decode one or more previously processed TUs of the same CU, or based on whether the one or more previously processed TUs contain residual data.

5. The apparatus according to claim 4, wherein the at least one processor is configured to partition the CU according to a subblock transform (SBT) mode, and to dedicate the entire budget of the normal CABAC bin determined for the CU to decoding a particular TU among the plurality of TUs that is signaled to contain non-null residual data.

6. The apparatus according to claim 4, wherein the at least one processor is configured to decode the plurality of TUs by partitioning the CU according to an intra-subpartitioning (ISP) mode and sequentially carrying over and reusing unused portions of the budget from preceding TUs for decoding subsequent TUs within the same CU.

7. A video encoding method, Regarding the coding unit (CU), this involves determining the budget for a normal CABAC (Context-Adaptive Binary Arithmetic Coding) bin, The aforementioned CU is divided into multiple conversion units (TU), Encoding the syntactic elements of the plurality of TUs into a bitstream using a normal CABAC encoding mode, wherein the budget of normal CABAC bins available for encoding the current TU among the plurality of TUs is determined based on the number of normal CABAC bins used to encode one or more previously processed TUs of the same CU, or based on whether the one or more previously processed TUs were selected to include residual data. Methods that include...

8. The method according to claim 7, wherein the partitioning of the CU is performed according to a subblock transform (SBT) mode, and the entire budget of the normal CABAC bin determined for the CU is dedicated to the coding of a particular TU among the plurality of TUs that is selected to contain non-null residual data.

9. The method according to claim 7, wherein the partitioning of the CU is performed according to an intra-subpartitioning (ISP) mode, and the encoding of the plurality of TUs includes sequentially carrying over and reusing unused portions of the budget from preceding TUs for encoding syntactic elements of subsequent TUs within the same CU.

10. A video encoding device, Memory and At least one processor and The at least one processor is Regarding coded units (CUs), the budget for a standard CABAC bin is determined, The aforementioned CU is divided into multiple conversion units (TU), A video encoding device configured to encode the syntactic elements of a plurality of TUs into a bitstream using a normal CABAC encoding mode, wherein the budget of normal CABAC bins available for encoding the current TU among the plurality of TUs is determined based on the number of normal CABAC bins used to encode one or more previously processed TUs of the same CU, or based on whether the one or more previously processed TUs were selected to include residual data.

11. The apparatus according to claim 10, wherein the at least one processor is configured to partition the CU according to a subblock transform (SBT) mode, and to dedicate the entire budget of the normal CABAC bin determined for the CU to the coding of a particular TU among the plurality of TUs that is selected to include non-null residual data.

12. The apparatus according to claim 10, wherein the at least one processor is configured to encode the plurality of TUs by partitioning the CUs according to an intra-subpartitioning (ISP) mode and sequentially carrying over and reusing unused portions of the budget from preceding TUs for encoding subsequent TUs within the same CU.