Cabac update for minimal probability
By detecting convergence phases and adjusting probability updates in CABAC encoding, the patent addresses inefficiencies in entropy coding, enhancing video compression efficiency through optimized probability management and resolution-specific context initialization.
Patent Information
- Application Number
- PCT/EP2024/086525
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-10
AI Technical Summary
Existing video compression technologies face inefficiencies in entropy coding, particularly in contexts where sequences of zero bins lead to suboptimal CABAC cost due to convergence of probabilities to minimum values, impacting overall encoding efficiency.
Implementing methods to detect convergence phases in CABAC encoding and decoding, adjusting probability updates to maintain optimal encoding efficiency by bounding probabilities within specific intervals and decoupling convergence speed from clipping values, and initializing contexts based on image resolution.
Enhances CABAC encoding efficiency by reducing overhead costs in sequences of zero bins and maintaining low encoding costs for non-zero bins, thereby improving overall video compression performance.
Smart Images

Figure EP2024086525_10072025_PF_FP_ABST
Abstract
Description
[0001] CABAC UPDATE FOR MINIMAL PROBABILITY
[0002] This application claims the priority to European Application No. 24305009.3 filed on 3 January 2024, which is incorporated herein by reference in its entirety.
[0003] TECHNICAL FIELD
[0004] The present embodiments generally relate to video compression. The present embodiments relate to a method and an apparatus for encoding or decoding an image or a video. More particularly, the present embodiments relate to improving entropy coding in video compression system.
[0005] BACKGROUND
[0006] To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content. Generally, intra or inter prediction is used to exploit the intra or inter picture correlation, then the differences between the original block and the predicted block, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded. In inter prediction, motion vectors used in motion compensation are often predicted from motion vector predictor. To reconstruct the video, the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction.
[0007] SUMMARY
[0008] According to an aspect, a first method for encoding an image or a video is provided. The first method for encoding is for example implemented for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image. The first method for encoding comprises arithmetically encoding the at least one binary symbol using the one or more probability values, updating the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, setting the updated value of the at least one probability value to a first given probability value. i According to another aspect, a first apparatus for encoding an image or a video is provided. The first apparatus for encoding comprises one or more processors operable to, for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, arithmetically encode the at least one binary symbol using the one or more probability values, update the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, set the updated value of the at least one probability value to a first given probability value.
[0009] According to an aspect, a first method for decoding an image or a video is provided. The first method for decoding is for example implemented for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image. The method for decoding comprises arithmetically decoding the at least one binary symbol using the one or more probability values, updating the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, setting the updated value of the at least one probability value to a first given probability value.
[0010] According to another aspect, a first apparatus for decoding an image or a video is provided. The first apparatus for decoding comprises one or more processors operable to, for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, arithmetically decode the at least one binary symbol using the one or more probability values, update the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, set the updated value of the at least one probability value to a first given probability value.
[0011] According to another aspect, a second method for encoding an image or a video is provided wherein for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, the second method for encoding comprises arithmetically encoding the at least one binary symbol using the one or more probability values and updating the one or more probability values based on the at least one binary symbol, a window size and a first given probability value and wherein updating the one or more probability values comprises clipping the update value of the one or more probability values to the first given probability value if the update value is below the first given probability value.
[0012] According to another aspect, a second apparatus for encoding an image or a video is provided. The second apparatus for encoding comprises one or more processors operable to for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, arithmetically encode the at least one binary symbol using the one or more probability values and update the one or more probability values based on the at least one binary symbol, a window size and a first given probability value and wherein updating the one or more probability values comprises clipping the update value of the one or more probability values to the first given probability value if the update value is below the first given probability value.
[0013] According to an aspect, a second method for decoding an image or a video is provided wherein for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, the second method for decoding comprises arithmetically decoding the at least one binary symbol using the one or more probability values and updating the one or more probability values based on the at least one binary symbol, a window size and a first given probability value and wherein updating the one or more probability values comprises clipping the update value of the one or more probability values to the first given probability value if the update value is below the first given probability value. According to another aspect, a second apparatus for decoding an image or a video is provided. The second apparatus for decoding comprises one or more processors operable to for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, arithmetically decode the at least one binary symbol using the one or more probability values and update the one or more probability values based on the at least one binary symbol, a window size and a first given probability value and wherein updating the one or more probability values comprises clipping the update value of the one or more probability values to the first given probability value if the update value is above the first given probability value.
[0014] According to another aspect, a third method for encoding or decoding an image or a video is provided. The method comprises for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding or decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, arithmetically encoding or decoding the at least one binary symbol using the one or more probability values and updating the one or more probability values based on the at least one binary symbol and a window size, wherein initial values for at least one of the one or more probability values or of one or more window sizes used for udapting the one or more probability values depend on a resolution of the image. According to another aspect, an apparatus for encoding or decoding an image or a video is provided. The apparatus comprises one or more processors operable to for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding or decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image, arithmetically encode or decode the at least one binary symbol using the one or more probability values and update the one or more probability values based on the at least one binary symbol and a window size, wherein initial values for at least one of the one or more probability values or of one or more window sizes used for udapting the one or more probability values depend on a resolution of the image. One or more aspects mentionned herein can be used along or in combination. Also, further embodiments that can be used alone or in combination are described herein.
[0015] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform any one of the methods for encoding or decoding an image or a video according to any of the embodiments described herein. One or more of the present embodiments also provide a non- transitory computer readable medium and / or a computer readable storage medium having stored thereon instructions for encoding or decoding an image or a video according to the methods described herein.
[0016] One or more embodiments also provide a computer readable storage medium having stored thereon a bitstream generated according to the methods described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving the bitstream generated according to the methods described above.
[0017] BRIEF DESCRIPTION OF THE DRAWINGS
[0018] 1A illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to an embodiment.
[0019] FIG. IB illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to another embodiment.
[0020] FIG. 1C illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to another embodiment.
[0021] FIG. 2 illustrates a block diagram of an embodiment of a video encoder within which aspects of the present embodiments may be implemented.
[0022] FIG. 3 illustrates a block diagram of an embodiment of a video decoder within which aspects of the present embodiments may be implemented.
[0023] FIG. 4 illustrates an example of a context-based entropy coding scheme.
[0024] FIG. 5 illustrates an example of a parameter initialization for a context-based entropy coding scheme.
[0025] FIG. 6 illustrates an example of a CAB AC engine in VVC encoding scheme.
[0026] FIG. 7 illustrates an example of a method for decoding a bin. FIG. 8 illustrates an example of a update of probaiblity and associated CABAC cost for a sequence of zero bins.
[0027] FIG. 9 illustrates an example of a method for encoding or decoding a sequence of binary symbols according to an embodiment.
[0028] FIG. 10 illustrates an example of a method for encoding or decoding a sequence of binary symbols according to another embodiment.
[0029] FIG. 11 illustrates an example of coding costs when changing a ratel parameter while keeping a same clipping value.
[0030] FIG. 12 illustrates an example of a method for initializing probability models for encoding or decoding a sequence of binary symbols, according to an embodiment.
[0031] FIG. 13 illustrates an example of a method for signaling or decoding syntax elements used in some embodiments described herein.
[0032] FIG. 14 shows two remote devices communicating over a communication network in accordance with an example of the present principles.
[0033] FIG. 15 shows the syntax of a signal in accordance with an example of the present principles.
[0034] DETAILED DESCRIPTION
[0035] This application describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the application or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.
[0036] The aspects described and contemplated in this application can be implemented in many different forms. FIGs. 1A, IB, 1C, 2 and 3 below provide some embodiments, but other embodiments are contemplated and the discussion of FIGs. 1 A, IB, 1C, 2 and 3 does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.
[0037] In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, the terms “image,” “picture” and “frame” may be used interchangeably.
[0038] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[0039] The present aspects are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
[0040] FIG. 1A-1C illustrates block diagrams of examples of systems in which various aspects and embodiments can be implemented. Any one of the systems 100A, 100B or 100B may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. In various embodiments, the system 100A, 100B or 100C is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 100A, 100B or 100C is configured to implement one or more of the aspects described in this application.
[0041] FIG. 1A illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. The system 100A includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 100 A includes at least one memory 120, e.g., a volatile memory device, and / or a non-volatile memory device, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The memory 120 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples. The processor 110 may be interconnected to the memory 120 by an interconnection bus 115.
[0042] Program code to be loaded onto processor 110 to perform the various aspects described in this application is subsequently loaded onto memory 120 for execution by processor 110.
[0043] In some embodiments, memory inside of the processor 110 is used to store program code instructions and to provide working memory for processing that is needed during encoding or decoding. The input to the elements of system 100A may be provided through various input devices (not represented). Both Processor 110 and memory 120 can also have one or more additional interconnections to external connections.
[0044] FIG. IB illustrates a block diagram of an example of a system 100B in which various aspects and embodiments can be implemented. The system 100B includes the processor 110 and memory 120 as described in relation with FIG. 1A. The input to the elements of system 100B may be provided through various input devices as indicated in block 105 which is described further below with FIG. 1C. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. IB, include composite video.
[0045] The various elements may be interconnected and transmit data therebetween using suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards. The system 100B includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium.
[0046] The system 100B may provide an output signal to various output devices, including a display, speakers, and other peripheral devices. The output devices may be communicatively coupled to system 100B via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100B using the communications channel 190 via the communications interface 150.
[0047] FIG. 1C illustrates a block diagram of an example of a system 100C in which various aspects and embodiments can be implemented according to another embodiment. Elements of system 100C, singly or in combination, may be embodied in a single integrated circuit, multiple Ics, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100C are distributed across multiple Ics and / or discrete components.
[0048] The system 100C includes the processor 110 and memory 120 as described in relation with FIG. 1A or IB.
[0049] System 100C includes a storage device 140, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.
[0050] System 100C includes an encoder / decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents module(s) that may be included in a device to perform the encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100C or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art. Program code to be loaded onto processor 110 or encoder / decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input data (image, video, volumetric content), the decoded data (image, video, volumetric content) or portions of the decoded data, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0051] In some embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for data encoding and decoding operations, such as for MPEG-2, HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding also known as H.266, standard developed by JVET, the Joint Video Experts Team).
[0052] The input to the elements of system 100C may be provided through various input devices as indicated in block 105, also mentionned in FIG. IB. Such input devices of system 100B or 100C include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. IB or 1C, include composite video.
[0053] In various embodiments, the input devices of block 105 in system 100B or 100C have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band- limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band- limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to- digital converter. In various embodiments, the RF portion includes an antenna.
[0054] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100B or 100C to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed- Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device.
[0055] Various elements of the systems 100A, 100B or 100C may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using the suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards. Similarly as for the sytem 100B of FIG. IB, the system 100C includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium.
[0056] Data is streamed to the system 100B or 100C, in various embodiments, using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100B or 100C using a set- top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100B or 100C using the RF connection of the input block 105. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
[0057] The system 100C may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The display 165 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other devices. The display 165 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide a function based on the output of the system 100C. For example, a disk player performs the function of playing the output of the system 100C.
[0058] In various embodiments, control signals are communicated between the system 100C and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100C via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100C using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100C in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
[0059] The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0060] In any of the systems 100A, 100B or 100C, the embodiments can be carried out by computer program product comprising code instructions that implements any one of embodiments described herein. The computer program product may be computer software implemented by the processor 110 or by hardware, or by a combination of hardware and software. As a nonlimiting example, the embodiments can be implemented by one or more integrated circuits. The memory 120 of any one of the systems 100A, 100B or 100C can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 110 of any one of the systems 100A, 100B or 100C can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0061] FIG. 2 illustrates an example of a block-based hybrid video encoder 200. Variations of this encoder 200 are contemplated, but the encoder 200 is described below for purposes of clarity without describing all expected variations.
[0062] In some embodiments, FIG. 2 also illustrate an encoder in which improvements are made to the HEVC standard or a VVC standard (Versatile Video Coding, Standard ITU-T H.266, 1SO / 1EC 23090-3, 2020) or an encoder employing technologies similar to HEVC or VVC, such as an encoder ECM (Enhanced Compression Model) under development by JVET (Joint Video Exploration Team).
[0063] Before being encoded, the video sequence may go through pre-encoding processing (201), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of color components), or re-sizing the picture (ex: down-scaling). Metadata can be associated with the pre-processing and attached to the bitstream.
[0064] In the encoder 200, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units of, for example, CUs (Coding units) or blocks. In the disclosure, different expressions may be used to refer to such a unit or block resulting from a partitioning of the picture. Such wording may be coding unit or CU, coding block or CB, luminance CB, or block. A CTU (Coding Tree Unit) refers to a group of blocks or group of units or group of coding units (CUs). In some embodiments, a CTU may be considered as a block, or a unit as itself.
[0065] Each unit is encoded using, for example, either an intra or inter mode. When a unit is encoded in an intra mode, it performs intra prediction (260). In an inter mode, motion estimation (275) and compensation (270) are performed. The intra mode and / or the inter mode may comprise several distinct sub-modes. For example, the intra mode may comprise directional intra predictions, template-based intra mode derivation prediction, intra block copy prediction or others modes spatially predicting the samples values of the unit. The inter mode may comprise skip mode, merge mode according to which motion information is derived from a list of motion candidates and no motion vector prediction residual is encoded, an inter mode according to which motion information is derived from a list of motion candidates and further refined either by encoding motion vector prediction residual or by template-matching performed both at the encoder and the decoder, further inter modes are also possible. The encoder decides (205) which one of the intra mode or inter mode to use for encoding the unit. When different intra modes and / or inter modes are possible, the endoder decides (205) which of the intra modes or inter modes to use. The encoder indicates the intra / inter decision by, for example, one or more syntax element signaling the prediction mode. The encoder may also blend (205) intra prediction result and inter prediction result, or blend results from different intra / inter prediction methods. Prediction residuals are calculated, for example, by subtracting (210) the predicted block from the original image block.
[0066] The motion refinement module (272) uses already available reference picture in order to refine the motion field of a block without reference to the original block. A motion field for a region can be considered as a collection of motion vectors for all pixels with the region. If the motion vectors are sub-block-based, the motion field can also be represented as the collection of all sub-block motion vectors in the region (all pixels within a sub-block have the same motion vector, and the motion vectors may vary from sub-block to sub-block). If a single motion vector is used for the region, the motion field for the region can also be represented by the single motion vector (same motion vectors for all pixels in the region).
[0067] The prediction residuals are then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes.
[0068] The encoder decodes (reconstructs) an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized (240) and inverse transformed (250) to decode prediction residuals. Combining (255) the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters (265) are applied to the reconstructed picture to perform, for example, one or more of a deblocking filtering, an SAG (Sample Adaptive Offset) filtering or an ALF (Adaptive Loop Filter) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer (280). Such filtered image is also referred to as a reference image in the following.
[0069] FIG. 3 illustrates a block diagram of a video decoder 300. In the decoder 300, a bitstream is decoded by the decoder elements as described below. Video decoder 300 generally performs a decoding pass reciprocal to the encoding pass as described in FIG. 2. The encoder 200 also generally performs video decoding as part of encoding video data.
[0070] In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide (335) the picture according to the decoded picture partitioning information. The transform coefficients are de- quantized (340) and inverse transformed (350) to decode the prediction residuals. Combining (355) the decoded prediction residuals and the predicted block, an image block is reconstructed.
[0071] The predicted block can be obtained (370) from intra prediction (360) or motion- compensated prediction (i.e., inter prediction) (375). In a similar manner as in the encoder, intra prediction and / or inter prediction may comprise several distinct sub-modes. The decoder obtains (370) the predictor block based on one or more syntax elements signaling the prediction mode among the available intra modes and inter modes. The decoder may blend (370) the intra prediction result and inter prediction result, or blend results from multiple intra / inter prediction methods. Before motion compensation, the motion field may be refined (372) by using already available reference pictures. In-loop filters (365) are applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380). Note that, for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side is identical to the contents of the reference picture buffer 280 on the encoder 200 side for the same picture.
[0072] The decoded picture can further go through post-decoding processing (385), for example, an inverse color transform (e.g. conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing (201), or re-sizing the reconstructed pictures (ex: up-scaling). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0073] Some of the embodiments described herein relates to entropy encoding and entropy decoding of at least one part of an image or a video to encode or decode. Embodiments described herein could also apply to any kind of input data that is being encoded / decoded, such as an image, a video, a volumetric content, audio signals, ....
[0074] In the case of image or video input data, any one of the embodiments described herein can be implemented for instance in entropy coding module of a video encoder and entropy decoding module of a video decoder. For instance, the embodiments described herein can be implemented in the entropy coding module 245 of the video encoder 200 in FIG. 2 or the entropy decoding module 330 of the video decoder 300 in FIG. 3.
[0075] In VVC (Versatile video coding, ITU-T H.266, TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU, SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS Infrastructure of audiovisual services - Coding of moving video, April 2022) and a current vesion of the ECM ( “Algorithm description of Enhanced Compression Model 10 (ECM 10) ”, M. Coban, R.-L. Liao, K. Naser, J. Strom, L. Zhang, document JVET-AE2025, 29th JVET Meeting), a large part of the signaling is done using an entropy coding of the values to transmit. Especially, using CAB AC (Context Adaptive Binary Arithmetic Coding), only binary values are encoded / decoded and generic (i.e. non-binary) values should first undergo a binarization process. The binzarization process provides a binary representation of the generic values, that is a representation using only binary symbols.
[0076] In the present document, the terms “bin” or “binary symbol” may be used interchangeably.
[0077] For each individual bin (binary symbol) to encode, a probability model is attached, to represent the conditional probability of the bin being equal to 1 or 0. This probability model may depend on some contextual information. This allows reaching an average coding rate for the considered bin that is close to the theoretical lower bound defined by the conditional entropy of the binary symbol given the contextual information. This conditional entropy is known to be lower than the non-contextual entropy, thus leading to a lower coding rate.
[0078] In the following, the probability of a bin is meant to be the probability that the bin value is ‘ 1’ . But the principles described herein apply similarly to the case that the probability is the one that the bin value is ‘O’.
[0079] The probability of each bin to encode is updated after each encoding / decoding, according to the value taken by the considered coded bin. The speed at which the probability is updated is a parameter of the model (i.e. the probability model associated to the bin). Another parameter is the initial probability used by the model, that is the probability value used for encoding / decoding the first bin that uses the considered context associated to the probability model.
[0080] An example of an entropy coding scheme is illustrated in FIG. 4:
[0081] For each bin to encode, at 410, a context is selected. Depending on the variant implemented, there may be a single context available for encoding the bin or multiple contexts available and a selection of a context among the multiple contexts can be made for example based on causal neighboring information of the block for which the bin is encoded / decoded.
[0082] At 420, it is determined if the bin is the first bin to encode using the selected context. If this is the case, at 430, an initial probability p is obtained for the selected context. The initial probability is known both at the encoder and decoder. Otherwise (no at 420), at 440, the current probability p associated to the context is obtained. At 450, the bin is encoded using the obtained probability value p. At 460, the current probability value is updated to provide an updated probability value p’ which then becomes the current probability value for the next bin to encode using the selected context. At 460, the current probability value is updated using a window size parameter w that is associated with the selected context.
[0083] As can be seen, each context is associated with a current probability p and a window size w corresponding to the update speed of the probability. Typically, the probability is updated (460 on FIG. 4) as follows: p' = a * b + (1 — a) * p, where p is the current probability, p’ is the updated probability, b the bin value encoded / decoded using the current probability p, and a = 1
[0084] — where w is the window size which determines the speed at which the probability is being updated.
[0085] In recent codecs like HEVC and VVC, a context is associated with two probabilities pO and pl and the two probabilities are maintained and updated during encoding / decoding. The two probabilities pO and pl are updated using two different window sizes, wO and wl. This allows to learn long term and short temr probabilities statsitics, for example wO is used to learn short term probabilities and wl is used to learn long term probabilities.
[0086] In the case of maintaining two probabilities, each bin is encoded / decoded by considering the weighted sum of the two probabilities pO and pl using a weight a. The windows sizes are updated each time a bin has been encoded / decoded depending on parameters dwOO, dwOl, dwlO, dwl 1 which are further explained below.
[0087] The current probability p or in case there are two probabilities pO and pl, are initialized using an initial probability value. Typically, pO and pl are initialized using the same initial probability value.
[0088] The parameters for a given context (initial probability, windows sizes, weight between probabilities) depend on external parameters, such as the type of slice to encode (intra I, biprediction, B uni-direction P), a swicth model (between B and P), a quantization parameter qp.
[0089] Each context uses parameters that are known between the encoder and the decoder, and these parameters are decided per context.
[0090] At the beginning of each slice, the parameters for each context are initialized as depicted in FIG. 5. Depending on a slice type of the slice being encoded / decoded, a slope and an offset are obtained, as well as window sizes wO and wl. The initial probability p is computed using the quantization parameter qp for the slice, and the obtained slope and offset. Moreover, a flag called sh_cabac_iniljl'ag is signaled in the slice header for non intra slices as shown in the table 1 below. This flag allows to switch the parameters to use for initialization: when the flag is true for a B slice, P slice parameters set are used, and the other way around.
[0091] Table 1
[0092] FIG. 6 shows an example of an overall CABAC engine for an inter slice. The parameters in dashed boxes are fixed parameters for a particular context. These parameters are chosen between 3 parameters set: the set defined for B slice, the set defined for P slice and the set from a previous model. The previous model is the one used for a previous coded slice which is carried over a following slice together with the probability values learned when encoding the previous slice. For Intra slices, only the I model is available.
[0093] Note that the dw values (dwOO, dwOl, dwlO, dwl l) are shared for all models (B, P, I and previous model).
[0094] As illustrated on FIG. 6, a bin is encoded / decoded (at 610) using probability value p which depends on probabilities p’0 and p’ 1 ans is determined as follows (at 620): p=a*p’0+(l-a)*p’l. The parameter a is tuned for each context and it allows to balance between the long term and short term probabilities independently for each context when encoding / decoding a bin.
[0095] These probabilities p’0 and p’ l are initialized with the same value as pO and pl (POi and Pli on FIG. 6), and p’0 and p’ l are updated using the previously used values of pO and pl, the observed bin and the modified update windows w0’ and wl’ (which also depend on the observed bin).
[0096] The probability pO and pl (at state t) are updated (respectively at 640 and 641) depending on the bin value, noted b in the formula, and window sizes wO and wl respectively: p0_next= p0_next and pl_next being the updated pO and pl values (at state t+1).
[0097] The windows size wO’ and wl’ are updated (respectively at 650 and 651) from wO and wl using the bin value b and the corresponding dw parameters as follows: w0’=w0+dw0[b] and wl’=wl+dwl[b], with dw0[b] and dwl[b] being dwOO and dwlO if b is 0, or dwOl and dwl l if b is 1. The probability p’O and p’ 1 (at state t) are then updated (respectively at 630 and 631) depending on the bin value b using the updated window sizes wO’ and wl’ as follows: pO’_next= aO’ * b with pO’ next and pl’ next being the updated p’O and p’ 1 values (at state t+1).
[0098] In the VVC standard, CABAC contains the following major changes compared to the design in HEVC: a Core CABAC engine, a separate residual coding structure for transform block and transform skip block and a context modeling for transform coefficients. The core CABAC engine is further described below.
[0099] The CABAC engine in HEVC uses a table -based probability transition process between 64 different representative probability states. In HEVC, the range ivlCurrRange representing the state of the coding engine is quantized to a set of 4 values prior to the calculation of the new interval range. The HEVC state transition can be implemented using a table containing all 64x4 8-bit pre-computed values to approximate the values of ivlCurrRange * pLPS( pStateldx ), where pLPS is the probability of the least probable symbol (LPS) and pStateldx is the index of the current state (this correspond for example to the probability p described in relation with FIG. 4). Also, a decode decision can be implemented using a pre-computed LUT. First ivlLpsRange is obtained using the LUT. Then, ivlLpsRange is used to update ivlCurrRange and calculate the output binV al. ivlLpsRange = rangeTabLps[ pStateldx ][ qRangeldx ] (3-1)
[0100] In VVC, the probability is linearly expressed by the probability index pStateldx. Therefore, all the calculation can be done with equations without LUT operation. To improve the accuracy of probability estimation, a multi-hypothesis probability update model is applied. The pStateldx used in the interval subdivision in the binary arithmetic coder is a combination of two probabilities pStateldxO and pStateldxl. The two probabilities pStateldxO and pStateldxlax associated with each context model and are updated independently with different adaptation rates. The two probabilities pStateldxO and pStateldxl correspond respectively for example to pO and pldescribed above in relation with FIG. 4 or FIG. 6 and adaptation rates respectively correspond to wO and wl. The adaptation rates of pStateldxO and pStateldxl for each context model are pre-trained based on the statistics of the associated bins. The probability estimate pStateldx is the weighted average of the estimates from the two hypotheses pStateldxO and pStateldxl . FIG. 7 shows a flowchart for decoding a single binary decision in VVC, that is FIG. 7 illustrates an example of decoding a binary symbol binVal. As done in HEVC, VVC CABAC also has a QP dependent initialization process invoked at the beginning of each slice. Given the initial value of a luma QP for the slice, the initial probability state of a context model, denoted as preCtxState, is derived as follows: m = slopeldx x 5 45 (3-2) n = (offsetldx << 3) +7 (3-3) preCtxState = Clip3(l, 127, ((m x (QP - 32)) » 4) + n) (3-4) where slopeldx and offsetldx are restricted to 3 bits, and total initialization values are represented by 6-bit precision. The probability state preCtxState represents the probability in the linear domain directly. Hence, preCtxState only needs proper shifting operations before input to arithmetic coding engine, and the logarithmic to linear domain mapping as well as the 256-byte table is saved. pStateldxO = preCtxState << 3 (3-5) pStateldxl = preCtxState << 7 (3-6)
[0101] In ECM (Algorithm description of Enhanced Compression Model 3 (ECM 3) JVET-X2025-v2 Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and 1SO / 1EC JTC 1 / SC 29
[0102] 23rd Meeting, by teleconference, 7-16 July 202 / ) extended precision and slice-type-based window size are proposed. The intermediate precision used in the arithmetic coding engine is increased, including three elements. First, the precisions for two probability states pStateldxO and pStateldxl are both increased to 15 bits, in comparison to 10 bits and 14 bits in VVC. Second, the LPS range update process is modified as follows:
[0103] If q>= 16384, q = 215-l-q
[0104] RLPS=((range*(q>>6))>>9)+l, where range is a 9-bit variable representing the width of the current interval, q is a 15 -bit variable representing the probability state of the current context model, and RLPS is the updated range for LPS. This operation can also be realized by looking up a 512x256-entry in 9-bit look-up table. Third, at the encoder side, the 256-entry look-up table used for bits estimation in VTM is extended to 512 entries.
[0105] Regarding the slice-type-based window size, since statistics are different with different slice types, it is beneficial to have a context probability state updated at a rate that is optimal under the given slice type. Therefore, for each context model, three window sizes are pre-defined for I-, B-, and P-slices, respectively, as the initialization parameters. The context initialization parameters and window sizes are retrained.
[0106] Study in ECM have been proposed for adaptive update rates, weighted average and state carryover (SEREGIN ET AL., "EE2-Test4.3: Combined, tests ofEE2-4.1 and EE2-4.2", Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, Document JVET- Z0135-vl, 26th Meeting, by teleconference, 20 April 2022). While the update rules for the two probability states pStateldxO and pStateldxl stay the same, encoding of the n-th bit (bin bn) with a given context is performed using probability states pStateldxO ’ and pStateldxl ’ as is described with FIG. 6. The probability states pStateldxO ’ and pStateldxl ’ are obtained from pStateldxO and pStateldxl using window sizes wO’ and wl’ which depend on the (n-l)-th bin noted bn-i (that is the previous encoded / decoded bin using the same given context), where: wO ’(n) = wO+dO(bn-i) wl ’(n) = wl+dl(bn-i) and dO(bi) and dl(bi) are look-up tables on offset values dOO, dOl, dlO and dl l defined at initialization for the considered context and added to the window size wO and wl.
[0107] Furthermore, the probability state used for encoding or decoding is a weighted average of pStateldxO ’ and pStateldxl ’ (instead of a simple average as before): pStateldx(n) = a*pState!dxO ’(n) + (1 -a)*pState!dx 1 ’(n), (3-55) where the parameter a (same as a in FIG. 6) is obtained from a EUT and depends on the context and on the slice type.
[0108] Finally, the initial state of some B or P-slices are inherited from previous slices instead of being re-initialized at each slice: more precisely, the two final states of each B-slice (or P-slice) are stored and used to initialize the next B-slice (or P-slice) in the same intra-period sharing the same temporal level and qp.
[0109] As explained above, for each slice, CAB AC encodes / decodes a series of bins bi, in specific contexts with distinct probability model. It may happen that a subset of these bin streams is exclusively composed of ‘0’ bins (Resp. ‘ 1’ bins). In the following, the trend of probability pStateldxl’ or pl’ as described in relation with FIG. 4-7 is considered, but the principles described herein can also be generalized to pStateldxO’ or pO’ and the global weighted probability pStateldx or p. p(i) is denoted the current probability pStateldx l’(i) in the probability model to have a ‘ 1’ bin (a binary symbol with a value 1). When coding N zero bins, p(i) progressively converges towards a small value. This progression follows the formula presented in (3-54) above and can be simplified when all bins bi = 0, and when adaptive update rates are null or negligible to: p(i+l) = (l-al’)*p(i) = (1-1 / 2ratel)*p(i), with ratel being the window size used for updating the probability. For example, when p(i) stands for the probability pStateldxl ’, ratel = wl’. p(i+l) / p(i) = 1- l / 2ratelp(i) = p(0)(l- 1 / 2iale l)' which is a geometric sequence of ratio r = (1- l / 2ratel) < 1, converging toward zero.
[0110] However, as illustrated in the update method in FIG. 7, probabilities are internally coded on 15 bits integer variables, e.g., p(i) values are comprised between [0, 32767] (215=32768). In VVC or ECM implementation, the above equation is not done using division with double precision, but binary shift operator in the integer domain: p(i+l)= p(i) - ( p(i) » ratel). This quantified formula keeps the same global geometric progression, with the artefact of clipping p(i) values whenever they are below a threshold.
[0111] It can be seen that whenever p(i) values can be coded on less than ratel bits, (e.g p below 2ratel), the term (p(i) » ratel) equals zero, and p(i+l) = p(i). As an effect, p(i) can stay above a minimum probability value pmin that is not null and depends on the initial windows size wl of the probability model.
[0112] A practical example of probability clamping and effect on CABAC cost is further explained below. FIG. 8 illustrates the updates of CABAC probability on a sequence of zero bins in ECM10, (top of FIG. 8) for sequence RitualDance, QP 37, image POC = 392, and CABAC context TransformSkipFlag(706). For this image, each transform skip flag for all CUs of the image is false, and a series of N zeros bins are encoded using a CABAC context with params: N zero bins = 584, p(0) = 1536, ratel=7. The two following equations are plot in FIG. 8: using a geometric progression of ratio r = (1-1 / 27) and using the update formula in ECM-10, that clampsPiif < 2ratel= 128. On the bottom of FIG. 8, the curves of the corresponding CAB AC cost are shown, that can be estimated with the following function: EstimatedCost(i)= -log2(l-p(i) / 32768). The total cost in bits of the encoded sequence is the sum of estimated costs for each bin, e.g. sum(EstimatedCost(i)). Visually, this can be interpreted as the integral or the area below the CAB AC curves. Faster the CAB AC curve converges to pmin and lower is pmin, smaller is the area surface, lower is the cumulated CAB AC cost for the sequence of bins.
[0113] The quantification on 15 bits and the clamping introduced by the ECM update method results in a total cost for the sequence of 10.2 bits, versus 8.6 bits for the geometric sequence, and therefore costs an overhead of 18 percents on the total encoded bits.
[0114] In the following, minimal CAB AC cost for encoding symbols is considered.
[0115] Minimal CAB AC cost for encoding a bin ‘O’:
[0116] As explained above regarding the entropy coding precision, the probabilities used for encoding are coded on 9 bits, so there are 29= 512 different encoding probabilities. By convention, each probability value is the middle of its interval. Therefore, the minimal encoding probability, called peiic min is Yi of 1 / 512 = 1 / 1024, and the minimal cost to encode a symbol zero is -log2(l- Penc min) = 0.0014 bits. This will be the first entry in the lookup table of encoding probabilities having 512 entries.
[0117] On the contrary the pStateldxl’ values used by the update method are coded on 15 bit, and must be normalized on 9 bits to find the corresponding CAB AC encoding probability in this lookup table of 512 elements.
[0118] In the above example, with ratel=7, due to clipping, pmin = 127 / 32768 which is comprised in the interval [1 / 512,2 / 512] corresponding to the second penc entry in the lookup table. So Penc=3 / 1024 is greater than penc_min and the CAB AC cost to encode symbol zero is not anymore miminal. It can be generalized that more ratel is high, more the penc due to clipping will be above penc_min, and less the CAB AC cost is minimal.
[0119] Practically, this reasoning concerns every context, for which pmin is clipped above 64 (26) because once normalized to 9 bits, the associated penc = Pmin » 6 is greater than 1. So every context having a parameter ratel > 6 is exposed to this issue.
[0120] Minimal CABAC cost for encoding a symbol ‘ 1” following a sequence of ‘0’ bins:
[0121] It must be observed that the p(i) value after a convergence to pmin on N zero bins bi also influences the cost to encode the next ‘ 1’ bin. If it is assumed that p(i) has converged to pmin equal to penc_min, then the cost to encode the next bin bn+i equal to 1 is : -log2(penc_min) = - log2( 1 / 1024) = 10.0 bits.
[0122] The total cost to encode such sequence bo . . . .bnbn+i = 0 01 is therefore not minimized by only decreasing pmin because doing so, the encoding cost of the last bin equal to 1 can be prohibitive, and the average cost for the wole sequence greater than with a higher pmin.
[0123] Some embodiments provide for optimizing the cumulated CABAC cost to encode a sequence of zero bins for which the pStateldxl’ probability (resp. the pStateldxO’, pStateldx’) converges toward a threshold value pmin. The fact that pmin is not the minimal encoding probability of the CABAC engine introduces a coding overhead that it is desirable to mitigate.
[0124] Some embodiments further provide for not deteriorating the CABAC encoding cost when a ‘ 1’ bin follows such a sequence of zero bins, for example by decreasing too much pmin that benefits for encoding ‘0’ bins, but presents a symmetric counter cost to encode next ‘ 1’ bin.
[0125] A further aim of some embodiments is to enforce the cabac encoder engine to internally use the lowest encoding probability as soon as such a convergence is detected. This is particulary significant for video sequences with small resolution (typically class C or D in CTC - Common Test Conditions), where a number of bins to encode in a picture can be low. For example in FIG. 8, on a serie of 584 bins, the update method reaches pmin only after 400 bins, e.g. 80 percent of the bins sequence, which is not enough reactive to adapt to the repartition probability to encounter a bin ‘O’.
[0126] According to a first embodiment, a detection of a convergence phase is added in the CABAC encoder and decoder. This additional step detects a convergence of pstates probabilities toward a minimum value, after a series of zero bins. When this convergence is detected, the pstate is updated so that it is bound in an interval [pl,p2] below the minimum value. Updating probabilities in this interval provides for lowering the probability state below an upper bound p2 that decreases the entropy coding cost on the series of ‘0’ bins and for limiting the probability state to be above a lower bound pl that limits the entropy coding cost for the first non-zero bin.
[0127] According to a second embodiment, in order to increase the rate of convergence while preserving the impact of changes on the threshold value of pmin, more flexibility in added in the upate process by decoupling two parameters: the window offset, and the clip value, which in the current implementation of ECM are tied in a mathematic relationship. This decoupling allows to increase the pstates precision, optimize the clip value, while increasing the convergence speed.
[0128] According to a third embodiment, the initialization parameters of the CABAC contexts are specialized per image resolution.
[0129] The mentionned embodiments are further described below. The embodiments are described for a given context denoted Ck in the following. They can apply to any of the CABAC contexts used in VVC or ECM or any arithmetic coding engine. Each context is identified by its context index and has its own initialization parameters defined per QP, and one or more internal states, e.g.: Stated, State 1, State which correspond respectively to the probability pStateldxO, pStateldxl, pStateldx described above. Embodiments described below relates to encoding or decoding a sequence of binary symbols or bins that uses the given context Ck. The sequence of bins corresponds to bins that have been generated for encoding at least one part of an image where the at least one part of the image can comprise: one or more blocks, one or more groups of blocks, a slice, one or more lines of groups of blocks, the image, or even more than an image. For example, embodiments described below are implemented in the entropy encoding module of a video encoder or the entropy decoding module of a video decoder. As described above, a binary symbol of the sequence represents either a syntax element to encode or decoder when the syntax element is a binary element such as a flag, or one bit from a binarization of a syntax element when the syntax element is not a binary element. In either cases, on the encoder side, the syntax elements have been obtained when encoding the at least one part of the image while on the decoder side, the syntax elements are to decoded from the decoded binary symbols to reconstruct the at least one part of the image.
[0130] For example, the binary symbols are the symbols that are provided as input to a context-based arithmetic encoding module of an encoder and as output of a context-based arithmetic decoding module of a decoder. As described above, the at least one context is associated to one or more probability values used for arithmetically encoding or decoding the sequence of binary symbols.
[0131] In a first embodiment, the update of the probabilities used for encoding or decoding a bin is modified when a convergence phase is detected for the the given context Ck. On both encoder and decoder sides, the update method to compute the new value p(i+l) of the probability is modified when it observes that p(i) probabilities converge to a minimum value. For that purpose, a new internal state is introduced, a so called “convergence state”. The engine enters this state for example after reception of bins when the updated probability p(i+ 1 ) is close to a pmin value for N consecutive bins. The engine stays in this state as long as bins bi are zero and leaves this state for example at first non-zero bin, or in in a variant, after a sufficient number M of non-zero bins has been observed.
[0132] When the engine enters the convergence phase, on each bin bi, the updated value p(i+ 1) of the probability is derived from the current value of the probability p(i) using a capped function in an interval [plk, p2k], below pmin.
[0133] FIG. 9 illustrates an example of a method 900 for encoding or decoding binary symbols that uses the context Ck according to this embodiment. As explained above, the method 900 is implemented for at least one part of an image to encode or decode and for at least one context Ck.
[0134] A counter nbZeroBin is initialized to 0, the counter counts the number of zero bin value encoded or decoded using the context Ck. A boolean IsInConvergenceState for the context Ck can also be initialized to false, this boolean value indicates when the context Ck is in a convergence state.
[0135] At 910, an ithbin is encoded or decoded using the probabilities associated to the context Ck, for example using equation 3-55 explained above. If the i'11bin is a 0, the counter nbZeroBin is incremented by 1. At 920, the probabilities are updated for example using equations 3-53 and 3-54 explained above. At 930, it is determined if the context Ck is in a convergence state. In other words, at 930, it is determined if one or more of the probabilities associated with the context Ck converge to a minimum probability value pmin on the N last encoded or decoded bins. The CAB AC engine can detect the entry in the convergence state in different ways. In a first variant, the offset added to the current probability p(i) to compute the new one (pi+1) is 0, which means that the probability keeps the same value. In a further variant, it is checked that the update states probabilities are equal to the previous ones, Nk times. In this further variant, it is checked that the probability stays the same for a number Nkof encoded / decoded bins.
[0136] In another variant, it is checked if the updated probability equals to or is below a minimum probability value. In a further variant, it is checked that the updated probability is in a window distance w2 to a minimum probability value pmin Nk times, e.g. [pmin -w2 / 2, pmin +w2 / 2], where Pmin and w2 are new CAB AC parameters added to each context Ck. For example, the minimum probability value pmin can be chosen as pmin = 2rate1 / 32768, and w2 a distance inversely proportional to the speed of convergence ratel. In another variant, the above criteria are used and a new initialization parameter Nk is added to the context Ck specifying a number of bins that have to be encoded or decoded before the detection of the convergence is performed.
[0137] In another variant, at 930, the convergence state is detected by checking in the history of bins encoded / decoded that the last Nk bins are equal to zero, for example by comparing the value of the counter nbZeroBin with Nk.
[0138] In another variant, a new State variable State” (i) is added to the context in addition to the stateO, Statel and State. The new State variable State”(i) is computed in the same way as Statel(i), but using a window coefficient w3, and it is checked if State”(i) converges towards zero.
[0139] If at 930, it is determined that the context Ck is in a convergence state, then at 940, it is determined whether it is the first bin for which the context Ck enters the convergence state or if the context Ck was already in a convergence state, for example the value of the boolean IsInConvergenceState is checked or the value of the counter nbZeroBin. If at 940, it is determined that this is the first bin bi for which a convergence is detected for the context Ck, then at 950, the updated probability p(i+ 1) is set to an upper bound probability value p2k that is below pmin. The value of the boolean IsInConvergenceState is set to true.
[0140] In a variant, the probability value p2k is a new initialization parameter of the context Ck that is known both at the encoder and decoder. In another variant, the probability value p2k is derived from p(i), for example with a linear function, p2k = a2k* p(i)+ b2k., with a2k and b2k being parameters defined for the context Ck.
[0141] If at 940, it is determined that this is not the first bin bi for which a convergence is detected for the context Ck, that is the context Ck progresses in the convergence state then at 960, the updated probability value p(i+ 1 ) is derived from p(i) using a function that bounds values inside an interval [plk,p2k], for example by clipping values above a lower bound plk and below an upper bound p2k.
[0142] This updating of the probability value is done while in the convergence state between bins i and j, with j being the bin for which the context Ck exits the convergence state.
[0143] In a variant, the updated probability value p(i+l) is derived from p(i) by decreasing slowly toward the lower bound plk, for example by subtracting a constant value at each bin, or applying a linear function p(i+l ) = alk* (i+l)+ blk. In another variant, the updated probability value p(i+l) is derived from p(i) by following a geometric progression toward the lower bound plk, at a convergence rate rk that depends on the context Ck: p(i+l ) = rk* p(i).
[0144] In another variant, the updated probability value p(i+l) is derived from p(i) by applying the same value: p(i+l ) = p(i) = p2k. If at 930, it is determined that the context Ck is not in a convergence state, then at 970, it is determined if the context Ckis exiting the convergence state. The engine exits the convergence state for example when at least one of the following conditions is true: after the first non-zero bin is observed or after a certain number Mk of consecutive non-zero bins is observed, or after a certain number Mk of not necessarily consecutive non-zero bins are observed or after a certain number Mk of non-zero bins are observed in a sliding window of fixed size Sk of bins. Or the value of the boolean IsInConvergenceState can be checked.
[0145] If it is determined that the context Ck is exiting the convergence state, then at 980, the context Ck leaves the convergence state by setting the updated value of the probability p(i+l ) to an exit convergence probability p3k. The counter nbZeroBin is also reset to 0 and the value of the boolean IsInConvergenceState is set to false.
[0146] In a variant, the p(i+ 1) probability is a fix value p3k which is a new initialization parameter of the context Ckthat can be known by the encoder and decoder. In another variant, p3k is derived from p(i), for example with a linear function, p3k = a3k* p(i)+ b3k. In another variant, the p(i+ 1 ) probability is derived from the last p(i) values, for example the p(i+ 1 ) probability is a weighted sum of the last L p(i) probabilities.
[0147] Otherwise (no at 970), when the engine is not in a convergence state, the update method computing p(i+l) is unchanged and the updated probalitiy p(i+l) determined at 920 is not changed.
[0148] In a second embodiment, the updating of the probabilites after encoding or decoding the binary symbol is modified with respect to conventional methods. This second embodiment does not involve the detection of any convergence phase.
[0149] The update of probabilities associated with a given context Ck after encoding / decoding a binary symbol using this given context Ck according for example to the current implementation of ECM is as follows: p(i+ 1 ) = p(i) - (p(i) > > rate 1 ) Equation( 1 ) with p(i) being one of the probabilities described above, for example pStateldxl’ and ratel then being wT.
[0150] Using Equation (1), it can be seen that ratel impacts the speed of convergence (increase the geometric ratio r = 1- l / 2ratel) and implicitly sets a clipping value, pmin equal to 2ratel, and stops the decrease of p(i) passed this value. In Equation (1), the clipping value and speed of convergence are tightly coupled which impacts any further training of the parameters, such as the speed of convergence. An aim of the embodiment described here is to advantageously decouple these two parameters: speed of convergence and clipping.
[0151] Equation(l) is thus modified to introduce a clipping parameter pciiP. An aim of this variable pciiPis to be able to tune separately in the update equation the speed convergence with variable ratel, and the pciiPvalue. The effect of changing these parameters independently allows reducing the ratel parameter (or, in other words, raising the convergence speed) which results in faster convergence, at a cost of a less stable probability model. However, unlike the current implementation, the choice of ratel no longer impacts the minimum probability state. Also, this change allows providing a higher pciiPthan the original clipping value of Equation ( 1 ) which results in lower coding cost for the first non-zero bin after a long series of zero bins but, at the same time, in a higher coding cost of long series of zero bins. This tradeoff can be optimized via training on a suitable dataset.
[0152] According to this embodiment, Equation (1) now becomes:
[0153] This modification can apply to the update of the one or more of the probabilities associated to the considered context. A pciiPvalue is thus defined for the context for each one of the probabilities which are to be updated using the modified Equation (1) above.
[0154] In a variant, the precision of the probability p(i) is also increased, for example from 15 bits to 32, to decrease the approximation introduced by the normalization on 15 bits, since the clipping is now ensured with pciiPand not by the normalization itself.
[0155] FIG. 10 illustrates an example of a method 1000 for encoding or decoding a sequence of binary symbols according to this embodiment. At 1010, a context Ck for encoding or decoding a binary symbol bi is obtained. As explained above, the method 1000 is implemented for at least one part of an image to encode or decode and for at least one context Ck. The method can be implemented or all contexts used by the arithmetic coding engine or for a subset of contexts. At 1020, a binary symbol bi is encoded or decoded using the probabilities associated to the context Ck, for example using equation 3-55 explained above. At 1030, the probability model is updated, as in the conventional manner, the window sizes are updated and the probabilities values are updated. When a probability p(i) is updated using the modified Equation (1), an update value p(i+l) of probability p(i) is determined as p(i)* (1- l / 2ratel). At 1040, it is checked whether the update value of p(i) satisfies or not a first criteria. In this embodiment, at 1040, it is checked if the update value p(i)*(l-l / 2ratel) is below the clipping value pciiPset for this probability, corresponding to max(p(i)*(l-l / 2rate1)), pciiP). If p(i)*( 1- l / 2ratel) is below pciiP, then at 1050, the update value assigned to p(i+l ) is pciiP, otherwise the update value assigned to p(i+l) remains p(i)*(l-l / 2ratel).
[0156] In a variant, in the method 1000, the updating of the probability value (1030), checking of the first criteria (1040) and assigning a new update value (1050) is done in a single operation using the max function in the updating step (1030), using the modified Equation (1):
[0157] FIG. 11 illustrates an example of coding costs when changing the ratel parameter while keeping a same clipping value pciiPfor the probability. The lower curve shows the use of the modified Equation (1), keeping the same value pciiP= pmin = 127, but with a modified ratel=6 instead of 7.
[0158] The upper curve shows the use of Equation (1). It can be seen that when using the modified Equation (1), the curve converges more rapidly toward pmin. Note that when using Equation (1) not modified, for ratel=6, the clipping will be much lower (pmin = 26-l = 63 in this case).
[0159] As other initializations parameters, pciiPfor each context Ck should be set to an optimal value learnt during a training phase. That is, instead of tuning (stateO, statel, rateO, ratel, ...) the training program now have to tune (stateO, statel, rateO, ratel, clipO, clipl . . .) resulting in different optimum values (ratel, clipl) than in the current implementation of ECM for example. In a third embodiment, the initialization parameters of the contexts used by the arithmetic coder are specialized per image resolution.
[0160] One limitation of the actual probability model for CAB AC context is that the same initial states and rates are used whatever the resolution of the picture is. The bins to encode at each slice corresponds to semantic flags in each CU, which number highly depends on the resolution (Typically N around 20000-40000 for sequences of class A(4K), and can be below 1000 for sequence of class C or D).
[0161] Obviously, as in example in FIG. 8, to converge toward a pmin after 400 bins has not the same impact on series of bins of length 1000, or for a few thousands, where the convergence phase is negligible relatively to the total length of the bins sequence.
[0162] In this third embodiment, for each context, different probability models are initialized with parameters (stateO, state 1, rateO, ratel, . . .) defined for each image resolution.
[0163] In a variant, a single probability model can be used for all resolutions, but some parameters are dervied from this common model (e.g: ratelcommon) depending on the current resolution R of the image, in the same way the initial state (initial probability value) is derived from the slice QP.
[0164] For example: for rate 1 parameter: ratel(R) = rate Lemmon * Weights(R), where WeightsQ is a lookup table of coefficients to multiply to the base rate of the model. Similar derivation can be done for the other parameters of the probability model of a context.
[0165] FIG. 12 illustrates an example of a method 1200 for initializing probability models for encoding or decoding a sequence of binary symbols, according to this embodiment. The method is for example performed for one or more slices of a video to encode or decode, and for each context Ck used by the arithmetic coder or for a subset of contexts. At 12010, the slice or image resolution R is read, for example in a SPS message accompanying the data encoding the image. At 1220, default probability models are read for the context Ck (slope value and offset value for each probability used in the model, rate for each probability used in the model,). At 1230, initial probability values are determined from the QP of the slice using the default slopes and offsets, for example as described with FIG. 5. At 1240, the window sizes or rates are derived from the default rates based on the image resolution for example by multiplying the default rate with a weight depending on the resolution. Embodiments provided herein can be implemented alone or in combination with one another. For example, in a further embodiment, the embodiments described in relation with FIG. 9 or 10 can be implemented in a combined embodiment wherein the updating of the probability at 920 in FIG. 9 is done using the embodiment described in FIG. 10 (1030-1040 and 1050), that is using the modified Equation (1).
[0166] In a variant, the initialization of the probability models used in the embodiments described in relation with FIG. 9 or FIG. 10 or in the combined embodiment mentioned above can be implemented according to the third embodiment with the initial values of the probability model that depends on the image resolution.
[0167] Any one of above-mentioned embodiments can be enabled or disabled for one or more contexts of the arithmetic coder, either enabled by default in a same manner at the encoder and the decoder, or enabled dynamically using appropriate signaling. In a variant, when enabled by default on the encoder and decoder, embodiments can also be disabled dynamically using appropriate signaling.
[0168] The different embodiments described herein decrease the CAB AC cost for a subset of contexts Ck, more specifically all contexts for which the series of bins follows a certain statistic distribution: where the proportion of sequence of zero bins is important, or where the average number of consecutive zero is important, so that a convergence happens.
[0169] As these characteristics are context and codec dependent, all the parameters of these embodiments are tuned differently for each context to adapt to the context’s statistic distribution, for example by training on a suitable dataset. These parameters include Mk,Nk. the bound interval [plk,p2k], the exit probability p3k, the coefficients of the update function when in convergence state, the coefficients alk,b Ik . . . a3k,b3k. the initialization parameter pciiP, the list of context Ck for which the use of one or more of the embodiments is enabled.
[0170] In a variant, all parameters of the invention are normative in the codec specification, and shared by the encoder and decoder. The specification gives the list of contexts for which a given embodiment described herein is enabled.
[0171] In a variant, appropriate signaling is used for enabling the use of a given embodiment described herein. FIG. 13 illustrates an example of a method 1300 for signaling or decoding syntax elements providing such appropriate signaling. At 1310, the encoder signals to the decoder a first activation list, in a metadata NAEU header, for example in a SPS, specifying the list of identifiers of the contexts or a group of contexts, for which the embodiment is active on a sequence of pictures. In a variant, in addition, at 1320, the encoder signals to the decoder for each slice, a second activation flag indicating if the embodiment is active or enabled on the current picture for a given context, for example in a slice header. This second activation flag can be signaled only for the contexts for which the embodiment is signaled as enabled at 1310. Or in a variant where 1310 is skipped, at 1320 the second activation flag is signaled for all contexts or a subset of contexts known to the decoder.
[0172] In another variant, the flags signaled at 1310 or 1320 can be disabling flag to disable the use of the embodiment for one or more contexts when the use of the embodiment is enabled by default at the encoder / decoder.
[0173] In a variant, all parameters, like plk, p2k, p3k, Pciip, are dynamically deduced from the current context parameters (ex: ratel) and / or the current probability p(i) by the encoder and the decoder. In a variant, a subset of these parameters are optimally computed by the encoder and signaled to the decoder, for example in a SPS header.
[0174] In an embodiment, illustrated in FIG. 14, in a transmission context between two remote devices A and B over a communication network NET, the device A comprises a processor in relation with memory RAM and ROM which are configured to implement a method for encoding a sequence of binary symbols according to any one of the embodiments described herein and the device B comprises a processor in relation with memory RAM and ROM which are configured to implement a method for decoding a sequence of binary symbols according to any one of the embodiments described herein. In accordance with an example, the network is a broadcast network, adapted to broadcast / transmit a coded video from device A to decoding devices including the device B.
[0175] FIG. 15 shows an example of the syntax of a signal transmitted over a packet-based transmission protocol. Each transmitted packet P comprises a header H and a payload PAYLOAD. In some embodiments, the pay load PAYLOAD may comprise data representative of at least one part of an image encoded according to any one of the embodiments described above. The payload can also comprise any signaling as described above. For example, the signal comprises the first or second activation flags mentioned above and / or new parameter values such as plk, p2k, p3k, Pciip mentioned above. Various implementations involve decoding. “Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this application, for example, entropy decoding a sequence of binary symbols to reconstruct image or video data.
[0176] As further examples, in one embodiment “decoding” refers only to entropy decoding, in another embodiment “decoding” refers only to differential decoding, and in another embodiment “decoding” refers to a combination of entropy decoding and differential decoding, and in another embodiment “decoding” refers to the whole reconstructing picture process including entropy decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0177] Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application, for example, determining re-sampling filter coefficients, re-sampling a decoded picture.
[0178] As further examples, in one embodiment “encoding” refers only to entropy encoding, in another embodiment “encoding” refers only to differential encoding, and in another embodiment “encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0179] Note that the syntax elements as used herein, are descriptive terms. As such, they do not preclude the use of other syntax element names.
[0180] This disclosure has described various pieces of information, such as for example syntax, that can be transmitted or stored, for example. This information can be packaged or arranged in a variety of manners, including for example manners common in video standards such as putting the information into an SPS, a PPS, a NAL unit, a header (for example, a NAL unit header, picture header or a slice header), or an SEI message. Other manners are also available, including for example manners common for system level or application level standards such as putting the information into one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for the purposes of session announcement and session invitation, for example as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. DASH MPD (Media Presentation Description) Descriptors, for example as used in DASH and transmitted over HTTP, a Descriptor is associated to a Representation or collection of Representations to provide additional characteristic to the content Representation. c. RTP header extensions, for example as used during RTP streaming. d. ISO Base Media File Format, for example as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and length also known as 'atoms' in some specifications. e. HES (HTTP live Streaming) manifest transmitted over HTTP. A manifest can be associated, for example, to a version or collection of versions of a content to provide characteristics of the version or collection of versions.
[0181] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process.
[0182] Some embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between the rate and distortion is usually considered, often given the constraints of computational complexity. The rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are different approaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one. Mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion.
[0183] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0184] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
[0185] Additionally, this application may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0186] Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0187] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0188] It is to be appreciated that the use of any of the following “and / or”, and “at least one of’, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[0189] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
[0190] As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0191] A number of embodiments has been described above. Features of these embodiments can be provided alone or in any combination, across various claim categories and types.
Claims
CLAIMS1. A method, comprising, for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image: arithmetically encoding the at least one binary symbol using the one or more probability values, updating the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, setting the updated value of the at least one probability value to a first given probability value.
2. An apparatus comprising one or more processors operable to, for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically encoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image: arithmetically encode the at least one binary symbol using the one or more probability values, update the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, set the updated value of the at least one probability value to a first given probability value.
3. A method comprising, for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image:arithmetically decoding the at least one binary symbol using the one or more probability values, updating the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, setting the updated value of the at least one probability value to a first given probability value.
4. An apparatus comprising one or more processors operable to, for at least one part of an image and for at least one context associated to one or more probability values used for arithmetically decoding at least one binary symbol, the at least one binary symbol being part of a binarization of a syntax element relating to the at least one part of the image: arithmetically decode the at least one binary symbol using the one or more probability values, update the one or more probability values based on the at least one binary symbol and a window size, and for at least one probability value of the one or more probability values, responsive to a determination that an updated value of the at least one probability value satisfies a first criteria, set the updated value of the at least one probability value to a first given probability value.
5. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein the determination that the updated value of the at least one probability value satisfies the first criteria includes comparing the updated value with the first given probability value, the first criteria being satisfied if the updated value is below the first given probability value.
6. The method or the apparatus of claim 5, wherein the one or more probability values are represented on more than 15 bits.
7. The method of any one of claims 1, 3 or 5-6 or the apparatus of any one of claims 2 or 4-6, wherein the updated value satisfies the first criteria if at least one of the following conditions is true:the updated value is the same as the at least one probability value before the updating, for a given number N of binary symbols using the at least one context, N being at least 1 , or the updated value is below a second given probability value, or the updated value is in a given range centered on the second given probability value, for the given number N of binary symbols using the at least one context, N being at least 1 , or a same value of binary symbols is arithmetically encoded or decoded for a given number of binary symbols using the at least one context.
8. The method or the apparatus of claim 7, wherein the first given probability value is an upper bound of a given probability range defined for the at least one context.
9. The method or the apparatus of claim 8, wherein the first given probability value is below the second given probability value.
10. The method or the apparatus of claim 8 or 9, wherein the first given probability value is derived from the at least one probability value.
11. The method or the apparatus of any one of claims 7-10, wherein for a subsequent binary symbol that uses the at least one context, the at least one probability value is updated using a function that bounds the updated value of the at least one probability value to a given probability range.
12. The method or the apparatus of claim 11, wherein the function comprises substracting a given offset or applying a linear function whose parameters are defined for the at least one context.
13. The method or the apparatus of any one of claims 7-12, wherein responsive to a determination that the updated value of the at least one probability value does not satisfy the first criteria and that a number M of binary symbols distinct from a binary symbol for which the first criteria was satisfied have been arithmetically encoded or decoded, the at least one probability value is assigned a third given probability value, with M being at least 1.
14. The method or the apparatus of claim 13, wherein the third given probability value is obtained based on parameters defined for the at least one context.
15. The method of any one of claims 1, 3, 5-14 or the apparatus of any one of claims 2 or 4-14, wherein initial values for at least one of the one or more probability values or of one or more window sizes used for updating the one or more probability values depend on a resolution of the image.
16. The method or apparatus of claim 15, wherein the one or more window sizes are derived for the at least one context using a base window size obtained for the at least context and a weight coefficient to apply to the base window size wherein the weight coefficient depends on the resolution of the image.
17. The method of any one of claims 1, 3, or 5-16 or the apparatus of any one of claims 2 or 4-16, wherein determining whether or not the updated value of the at least one probability value satisfies the first criteria is responsive to a determination that a first flag signaled at a sequence level indicates that the determination is enabled for the at least one context.
18. The method or the apparatus of claim 17, wherein determining whether or not the updated value of the at least one probability value satisfies the first criteria is further responsive to a determination that a second flag signaled at a slice level indicates that the determination is enabled for the at least one context.
19. A non-transitory computer readable medium storing coded data representative of at least one part of an image encoded according to the method of any one of claims 1 or 5-18.
20. The non-transitory computer readable medium of claim 19, wherein the coded data comprises a first flag signaled at a sequence level indicating whether the determination that the updated at least one probability value satisfies the first criteria is enabled or disabled for the at least one context.
21. The non-transitory computer readable medium of claim 19 or 20, wherein the codeddata comprises a second flag signaled at a slice level indicating whether the determination that the updated at least one probability value satisfies the first criteria is enabled or disabled for the at least one context.
22. A computer program product including instructions for causing one or more processors to carry out the method of any of claims 1, 3 or 5-21.
23. A non-transitory computer readable medium storing executable program instructions to cause a computer executing the program instructions to perform the method of any of claims 1, 3 or 5-21.
24. A device comprising: an apparatus according to claim 4; and at least one of (i) an antenna configured to receive or transmit a signal, the signal including data representative of the at least one part of the image, (ii) a band limiter configured to limit the signal to a band of frequencies that includes the data representative of the at least one part of the image, or (iii) a display configured to display the image.
25. A device according to claim 27, wherein the device comprises at least one of a television, a cell phone, a tablet, a set-top box.
Citation Information
Patent Citations
Context adaptive binary arithmetic coding method and device
US20220408091A1
Systems and methods for regularization-free multi-hypothesis arithmetic coding
US20230308651A1
Regularization of a probability model for entropy coding
WO2023163808A1
EP24305009A