Template-based reordering of GPM candidates
By reordering template-based costs and encoding syntax elements for pairs of candidates in video compression systems, the method addresses the challenge of achieving high compression efficiency with geometric partitioning modes, resulting in improved video encoding and decoding performance.
Patent Information
- Application Number
- PCT/EP2024/083231
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-05
AI Technical Summary
Existing video compression systems face challenges in achieving high compression efficiency when using geometric partitioning modes, particularly in accurately predicting and encoding geometric partitions within video blocks.
The method involves obtaining a list of template-based costs for pairs of candidates associated with splitting modes, reordering these costs, and encoding a syntax element that identifies the pair of candidates in the reordered list, allowing for efficient encoding and decoding of video blocks using geometric partitioning modes.
This approach enhances compression efficiency by optimizing the prediction and encoding of geometric partitions, leading to improved video encoding and decoding performance.
Smart Images

Figure EP2024083231_05062025_PF_FP_ABST
Abstract
Description
[0001] TEMPLATE-BASED REORDERING OF GPM CANDIDATES
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the priority to European Patent Application No.23307083.8, filed on November 29th, 2023, the entire disclosure of which is incorporated herein by reference.
[0004] TECHNICAL FIELD
[0005] The present embodiments generally relate to video compression. The present embodiments relate to a method and an apparatus for encoding or decoding an image or a video. More particularly, the present embodiments relate to improving geometric partionning modes in video compression system.
[0006] BACKGROUND
[0007] To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content. Generally, intra or inter prediction is used to exploit the intra or inter picture correlation, then the differences between the original block and the predicted block, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded. In inter prediction, motion vectors used in motion compensation are often predicted from motion vector predictor. To reconstruct the video, the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction.
[0008] SUMMARY
[0009] According to an aspect, a method for encoding a video is provided. The method comprises for at least one block of a video, that is intended to be encoded using a geometric partitioning mode, obtaining at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reordering the at least one template-based cost of at least one of the at least one list in a given order, encoding at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, encoding the at least one block using the pair of candidates identified and using a given splitting mode.
[0010] According to another aspect, an apparatus for encoding a video is provided. The apparatus comprises one or more processors operable to for at least one block of a video, that is intended to be encoded using a geometric partitioning mode, obtain at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reorder the at least one template-based cost of at least one of the at least one list in a given order, encode at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, encode the at least one block using the pair of candidates identified and using a given splitting mode.
[0011] According to an aspect, a method for decoding a video is provided. The method comprises for at least one block of a video, the at least one block being encoded using a geometric partitioning mode, obtaining at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reordering the at least one template-based cost of at least one of the at least one list in a given order, decoding at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, decoding the at least one block using the pair of candidates identified and using a given splitting mode.
[0012] According to another aspect, an apparatus for decoding a video is provided. The apparatus comprises one or more processors operable to for at least one block of a video, the at least one block being encoded using a geometric partitioning mode, obtain at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reorder the at least one template-based cost of at least one of the at least one list in a given order, decode at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, decode the at least one block using the pair of candidates identified and using a given splitting mode.
[0013] Further embodiments that can be used alone or in combination are described herein. One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform any one of the methods for encoding or decoding a video according to any of the embodiments described herein. One or more of the present embodiments also provide a non-transitory computer readable medium and / or a computer readable storage medium having stored thereon instructions for encoding or decoding a video according to the methods described herein.
[0014] One or more embodiments also provide a computer readable storage medium having stored thereon a bitstream generated according to the methods described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving the bitstream generated according to the methods described above.
[0015] BRIEF DESCRIPTION OF THE DRAWINGS
[0016] FIG. 1A illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to an embodiment.
[0017] FIG. 1 B illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to another embodiment.
[0018] FIG. 1 C illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to another embodiment.
[0019] FIG. 2 illustrates a block diagram of an embodiment of a video encoder within which aspects of the present embodiments may be implemented.
[0020] FIG. 3 illustrates a block diagram of an embodiment of a video decoder within which aspects of the present embodiments may be implemented.
[0021] FIG. 4 illustrates an example of a prediction process of a geometric partitioning mode (GPM) FIG. 5 illustrates an example of a geometric split line description.
[0022] FIG. 6 illustrates an example of a ramp function of the weights of the GPM blending based on the displacement (d) from a predicted sample position to the GPM partitioning boundary and the blending area size ( r).
[0023] FIG. 7 illustrates an example of a GPM signaling process.
[0024] FIG. 8A illustrates an example of a template of a block predicted using GPM.
[0025] FIG. 8B illustrates examples of blocks predicted using GPM with intra prediction.
[0026] FIG. 8C illustrates examples of angular and wide angular intra prediction modes from VVC.
[0027] FIG. 9 illustrates an example of a method for encoding a block of a video according to an embodiment.
[0028] FIG. 10 illustrates an example of a method for decoding a block of a video according to an embodiment. FIG. 11 illustrates an example of a method for encoding a block of a video according to another embodiment.
[0029] FIG. 12 illustrates an example of a method for decoding a block of a video according to another embodiment.
[0030] FIG. 13 illustrates an example of a method for encoding a block of a video according to another embodiment.
[0031] FIG. 14 illustrates an example of a method for decoding a block of a video according to another embodiment.
[0032] FIG. 15 illustrates an example of template-based costs determined for a set of splitting modes and a set of pairs of candidates.
[0033] FIG. 16 illustrates an example of a list of pairs of candidates ordered according to their second- best template-based costs.
[0034] FIG. 17 illustrates an example of a method for encoding a block of a video according to another embodiment.
[0035] FIG. 18 illustrates an example of a method for decoding a block of a video according to another embodiment.
[0036] FIG. 19 shows two remote devices communicating over a communication network in accordance with an example of the present principles.
[0037] FIG. 20 shows the syntax of a signal in accordance with an example of the present principles.
[0038] DETAILED DESCRIPTION
[0039] This application describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the application or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.
[0040] The aspects described and contemplated in this application can be implemented in many different forms. FIGs. 1A, 1 B, 1 C, 2 and 3 below provide some embodiments, but other embodiments are contemplated and the discussion of FIGs. 1 A, 1 B, 1 C, 2 and 3 does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.
[0041] In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, the terms “image,” “picture” and “frame” may be used interchangeably.
[0042] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[0043] The present aspects are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
[0044] FIG. 1A-1 C illustrates block diagrams of examples of systems in which various aspects and embodiments can be implemented. Any one of the systems 100A, 100B or 100B may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. In various embodiments, the system 100A, 100B or 100C is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 100A, 100B or 100C is configured to implement one or more of the aspects described in this application.
[0045] FIG. 1A illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. The system 100A includes at least one processor 1 10 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 1 10 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 100A includes at least one memory 120, e.g., a volatile memory device, and / or a non-volatile memory device, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The memory 120 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples. The processor 1 10 may be interconnected to the memory 120 by an interconnection bus 1 15.
[0046] Program code to be loaded onto processor 1 10 to perform the various aspects described in this application subsequently loaded onto memory 120 for execution by processor 110.
[0047] In some embodiments, memory inside of the processor 110 is used to store program code instructions and to provide working memory for processing that is needed during encoding or decoding. The input to the elements of system 100A may be provided through various input devices (non represented). Both Processor 1 10 and memory 120 can also have one or more additional interconnections to external connections.
[0048] FIG. 1 B illustrates a block diagram of an example of a system 100B in which various aspects and embodiments can be implemented. The system 100B includes the processor 110 and memory 120 described in relation with FIG. 1 A. The input to the elements of system 100B may be provided through various input devices as indicated in block 105 which is described further below with FIG. 1 C. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. 1 B, include composite video.
[0049] The various elements may be interconnected and transmit data therebetween using suitable connection arrangement 1 15, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
[0050] The system 100B includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium.
[0051] The system 100B may provide an output signal to various output devices, including a display, speakers, and other peripheral devices. The output devices may be communicatively coupled to system 100B via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100B using the communications channel 190 via the communications interface 150.
[0052] FIG. 1 C illustrates a block diagram of an example of a system 100C in which various aspects and embodiments can be implemented according to another embodiment. Elements of system 100C, singly or in combination, may be embodied in a single integrated circuit, multiple les, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100C are distributed across multiple les and / or discrete components.
[0053] The system 100C includes the processor 1 10 and memory 120 described in relation with FIG. 1 A or 1 B.
[0054] System 100C includes a storage device 140, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.
[0055] System 100C includes an encoder / decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents module(s) that may be included in a device to perform the encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100C or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
[0056] Program code to be loaded onto processor 1 10 or encoder / decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 1 10. In accordance with various embodiments, one or more of processor 1 10, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0057] In some embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 1 10 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2, HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding also known as H.266, standard developed by JVET, the Joint Video Experts Team).
[0058] The input to the elements of system 100C may be provided through various input devices as indicated in block 105, also mentionned in FIG. 1 B. Such input devices of system 100B or 100C include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. 1 B or 1 C, include composite video.
[0059] In various embodiments, the input devices of block 105 in system 100B or 100C have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna. Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100B or 100C to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 1 10 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device.
[0060] Various elements of system 100C may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
[0061] Similarly as for the sytem 100B of FIG. 1 B, the system 100C includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium.
[0062] Data is streamed to the system 100B or 100C, in various embodiments, using a Wi-Fi network such as IEEE 802.1 1 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100B or 100C using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100B or 100C using the RF connection of the input block 105. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
[0063] The system 100C may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The display 165 of various embodiments includes one or more of, for example, a touchscreen display, an organic lightemitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other devices. The display 165 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide a function based on the output of the system 100C. For example, a disk player performs the function of playing the output of the system 100C.
[0064] In various embodiments, control signals are communicated between the system 100C and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100C via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100C using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100C in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
[0065] The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0066] In any of the systems 100A, 100B or 100C, the embodiments can be carried out by computer program product comprising code instructions that implements any one of embodiments described herein. The computer program product may be computer software implemented by the processor 1 10 or by hardware, or by a combination of hardware and software. As a nonlimiting example, the embodiments can be implemented by one or more integrated circuits. The memory 120 of any one of the systems 100A, 100B or 100C can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as nonlimiting examples. The processor 1 10 of any one of the system 100A, 100B or 100C can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0067] FIG. 2 illustrates an example of a block-based hybrid video encoder 200. Variations of this encoder 200 are contemplated, but the encoder 200 is described below for purposes of clarity without describing all expected variations.
[0068] In some embodiments, FIG. 2 also illustrate an encoder in which improvements are made to the HEVC standard or a VVC standard Versatile Video Coding, Standard ITU-T H.266, ISO / IEC 23090-3, 2020) or an encoder employing technologies similar to HEVC or VVC, such as an encoder ECM (Enhanced Compression Model) under development by JVET (Joint Video Exploration Team).
[0069] Before being encoded, the video sequence may go through pre-encoding processing (201 ), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of color components), or re-sizing the picture (ex: down-scaling). Metadata can be associated with the pre-processing and attached to the bitstream.
[0070] In the encoder 200, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units of, for example, CUs (Coding units) or blocks. In the disclosure, different expressions may be used to refer to such a unit or block resulting from a partitioning of the picture. Such wording may be coding unit or CU, coding block or CB, luminance CB, or block. A CTU (Coding Tree Unit) refers to a group of blocks or group of units or group of coding units (CUs). In some embodiments, a CTU may be considered as a block, or a unit as itself.
[0071] Each unit is encoded using, for example, either an intra or inter mode. When a unit is encoded in an intra mode, it performs intra prediction (260). In an inter mode, motion estimation (275) and compensation (270) are performed. The intra mode and / or the inter mode may comprise several distinct sub-modes. For example, the intra mode may comprise directional intra predictions, template-based intra mode derivation prediction, intra block copy prediction or others modes spatially predicting the samples values of the unit. The inter mode may comprise skip mode, merge mode according to which motion information is derived from a list of motion candidates and no motion vector prediction residual is encoded, an inter mode according to which motion information is derived from a list of motion candidates and further refined either by encoding motion vector prediction residual or by template-matching performed both at the encoder and the decoder, further inter modes are also possible, such as the GPM mode described further below. The encoder decides (205) which one of the intra mode or inter mode to use for encoding the unit. When different intra modes and / or inter modes are possible, the endoder decides (205) which of the intra modes or inter modes to use. The encoder indicates the intra / inter decision by, for example, one or more syntax element signaling the prediction mode. The encoder may also blend (205) intra prediction result and inter prediction result, or blend results from different intra / inter prediction methods. Prediction residuals are calculated, for example, by subtracting (210) the predicted block from the original image block.
[0072] The motion refinement module (272) uses already available reference picture in order to refine the motion field of a block without reference to the original block. A motion field for a region can be considered as a collection of motion vectors for all pixels with the region. If the motion vectors are sub-block-based, the motion field can also be represented as the collection of all sub-block motion vectors in the region (all pixels within a sub-block have the same motion vector, and the motion vectors may vary from sub-block to sub-block). If a single motion vector is used for the region, the motion field for the region can also be represented by the single motion vector (same motion vectors for all pixels in the region).
[0073] The prediction residuals are then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes.
[0074] The encoder decodes (reconstructs) an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized (240) and inverse transformed (250) to decode prediction residuals. Combining (255) the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters (265) are applied to the reconstructed picture to perform, for example, one or more of a deblocking filtering, an SAO (Sample Adaptive Offset) filtering or an ALF (Adaptive Loop Filter) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer (280). Such filtered image is also referred to as a reference image in the following.
[0075] FIG. 3 illustrates a block diagram of a video decoder 300. In the decoder 300, a bitstream is decoded by the decoder elements as described below. Video decoder 300 generally performs a decoding pass reciprocal to the encoding pass as described in FIG. 2. The encoder 200 also generally performs video decoding as part of encoding video data.
[0076] In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide (335) the picture according to the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. Combining (355) the decoded prediction residuals and the predicted block, an image block is reconstructed. The predicted block can be obtained (370) from intra prediction (360) or motion- compensated prediction (i.e., inter prediction) (375). In a similar manner as in the encoder, intra prediction and / or inter prediction may comprise several distinct sub-modes. The decoder obtains (370) the predictor block based on one or more syntax elements signaling the prediction mode among the available intra modes and inter modes. The decoder may blend (370) the intra prediction result and inter prediction result, or blend results from multiple intra / inter prediction methods. Before motion compensation, the motion field may be refined (372) by using already available reference pictures. In-loop filters (365) are applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380). Note that, for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side is identical to the contents of the reference picture buffer 280 on the encoder 200 side for the same picture.
[0077] The decoded picture can further go through post-decoding processing (385), for example, an inverse color transform (e.g. conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing (201 ), or re-sizing the reconstructed pictures (ex: up-scaling). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0078] Some of the embodiments described herein relates to geometric partitionning mode used for encoding or decoding a block of the video to encode or decode. Embodiments provided herein aim at improving compression efficiency when using geometric partitionning mode.
[0079] Any one of the embodiments described herein can be implemented for instance in the prediction mode module of a video encoder and prediction mode module of a video decoder. For instance, the embodiments described herein can be implemented in the prediction mode module 205 of the video encoder 200 in FIG. 2 or the prediction mode module 370 of the video decoder 300 in FIG. 3.
[0080] Geometric Partition Mode (GPM) is also known as GEO mode or geometric split mode. GPM and GEO may be used interchangeably throughout the document, and both refer to the Geometric Partition Mode.
[0081] In VVC, a geometric partitioning mode (GPM) is supported for inter prediction. The GPM aims to increase the partition precision and to better fit the moving objects boundaries using a geometrical partition of a block. The geometric partitioning mode is signaled in VVC using a CU-level (block-level) flag as one kind of merge mode, with other merge modes.
[0082] FIG. 4 shows an example of the prediction process of GPM. In VVC, GPM is designed for block with size w x h = 2k x 21 (in terms of luma samples) with k, I e {3,...,6}. Moreover, GPM is disabled for a block that has an aspect ratio larger than 4:1 or smaller than 1 :4, considering narrow blocks rarely contain geometrically separated patterns.
[0083] When GPM is applied, the block is split into two parts, also called partitions in the following, which may be non-rectangular or non asymetric rectangular. The block is split by a straight partitioning boundary or edge, parameterized by an angle q>i and an offset pi. The partitioning boundary is also referred in the following as splitting line, or partitioning edge or partitioning line.
[0084] In total, 64 partitioning lines are supported in VVC and indexed by GPM partition index. Each part of the block associates a unidirectional MV that is coded in VVC with merge mode.
[0085] In VVC, a regular merge list is populated with spatial candidates which utilize motion information from spatially adjacent blocks, temporal candidates which utilize motion information from temporal blocks, HMVP (History-Based Motion Vector Predictor) candidates, paiwise average candidate and zero Motion Vector candidates. HMVP candidates are candidates that contain motion information that was previously coded and which is associated with adjacent and non-adjacent blocks relative to the block. A pairwise average candidate is generated by averaging the motion vectors of the first two available candidates of the merge candidate list.
[0086] The GPM merge list (also referred as geometric uni-prediction candidate list for VVC) is directly derived from this regular merge list using the parity of the indices Denote n as the index of the uni-prediction motion in the geometric uni-prediction candidate list. The LX motion vector of the n-th extended merge candidate, with X equal to the parity of n, is used as the n- th uni-prediction motion vector for geometric partitioning mode. In case a corresponding LX motion vector of the n-th extended merge candidate does not exist, the L(1 - X) motion vector of the same candidate is used instead as the uni-prediction motion vector for geometric partitioning mode.
[0087] In ECM, this process is invoked only for small blocks 8x8, 16x8 and 8x16. For larger blocks, the extraction process is bypassed, so the initial merge list as defined for a merge mode without GPM (and which may contain merged Bi-MVs candidates) is directly used as the GPM merge list.
[0088] A single GPM merge list is generated, and each GPM partition selects a candidate in this list. The GPM partition index signaling the splitting mode and the two GPM merge indices signaling respectively the candidates in the GPM merge list for each partition are coded into the bitstream. For each part, a block-based motion compensation prediction (MCP) is performed, resulting in two intermediate prediction blocks Po and Pi. A GPM prediction block PG is generated by performing a blending process using integer blending matrices Woand Wi, containing weights in the value range of [0, 8]. This can be expressed as PG = (Wo ° Po + Wi° Pi+ 4) » 3 (1 )
[0089] With Wo + Wi= 8Jw,h (2) where in (1 ) denotes the Hadamard product and Jw,h denotes a matrix of ones with size of the current block. The generated GPM prediction PGis subtracted from the original signal to generate residuals, which is transformed and coded into the bitstream using the regular VVC transform coding and CABAC engine.
[0090] The weights in the blending matrices of GPM are derived based on the displacement from a sample location to the partitioning boundary. The location of the splitting line is mathematically derived from the angle and offset parameters of a specific split as illustrated on FIG. 5. The angle q>i is quantized from between 0 and 360 degrees with a step equal to 11.25 degree (total 32). The description of a geometric split with angle <p and distance pi is depicted in FIG. 5. In the following, a geometric split with a given angle and a given offset is referred to as a split mode or a splitting mode.
[0091] Each partition of a geometric split mode in the block is inter-predicted using its own motion; only uni-prediction is allowed for each partition, that is, each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure that same as the conventional bi-prediction, only two motion compensated prediction are needed for each block. A blending of the two predictions is performed for samples in a blending are on each part of the splitting line.
[0092] In ECM (M. Coban, R-L. Liao, K.Naser, J. Strom, L.Zhang, "Algorithm description of Enhanced Compression Model 9 (ECM 9), " document JVET-AD2025, 30th Meeting, Antalya, TR, 21-28 April 2023), bi-prediction is allowed for each partition of a block split using GPM. In ECM, the blending area corresponding to the ramp strength can be selected among 5 values as depicted in FIG. 6. FIG. 6 shows the ramp function for the weights for GPM blending based on the displacement (d) from a predicted sample position to the GPM partitioning boundary and the blending area size (T).
[0093] GPM from VVC is extended in ECM by applying motion vector refinement on top of the existing GPM uni-directional MVs. Each geometric partition of a GPM block can decide whether to signal MVD (Motion Vector Difference) or not. If MVD is signalled for a geometric partition, after a GPM merge candidate is selected for a partition, the motion of the partition is further refined by the signalled MVD information. This is also referred as MMVD (Merge Motion Vector Differences).
[0094] FIG. 7 describes the signaling (700) of GPM in ECM. A block-level flag (geoFlag) indicates (710) whether the block is coded using GPM. Here, it is assumed the flag is set to true. Then, a flag is signaled (720) indicating whether the first partition of the block using GPM uses MMVD. If this is is the case, an index is signaled (730) to indicate the MVD. Otherwise, it is determined (740) whether the first partition is coded in an intra mode. If this is the case, syntax elements are signaled (750) indicating: a splitting mode indicating the split line, an index indicating an intra candidate used for the first partition and an index indicating an inter candidate used for the second partition. Otherwise, a flag is signaled (760) indicating whether the second partition of the block using GPM uses MMVD. If this is the case, an index is signaled (770) to indicate the MVD for the second partition, then the process goes to step 790. Otherwise, it is determined (780) whether the second partition is coded in an intra mode. If this is the case, syntax elements are signaled (750) indicating a splitting mode indicating the split line, an index indicating an intra candidate used for the second partition and an index indicating an inter candidate for the first partition. Otherwise, a flag is signaled (790) to indicate whether template matching (TM) is applied to both partitions of the block in GPM mode. When the template matching used, motion information for each geometric partition is refined using TM. When TM is chosen, a template is constructed using left, above or left and above neighboring samples according to the partition angle. The motion is then refined by minimizing the difference between the current template and the template in the reference picture.
[0095] When both partitions are inter-coded, as can be seen in the bottom left rectangle part (795), the split line used is signaled (out of 64 possible split lines), and each merge candidate used to predict one of the two partitions is signaled. A first index (mergeCandPartO) is signaled to indicate the merge candidate used for predicting the first partition of the block and a second index (mergeCandPartl ) is signaled to indicate the merge candidate used for predicting the second partition of the block.
[0096] In template matching (TM) based reordering for GPM split modes, given the motion information of the current GPM block (pair of inter motion candidates), the respective TM cost values of GPM split modes are computed. Then, all GPM split modes are reordered in ascending ordering based on the template-based cost values. For example, template-based costs are determined using SAD (Sum of Absolute Difference) but other metrics can be used for determining a cost using the template of a geometric partitioned block.
[0097] Instead of sending the GPM split mode, an index is signaled using Golomb-Rice code to indicate where the exact GPM split mode is located in the reordered list.
[0098] The reordering method for GPM split modes is a two-step process performed after the respective reference templates of the two GPM partitions in a block are generated, as follows. The GPM partition edge is extended into the reference templates of the two GPM partitions as illustrated on FIG. 8 showing the splitting edge of the block (current CU) extending in the template of the block (top template and left template). This results in 64 reference templates and the respective template-based costs are computed for each of the 64 reference templates. The GPM split modes are reordered based on their template-based cost values in ascending order and the best 32 split modes are marked as available split modes.
[0099] When determining template-based costs, the edge on the template is extended from that of the block as shown in FIG. 8, but GPM blending process is not used in the template area across the edge. The template-based cost is determined by computing a metric, for example the SAD, between samples values of the template of the first partition of the block and samples values of the corresponding template of the candidate for the first partition (grey area in FIG. 8) and between samples values of the template of the second partition of the block and samples values of the corresponding template of the candidate for the second partition (grey area in FIG. 8).
[0100] After ascending reordering using template-based cost, an index is signaled to indicate which splitting mode to use for predicting the block. The index signaling the splitting mode to use for encoding or decoding the block is signaled as the position of the splitting mode in the reordered list. For example, the splitting mode to use for predicting the block is determined by comparing a distortion cost or a rate-distortion cost determined for the current block and testing each one of the splitting modes that are marked as available. For example, the distortion cost is determined using a metric computed between the original samples values of the current block and the samples values of a reconstructed version of the current block when predicted using a given combination of a pair of candidates and a splitting mode. The ratedistortion cost further takes into account the cost for signaling the selected combination.
[0101] In ECM, in GPM with inter and intra prediction, the final prediction samples are generated by weighting inter predicted samples and intra predicted samples for each GPM-separated region. The inter predicted samples are derived by inter GPM whereas the intra predicted samples are derived by an intra prediction mode (IPM) candidate list and an index signaled from the encoder as described with FIG. 7. The IPM candidate list size is pre-defined as 3. The available IPM candidates are the parallel angular mode against the GPM block boundary (Parallel mode), the perpendicular angular mode against the GPM block boundary (Perpendicular mode), and the Planar mode as shown FIG. 8B (a-c) respectively. Furthermore, GPM with intra and intra prediction as shown FIG. 8B (d) is restricted to reduce the signalling overhead for IPMs and avoid an increase in the size of the intra prediction circuit on the hardware decoder. In addition, a direct motion vector and IPM storage on the GPM-blending area is introduced to further improve the coding performance.
[0102] In DIMD and neighboring mode based IPM derivation, Parallel mode is registered first. Therefore, at most two IPM candidates are derived from the decoder-side intra mode derivation (DIMD) method and / or the neighboring blocks can be registered if there is not the same IPM candidate in the list. As for the neighboring mode derivation, there are five positions for available neighboring blocks at most, but they are restricted by the angle of GPM block boundary as shown in Table 1 below, which are already used for GPM with template matching (GPM-TM).
[0103] Table 1 : The position of available neighboring blocks for IPM candidate derivation based on the angle of GPM block boundary. A and L denotes the above and left side of the prediction block.
[0104] GPM-intra can be combined with GPM with merge with motion vector difference (GPM-MMVD) TIMD is used on IPM candidates of GPM-intra to further improve the coding performance. The Parallel mode can be registered first, then IPM candidates of TIMD, DIMD, and neighboring blocks.
[0105] As can be seen from above, in ECM or VVC, the signaling of the merge candidates used for predicting each partition of a block coded using GPM is done independently. Also, the signaling of the split line is done independently from the signaling of the merge candidates used for predicting each partition of a block coded using GPM.
[0106] From the foregoing, it appears that coding cost of GPM parameters could be reduced, for example by using template information.
[0107] At least one embodiment relates to a method for encoding or decoding a video wherein one or more blocks of the video are encoded using GPM and wherein a same information is signaled to indicate a first candidate used for predicting a first partition of the block and a second candidate used for predicting a second partition of the block. The same information allows to identify a pair of candidates that includes the first candidate and the second candidate. In some variants, the pair of candidates is identified in an ordered list of pairs of candidates, the list being ordered according to template-based costs determined for each pair of candidates of the list.
[0108] In a variant, the same information is also used for identifying the split mode of the block in GPM.
[0109] In the document, the wording candidate refers to a predictor that can used for predicting the block. A predictor is a combination of specified values or data elements (e.g., sample value or motion vector) that may have been previously used for encoding or decoding a previous block.
[0110] In one variant, the candidates are merge candidates and the signaling provided in the embodiments described herein modifies the signaling of the first candidate and the second candidate as described at 795 in FIG. 7.
[0111] In another variant, the candidates of the pair of candidates that are signaled according to the embodiments described herein can be any of possible candidates for a geometric partition. For example, a candidate can be a motion candidate (from a merge mode, with or without mvd, or template-matching motion refinement), or an intra candidate obtained from an intra prediction mode. In this variant, the signaling provided in the embodiments described herein modifies the signaling of the pair of candidates signaled for example at 795 and 750 from FIG. 7.
[0112] It is assumed here that candidates are provided from a candidate list that is propulated in a similar manner at the encoder and the decoder. For example, the candidate list can be populated in as in VVC or ECM or according to any variation of the process done in VVC or ECM. The populating of the candidate list as described above is provided for description purposes, but the content of the candidate list is not limited to the sole examples of candidates described herein.
[0113] FIG.9 illustrates an embodiment of a method 900 for encoding a block of the video. At 910, one or more lists comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode is obtained. As discussed above, the splitting mode is part of a set of given splitting modes defined by an angle and an offset as described with FIG. 5. The splitting mode defines a splitting edge that splits the at least one block into two geometric partitions, and each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block. At 920, a combination of a pair of candidates and its associated splitting mode is selected for predicting the at least one block, for example using a RDO process. The selected combination is the one that has to be signaled to a decoder for the block to be correctly decoded. At 950, the block is encoded using the selected combination of pair of candidates and splitting mode. That is, each partition of the block is predicted from one of the candidates of the pair. Prediction residual is determined for the block and encoded.
[0114] For the signaling of the selected combination, at 930, at least one of the one or more lists is reordered in a given order, for example in ascending order of template-based costs considered. At 940, an indication for identifying the selected combination is signaled. Depending on the variants used, one or more syntax elements can be used. Some variants are described further below for the reordering of the at least one of the one or more lists at 930.
[0115] FIG. 10 illustrates an embodiment of a method 1000 for decoding a block of the video. For example, the block has been encoded using the method 900 described in relation with FIG. 10. At 1010, one or more lists comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode is obtained. At 1020, in a similar manner as what is done on the encoder side, at least one of the one or more lists is reordered in a given order, for example in ascending order of template-based costs considered and at 1030, an indication for identifying in the reordered list a combination of a pair of candidates and a splitting mode is decoded. Depending on the variants used, one or more syntax elements can be used for decoding the indication. Variants are described further below for the reordering of the at least one of the one or more lists at 1020. At 1040, the block is decoded, i.e. reconstructed, using the decoded combination of pair of candidates and splitting mode. That is, the block is partitioned according to the splitting mode and each partition of the block is predicted from one of the candidates of the identified pair. Prediction residual is decoded for the block and added to the prediction to reconstruct the block.
[0116] In a first embodiment, an index is used to signal the pair of candidates in the reordered pair of candidates. In this embodiment, the splitting mode is signaled by an index as in VVC or ECM. The signaling of the candidates in GPM is changed from what is done in ECM, by reordering the pairs of candidates that are associated to the signaled splitting mode based on their template-based costs. In this embodiment, instead of signaling each candidate separately, one index is signaled that corresponds to the ranking of the pairs of candidates in the reordered list based on their template-based costs.
[0117] In a second embodiment, a split index is signaled that corresponds to the ranking of the splitting mode for a specific pair of candidates. In this embodiment, the template costs are computed of each combination of pair of candidates and splitting modes. The signaled split is the ranking of the split for the chosen / selected pair of candidates. All pairs of candidates are reordered based on the template costs of the signaled rank, e.g., if rank 5 is signaled, all pairs of candidates are reordered based on the template cost of their 5thbest splitting mode, and an index signals which pair of candidates should be used in this reordered list.
[0118] In a third embodiment, a global index is signaled that identifies both the selected splitting mode and the selected pair of candidates. In this embodiment, only one index is required for signaling both a pair of candidates and a splitting mode.
[0119] FIG. 11 illustrates an example of a method 1100 for encoding a block of a video according to the first embodiment. At 11 10, one list per splitting mode is obtained. Each list comprises template-based costs determined for pairs of candidates associated to the same splitting mode (the splitting mode of the list). At 1 120, a best combination of a pair of candidates and a splitting mode is selected, for example based on RDO performed on the sample values of the block and on the reconstructed version of the block when predicted using the combination of the pair of candidates and the splitting mode. It is understood that the best combination is selected among all the determined combination of pairs of candidates and splitting mode determined at 1 110. At 1130, the splitting mode corresponding to the splitting mode of the combination selected is signaled, that is encoded in a bitstream. At 1 140, the list of templatebased costs obtained for the signaled splitting mode is reordered, for example in ascending order. At 1 150, the location in the reordered list of the pair of candidates corresponding to the pair of candidates of the selected combination is signaled. At 1160, the block is encoded as in step 950 described in relation with FIG. 9.
[0120] FIG. 12 illustrates an example of a method 1200 for decoding a block of a video according to the first embodiment. At 1210, the splitting mode is decoded from the bitstream. At 1220, a list of template-based costs is obtained where the template-based costs are determined for each available pair of candidates and the decoded splitting mode. At 1230, the list of templatebased costs obtained for the decoded splitting mode is reordered, for example in ascending order. At 1240, a syntax element indicating a location in the reordered list is decoded. The syntax element indicates the location of the pair of candidates in the reordered list from which the block shall be predicted at 1250. At 1250, the block is reconstructed as in step 1040 described in relation with FIG. 10.
[0121] In this embodiment, the template-based costs of each pair of merge candidates is evaluated. The split that shall be used for the block at the decoder is signaled with an index (1 130, 1210). In a variant, the splitting mode can be either not reordered. They are in a predefined order identical for all blocks. In another variant, the splitting modes are reordered based on parameters that are not the pairs of candidates as the same reordering shall be done on the encoder and decoder side. For example, the splitting modes may be ordered differently depending on block size and shape of a block to split (e.g., the diagonal and anti-diagonal splits of each splitting mode may be coded on fewer bits).
[0122] In another example, the splitting modes can be reordered based on template information that does not use the pairs of candidates.
[0123] For example, the splitting modes can be reordered based on a cost obtained for predicting samples of the template of the block using directional intra prediction modes. FIG. 8C ilustrates examples of directional intra prediction modes. An intra prediction angle of a directional intra prediction mode can be rounded to a closest angle that is parallel to a split line of a splitting mode, using the logic of GPM with inter and intra prediction, which uses the intra modes parallel to split lines (FIG. 8B (a)). In an example, splitting modes are reordered based on the cost of the corresponding directional intra prediction modes. In another example, only splitting modes having an angle corresponding to an angle of a directional intra prediction mode that provides a best cost are considered. And the splitting mode is selected among the set of splitting modes having this angle.
[0124] In another example, DIMD (Decoder-side Intra Mode Derivation) can be used to derive an intra prediction angle. Decoder side intra mode derivation (DIMD) is a prediction mode for coding a block in intra in ECM. When DIMD is applied to a block, up to five intra modes are derived from the reconstructed neighbor samples, and these five predictors are combined with the planar mode predictor with the weights derived from a histogram of gradients.
[0125] The DIMD process can be used to derive angles for the split lines based on the intra prediction angle from the derived intra modes provided by DIMD. In a similar manner as in the examples above, an intra prediction angle can be rounded to the closest angle that is parallel to a split line of a splitting mode, using the logic of GPM with inter and intra prediction, which uses the intra modes parallel to split lines (FIG. 8B (a)). In an example, only the splitting modes having an angle derived using the DIMD are considered.
[0126] At 1 140 and 1230, the pairs of candidates are then ordered based on their template costs when used with the signaled splitting mode, and the index of the pair in the ordered list is coded. In this way, the pairs of candidates with the lowest template cost (i.e. that provide better predictions of the template) may be coded on fewer bits than candidates with higher template costs. For example, in some variants, the code used may be a unary code. In some variants, the first pair of candidates (which has the lowest cost) may be signaled with a CABAC coded flag (e.g., the merge flag, to use the same contexts as the original GPM process) and the remaining pairs of candidates use a unary code. In this variant, the flag indicates whether the pair of candidates at the first location in the reordered list is the selected pair of candidates (1120) or not. If not, an index is then coded to signal the location of the selected pair of candidates between the second location in reordered list and the end of the list. In some variants, the reordered list is truncated to a maximum number of pairs of candidates to reduce the size of the list and the average cost of transmission of the location of the selected pair of candidates. For example, a maximum of 10 pairs may be used. In some variants, the list is truncated based on another parameter. For example, in some variants, only the pairs with a template cost smaller than K times the cost of the lowest template cost are selected (for example with K=1.25). In some variants the list is limited based on max_num_merge_cand which the maximum number of merge candidates. For example, only the first 2*max_num_merge_cand pairs are allowed to be used.
[0127] FIG. 13 illustrates an example of a method 1300 for encoding a block of a video according to the second embodiment. At 1310, one list per pair of candidates is obtained. Each list comprises template-based costs determined for splitting modes associated to a same pair of candidates and each list is ordered according to a given order, for example in ascending order. At 1320, a best combination of a pair of candidates and a splitting mode is selected, for example based on RDO performed on the original samples values of the block and on the samples values of the reconstructed version of the block when predicted using the combination of the pair of candidates and the splitting mode. It is understood that the best combination is selected among all the determined combination of pairs of candidates and splitting mode determined at 1310. At 1330, a ranking is signaled or encoded in a bitstream, the ranking is the ranking the selected splitting mode in the list associated to the selected pair of candidates. At 1340, the template-based costs of each list obtained at 1310 that are obtained for the signaled ranking are ordered, for example in ascending order, in a list. In other words, each template-based costs at the signaled ranking are put in a new list and ordered in ascending order, thus resulting in an ordered list of template-based costs for different pairs of candidates and splitting modes. At 1350, the location in the new ordered list of the selected pair of candidates is signaled. At 1360, the block is encoded as in step 950 described in relation with FIG. 9.
[0128] FIG. 14 illustrates an example of a method for decoding a block of a video according to the second embodiment. At 1410, a ranking of a splitting mode is decoded from in the bitstream. At 1420, in a similar manner as in the encoder at 1310, one list per available pair of candidates is obtained. Each list comprises template-based costs determined for splitting modes associated to a same pair of candidates and each list is ordered according to a given order, for example in ascending order. At 1430, as in step 1340 of the encoding method, the template-based costs of each ordered list obtained at 1420 that are ranked at the decoded ranking are ordered, for example in ascending order, in a second list. At 1440, a syntax element indicating a location in the ordered second list is decoded. The syntax element indicates the location of the pair of candidates in the ordered second list from which the block shall be predicted at 1450. At 1450, the block is reconstructed as in step 1040 described in relation with FIG. 10.
[0129] In this second embodiment, all combinations of pairs of candidates and splitting mode have their template cost computed. The splitting mode used for encoding / decoding the block is signaled as its ranking R in the ordered list of splitting modes for the pair of candidates selected. The pairs of candidates are then reordered based on the template-based cost of their Rthbest splitting mode. An index signals the ranking of the selected pair of candidates in this reordered list.
[0130] An example is provided below to describe the coding process. In this example, templatebased costs that have been found (1310, 1420) for using one of 6 splitting modes SO to S5 along with one of 4 pairs P0 to P3 are shown in FIG. 15. To signal the use of the splitting mode SO with the pair of candidates P3 (for example the combination selected at 1320), the splitting mode SO is the second-best splitting mode when used with the pair of candidates P3 (with a cost of 34, second only to splitting mode S5 with a cost of 19 for the same pair P3 of candidates). A first index of value 1 is signaled (1330, 1410). An index of 0 would have been signaled for the best splitting mode, i..e the first best splitting mode. All pairs (P0-P3) are reordered based on the template-based cost of their second-best splitting mode (1340, 1430). See the table illustrated on FIG. 16 showing the template-based costs of the second-best splitting mode for each pair, and the corresponding splitting mode indicated in parenthesis. The selected pair of candidates P3 is at the first location in this ordered list. A second index of value 0 is signaled (1350, 1440) to indicate that the pair of candidates P3 is used.
[0131] In some variants, the number of splitting modes that can be used is reduced from the number of all splitting modes available. This allows to reduce the complexity of the method. For example, at 1310 and 1420, the number of splitting modes for which template-based costs are obtained is reduced compared to the number of available splitting modes. This allows to avoid determining template-based costs for all possible combination of pair of candidates and splitting mode. In some variants, a fix number of N splitting modes are allowed (e.g., N=48).
[0132] In other variants, the number of splitting modes allowed is reduced but may vary, for example by reducing the number of splitting modes based on the template-based cost of the splitting mode for the pair of candidates that has the lowest template-based cost. For example, let C be the lowest template-based cost found for a combination of a splitting mode and a pair of candidates. For this pair of candidates, let N be the number of splitting modes that have a template-based cost strictly under K*C, with K=1 .25 for example. In this example, in some variants, only N splitting modes are used. Other values of K are also possible. This variant allows to reduce the signaling cost of the first ranking (1330, 1410) and the second ranking (1350, 1440). At 1320, the combination of splitting mode and pair of candidates used for predicting the block is thus selected among the reduced number of template-based costs that have a template-based cost strictly under K*C.
[0133] FIG. 17 illustrates an example of a method 1700 for encoding a block of a video according to the third embodiment. At 1710, one list of template-based costs is obtained that comprises template-based costs determined for pairs of candidates and splitting modes. In this variant, a single list is obtained that comprises all template-based costs available. The list is ordered in a given order, for example in ascending order of template-based costs. At 1720, a best combination of a pair of candidates and a splitting mode is selected, for example based on RDO performed on the samples values of the block and on the reconstructed version of the block when predicted using the combination of the pair of candidates and the splitting mode. It is understood that the best combination is selected among all the determined combination of pairs of candidates and splitting mode determined at 1710. At 1730, the location in the ordered list of the selected combination of the pair of candidates and the splitting mode is signaled. At 1740, the block is encoded as in step 950 described in relation with FIG. 9.
[0134] FIG. 18 illustrates an example of a method 1800 for decoding a block of a video according to the third embodiment. At 1810, in a similar manner as in the encoder (at 1710) one list of template-based costs is obtained that comprises template-based costs determined for pairs of candidates and splitting modes. In this variant, a single list is obtained that comprises template-based costs for all combination of pair of cnadidates and splitting modes that are available. The list is ordered in a given order, for example in ascending order of templatebased costs. At 1820, a syntax element indicating a location in the ordered list is decoded. The syntax element indicates the location of combination of the pair of candidates and the splitting mode that shall be used for predicting the block at 1830. At 1830, the block is reconstructed as in step 1040 described in relation with FIG. 10.
[0135] In this third embodiment, one index is used to signal both the splitting mode and the pair of candidates. In this embodiment, the template-based costs of all the combinations pairs of candidates and splitting modes are computed. All combinations are then ordered (1710, 1810) based on their template-based cost and an index transmits (1730, 1820) the ranking of the selected combination in the ordered list of combinations of splitting modes and pairs of candidates. Combinations with a lower template-based cost (and a lower rank in the ordered list) can then be coded on fewer bins than combinations with a higher template-based cost.
[0136] In some variants, the list obtained at 1710 and 1810 is truncated to limit the overall number of candidates (e.g., to a fixed number of candidates, or to a number of candidates that depends on parameters such as block size or shape, for example). In some variants, the list is truncated based on the template-based costs of the combinations of splitting modes and pairs of candidates. For example, only the combinations with a template-based cost smaller than or equal to K times the template-based cost of the best found combination (with lowest tempalte- based cost) are kept, with K =1 .25 for example (other values of K can be used).
[0137] In some variants, the list of combinations of pairs of candidates and splitting modes is limited but only after the list contains at least N different splitting modes (e.g., N=8). In other variants, the list is limited but only after it contains at least M different pairs of candidates (e.g., M=8, or M=max_num_merge_cand). In further other variants, the list is truncated so that there are not too many candidates with the same splitting modes. For example, only the first 4 combinations of splitting mode and pair of candidatess that have a given splitting mode are kept, and the fifth and subsequent combinations having the given splitting mode are not allowed.
[0138] In an embodiment, illustrated in FIG. 19, in a transmission context between two remote devices A and B over a communication network NET, the device A comprises a processor in relation with memory RAM and ROM which are configured to implement a method for encoding a video according to any one of the embodiments described herein and the device B comprises a processor in relation with memory RAM and ROM which are configured to implement a method for decoding a video according to any one of the embodiments described herein. In accordance with an example, the network is a broadcast network, adapted to broadcast / transmit a coded video from device A to decoding devices including the device B. FIG. 20 shows an example of the syntax of a signal transmitted over a packet-based transmission protocol. Each transmitted packet P comprises a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may comprise video data encoded according to any one of the embodiments described above. The payload can also comprise any signaling as described above. For example, the signal comprises coded data representative of at least one block of the video, the at least one block being encoded using a geometric partitioning mode and the coded data comprises at least one first syntax element that at least identifies a given pair of candidates in an ordered list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from the set of splitting modes. In some variants, the coded data further comprises a second syntax element indicating a given splitting mode in the set of splitting modes or indicating a rank order of the given splitting mode in an ordered list of splitting modes for the given candidates pair.
[0139] Various implementations involve decoding. “Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this application, for example, entropy decoding a sequence of binary symbols to reconstruct image or video data.
[0140] As further examples, in one embodiment “decoding” refers only to entropy decoding, in another embodiment “decoding” refers only to differential decoding, and in another embodiment “decoding” refers to a combination of entropy decoding and differential decoding, and in another embodiment “decoding” refers to the whole reconstructing picture process including entropy decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0141] Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application, for example, determining re-sampling filter coefficients, resampling a decoded picture.
[0142] As further examples, in one embodiment “encoding” refers only to entropy encoding, in another embodiment “encoding” refers only to differential encoding, and in another embodiment “encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0143] Note that the syntax elements as used herein, are descriptive terms. As such, they do not preclude the use of other syntax element names.
[0144] This disclosure has described various pieces of information, such as for example syntax, that can be transmitted or stored, for example. This information can be packaged or arranged in a variety of manners, including for example manners common in video standards such as putting the information into an SPS, a PPS, a NAL unit, a header (for example, a NAL unit header, picture header or a slice header), or an SEI message. Other manners are also available, including for example manners common for system level or application level standards such as putting the information into one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for the purposes of session announcement and session invitation, for example as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. DASH MPD (Media Presentation Description) Descriptors, for example as used in DASH and transmitted over HTTP, a Descriptor is associated to a Representation or collection of Representations to provide additional characteristic to the content Representation. c. RTP header extensions, for example as used during RTP streaming. d. ISO Base Media File Format, for example as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and length also known as 'atoms' in some specifications. e. HLS (HTTP live Streaming) manifest transmitted over HTTP. A manifest can be associated, for example, to a version or collection of versions of a content to provide characteristics of the version or collection of versions.
[0145] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process.
[0146] Some embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between the rate and distortion is usually considered, often given the constraints of computational complexity. The rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are different approaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one. Mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion.
[0147] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0148] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
[0149] Additionally, this application may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0150] Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0151] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information. It is to be appreciated that the use of any of the following “ / ”, “and / or”, and “at least one of”, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[0152] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
[0153] As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor- readable medium.
[0154] A number of embodiments has been described above. Features of these embodiments can be provided alone or in any combination, across various claim categories and types.
Claims
CLAIMS1 . A method, comprising, for at least one block of a video, the at least one block being encoded using a geometric partitioning mode: obtaining at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reordering the at least one template-based cost of at least one of the at least one list in a given order, decoding at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, decoding the at least one block using the pair of candidates identified and using a given splitting mode.
2. An apparatus comprising one or more processors operable to, for at least one block of a video, the at least one block being encoded using a geometric partitioning mode: obtain at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reorder the at least one template-based cost of at least one of the at least one list in a given order, decode at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, decode the at least one block using the pair of candidates identified and using a given splitting mode.
3. A method comprising, for at least one block of a video, the at least one block being intended to be encoded using a geometric partitioning mode:obtaining at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reordering the at least one template-based cost of at least one of the at least one list in a given order, encoding at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, encoding the at least one block using the pair of candidates identified and using a given splitting mode.
4. An apparatus comprising one or more processors operable to, for at least one block of a video, the at least one block being intended to be encoded using a geometric partitioning mode: obtain at least one list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from a set of splitting modes, the splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each candidate of the pair of candidates respectively predicts one of the two geometric partitions of the at least one block, reorder the at least one template-based cost of at least one of the at least one list in a given order, encode at least one first syntax element that at least identifies a pair of candidates in the re-ordered at least one list, encoding the at least one block using the pair of candidates identified and using a given splitting mode.
5. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein the given splitting mode is the splitting mode associated with the pair of candidates identified in the reordered list.
6. The method of any one of claims 1 , 3 or 5 or the apparatus of any one of claims 2 or 4-5, wherein the at least one list obtained is a single list wherein each one of the pair of candidates in the single list is associated to a splitting mode from the set of splittingmodes.
7. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein each one of the at least one list is associated to one splitting mode, the method further comprises or the one or more processors are further configured to decoding or encoding a second syntax element indicating the given splitting mode in the set of splitting modes, and wherein only the list associated to the given splitting mode is reordered.
8. The method of any one of claims 1 , 3 or 5 or the apparatus of any one of claims 2 or 4-5, wherein each list of the at least one list comprises template costs associated to a same pair of candidates, the same pair of candidates associated to each list being distinct, and each list of the at least one list is an ordered list wherein template costs of said list are ordered in the given order, the method further comprises or the one or more processors are further configured to decoding or encoding a third syntax element indicating a rank order of the given splitting mode in an ordered list of splitting modes for a given candidates pair, and the reordering of the at least one list is a reordering in the given order of template costs associated with pairs of candidates that are associated with a splitting mode that is at the indicated rank order in each of the ordered lists respectively associated with the pairs of candidates.
9. The method of any one of claims 1 , 3 or 5-8 or the apparatus of any one of claims 2 or 4-8, wherein the at least one first syntax element comprises a flag that indicates whether or not the pair of candidates to identify in the reordered at least one list is at a first position in the reordered at least one list.
10. The method or the apparatus of claim 9, wherein the at least one first syntax element comprises an index indicating a position of the pair of candidates to identify in the reordered at least one list after the first position.1 1 . The method of any one of claims 1 , 3 or 5-10 or the apparatus of any one of claims 2 or 4-10, wherein the reordered at least one list is truncated to a given number of template costs.
12. The method of any one of claims 1 , 3 or 5-11 or the apparatus of any one of claims 2 or 4-11 , wherein the reordered at least one list only comprises template costs that are lower than a given percentage of a lowest template cost.
13. The method or the apparatus of claim 9, wherein a number of splitting modes allowed for a pair of candidates is set to a given number.
14. The method or the apparatus of claim 9, wherein each list of the at least one list is associated to one pair of candidates, and each list only comprises template costs that are lower than a given percentage of a lowest template cost associated to the pair of candidates associated to said list.
15. The method or the apparatus of claim 11 , wherein the reordered at least one list is truncated only when the reordered at least one list comprises template costs for a given number of different splitting modes.
16. The method or the apparatus of claim 11 , wherein the reordered at least one list is truncated only when the reordered at least one list comprises template costs for a given number of different pairs of candidates.
17. The method of any one of claims 1 , 3 or 5-10 or the apparatus of any one of claims 2 or 4-10, wherein the reordered at least one list comprises less than a given number of template costs associated with a same splitting mode.
18. A non-transitory computer readable medium storing coded data representative of at least one block of a video, the at least one block being encoded using a geometric partitioning mode comprising a given splitting mode from a set of splitting modes, the given splitting mode defining a splitting edge that splits the at least one block into two geometric partitions, and wherein each one of the two geometric partitions is respectively predicted by one candidate of a given pair of candidates, the coded data comprising at least one first syntax element that at least identifies the given pair of candidates in an ordered list comprising at least one template-based cost determined for a pair of candidates associated with a splitting mode from the set of splitting modes.
19. The non-transitory computer readable medium of claim 18, wherein the coded data further comprises a second syntax element indicating the given splitting mode in the set of splitting modes.
20. The non-transitory computer readable medium of claim 18, wherein the given splitting mode is the splitting mode associated with the pair of candidates identified in the reordered list.21 . The non-transitory computer readable medium of claim 20, wherein the coded data further comprises a second syntax element indicating a rank order of the given splitting mode in an ordered list of splitting modes for the given pair of candidates.
22. The non-transitory computer readable medium of any one of claims 18-21 , wherein the at least one first syntax element comprises a flag that indicates whether or not the given pair of candidates to identify in the reordered at least one list is at a first position in the reordered at least one list.
23. The non-transitory computer readable medium of claim 22, wherein the at least one first syntax element comprises an index indicating a position of the given pair of candidates to identify in the reordered at least one list after the first position.
24. A computer program product including instructions for causing one or more processors to carry out the method of any of claims 1 , 3 or 5-17.
25. A non-transitory computer readable medium storing executable program instructions to cause a computer executing the program instructions to perform the method of any of claims 1 , 3 or 5-17.
26. A device comprising: an apparatus according to claim 2; and at least one of (i) an antenna configured to receive or transmit a signal, the signal including data representative of a video, (ii) a band limiter configured to limit the signal to a band of frequencies that includes the data representative of the video, or (iii) a display configured to display the video.
27. A device according to claim 28, wherein the device comprises at least one of a television, a cell phone, a tablet, a set-top box.