Template-based reordering of gpm candidates
By optimizing the candidate pair reordering and encoding of the geometric partitioning pattern, the problem of low video coding efficiency in existing technologies is solved, and more efficient video compression is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTERDIGITAL CE PATENT HOLDINGS SAS
- Filing Date
- 2024-11-22
- Publication Date
- 2026-06-26
Smart Images

Figure CN122295933A_ABST
Abstract
Description
[0001] Cross-reference to related applications This application claims priority to European Patent Application No. 23307083.8, filed on 29 November 2023, the entire disclosure of which is incorporated herein by reference. Technical Field
[0002] This embodiment generally relates to video compression. Specifically, it relates to methods and apparatus for encoding or decoding images or videos. More specifically, this embodiment relates to improving the geometric partitioning pattern in a video compression system. Background Technology
[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to utilize spatial and temporal redundancy in video content. Intra-frame or inter-frame prediction is usually used to leverage intra-frame or inter-frame image correlations, followed by transform, quantization, and entropy coding of the difference between the original block and the predicted block, typically represented as prediction error or prediction residual. In inter-frame prediction, the motion vectors used for motion compensation are usually predicted from a motion vector predictor. To reconstruct the video, the compressed data is decoded through the inverse process corresponding to entropy coding, quantization, transform, and prediction. Summary of the Invention
[0004] According to one aspect, a method for encoding video is provided. The method includes, for at least one block of a video intended to be encoded using a geometric partitioning pattern, obtaining at least one list comprising at least one template-based cost determined for at least one pair of candidate partitions associated with a partitioning pattern from a set of partitioning patterns, the partitioning pattern defining partitioning edges that divide the at least one block into two geometric partitions, and wherein each candidate in the candidate pair predicts one of the two geometric partitions of the at least one block, reordering at least one template-based cost of at least one in the at least one list in a given order, encoding at least one first syntax element that at least identifies the candidate pair in the reordered at least one list, and encoding the at least one block using the identified candidate pair and the given partitioning pattern.
[0005] According to another aspect, an apparatus for encoding video is provided. The apparatus includes one or more processors operable to, for at least one block of a video intended to be encoded using a geometric partitioning pattern, obtain at least one list comprising at least one template-based cost determined for at least one pair of candidate partitions associated with a partitioning pattern from a set of partitioning patterns, the partitioning pattern defining partitioning edges that divide the at least one block into two geometric partitions, and wherein each candidate in the candidate pair predicts one of the two geometric partitions of the at least one block, reorders at least one template-based cost of at least one in the at least one list in a given order, encodes at least one first syntax element that at least identifies the candidate pair in the reordered at least one list, and encodes the at least one block using the identified candidate pair and the given partitioning pattern.
[0006] According to one aspect, a method for decoding video is provided. The method includes, for at least one block of video, encoding the at least one block using a geometric partitioning pattern, obtaining at least one list, the at least one list including at least one template-based cost determined for candidate pairs associated with a partitioning pattern from a set of partitioning patterns, the partitioning pattern defining partitioning edges that divide the at least one block into two geometric partitions, and wherein each candidate in the candidate pair predicts one of the two geometric partitions of the at least one block, reordering at least one template-based cost of the at least one list in a given order, decoding at least one first syntax element that at least identifies the candidate pairs in the reordered at least one list, and decoding the at least one block using the identified candidate pairs and the given partitioning pattern.
[0007] According to another aspect, an apparatus for decoding video is provided. The apparatus includes one or more processors operable to, for at least one block of video encoded using a geometric partitioning pattern, obtain at least one list comprising at least one template-based cost determined for candidate pairs associated with a partitioning pattern from a set of partitioning patterns, the partitioning pattern defining partitioning edges that divide the at least one block into two geometric partitions, and wherein each candidate in the candidate pair predicts one of the two geometric partitions of the at least one block, reorders at least one template-based cost of at least one in the at least one list in a given order, decodes at least one first syntax element that at least identifies the candidate pair in the reordered at least one list, and decodes the at least one block using the identified candidate pair and the given partitioning pattern.
[0008] This document describes other embodiments that can be used alone or in combination.
[0009] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause one or more processors to perform any of the methods for encoding or decoding video according to any embodiment described herein. One or more of these embodiments also provide a non-transitory computer-readable medium and / or a computer-readable storage medium having instructions stored thereon for encoding or decoding video according to the methods described herein.
[0010] One or more embodiments also provide a computer-readable storage medium having a bit stream generated according to the method described herein stored thereon. One or more embodiments also provide methods and apparatus for transmitting or receiving bit streams generated according to the methods described above. Attached Figure Description
[0011] Figure 1A A block diagram of a system in which aspects of this embodiment can be implemented according to an embodiment is shown.
[0012] Figure 1B A block diagram of a system according to another embodiment in which aspects of this embodiment can be implemented is shown.
[0013] Figure 1C A block diagram of a system according to another embodiment in which aspects of this embodiment can be implemented is shown.
[0014] Figure 2 A block diagram of an embodiment of a video encoder in which various aspects of this embodiment can be implemented is shown.
[0015] Figure 3 A block diagram of an embodiment of a video decoder in which various aspects of this embodiment can be implemented is shown.
[0016] Figure 4 An example of the prediction process for the Geometric Partitioning Pattern (GPM) is shown. Figure 5 An example of a geometric dividing line description is shown.
[0017] Figure 6 An example of a ramp function is shown, which calculates the weights of GPM mixing based on the displacement (d) from the predicted sample location to the GPM partition boundary and the size (r) of the mixing region.
[0018] Figure 7 An example of the GPM signaling process is shown.
[0019] Figure 8A An example template for a block predicted using GPM is shown.
[0020] Figure 8BAn example of using GPM prediction with intra-frame prediction is shown.
[0021] Figure 8C Examples of angular and wide-angle intra-frame prediction modes from VVC are shown.
[0022] Figure 9 An example of a method for encoding blocks of video according to an embodiment is shown.
[0023] Figure 10 An example of a method for decoding blocks of video according to an embodiment is shown.
[0024] Figure 11 An example of a method for encoding blocks of video according to another embodiment is shown.
[0025] Figure 12 An example of a method for decoding blocks of video according to another embodiment is shown.
[0026] Figure 13 An example of a method for encoding blocks of video according to another embodiment is shown.
[0027] Figure 14 An example of a method for decoding blocks of video according to another embodiment is shown.
[0028] Figure 15 An example of template-based cost determination for a set of partitioning patterns and a set of candidate pairs is shown.
[0029] Figure 16 An example is shown of a list of candidate pairs sorted by their second-best template-based cost.
[0030] Figure 17 An example of a method for encoding blocks of video according to another embodiment is shown.
[0031] Figure 18 An example of a method for decoding blocks of video according to another embodiment is shown.
[0032] Figure 19 An example of two remote devices communicating via a communication network is shown, based on this principle.
[0033] Figure 20 The syntax of an example signal based on this principle is shown. Detailed Implementation
[0034] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in a specific manner and are generally described in a way that may sound restrictive, at least to illustrate individual characteristics. However, this is for the purpose of clarity and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in previous filings.
[0035] The aspects described and envisioned in this application can be implemented in many different forms. The following... Figure 1A , 1B 1C, 2, and 3 provide some embodiments, but other embodiments are contemplated, and... Figure 1A , 1B The discussions in 1C, 2, and 3 do not limit the breadth of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of generated or encoded bitstreams. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.
[0036] In this application, the terms “reconstructed” and “decoded” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0037] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first," "second," etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and may occur, for example, before, during, or in a time period overlapping with the second decoding.
[0038] This aspect is not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations, whether pre-existing or developed in the future, as well as any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.
[0039] Figure 1A-1C A block diagram illustrating examples of systems in which various aspects and embodiments can be implemented is shown. Any of systems 100A, 100B, or 100C can be embodied as a device including the various components described below and configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. In various embodiments, systems 100A, 100B, or 100C are communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, systems 100A, 100B, or 100C are configured to implement one or more aspects described in this application.
[0040] Figure 1A A block diagram illustrating an example system in which various aspects and embodiments can be implemented is shown. System 100A includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the aspects described herein. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100A includes at least one memory 120, such as a volatile memory device and / or a non-volatile memory device, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drive, and / or optical disk drive. As a non-limiting example, memory 120 may include an internal storage device, an attached storage device, and / or a network-accessible storage device. Processor 110 may be interconnected to memory 120 via interconnect bus 115.
[0041] The program code to be loaded onto processor 110 to execute the various aspects described in this application is then loaded onto memory 120 for execution by processor 110.
[0042] In some embodiments, the memory within processor 110 is used to store program code instructions and provide working memory for processing required during encoding or decoding. Input to the components of system 100A can be provided through various input devices (not shown). Both processor 110 and memory 120 may also have one or more additional interconnects to external connections.
[0043] Figure 1B A block diagram illustrating an example of a system 100B in which various aspects and embodiments can be implemented is shown. System 100B includes... Figure 1AThe processor 110 and memory 120 are described. Inputs to the components of system 100B can be provided through various input devices, as indicated in box 105, which are utilized below. Figure 1C Further description. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a collection of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1B Other examples not shown include composite video.
[0044] Various components can be interconnected and data can be transferred between them using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit board.
[0045] System 100B includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.
[0046] System 100B can provide output signals to various output devices, including displays, speakers, and other peripherals. Output devices can be communicatively coupled to system 100B via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100B via communication interface 150 using communication channel 190.
[0047] Figure 1C A block diagram of an example of a system 100C according to another embodiment, in which various aspects and embodiments may be implemented, is shown. The elements of the system 100C may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of the system 100C are distributed across multiple ICs and / or discrete components.
[0048] System 100C includes information about Figure 1A or Figure 1B The processor 110 and memory 120 are described.
[0049] System 100C includes storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drive, and / or optical disk drive. As a non-limiting example, storage device 140 may include internal storage device, attached storage device, and / or network-accessible storage device.
[0050] System 100C includes an encoder / decoder module 130 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 130 may be implemented as a separate element of system 100C, or it may be incorporated into processor 110 as a combination of hardware and software known to those skilled in the art.
[0051] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.
[0052] In some embodiments, memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, for example for MPEG-2, HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Coding, also known as H.266, a standard developed by JVET, the Joint Video Experts Group).
[0053] Inputs to the components of system 100C can be provided through various input devices, as indicated in box 105, and also... Figure 1B As mentioned above, such input devices of system 100B or 100C include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) component (COMP) input terminals (or a collection of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1B or Figure 1C Other examples not shown include composite video.
[0054] In various embodiments, the input device of block 105 in system 100B or 100C has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting it again to a narrower frequency band to select, for example, a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, downconverter, demodulator, error corrector, and demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0055] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 100B or 100C to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, within a separate input processing IC or within processor 110. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processor 110. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in combination with memory and storage elements to process the data streams as needed for presentation on the output device.
[0056] Various components of system 100C can be provided within an integrated housing, in which various components can be interconnected and transmit data between them using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit board.
[0057] Similar to Figure 1BSystem 100B and System 100C include a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.
[0058] In various embodiments, data is streamed to system 100B or 100C using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). Wi-Fi signals in these embodiments are received via a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 100B or 100C, delivering data via an HDMI connection to input block 105. Still other embodiments use an RF connection to input block 105 to provide streaming data to system 100B or 100C. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0059] System 100C can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. The display 165 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a flexible display, and / or a foldable display. The display 165 can be used in a television, tablet computer, laptop computer, cellular phone (mobile phone), or other device. The display 165 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 185 include one or more of a standalone digital video disc (or digital multifunction disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments utilize one or more peripheral devices 185 that provide functionality based on the output of system 100C. For example, a disc player performs the function of playing the output of system 100C.
[0060] In various embodiments, signaling is used to transmit control signals between system 100C and display 165, speaker 175, or other peripheral devices 185. This signaling may be AV.Link, CEC, or other communication protocols enabling device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 100C via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100C via communication interface 150 using communication channel 190. Display 165 and speaker 175 may be integrated into a single unit along with other components of system 100C in an electronic device such as a television. In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0061] For example, if the RF portion of input 105 is part of a separate set-top box, then display 165 and speaker 175 may alternatively be separate from one or more other components. In various embodiments where display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0062] In any of systems 100A, 100B, or 100C, the embodiments may be implemented by a computer program product including code instructions that implement any of the embodiments described herein. The computer program product may be computer software implemented via processor 110 or via hardware or a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 120 of any of systems 100A, 100B, or 100C may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. As a non-limiting example, the processor 110 of any of systems 100A, 100B, or 100C may be of any type suitable for the technical environment and may include one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0063] Figure 2 An example of a block-based hybrid video encoder 200 is shown. Variations of this encoder 200 are envisioned, but for clarity, the encoder 200 is described below without describing all anticipated variations.
[0064] In some embodiments, Figure 2 It also shows the HEVC standard or VVC standard ( Universal Video Coding, Standard ITU-T H.266, ISO / IEC 23090-3, 2020 An improved encoder, or an encoder that uses a technology similar to HEVC or VVC, such as the ECM (Enhanced Compression Model) encoder developed by JVET (Joint Video Exploration Group).
[0065] Before being encoded, the video sequence may undergo pre-coding (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), performing remapping of the input image components to obtain a signal distribution more resilient to compression (e.g., using histogram equalization of the color components), or resizing the image (e.g., reducing its size). Metadata may be associated with pre-processing and attached to the bitstream.
[0066] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is segmented (202) and processed in units such as CUs (coding units) or blocks. In this disclosure, different expressions may be used to refer to such units or blocks resulting from the segmentation of the image. Such terms may be coding unit or CU, coding block or CB, luminance CB or block. CTU (coding tree unit) refers to a group of blocks or a group of units or a group of coding units (CUs). In some embodiments, a CTU may be considered as a block or a unit itself.
[0067] Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in an intra-frame mode, intra-frame prediction (260) is performed. In an inter-frame mode, motion estimation (275) and compensation (270) are performed. Intra-frame and / or inter-frame modes may include several different sub-modes. For example, intra-frame modes may include directional intra-frame prediction, template-based intra-frame mode derived prediction, intra-frame block copy prediction, or other modes of spatial prediction of sample values of the unit. Inter-frame modes may include skip mode, merge mode, inter-frame mode, according to which motion information is derived from a motion candidate list and the motion vector prediction residual is not encoded, according to which motion information is derived from a motion candidate list and the motion information is further refined by encoding the motion vector prediction residual or by template matching performed at both the encoder and decoder, and other inter-frame modes are also possible, such as the GPM mode described further below. The encoder determines (205) which of the intra-frame or inter-frame modes is used to encode the unit. When different intra-frame modes and / or inter-frame modes are possible, the encoder decides (205) which of the intra-frame or inter-frame modes to use. The encoder indicates the intra-frame / inter-frame decision by, for example, signaling one or more syntax elements of the prediction mode. The encoder can also mix (205) intra-frame prediction results and inter-frame prediction results, or mix results from different intra-frame / inter-frame prediction methods. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0068] The motion refinement module (272) uses an already available reference image to refine the motion field of a block without referencing the original block. The motion field of a region can be considered as the set of motion vectors for all pixels in that region. If the motion vectors are based on sub-blocks, the motion field can also be represented as the set of motion vectors for all sub-blocks in the region (all pixels within a sub-block have the same motion vector, and the motion vectors can vary between sub-blocks). If a single motion vector is used for the region, the motion field of that region can also be represented by a single motion vector (the same motion vector for all pixels in that region).
[0069] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy-encoded (245) to output a bitstream. The encoder can skip the transform and apply the quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is directly encoded without applying either the transform or quantization process.
[0070] The encoder decodes (reconstructs) the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed image to perform one or more of, for example, unblocking filtering, SAO (Sample Adaptive Shift) filtering, or ALF (Adaptive Loop Filter) filtering to reduce coding artifacts. The filtered image is stored at the reference image buffer (280). This filtered image is also referred to below as the reference image.
[0071] Figure 3 A block diagram of a video decoder 300 is shown. In decoder 300, the bitstream is decoded by decoder elements, as described below. Video decoder 300 typically performs operations similar to... Figure 2 The decoding passes are the inverse of the encoding passes described in the document. Encoder 200 typically also performs video decoding as part of the encoded video data.
[0072] Specifically, the decoder's input includes a video bitstream that can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. Image segmentation information indicates how to segment the image. Therefore, the decoder can segment (335) the image based on the decoded image segmentation information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct image blocks.
[0073] The (370) prediction block can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). In a similar manner to that in the encoder, intra-frame prediction and / or inter-frame prediction can include several different sub-modes. The decoder obtains the (370) predictor block based on one or more syntax elements that signal the available prediction modes in the intra-frame and inter-frame modes. The decoder can mix the (370) intra-frame prediction results and inter-frame prediction results, or mix results from multiple intra-frame / inter-frame prediction methods. Before motion compensation, the motion field can be refined (372) by using an already available reference image. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored at the reference image buffer (380). Note that for a given image, the contents of the reference image buffer 380 on the decoder 300 side are the same as the contents of the reference image buffer 280 on the encoder 200 side for the same image.
[0074] The decoded image can also undergo post-decoding processing (385), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping of the remapping process performed in pre-encoding processing (201), or resizing the reconstructed image (e.g., enlarging the size). Post-decoding processing can utilize metadata exported in pre-encoding processing and signaled in the bitstream.
[0075] Some embodiments described herein relate to geometric partitioning patterns for encoding or decoding video blocks to be encoded or decoded. The embodiments provided herein are intended to improve compression efficiency when using geometric partitioning patterns.
[0076] Any of the embodiments described herein can be implemented, for example, in the prediction mode module of a video encoder and the prediction mode module of a video decoder. For example, the embodiments described herein can be implemented in... Figure 2 The prediction mode module 205 of the video encoder 200 or Figure 3 It is implemented in the prediction mode module 370 of the video decoder 300.
[0077] Geometric Partitioning (GPM) is also known as GEO mode or geometric partitioning mode. GPM and GEO are used interchangeably throughout the document, and both refer to geometric partitioning mode.
[0078] In VVC, Geometric Partitioning (GPM) is supported for inter-frame prediction. GPM aims to increase partitioning accuracy and better fit moving object boundaries using block-level geometric partitioning. In VVC, CU-level (block-level) flags are used as a merge mode to signal the geometric partitioning mode to other merge modes.
[0079] Figure 4 An example of the prediction process for GPM is shown. In VVC, GPM is designed for blocks with a size w×h=2k×2l (in terms of luminance samples), where k, l∈{3,...,6}. Furthermore, considering that narrow blocks rarely contain geometrically separated patterns, GPM is disabled for blocks with aspect ratios greater than 4:1 or less than 1:4.
[0080] When GPM is applied, the block is divided into two parts, referred to below as partitions, which can be non-rectangular or asymmetric rectangular. The block is divided by straight partition boundaries or edges, and by angle φ. i and offset ρ i Parameterization. The partition boundary is also referred to below as the dividing line, partition edge, or partition line.
[0081] In total, VVC supports 64 partition lines, and these 64 partition lines are indexed by the GPM partition index. Each part of a block will be associated with a merge pattern in a one-way MV encoded in VVC.
[0082] In VVC, the regular merge list is populated with spatial candidates utilizing motion information from spatially neighboring blocks, temporal candidates utilizing motion information from temporal blocks, HMVP (History-Based Motion Vector Predictor) candidates, pairwise average candidates, and zero motion vector candidates. HMVP candidates are those containing previously encoded motion information associated with neighboring and non-neighboring blocks relative to the block. Pairwise average candidates are generated by averaging the motion vectors of the first two available candidates in the merge candidate list.
[0083] The GPM merge list (also known as the geometric one-way prediction candidate list for VVC) is derived directly from the regular merge list using the parity of the index. Let n denote the index of the one-way predicted motion in the geometric one-way prediction candidate list. The LX motion vector of the nth extended merge candidate, where X equals the parity of n, is used as the nth one-way predicted motion vector for the geometric partitioning pattern. If the corresponding LX motion vector of the nth extended merge candidate does not exist, the L(1-X) motion vector of the same candidate is used instead as the one-way predicted motion vector for the geometric partitioning pattern.
[0084] In ECM, this procedure is invoked only for small blocks of 8x8, 16x8, and 8x16. For larger blocks, the extraction procedure is bypassed, so the initial merge list defined for merge modes without GPM (and which may contain Bi-MV candidates for merging) is directly used as the GPM merge list.
[0085] A single GPM merge list is generated, and candidates from this list are selected for each GPM partition. The GPM partition index, which signals the partitioning pattern, and two GPM merge indices, which signal the candidates from the GPM merge list for each partition respectively, are encoded into a bitstream. For each partition, block-based motion compensation prediction (MCP) is performed, resulting in two intermediate prediction blocks P0 and P1. A mixing process is performed using integer mixing matrices W0 and W1 to generate the GPM prediction block P. G This includes weights in the value range [0, 8]. This can be represented as... in Among them (1) Denotes the Hadamard product, and J w、h This represents a matrix of 1s with the current block size. The generated GPM prediction P is subtracted from the original signal. GTo generate residuals, the residuals are transformed and encoded into a bitstream using conventional VVC transform coding and the CABAC engine.
[0086] The weights in the GPM mixing matrix are derived based on the displacement from the sample location to the partition boundary. The position of the partition line is mathematically derived from... Figure 5 The angle and offset parameters for the specific division shown above are derived. Angle φ i Quantization is performed between 0 and 360 degrees, with a step size of 11.25 degrees (32 steps in total). Figure 5 The text describes an angle φ. i and distance ρ i The description of the geometric division. In the following text, a geometric division with a given angle and a given offset is referred to as a division pattern.
[0087] Inter-frame predictions are performed on each partition of the geometric partitioning pattern within the block using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with regular bidirectional prediction, only two motion-compensated predictions are needed per block. The two predictions are then blended for samples in the blending region on each part of the partition line.
[0088] In ECM ( M.Coban, RL. Liao, K.Naser, J.Ström, L.Zhang, "Algorithm Description of Enhanced Compression Model 9 (ECM 9), "Document JVET-AD2025, 30th Meeting" Parliament, Antalya, TR, April 21–28, 2023 In GPM, each partition of the block allows for bidirectional prediction. In ECM, a mixing region corresponding to the slope intensity can be selected from five values, such as... Figure 6 As shown in the image. Figure 6 The ramp function is shown to weight the GPM mixing based on the displacement (d) from the predicted sample location to the GPM partition boundary and the mixing region size (τ).
[0089] By applying motion vector refinement on top of the existing GPM unidirectional MV, the GPM from the VVC is extended in the ECM. Each geometric partition of the GPM block can decide whether to signal the MVD (Motion Vector Difference). If the MVD is signaled for the geometric partition, the motion of the partition is further refined using the signaled MVD information after selecting GPM merging candidates for the partition. This is also known as MMVD (Merged Motion Vector Difference).
[0090] Figure 7The signaling of GPM in the ECM is described (700). A block-level flag (geoFlag) indicates (710) whether GPM is used to encode the block. Here, it is assumed that the flag is set to true. Then, a flag indicating whether the first partition of the block using GPM uses MMVD is signaled (720). If so, an index is signaled (730) to indicate MVD. Otherwise, it is determined (740) whether the first partition is encoded in intra-frame mode. If so, a syntax element is signaled (750) indicating: the partition mode of the partition line, the index of the intra-frame candidate for the first partition, and the index of the inter-frame candidate for the second partition. Otherwise, a flag indicating whether the second partition of the block using GPM uses MMVD is signaled (760). If so, an index is signaled (770) to indicate MVD for the second partition, and then the process proceeds to step 790. Otherwise, it is determined (780) whether the second partition is encoded in intra-frame mode. If this is the case, a signal notification (750) syntax element is used, indicating the partitioning mode of the dividing line, the index of the intra-candidate for the second partition, and the index of the inter-candidate for the first partition. Otherwise, a signal notification (790) flag is used to indicate whether template matching (TM) is applied to both partitions of a block in GPM mode. When template matching is used, TM is used to refine the motion information of each geometric partition. When TM is selected, a template is constructed using left, top, or top-left adjacent samples based on the partition angle. Motion is then refined by minimizing the difference between the current template and the template in the reference image.
[0091] When both partitions are inter-coded, as can be seen in the lower left rectangular section (795), the partition line used (out of 64 possible partition lines) is signaled, and each merge candidate used to predict one of the two partitions is signaled. The first index (mergeCandPart0) is signaled to indicate the merge candidate for the first partition used to predict the block, and the second index (mergeCandPart1) is signaled to indicate the merge candidate for the second partition used to predict the block.
[0092] In template-matching (TM) reordering for GPM partitioning patterns, given motion information of the current GPM block (a pair of inter-frame motion candidates), the corresponding TM cost value for the GPM partitioning pattern is calculated. Then, all GPM partitioning patterns are reordered in ascending order based on the template-based cost values. For example, SAD (Sum of Absolute Differences) is used to determine the template-based cost, but other metrics can be used to determine the cost using templates for geometric partitioning blocks.
[0093] Instead of sending the GPM partition pattern, the index is signaled using Glenn-Less code to indicate where the exact GPM partition pattern is located in the reordered list.
[0094] The reordering method for GPM partitioning patterns is a two-step process performed after generating corresponding reference templates for the two GPM partitions in the block, as shown below. As illustrated in Figure 8, the GPM partition edges are extended into the reference templates of the two GPM partitions. Figure 8 shows the partitioning edges of the block (current CU) extended within the templates of the block (top template and left template). This results in 64 reference templates, and a corresponding template-based cost is calculated for each of these 64 templates. The GPM partitioning patterns are reordered in ascending order based on their template-based cost values, and the optimal 32 partitioning patterns are marked as available partitioning patterns.
[0095] When determining template-based costs, the edges on the template extend from the edges of the block as shown in Figure 8, but the GPM blending process is not used in template regions that cross the edges. Template-based costs are determined by calculating a metric such as SAD between the sample values of the template in the first partition of the block and the sample values of the corresponding candidate templates in the first partition (gray area in Figure 8), and between the sample values of the template in the second partition of the block and the sample values of the corresponding candidate templates in the second partition (gray area in Figure 8).
[0096] After ascending reordering using template-based costs, an index is signaled to indicate which partitioning mode to use for block prediction. The index signaled for the partitioning mode used to encode or decode the block is signaled as the position of the partitioning mode in the reordered list. For example, the partitioning mode used for block prediction is determined by comparing the distortion cost or rate distortion cost determined for the current block and testing each of the partitioning modes marked as available. For example, when prediction is performed using a given combination of candidate pairs and partitioning modes, the distortion cost is determined using a metric calculated between the original sample values of the current block and the sample values of the reconstructed version of the current block. The rate distortion cost also takes into account the cost signaled for the selected combination.
[0097] In ECM, within a GPM with inter-frame and intra-frame prediction, the final prediction sample is generated by weighting the inter-frame and intra-frame prediction samples from the regions separated by each GPM. Inter-frame prediction samples are derived from the inter-frame GPM, while intra-frame prediction samples are derived from the intra-frame prediction mode (IPM) candidate list and from the index provided by the encoder using signals, such as... Figure 7 As described. The IPM candidate list size is predefined as 3. The available IPM candidates are respectively as follows: Figure 8B(ac) shows the parallel angle mode (parallel mode), the perpendicular angle mode (perpendicular mode), and the planar mode relative to the GPM block boundary. Additionally, as... Figure 8B (d) shows that the GPM with intra-frame and intra-frame prediction is limited to reduce the signaling overhead for IPM and avoid increasing the size of the intra-frame prediction circuitry on the hardware decoder. Furthermore, direct motion vectors and IPM storage on the GPM mixing region are introduced to further improve coding performance.
[0098] In DIMD-based and adjacent-mode IPM export, parallel modes are registered first. Therefore, at most two IPM candidates are exported from the decoder-side Intra-Mode Export (DIMD) method, and / or adjacent blocks can be registered if no identical IPM candidate exists in the list. For adjacent-mode export, there are a maximum of five locations for available adjacent blocks, but these are limited by the angle of the GPM block boundary, as shown in Table 1 below, and have been used for GPM with template matching (GPM-TM).
[0099] Table 1: Positions of available neighboring blocks derived from IPM candidates, based on the angles of the GPM block boundaries. A and L represent the top and left sides of the predicted block. Angle of dividing line 0 2 3 4 5 8 11 12 13 14 Section 1 A A A A L+A L+A L+A L+A A A Division 2 L+A L+A L+A L L L L L+A L+A L+A Angle of dividing line 16 18 19 20 21 24 27 28 29 30 Section 1 A A A A L+A L+A L+A L+A A A Division 2 L+A L+A L+A L L L L L+A L+A L+A
[0100] GPM can be combined with other GPMs within a frame, with GPM having a combination with motion vector difference (GPM-MMVD). TIMD is used as an IPM candidate within a GPM frame to further improve coding performance. Parallel modes can be registered first, followed by TIMD, DIMD, and IPM candidates from adjacent blocks.
[0101] As can be seen from the above, in ECM or VVC, the signaling used to predict merge candidates for each partition of a block using GPM encoding is performed independently. Furthermore, the signaling for the partition line is performed independently of the signaling used to predict merge candidates for each partition of a block using GPM encoding.
[0102] Based on the foregoing, it appears that the encoding cost of GPM parameters could be reduced, for example, by using template information.
[0103] At least one embodiment relates to a method for encoding or decoding video, wherein one or more blocks of the video are encoded using GPM, and wherein the same information is signaled to indicate a first candidate for a first partition of the prediction block and a second candidate for a second partition of the prediction block. The same information allows identification of pairs of candidates that include both the first and second candidates. In some variations, pairs of candidates are identified in an ordered list, which is sorted according to a template-based cost determined for each pair of candidates in the list.
[0104] In the variant, the same information is also used to identify the block partitioning pattern in the GPM.
[0105] In this paper, word candidates refer to predictors that can be used to predict blocks. A predictor is a combination of specified values or data elements (e.g., sample values or motion vectors) that may have been previously used to encode or decode previous blocks.
[0106] In one variant, the candidate is a merged candidate, and the signaling provided in the embodiments described herein modifies the signaling of the first and second candidates, such as... Figure 7 The 795th description in the text.
[0107] In another variation, the candidates in the pair of candidates signaled according to the embodiments described herein can be any possible candidates for geometric partitioning. For example, the candidates can be motion candidates (from a merging mode, with or without MVD, or template matching motion refinement), or intra-frame candidates obtained from an intra-prediction mode. In this variation, the signaling provided in the embodiments described herein modifies, for example, signals from... Figure 7 Signals at positions 795 and 750 are used to notify the candidate pair of signaling.
[0108] This assumes that candidates are provided from a candidate list that is populated in a similar manner at the encoder and decoder. For example, the candidate list can be populated as in the VVC or ECM, or according to any variation of the process performed in the VVC or ECM. The population of the candidate list as described above is provided for descriptive purposes, but the contents of the candidate list are not limited to the only examples of candidates described herein.
[0109] Figure 9 An embodiment of a method 900 for encoding blocks of video is shown. At 910, one or more lists are obtained, including costs for at least one template-based pair determined for candidates associated with a partitioning pattern. As described above, the partitioning pattern is determined by using... Figure 5 The described angles and offsets define a portion of a set of given partitioning patterns. A partitioning pattern defines a partitioning edge that divides at least one block into two geometric partitions, and each candidate in the candidate pair predicts one of the two geometric partitions of at least one block. At 920, for example, an RDO procedure is used to select a combination of candidate pairs and their associated partitioning patterns for predicting at least one block. The selected combination is one that must be signaled to the decoder for correct block decoding.
[0110] At position 950, the block is encoded using a combination of candidate pairs and the selection of a partitioning pattern. That is, each partition of the block is predicted from one of the candidates in that pair. Prediction residuals are determined for each block and encoded.
[0111] For the selected combination, at 930, at least one of the lists is reordered in a given order, for example, in ascending order of the considered template-based costs. At 940, an indication is given to identify the selected combination. Depending on the variant used, one or more syntax elements may be used. Some variants for reordering at least one of the lists at 930 are further described below.
[0112] Figure 10 An embodiment of a method 1000 for decoding blocks of video is shown. For example, methods related to... Figure 10 The described method 900 encodes the block. At 1010, one or more lists are obtained, including at least one template-based cost determined for candidate pairs associated with the partitioning pattern. At 1020, at least one of the one or more lists is reordered in a given order, e.g., in ascending order of the considered template-based costs, in a manner similar to that performed on the encoder side, and at 1030, an indication used to identify combinations of candidate pairs and partitioning patterns in the reordered lists is decoded. Depending on the variant used, one or more syntax elements may be used to decode the indication. A variant for reordering at least one of the one or more lists at 1020 is further described below. At 1040, the block is decoded, i.e., reconstructed, using a combination of the decoded candidate pairs and partitioning patterns. That is, the block is partitioned according to the partitioning pattern, and each partition of the block is predicted according to one of the identified candidates of the pair. The prediction residuals for the block are decoded and added to the prediction to reconstruct the block.
[0113] In a first embodiment, an index is used to signal the pairs of candidates in the reordered candidate pair. In this embodiment, the partitioning pattern is signaled by an index in the VVC or ECM. The changes made from the ECM to the candidates in the GPM are signaled and reordered by the candidate pairs associated with the signaled partitioning pattern based on their template-based cost pairs. In this embodiment, instead of signaling each candidate individually, an index is signaled that corresponds to the rank of the candidate pairs in the reordered list based on their template-based cost.
[0114] In the second embodiment, a partition index corresponding to the ranking of the partitioning patterns of a specific pair of candidates is signaled. In this embodiment, the template cost of each combination of candidate pairs and partitioning patterns is calculated. The partition signaled is the ranking of the partition of the selected / chosen pair of candidates. All candidate pairs are reordered based on the template cost of the ranking signaled; for example, if the ranking signaled is 5, then all candidate pairs are reordered based on the template cost of the 5th best partitioning pattern of all candidate pairs, and the index signaled which candidate pair should be used in this reordered list.
[0115] In the third embodiment, a global index is signaled, which identifies the selected partitioning pattern and the candidate selected pairs. In this embodiment, only one index is needed to signal both the candidate pairs and the partitioning pattern.
[0116] Figure 11 An example of a method 1100 for encoding blocks of video according to a first embodiment is shown. At 1110, a list is obtained for each partitioning pattern. Each list includes template-based costs determined for candidate pairs associated with the same partitioning pattern (partitioning patterns in the list). At 1120, when prediction is performed using a combination of candidate pairs and partitioning patterns, for example, RDO performed based on sample values of the blocks and reconstructed versions of the blocks, the optimal combination of candidate pairs and partitioning patterns is selected. It should be understood that the optimal combination is selected from all determined combinations of candidate pairs and partitioning patterns determined at 1110. At 1130, the partitioning pattern corresponding to the selected combination is signaled, which is encoded in the bitstream. At 1140, the list of template-based costs obtained for the signaled partitioning patterns is reordered, for example, in ascending order. At 1150, the position of the candidate pairs corresponding to the candidate pairs of the selected combination in the reordered list is signaled. At 1160, as per... Figure 9 The block is encoded as described in step 950.
[0117] Figure 12 An example of a method 1200 for decoding blocks of video according to a first embodiment is shown. At 1210, a partitioning pattern is decoded from the bitstream. At 1220, a list of template-based costs is obtained, wherein a template-based cost is determined for each available candidate pair and the decoded partitioning pattern. At 1230, the list of template-based costs obtained for the decoded partitioning patterns is reordered, for example, in ascending order. At 1240, syntax elements indicating positions in the reordered list are decoded. The syntax elements indicate the positions of candidate pairs from the reordered list of their predicted blocks at 1250. At 1250, as in the section on... Figure 10Rebuild the block as described in step 1040.
[0118] In this embodiment, the template-based cost of each pair of merge candidates is evaluated. The partitioning to be used for the block is signaled at the decoder using indices (1130, 1210). In a variant, the partitioning patterns may not be reordered. They are in a predetermined order that is the same for all blocks. In another variant, the partitioning patterns are reordered based on parameters of pairs that are not candidates, since the same reordering should be done on both the encoder and decoder sides. For example, the partitioning patterns may be ordered differently depending on the block size and shape of the block to be partitioned (e.g., the diagonal and anti-diagonal partitions of each partitioning pattern may be encoded on fewer bits).
[0119] In another example, the partitioning patterns can be reordered based on template information that does not use candidate pairs.
[0120] For example, partitioning patterns can be reordered based on the cost obtained from samples that predict blocks using templates from directional intra-prediction patterns. Figure 8C An example of a directional intra-prediction mode is shown. The intra-prediction angle of a directional intra-prediction mode can be rounded to the nearest angle parallel to the partition line of the partition mode using the intra-prediction mode parallel to the partition line (FIG.8B(a)). In the example, the partition modes are reordered based on the cost of the corresponding directional intra-prediction mode. In another example, only partition modes with angles corresponding to the angles of the directional intra-prediction mode that provides the best cost are considered. And the partition mode is selected from the set of partition modes with that angle.
[0121] In another example, DIMD (decoder-side intra-frame mode derivation) can be used to derive intra-frame prediction angles. Decoder-side intra-frame mode derivation (DIMD) is a prediction mode used to encode intra-frame blocks in the ECM. When DIMD is applied to a block, up to five intra-frame modes are derived from the reconstructed neighbor samples, and these five predictors are combined with a planar mode predictor with weights derived from the gradient histogram.
[0122] The DIMD procedure can be used to derive the dividing line angle based on the intra-prediction angle of the intra-frame mode derived from DIMD. Similar to the example above, the logic of GPM with inter-frame and intra-frame prediction can be used to round the intra-prediction angle to the nearest angle parallel to the dividing line of the dividing pattern, which uses the intra-frame mode parallel to the dividing line (FIG.8B(a)). In the example, only dividing patterns with angles derived using DIMD are considered.
[0123] At 1140 and 1230, then, when used with a partitioning mode notified by signaling, candidate pairs are sorted based on their template costs, and the indices of pairs in the ordered list are encoded. In this way, candidate pairs with the lowest template costs (i.e., providing better predictions of the template) can be encoded in fewer bits compared to candidates with higher template costs. For example, in some variations, the code used can be unary. In some variations, the first candidate pair (with the lowest cost) can be signaled using a CABAC-encoded flag (e.g., a merge flag to use the same context as the original GPM process), and the remaining candidate pairs use unary codes. In this variation, the flag indicates whether the candidate pair at the first position in the reordered list is the candidate's chosen pair (1120). If not, the index is encoded to signal the position of the candidate's chosen pair between the second position in the reordered list and the end of the list.
[0124] In some variations, the reordered list is truncated to a maximum number of candidate pairs to reduce the list size and the average cost of transferring the selected candidate pairs. For example, a maximum of 10 pairs could be used. In some variations, the list is truncated based on another parameter. For example, in some variations, only pairs with a template cost less than K times the cost of the lowest template cost are selected (e.g., K=1.25). In some variations, the list is limited based on max_num_merge_cand, which is the maximum number of merge candidates. For example, only the first 2*max_num_merge_cand pairs are allowed.
[0125] Figure 13An example of method 1300 for encoding blocks of video according to a second embodiment is shown. At 1310, a list is obtained for each pair of candidates. Each list includes a template-based cost determined for the partitioning patterns associated with the same candidate pair, and each list is sorted according to a given order, e.g., ascending order. At 1320, when prediction is performed using a combination of candidate pairs and partitioning patterns, e.g., RDO performed based on sample values of the original sample values of the block and the reconstructed version of the block, the best combination of candidate pairs and partitioning patterns is selected. It should be understood that the best combination is selected from all determined combinations of candidate pairs and partitioning patterns determined at 1310. At 1330, a ranking is signaled or encoded in the bitstream, the ranking being a ranking of the selected partitioning patterns in the list associated with the selected candidate pairs. At 1340, the template-based costs for each list obtained at 1310 for the ranking signaled are sorted in the list, e.g., ascending order. In other words, the template-based cost at each ranking position, signaled by a signal, is placed into a new list and sorted in ascending order, thus resulting in an ordered list of template-based costs for candidates and splitting patterns for different pairs. At 1350, the selected pair of candidates is signaled to its position in the new ordered list. At 1360, as per... Figure 9 The block is encoded as described in step 950.
[0126] Figure 14 An example of a method for decoding blocks of video according to a second embodiment is shown. At 1410, a ranking of partitioning patterns is decoded from the bitstream. At 1420, a list of available pairs for each candidate is obtained in a manner similar to that in the encoder at 1310. Each list includes a template-based cost determined for the partitioning pattern associated with the same candidate pair, and each list is sorted according to a given order, e.g., ascending order. At 1430, as in step 1340 of the encoding method, the template-based costs of each ordered list obtained at 1420, ranked by the decoded ranking, are sorted, for example, in a second list in ascending order. At 1440, syntax elements indicating positions in the ordered second list are decoded. The syntax elements indicate the positions of candidate pairs from the ordered second list of their predicted blocks at 1450. At 1450, as in the section on... Figure 10 Rebuild the block as described in step 1040.
[0127] In this second embodiment, all combinations of candidate pairs and partitioning patterns have their calculated template costs. The partitioning pattern used to encode / decode the block is signaled as its rank R in an ordered list of partitioning patterns for the selected candidate pair. The candidate pairs are then reordered based on the template-based cost of the Rth best partitioning pattern of the candidate pair. An index is signaled to indicate the rank of the selected candidate pair in this reordered list.
[0128] The following example illustrates the encoding process. In this example, Figure 15 The diagram shows the template-based costs of the discovered (1310, 1420) partitioning pattern (S0-S5) using one of six partitioning patterns S0 to S5 along with four pairs of P0 to P3. To signal the use of partitioning pattern S0 with candidate pair P3 (e.g., the combination selected at 1320), partitioning pattern S0 is the second best partitioning pattern when used with candidate pair P3 (cost 34 for the same candidate pair P3, second only to partitioning pattern S5 with cost 19). The first index (1330, 1410) with a value of 1 is signaled. Index 0, which has been signaled for the best partitioning pattern, is then signaled as the first best partitioning pattern. The template-based costs of all pairs (P0-P3) based on their second best partitioning patterns are reordered (1340, 1430). See also... Figure 16 The table above shows the template-based cost of the second optimal partitioning pattern for each pair, along with the corresponding partitioning pattern indicated in parentheses. The selected candidate pair P3 is located at the first position in this ordered list. The second index (1350, 1440) with a value of 0 is signaled to indicate the use of candidate pair P3.
[0129] In some variations, the number of available partitioning patterns can be reduced from the total number of available partitioning patterns. This allows for a reduction in the complexity of the method. For example, at 1310 and 1420, the number of partitioning patterns for obtaining template-based costs is reduced compared to the number of available partitioning patterns. This allows for avoiding template-based costs in determining candidate pairs and all possible combinations of partitioning patterns. In some variations, a fixed number of N partitioning patterns is allowed (e.g., N=48).
[0130] In other variations, the number of allowed partitioning patterns is reduced, but this can be varied, for example, by reducing the number of partitioning patterns based on the template-based cost of the partitioning pattern of the candidate pair with the lowest template-based cost. For example, let C be the lowest template-based cost found for a combination of partitioning patterns and candidate pairs. For this candidate pair, let N be the number of partitioning patterns with template-based costs strictly below K*C, for example, K=1.25. In this example, in some variations, only N partitioning patterns are used. Other values of K are also possible. This variation allows for a reduction in signaling costs for the first rank (1330, 1410) and the second rank (1350, 1440). At 1320, the combination of partitioning patterns and candidate pairs for predicting the block is thus selected from the reduced number of template-based costs with template-based costs strictly below K*C.
[0131] Figure 17 An example of method 1700 for encoding blocks of video according to a third embodiment is shown. At 1710, a list of template-based costs is obtained, which includes template-based costs determined for candidate pairs and partitioning patterns. In this variation, a single list including all available template-based costs is obtained. This list is sorted in a given order, for example, in ascending order of template-based costs. At 1720, when prediction is performed using combinations of candidate pairs and partitioning patterns, for example, RDO performed based on sample values of the blocks and reconstructed versions of the blocks, the best combination of candidate pairs and partitioning patterns is selected. It should be understood that the best combination is selected from all determined combinations of candidate pairs and partitioning patterns determined at 1710. At 1730, the position of the selected combination of candidate pairs and partitioning patterns in the ordered list is signaled. At 1740, as per... Figure 9 The block is encoded as described in step 950.
[0132] Figure 18 An example of method 1800 for decoding a block of video according to a third embodiment is shown. At 1810, a list of template-based costs is obtained in a manner similar to that in the encoder (at 1710), which includes template-based costs determined for candidate pairs and partitioning patterns. In this variation, a single list is obtained, which includes template-based costs for all combinations of available partitioning patterns and candidate pairs. This list is sorted in a given order, for example, in ascending order of template-based costs. At 1820, syntax elements indicating positions in the ordered list are decoded. Syntax elements indicate positions at 1830 that will be used to predict combinations of candidate pairs and partitioning patterns for the block. At 1830, as in the section on... Figure 10 Rebuild the block as described in step 1040.
[0133] In this third embodiment, an index is used to signal both the partitioning pattern and the candidate pairs. In this embodiment, a template-based cost is calculated for all combinations of candidate pairs and partitioning patterns. All combinations are then sorted (1710, 1810) based on their template-based costs, and the index is transferred (1730, 1820) to indicate the rank of the selected combination in the ordered list of combinations of partitioning patterns and candidate pairs. Combinations with lower template-based costs (and lower rank in the ordered list) can then be encoded on fewer bins compared to combinations with higher template-based costs.
[0134] In some variations, the list obtained at 1710 and 1810 is truncated to limit the total number of candidates (e.g., to a fixed number of candidates, or to a number of candidates that depends on parameters such as block size or shape). In some variations, the list is truncated based on the template-based cost of combinations of partitioning patterns and candidate pairs. For example, only combinations with a template-based cost less than or equal to K times the cost of the best-found combination (with the lowest template-based cost) are retained, such as K=1.25 (other values of K can be used).
[0135] In some variations, the list of candidate pairs and partitioning patterns is finite, but only after the list contains at least N distinct partitioning patterns (e.g., N=8). In other variations, the list is finite, but only after it contains at least M distinct candidate pairs (e.g., M=8, or M=max_num_merge_cand). In still other variations, the list is truncated so that not too many candidates have the same partitioning pattern. For example, only the first four combinations of partitioning patterns and candidate pairs with a given partitioning pattern are retained, and fifth and subsequent combinations with the given partitioning pattern are not allowed.
[0136] exist Figure 19 In the illustrated embodiment, within the transmission context of a communication network NET between two remote devices A and B, device A includes a processor associated with RAM and ROM, configured to implement a method for encoding video according to any of the embodiments described herein, and device B includes a processor associated with RAM and ROM, configured to implement a method for decoding video according to any of the embodiments described herein. According to the example, the network is a broadcast network adapted to broadcast / transmit the encoded video from device A to a decoding device including device B.
[0137] Figure 20An example of the syntax for a signal transmitted via a packet-based transport protocol is shown. Each transmitted packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include video data encoded according to any of the embodiments described above. The payload may also include any signaling as described above. For example, the signal includes encoded data representing at least one block of video, said at least one block being encoded using a geometric partitioning pattern, and said encoded data including at least one first syntax element that identifies at least one pair of given candidate pairs in an ordered list, said ordered list including at least one template-based cost determined for candidate pairs associated with partitioning patterns from a set of partitioning patterns. In some variations, the encoded data also includes a second syntax element that indicates the ranking order of a given partitioning pattern in a set of partitioning patterns or in an ordered list of partitioning patterns for a given candidate pair.
[0138] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes, or alternatively includes, processes performed by decoders of the various implementations described herein, such as entropy decoding of a sequence of binary symbols to reconstruct image or video data.
[0139] As further examples, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding; and in yet another embodiment, "decoding" refers to the entire image reconstruction process including entropy decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and will be considered well understood by those skilled in the art.
[0140] Various implementations involve encoding. In a manner similar to the above discussion of “decoding,” the term “encoding” as used herein can encompass all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various examples, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, such processes also include, or alternatively include, processes performed by the encoders of the various implementations described herein, such as determining resampling filter coefficients, resampling the decoded picture.
[0141] As another example, in one embodiment, "encoding" refers only to entropy encoding; in another embodiment, "encoding" refers only to differential encoding; and in yet another embodiment, "encoding" refers to a combination of differential and entropy encoding. It will be clear, and is considered well understood by those skilled in the art, whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, based on the context of the specific description.
[0142] Note that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.
[0143] This disclosure describes various information fragments that can be transmitted or stored, such as syntax. This information can be encapsulated or arranged in a variety of ways, including those common in video standards, such as placing the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers, picture headers, or slice headers), or SEI messages. Other methods are also available, including those common in system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (Session Description Protocol), a format for describing multimedia communication sessions for session announcement and session invitation purposes, for example, as described in the RFC and used in conjunction with RTP (Real-Time Transport Protocol) transmission.
[0144] b. For example, the DASH MPD (Media Presentation Description) descriptor used in DASH and transmitted over HTTP, which is associated with a representation or set of representations to provide additional features to the content representation.
[0145] c. RTP header extensions, for example, used during RTP streaming.
[0146] d. ISO Basic Media File Format, such as that used in OMAF, and uses boxes as object-oriented building blocks defined by unique type identifiers and lengths, also referred to as "atoms" in some specifications.
[0147] e. An HLS (HTTP Live Stream) manifest transmitted via HTTP. The manifest can be associated with, for example, a version or set of versions of the content to provide characteristics of the version or set of versions.
[0148] When the accompanying drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding apparatus. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.
[0149] Some embodiments involve rate distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account computational complexity constraints. Rate distortion optimization is generally formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate distortion optimization problem. For example, these approaches can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, with a complete evaluation of the encoding cost and associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save encoding complexity, particularly by utilizing computations based on approximate distortion based on predicting or predicting the residual signal rather than the reconstructed signal. A hybrid of these approaches can also be used, such as by using approximate distortion only for some possible encoding options and full distortion for others. Other methods evaluate only a subset of possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding cost and associated distortion.
[0150] The implementations and aspects described herein can be implemented, for example, in a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method), the implementation of the features in question can be implemented in other forms (e.g., apparatus or program). An apparatus can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0151] References to “an embodiment” or “an embodiment” or “an implementation” or “an implementation” and their variations mean that a particular feature, structure, characteristic, and such is included in at least one embodiment. Therefore, the phrases “in an embodiment” or “in an embodiment” or “in an implementation” or “in an implementation” appearing throughout this application and any other variations do not necessarily refer to the same embodiment.
[0152] Additionally, this application may involve "determining" various pieces of information. Determining information may include, for example, one or more of the following: estimated information, calculated information, predicted information, or information retrieved from memory.
[0153] Furthermore, this application may involve "accessing" various information fragments. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.
[0154] Additionally, this application may involve "receiving" various pieces of information. Like "access," receiving is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) from one or more sources. Furthermore, "receiving" is generally involved in one or more ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0155] It should be understood that, for example, the use of any of the above " / ", "and / or", and "...at least one of A and B" in the cases of "A / B", "A and / or B", and "at least one of A and B" is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the cases of "A, B and / or C" and "at least one of A, B, and C", this wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as listed.
[0156] Furthermore, as used herein, the term "signal" refers, among other things, to instructing the corresponding decoder to something. In this way, in one embodiment, the same parameter is used on both the encoder and decoder sides. Thus, for example, the encoder can transmit (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has said specific parameter as well as other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select said specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, and the like are used to signal information to the corresponding decoder. While the foregoing refers to the verb form of the term "signal," the term "signal" can also be used as a noun herein.
[0157] It will be apparent to a person skilled in the art that implementations can generate various signals that are formatted to carry information, for example, that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0158] Several embodiments have been described above. The features of these embodiments may be provided individually or in any combination across various claim classes and types.
Claims
1. A method comprising encoding at least one block of a video using a geometric partitioning pattern: Obtain at least one list, which includes at least one template-based cost determined for a pair of candidates associated with a partitioning pattern from a set of partitioning patterns, wherein the partitioning pattern defines a partitioning edge that divides at least one block into two geometric partitions, and wherein each candidate in the pair of candidates predicts one of the two geometric partitions of at least one block. Reorder at least one template-based cost of at least one of at least one of at least one of the at least one lists in a given order. Decode at least one first syntactic element that identifies at least one candidate pair in at least one reordered list. Decode at least one block using the identified candidate pairs and the given partitioning pattern.
2. An apparatus comprising one or more processors operable to target at least one block of video, the at least one block being encoded using a geometric partitioning pattern: Obtain at least one list, which includes at least one template-based cost determined for a pair of candidates associated with a partitioning pattern from a set of partitioning patterns, wherein the partitioning pattern defines a partitioning edge that divides at least one block into two geometric partitions, and wherein each candidate in the pair of candidates predicts one of the two geometric partitions of at least one block. Reorder at least one template-based cost of at least one of at least one of at least one of the at least one lists in a given order. Decode at least one first syntactic element that identifies at least one candidate pair in at least one reordered list. Use the identified candidate pairs and the given partitioning pattern to decode at least one block.
3. A method comprising: For at least one block of the video, the at least one block is intended to be encoded using a geometric partitioning pattern: Obtain at least one list, which includes at least one template-based cost determined for a pair of candidates associated with a partitioning pattern from a set of partitioning patterns, wherein the partitioning pattern defines a partitioning edge that divides at least one block into two geometric partitions, and wherein each candidate in the pair of candidates predicts one of the two geometric partitions of at least one block. Reorder at least one template-based cost of at least one of at least one of at least one of the at least one lists in a given order. The encoding identifies at least one first syntactic element of a candidate pair in at least one reordered list. Encode at least one block using the identified candidate pairs and the given partitioning pattern.
4. An apparatus comprising one or more processors operable to encode at least one block of video, the at least one block being designed to use a geometric partitioning pattern: Obtain at least one list, which includes at least one template-based cost determined for a pair of candidates associated with a partitioning pattern from a set of partitioning patterns, wherein the partitioning pattern defines a partitioning edge that divides at least one block into two geometric partitions, and wherein each candidate in the pair of candidates predicts one of the two geometric partitions of at least one block. Reorder at least one template-based cost of at least one of at least one of at least one of the at least one lists in a given order. The encoding identifies at least one first syntactic element of a candidate pair in at least one reordered list. Encode at least one block using the identified candidate pairs and the given partitioning pattern.
5. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein the given partitioning pattern is a partitioning pattern associated with a pair of candidates identified in the reordered list.
6. The method of any one of claims 1, 3 or 5, or the apparatus of any one of claims 2 or 4-5, wherein the at least one list obtained is a single list, wherein each candidate in a pair of candidates in the single list is associated with a partitioning pattern from a set of partitioning patterns.
7. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein each of the at least one list is associated with a partitioning pattern, the method further comprising, or one or more processors being configured to decode or encode a second syntax element of a given partitioning pattern in a set indicating partitioning patterns, and wherein only the list associated with the given partitioning pattern is reordered.
8. The method of any one of claims 1, 3, or 5, or the apparatus of any one of claims 2 or 4-5, wherein each of the at least one list includes template costs associated with the same candidate pair, the same candidate pair associated with each list is distinct, and each of the at least one list is an ordered list, wherein the template costs of the lists are ordered in a given order, the method further comprising, or one or more processors being further configured to decode or encode a third syntax element indicating the rank order of a given partitioning pattern in an ordered list of partitioning patterns for a given candidate pair, and the reordering of at least one list being a reordering of the template costs associated with the candidate pair, the candidate pair being associated with a partitioning pattern in the indicated rank order in each ordered list of the ordered lists associated with the candidate pair.
9. The method of any one of claims 1, 3, or 5-8, or the apparatus of any one of claims 2 or 4-8, wherein at least one first syntax element includes a flag indicating whether a candidate pair to be identified in at least one reordered list is in a first position in at least one reordered list.
10. The method or apparatus of claim 9, wherein at least one first syntax element includes an index indicating the position of a candidate pair to be identified in at least one reordered list after the first position.
11. The method according to any one of claims 1, 3 or 5-10, or the apparatus according to any one of claims 2 or 4-10, wherein at least one reordered list is truncated to a given number of template costs.
12. The method of any one of claims 1, 3, or 5-11, or the apparatus of any one of claims 2 or 4-11, wherein at least one reordered list includes only template costs that are less than a given percentage of the minimum template cost.
13. The method or apparatus of claim 9, wherein the number of partitioning patterns allowed for candidate pairs is set to a given number.
14. The method or apparatus of claim 9, wherein each of the at least one list is associated with a pair of candidates, and each list includes only a template cost that is less than a given percentage of the lowest template cost associated with the pair of candidates associated with the list.
15. The method or apparatus of claim 11, wherein the at least one reordered list is truncated only if the at least one reordered list includes a given number of template costs for different partitioning patterns.
16. The method or apparatus of claim 11, wherein the at least one reordered list is truncated only if the reordered list includes template costs for a given number of different candidate pairs.
17. The method of any one of claims 1, 3, or 5-10, or the apparatus of any one of claims 2, 4-10, wherein at least one reordered list includes fewer than a given number of template costs associated with the same partitioning pattern.
18. A non-transitory computer-readable medium storing encoded data representing at least one block of a video, the at least one block being encoded using a geometric partitioning pattern comprising a given partitioning pattern from a set of partitioning patterns, the given partitioning pattern defining partitioning edges that divide the at least one block into two geometric partitions, and wherein each of the two geometric partitions is predicted by a candidate from a given pair of candidates, the encoded data comprising at least one first syntax element that identifies at least one pair of given candidates in an ordered list, the ordered list comprising at least one template-based cost determined for the pair of candidates associated with the partitioning pattern from the set of partitioning patterns.
19. The non-transitory computer-readable medium of claim 18, wherein the encoded data further comprises a second syntax element indicating a given partition pattern in a set of partition patterns.
20. The non-transitory computer-readable medium of claim 18, wherein the given partitioning pattern is a partitioning pattern associated with a pair of candidates identified in a reordered list.
21. The non-transitory computer-readable medium of claim 20, wherein the encoded data further includes a second syntactic element indicating the rank order of a given partitioning pattern in an ordered list of partitioning patterns for a given candidate pair.
22. The non-transitory computer-readable medium according to any one of claims 18-21, wherein at least one first syntax element includes a flag indicating whether a given pair of candidates to be identified in at least one reordered list is at a first position in at least one reordered list.
23. The non-transitory computer-readable medium of claim 22, wherein at least one first syntax element includes an index indicating the position of a given pair of candidates to be identified in at least one reordered list after the first position.
24. A computer program product comprising instructions for causing one or more processors to perform the method according to any one of claims 1, 3, or 5-17.
25. A non-transitory computer-readable medium storing executable program instructions to cause a computer executing the program instructions to perform the method according to any one of claims 1, 3, or 5-17.
26. An apparatus comprising: The apparatus according to claim 2; as well as At least one of the following: (i) an antenna configured to receive or transmit a signal including data representing video; (ii) a band limiter configured to limit the signal to a frequency band including data representing video; or (iii) a display configured to display video.
27. The device of claim 28, wherein the device comprises at least one of a television, a cellular phone, a tablet computer, and a set-top box.