Deriving coding modes from one or more template-based costs

By using a template-based cost evaluation method to optimize the selection of video encoding modes and the merging of candidate lists, the problem of insufficient video encoding compression efficiency in existing technologies is solved, and more efficient video data compression is achieved.

CN121925841APending Publication Date: 2026-04-24INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-09-12
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing video coding schemes suffer from insufficient compression efficiency when utilizing spatial and temporal redundancy in video content, especially lacking effective cost assessment methods when determining coding modes.

Method used

A template-based cost evaluation method is adopted to determine the coding mode by calculating the difference between the template and the prediction block. The selection of the coding mode is optimized by using the template shape and parameters, and the candidate list is optimized by combining template matching and reordering algorithms.

Benefits of technology

It improves the compression efficiency of video encoding by more accurate encoding mode selection and candidate merging optimization, thereby enhancing the compression performance of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925841A_ABST
    Figure CN121925841A_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding or decoding a video are provided in which a block is encoded using one or more encoding modes using template-based cost. In one embodiment, the determination of the cost based on the template is adapted according to the encoding mode. In one embodiment, one or more template-based costs are determined using a first cost obtained for a first component (e.g., luminance) of a block and at least one additional cost determined for the block. In one variant, the additional cost is a template-based cost determined for another component of the block, such as a chroma component. In another variation, the additional cost is the cost for signaling one or more syntax elements of the coding mode. In another embodiment, one or more template-based costs are determined using a template shape adapted according to a given criterion. Once the parameters of the encoding mode have been determined, a prediction of the block is obtained based on the encoding mode, and the block is encoded or decoded based on the prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to European Application No. 23306631.5, filed on 28 September 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This document generally relates to video compression. Specifically, it relates to a method and apparatus for encoding or decoding images or videos. More specifically, it relates to an improved encoding scheme for a template-based cost video compression system. Background Technology

[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to utilize spatial and temporal redundancy in video content. Intra-frame or inter-frame prediction is usually used to leverage intra-frame or inter-frame image correlations, and then the differences between the original block and the predicted block (typically represented as prediction error or prediction residual) are transformed, quantized, and entropy-coded. In inter-frame prediction, motion vectors used for motion compensation are typically predicted from motion vector prediction values. To reconstruct the video, the compressed data is decoded through the inverse process corresponding to entropy coding, quantization, transform, and prediction. Summary of the Invention

[0004] According to one aspect, a method for encoding or decoding blocks of video is provided. The method includes: obtaining a prediction of the block based on an encoding mode, wherein the encoding mode uses at least one template-based cost, the template-based cost being determined using at least one first cost obtained for a first component of the block; and encoding or decoding the block based on the prediction, wherein the at least one template-based cost is determined using at least one additional cost determined for the block.

[0005] According to another aspect, an apparatus is provided for encoding or decoding blocks of video. The apparatus includes one or more processors operable to: obtain a prediction of the block based on an encoding mode, wherein the encoding mode uses at least one template-based cost, the template-based cost being determined using at least one first cost obtained for a first component of the block; and encode or decode the block based on the prediction, wherein the at least one template-based cost is determined using at least one additional cost determined for the block.

[0006] In some embodiments, the additional cost is a template-based cost obtained for another component of the block. In other embodiments, the additional cost is the cost of signaling one or more syntax elements associated with the encoding pattern.

[0007] According to another aspect, a method for encoding or decoding blocks of video is provided. The method includes: obtaining a prediction of the block based on an encoding mode, wherein the encoding mode uses at least one template-based cost; and encoding or decoding the block based on the prediction, wherein the at least one template-based cost is determined using a template shape, the template shape being determined based on at least one parameter of the encoding mode.

[0008] According to another aspect, an apparatus is provided for encoding or decoding blocks of video. The apparatus includes one or more processors operable to: obtain a prediction of the blocks based on an encoding mode, wherein the encoding mode uses at least one template-based cost; and encode or decode the blocks based on the prediction, wherein the at least one template-based cost is determined using a template shape derived from at least one parameter of the encoding mode.

[0009] This document describes other embodiments that can be used alone or in combination.

[0010] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods for encoding or decoding video according to any of the embodiments described herein. One or more embodiments of the present invention also provide a non-transitory computer-readable medium and / or a computer-readable storage medium storing instructions thereon for encoding or decoding video according to the methods described herein.

[0011] One or more embodiments also provide a computer-readable storage medium storing a bit stream generated according to the method described herein. One or more embodiments also provide methods and apparatus for transmitting or receiving a bit stream generated according to the described method. Attached Figure Description

[0012] Figure 1 A block diagram of a system in which various aspects of the embodiments described herein can be implemented is shown.

[0013] Figure 2 A block diagram is shown of an embodiment of a video encoder in which various aspects of the embodiments described herein may be implemented.

[0014] Figure 3 A block diagram is shown of an embodiment of a video decoder in which aspects of the embodiments described herein may be implemented.

[0015] Figure 4 An example template is shown for determining the template cost used to reorder merge candidates in a list.

[0016] Figure 5An example is shown of a template of a block with sub-block motion information using the motion information of the current block, and a reference sample of the template.

[0017] Figure 6 An example template for a block used for geometric partitioning is shown.

[0018] Figure 7 An example of a prediction direction used for intra-frame prediction is shown.

[0019] Figure 8 An example flowchart for encoding or decoding at least one block of video is shown according to one embodiment.

[0020] Figure 9 An example flowchart for encoding or decoding at least one block of a video according to another embodiment is shown.

[0021] Figure 10 An example flowchart for encoding or decoding at least one block of a video according to another embodiment is shown.

[0022] Figure 11 An example flowchart for encoding or decoding at least one block of a video according to another embodiment is shown.

[0023] Figure 12 A block diagram of a system according to another embodiment, in which aspects of the embodiments herein may be implemented, is shown.

[0024] Figure 13 This paper illustrates an example of two remote devices communicating via a communication network, based on the principles outlined herein.

[0025] Figure 14 The syntax of a signal, based on the principles outlined in this paper, is shown.

[0026] Figure 15 An example is shown for deriving spatially adjacent blocks as candidates for spatial merging. Detailed Implementation

[0027] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in detail and, at least to illustrate individual characteristics, are generally described in a manner that may sound restrictive. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with those described in previous documents.

[0028] The aspects described and envisioned in this application can be implemented in many different forms. The following... Figure 1 , Figure 2 and Figure 3 Some embodiments have been provided, but other embodiments are contemplated, and... Figure 1 , Figure 2 and Figure 3 The discussion does not limit the breadth of implementations. At least one aspect relates generally to video encoding and decoding, and at least one other aspect relates generally to transmitting a generated or encoded bitstream. These and other aspects may be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having a bitstream generated according to any of the described methods stored thereon.

[0029] In this application, the terms “reconstruction” and “decoding” are used interchangeably, as are the terms “pixel” and “sample”, and the terms “image”, “picture” and “frame” are used interchangeably.

[0030] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first" and "second" may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of these terms does not imply an ordering of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and may occur, for example, before, during, or in a time period overlapping with the second decoding.

[0031] The aspects described herein are not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations, whether pre-existing or developed in the future, as well as any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically excluded, the aspects described in this application may be used individually or in combination.

[0032] Figure 1A block diagram illustrating an example of a system in which various aspects and embodiments can be implemented is shown. System 100 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100 may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in this application.

[0033] System 100 includes at least one processor 110 configured to execute instructions loaded therein, in accordance with various aspects described in this application. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0034] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents multiple modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, the encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated into processor 110 as a combination of hardware and software, as is known to those skilled in the art.

[0035] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 406 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0036] In some embodiments, memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2, HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Multi-Functional Video Coding, also known as H.266, a standard developed by JVET (Joint Video Experts Group)).

[0037] Inputs to the components of system 100 can be provided via various input devices indicated in block 105. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) a component input terminal (or a set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1 Other examples shown include composite videos.

[0038] In various embodiments, the input device of block 105 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal bandwidth to a frequency band), (ii) down-converting the selected signal, (iii) again bandwidth-limiting to a narrower frequency band to select (e.g.,) a signal frequency band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and bandwidth-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, bandwidth limiters, channel selectors, filters, down-converters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0039] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 100 to other electronic devices via USB and / or HDMI connections. It will be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 110 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 110. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in combination with memory and storage elements to process the data streams as needed for presentation on the output device.

[0040] Various components of system 100 can be housed within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data therebetween using a suitable connection arrangement 115 (e.g., internal buses known in the art, including I2C buses, wiring, and printed circuit boards).

[0041] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0042] In various embodiments, data is streamed to system 100 using a Wi-Fi network, such as IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers). Wi-Fi signals in these embodiments are received via communication channel 190 and communication interface 150, which are adapted for Wi-Fi communication. Communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that transmits data over an HDMI connection on input block 105 to provide streaming data to system 100. Further embodiments use an RF connection on input block 105 to provide streaming data to system 100. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0043] System 100 can provide output signals to various output devices, including display 165, speaker 175, and other peripheral devices 185. Display 165 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 165 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. Display 165 can also be integrated with other components (e.g., as in a smartphone) or can be a standalone device (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 185 include one or more of a standalone digital video disc (or digital multifunction disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments utilize one or more peripheral devices 185 that provide functionality based on the output of system 100. For example, a disc player performs the function of playing the output of system 100.

[0044] In various embodiments, signaling (such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention) is used to transmit control signals between system 100 and display 165, speaker 175, or other peripheral devices 185. Output devices may be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 via communication interface 150 using communication channel 190. In electronic devices such as televisions, display 165 and speaker 175 may be integrated into a single unit along with other components of system 100. In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0045] For example, if the RF portion of input 105 is part of a separate set-top box, display 165 and speaker 175 can alternatively be separated from one or more other components. In various embodiments where display 165 and speaker 175 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0046] The embodiments can be executed by processor 110 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, memory 120 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 110 can be of any type suitable for the technical environment and can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0047] Figure 2 An example of a block-based hybrid video encoder 200 is shown. Variations of this encoder 200 are envisioned, but for clarity, the encoder 200 is described below without describing all anticipated variations.

[0048] In some embodiments, Figure 2 Encoders are also shown, which are improvements on the HEVC or VVC standard (General Video Coding, standard ITU-T H.266, ISO / IEC 23090-3, 2020) or encoders using similar technologies to HEVC or VVC, such as the encoder ECM being developed by JVET (Joint Video Exploration Group).

[0049] Before being encoded, the video sequence may undergo pre-coding (201), for example, applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), performing remapping of the input image components to obtain a more compression-resistant signal distribution (e.g., using histogram equalization of the color components), or resizing the image (e.g., downsizing). Metadata may be associated with pre-processing and appended to the bitstream.

[0050] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed in units such as CUs (coding units) or blocks. In this disclosure, different expressions may be used to refer to such units or blocks generated from the partitioning of the image. Such terminology may be coding unit or CU, coding block or CB, luminance CB, or block. A CTU (coding tree unit) refers to a set of blocks or a set of units or a set of coding units (CUs). In some embodiments, a CTU may be considered a block or a unit itself. Figure 15 The above shows an example of partitioning using CTU and CU.

[0051] For example, in VVC, as in HEVC, the image is partitioned into multiple non-overlapping CTUs. The CTU size in VVC can be set up to 128x128 or 256x256 in luminance samples, while in HEVC it ​​can be set up to 64x64. In HEVC, recursive quadtree (QT) splitting can be applied to each CTU, resulting in one or more CUs, all of which have a square shape. In VVC, rectangular CUs are supported along with square CUs. Binary tree (BT) splitting and ternary tree (TT) splitting are also used in VVC. Figure 16 shows some examples of CTU partitioning. Further splitting of the resulting CUs is also possible.

[0052] Each unit is encoded using, for example, intra-frame or inter-frame modes. When a unit is encoded in intra-frame mode, intra-frame prediction (260) is performed. In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes is used for the encoded unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. The encoder can also mix (263) intra-frame prediction results and inter-frame prediction results, or mix results from different intra-frame / inter-frame prediction methods. For example, prediction residuals are calculated by subtracting (210) the prediction block from the original image block.

[0053] The motion refinement module (272) uses an already available reference image to refine the motion field of a block without referencing the original block. The motion field of a region can be considered as the set of motion vectors of all pixels in that region. If the motion vectors are based on sub-blocks, the motion field can also be represented as the set of motion vectors of all sub-blocks in the region (all pixels within a sub-block have the same motion vector, and the motion vector can vary depending on the sub-block). If a single motion vector is used for the region, the motion field of that region can also be represented by a single motion vector (the same motion vector for all pixels in that region).

[0054] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply the quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode the residual without applying the transform or quantization process.

[0055] The encoder decodes (reconstructs) the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inversely transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored at the reference image buffer (280).

[0056] Figure 3 A block diagram of a video decoder 300 is shown. In decoder 300, the bitstream is decoded by decoder elements as described below. Video decoder 300 typically performs operations similar to... Figure 2 The encoding process described herein is the inverse of the decoding process. Encoder 200 typically also performs video decoding as part of the encoded video data.

[0057] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. First, the bitstream is entropy-decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how to partition the image. Therefore, the decoder can partition (335) the image based on the decoded image partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct the image blocks.

[0058] The (370) prediction block can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). The decoder can mix (373) the intra-frame prediction results and the inter-frame prediction results, or mix the results from multiple intra-frame / inter-frame prediction methods. Before motion compensation, the motion field can be refined (372) by using an already available reference image. A loop filter (365) is applied to the reconstructed image. The filtered image is stored at the reference image buffer (380).

[0059] The decoded image can undergo further post-decoding processing (385), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping of the remapping process performed in the pre-encoding process (201), or resizing the reconstructed image (e.g., scaling up). Post-decoding processing can utilize metadata derived in the pre-encoding process and signaled in the bitstream.

[0060] Some of the embodiments described herein relate to template-based encoding patterns and to improving the determination of the cost of template-based encoding patterns used in these encoding patterns.

[0061] Any of the embodiments described herein can be implemented, for example, in the intra-prediction module or inter-prediction module of a video encoder or video decoder. For example, the embodiments described herein can be implemented in the intra-prediction module 260 or motion estimation module 275, or motion refinement module 272 or motion compensation module 270 of video encoder 200, or the intra-prediction module 360 ​​or motion refinement module 372, or motion compensation module 375 of video decoder 300.

[0062] In the ECM algorithm description (Muhamed Coban, Ru-Ling Liao, Karam Naser, Jacob Ström, Li Zhang, “Algorithm Description of Enhanced Compression Model 9 (ECM 9)”, 30th Session, Antalya, TR, April 21-28, 2023), several tools or coding schemes use template-based cost to determine or derive one or more parameters of the coding scheme. Some of these are described below.

[0063] Fusion for Template-Based Intra-Frame Mode Export (TIMD) In ECM, for each intra-prediction mode in the MPM (most likely mode list) and the wide-angle mode when upper right and / or lower left reference samples are available, the SATD (Sum of Absolute Transform Differences, e.g., Hadamard Transform) between the predicted and reconstructed samples of the template of the current block to be encoded or decoded is calculated.

[0064] Then, the two intra-prediction modes that provide the minimum SATD are selected as TIMD modes. These two TIMD modes are weighted and fused after applying the PDPC (Location-Associated Intra-Prediction Combination) process, and this weighted intra-prediction is used to encode the current CU / block. The Location-Associated Intra-Prediction Combination (PDPC) is included in the derivation of the TIMD modes.

[0065] The costs of the two selected modes (costMode1, costMode2) are compared with a threshold. In the comparison, cost factor 2 is applied as follows: costMode2 < 2 * costMode1.

[0066] If this condition is true, then the fusion of the two selected modes is applied; otherwise, only mode 1 (which provides CostMode1) is used.

[0067] When merging two selected patterns, the pattern weights (weight1, weight2) are calculated from their SATD costs, as follows: weight1=costMode2 / (cost Mode1+costMode2) weight2 = 1 - weight1.

[0068] Adaptive reordering of merge candidates with template matching (ARMC-TM) In ECM, in merge mode, the candidate motion list is populated with motion information from previous encoded blocks considered in a given order.

[0069] The candidates in the merge list (also known as merge candidates) are adaptively reordered using template matching (TM). This reordering method is applied to the following merge modes: regular merge mode, TM merge mode, and affine merge mode (excluding the SbTMVP candidate-sub-block temporal motion vector predictor). For TM merge mode, the merge candidates are reordered before the refinement process is applied to the selected merge candidates.

[0070] First, an initial list of candidates to be merged is constructed based on the given inspection order of the candidates. For example, the following candidates can be considered: spatial candidates, temporal motion vector predictions (TMVP), non-nearest candidates, historical motion vector predictions (HMVP), paired candidates, and virtual merge candidates. Then, the candidates in the initial list are divided into several subgroups.

[0071] For Template Matching (TM) merging mode and Adaptive DMVR (Decoder-Side Motion Vector Refinement) mode, TM or multiple passes of DMVR are first used to refine each merging candidate in the initial list.

[0072] The merge candidates in each subgroup are reordered to generate a reordered merge candidate list. This reordering is done based on a template-based cost determined for each candidate. The index of the selected merge candidate in the reordered merge candidate list is then signaled to the decoder.

[0073] For simplicity, merge candidates in the last but not the first subgroup are not reordered. All zero candidates from the ARMC reordering process are excluded during the construction of the merge motion vector candidate list. For regular merge mode and TM merge mode, the subgroup size is set to 5. For affine merge mode, the subgroup size is set to 3.

[0074] Template-based cost calculation: The template matching cost of merging candidates during the reordering process is measured by the SAD (Sum of Absolute Differences) between the samples of the current block's template and their corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. Reference (predicted) samples of the template are located using the motion information (motion compensation) of the merging candidates. When merging candidates utilize bidirectional prediction, the reference samples of the merging candidate's template are also generated through bidirectional prediction, such as... Figure 4 As shown in the image.

[0075] Refinement of the initial merge candidate list: When multiple passes of DMVR are used to export the refined motion to the initial merge candidate list, only the first pass of the multiple passes of DMVR (i.e., PU level) is applied during reordering. When template matching is used to export the refined motion, the template size is set to 1. When a block is a flat block whose width is greater than twice its height, or a narrow block whose height is greater than twice its width, only the top or left template is used respectively during TM motion refinement. TM is expanded to perform 1 / 16 pixel MVD precision. The refined motion is used to reorder the first four merge candidates in TM merge mode.

[0076] For sub-block-based merge candidates with a sub-block size equal to Wsub × Hsub, the above template includes several sub-templates of size Wsub × 1, and the left template includes several sub-templates of size 1 × Hsub. For example... Figure 5 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points for each sub-template.

[0077] Reordering criteria: During the reordering process, if the cost difference between a candidate and its predecessor is less than λ (e.g., |D1-D2|<λ), then the candidate is considered redundant, where D1 and D2 are the costs obtained during the first ARMC sorting, and λ is the Lagrange parameter used in the RD criterion on the encoder side.

[0078] The algorithm is defined as follows. Among all candidates in the list, determine the candidate and the minimum cost difference between them. If the minimum cost difference is greater than or equal to λ, the list is considered sufficiently diverse, and reordering stops. Otherwise (the minimum cost difference is less than λ), the candidate is considered redundant, and it is moved to another position in the list. This other position is the first position, where the candidate is sufficiently diverse compared to its predecessor. The algorithm stops after a finite number of iterations (if the minimum cost difference is not less than λ).

[0079] The aforementioned reordering algorithm is applied to conventional, TM, BM (bilateral matching), and affine merging modes. Similar algorithms have been applied to merged MMVD (merging with MVD) and symbolic MVD (motion vector difference) prediction methods, which also use ARMC for reordering.

[0080] The value of λ is set to be equal to the rate-distortion criterion λ used for selecting the best merging candidate on the encoder side for the low-latency configuration, and equal to the value λ corresponding to another QP for the random access configuration. A set of λ values ​​corresponding to each QP offset transmitted with a signal is provided in the SPS or in the strip header for QP offsets not present in the SPS.

[0081] Extensions to the AMVP pattern: The ARMC design also applies to the AMVP (Adaptive Motion Vector Prediction) mode, where AMVP candidates are reordered based on TM cost. For the Template Matching (TM-AMVP) mode used for advanced motion vector prediction, an initial AMVP candidate list is constructed, and then refined using TM to build a more refined AMVP candidate list. Furthermore, MVP candidates with TM costs greater than a threshold equal to five times the cost of the first MVP candidate are skipped.

[0082] Geometric Partitioning Mode with Template Matching (GPM) (GPM-TM) In ECM, template matching is applied to GPM (Geometric Partitioning Mode). When GPM mode is enabled for a CU / block, a CU-level flag (GPM-TM flag) is signaled to indicate whether TM is applied to two geometric partitions. TM is used to refine the motion information of each geometric partition. When TM is selected, a template is constructed using adjacent samples from the left and top, or left and top, depending on the partition angle, as shown in Table 1 below. The motion is then refined by minimizing the difference between the current template and the template in the reference image using the same search method in merge mode with the half-pixel interpolation filter disabled.

[0083] Table 1: Templates for the first and second geometric partitions, where A indicates the use of the top sample point, L indicates the use of the left sample point, and L+A indicates the use of both the left and top samples.

[0084] In addition to GPM-TM, GPM-MMVD allows MVD information to be signaled to correct selected motion vector candidates. GPM-MMVD and GPM-TM are exclusively enabled for a GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching should be applied to those two GPM partitions. Otherwise (if at least one GPM-MMVD flag is true), the GPM-TM flag is inferred to be false.

[0085] Template-match-based reordering for GPM splitting patterns (GPM-SPLIT-TM) In ECM, template matching can also be used to reorder different geometric partitions of blocks corresponding to different edge angles. These partitions are also known as GPM split patterns. There are 64 GPM split patterns.

[0086] In template-match-based reordering for GPM split patterns, given the motion information of the current GPM block, the corresponding TM cost value for the GPM split pattern is calculated. Then, all GPM split patterns are reordered in ascending order based on their TM cost values ​​(SAD). Instead of sending GPM split patterns, a signal is used to send an index indicating exactly where the GPM split pattern is located in the reordered list.

[0087] The reordering method for the GPM split pattern is a two-step process performed after generating corresponding reference templates for the two GPM partitions in the coding unit, as follows. The edges of the GPM partitions are extended into the reference templates of the two GPM partitions, resulting in 64 reference templates, and the corresponding TM cost is calculated for each of the 64 reference templates.

[0088] The GPM split patterns are then reordered in ascending order based on their TM cost values, and the top 32 split patterns are marked as available split patterns.

[0089] The edges on the template extend from the edges of the current CU, such as Figure 6 As shown, however, the GPM blending process is not used in template regions that cross edges. After reordering in ascending order using TM costs, the index of the selected splitting pattern is signaled.

[0090] Template-match-based BCW index export for merge patterns (BCW-TM) In BCW for the merge mode, weighted bidirectional prediction is used to predict blocks, given for example by: B = (W0xP0 + W1xP1) / 8, where B is the predicted block, P0 and P1 are reference blocks, and W0 and W1 are the corresponding weights applied to the reference blocks. Motion information from candidates in the merge list is inherited to obtain the reference blocks. In BCW, for the merge mode, weights W0 and W1 are selected from a given table using BCW indexes. An example of such a table could be: .

[0091] In BCW-TM mode, the BCW (with CU-based weighted bidirectional prediction) index of the merged encoded CU is derived based on template matching cost, instead of being derived from the BCW index of adjacent blocks.

[0092] Given the selected merge candidates, the TM cost value is calculated using different bidirectional forecast weights, and then the merge CU is predicted using the bidirectional forecast weight with the minimum TM cost value.

[0093] When calculating the TM cost for bidirectional prediction weights, the following rules apply. Since inherited bidirectional prediction weights may have higher accuracy than other weights, only inherited bidirectional prediction weights and their two adjacent weights (i.e., ±1) are considered. For example, if the inherited bidirectional prediction weight is 4, then only three weights {3, 4, 5} are involved in the TM cost calculation. The TM cost of the inherited BCW index is multiplied by 0.90625, i.e., a cost reduction of 3 / 32. The TM cost for equal weights is multiplied by 0.90625 because bidirectional prediction samples benefit BDOF, and BDOF is only applied to CUs with equal weights.

[0094] Template matching-based BCW index export was applied to CUs encoded with regular merging, template matching, adaptive decoder-side motion vector refinement, and MMVD mode.

[0095] Additionally, the bidirectional prediction weights used for the merged mode are expanded from {-2, 3, 4, 5, 10} to {1, 2, 3, 4, 5, 6, 7}. Furthermore, the negative bidirectional prediction weights {-2, 10} for the non-merged mode are replaced with positive weights {1, 7}.

[0096] In the example described above, it appears that the determination of the template cost can be improved. It can be observed that only a single component (e.g., the luminance component) is used to calculate the template cost or template matching cost. Furthermore, the template matching cost does not consider the mode signaling cost, which, in some coding schemes, may influence the rate / distortion cost of the selected candidate. Additionally, as seen in the example described above, the TM shape is fixed for all mode candidates, and it does not consider other mode parameters. Therefore, there is a need to improve template-based coding schemes.

[0097] In this document, the costs of wording based on templates and the costs of template matching can be used interchangeably.

[0098] This document provides embodiments for improving template-based coding modes, namely, using template-based costs to determine coding parameters such as motion vectors or intra-frame prediction modes, or refined derived coding parameters, or candidate prediction values ​​for reordered blocks.

[0099] Some of the embodiments described herein may include at least one of the following: determining a template matching cost using at least two components (e.g., luminance and chrominance components), determining a template matching cost that includes pattern signaling costs (e.g., candidate indices), or determining a template matching cost in which the shape of the template is adapted according to a given criterion. For example, such a criterion may be at least one of pattern candidates or encoding parameters of an encoding pattern. The TM shape may be implicitly derived based on some pattern parameters (e.g., spatial candidate positions relative to the current block). The TM cost may be appropriately renormalized when compared to template costs with different shape sizes.

[0100] According to a general embodiment, a method is provided for encoding or decoding blocks of video, wherein an encoding mode is used to encode / decode the block, the encoding mode using one or more template-based costing methods. This general embodiment, for example, in... Figure 8 As shown above, Figure 8 A method 800 for encoding / decoding video blocks is shown.

[0101] At point 810, for a given encoding pattern using template-based cost, encoding pattern parameters are determined based on one or more template-based costs. The determination of template-based costs is adapted according to the encoding pattern.

[0102] In one embodiment, one or more template-based costs are determined using a first cost obtained for a first component of the block (e.g., luminance) and at least one additional cost determined for the block. In a variant, the additional cost is a template-based cost determined for another component of the block (e.g., chroma component). In another variant, the additional cost is the cost of one or more syntax elements used to signal the encoded pattern.

[0103] In another embodiment, one or more template-based costs are determined using a template shape adapted according to a given criterion, such as an encoding pattern or a candidate used in the encoding pattern.

[0104] At 820, once the parameters of the encoding mode have been determined, the prediction of the block is obtained based on the encoding mode, and at 830, the block is encoded or decoded based on the prediction.

[0105] The embodiments described below, as well as other variations of these embodiments, at step 810 above. These embodiments / variations can be used individually or in combination to further enhance template-based cost determination or reduce the complexity of template-based cost determination. The embodiments / variations above and below are described with reference to a given coding mode for which template-based cost determination of coding mode parameters are considered. These parameters may be candidate motions, selected candidates from a candidate list, intra-frame mode prediction, geometric splitting modes, weighted prediction weights, etc.

[0106] For example, where applicable, the embodiments / variations described above and below are applied to any coding mode of the ECM algorithm mentioned above. For example, the coding mode is one of the following: an intra-frame mode that derives intra-prediction modes using template-based cost; a merge mode that reorders a candidate motion list using template-based cost; an adaptive motion vector prediction mode that reorders a candidate motion list using template-based cost; a geometric partitioning mode that determines motion for geometric partitions using template-based cost; a geometric partitioning mode that reorders a candidate splitting mode list for blocks using template-based cost; and a merge mode that determines weights for block-based weighted bidirectional prediction using template-based cost.

[0107] The embodiments / variations can also be applied to any coding scheme that requires determining template-based costs.

[0108] In ECM, several encoding tools (such as TIMD, ARMC-TM, GPM-TM, GPM-SPLIT-TM, BCW-TM) use template-based cost to estimate the cost of a set of candidate modes. The cost is measured as the difference between the reconstructed samples in the template and the predictions obtained using the considered candidate modes against the template. The difference measurement function can be, for example, SAD (Sum of Absolute Differences), SATD (Hadamard), MR-SAD (Mean Slower SAD), MR-SATD (Mean Slower SATD), or any quadratic error function, such as MSE (Sum of Squared Errors). Typically, the difference is measured using a single component, which is the luminance sample.

[0109] In the embodiment described herein, the difference is measured using several components, specifically at least two components, nc > 1, where nc is the number of components used in the template-based cost determination. These two components can be luminance and chrominance components. The final measurement result is then the sum of nc measurements.

[0110] In one variation, the final measurement result can be a weighted sum of nc measurements, where predefined weights are associated with each component. In another variation, the weights can depend on the chroma format.

[0111] For example, in the case of 4:4:4 RGB, equal weights can be used for each RGB component. In another example, in the case of 4:4:4 YCbCr, unequal weights (e.g., 4, 1, 1) can be used to emphasize distortion in the luminance component. In yet another example, in the case of 4:2:0 YCbCr, using equal weights will implicitly emphasize luminance because the number of samples in the template for the Y component is greater than the number of samples in the template for the Cb or Cr components.

[0112] In one variant, a signal for determining one or more components of the template-based cost can be sent to the decoder.

[0113] Alternatively, in another variation, the use of one or more components for determining template-based costs can be implicitly derived from the coding mode. For example, in cases where the prediction mode is a different intra-mode for luma and chroma or in the case of dual-tree, a regular difference measure is used, which uses a single component (nc=1), while in cases where the prediction mode is the same intra-mode for luma and chroma, a combination of several difference measures (nc>1) is used.

[0114] In another variant, certain components (e.g., nc=3) are used only in inter-frame mode, or in a subset of inter-frame modes (e.g., BCW-TM).

[0115] Figure 9 A method 900 for encoding / decoding video blocks according to this embodiment is illustrated. In this embodiment, at 910, parameters of the encoding mode for the block are determined based on one or more template-based costs, wherein the template-based costs are determined using at least two components according to any of the variations described above. Once the parameters of the encoding mode for the block have been determined, at 920, a prediction of the block is obtained based on the encoding mode, and at 930, the block is encoded or decoded based on the prediction.

[0116] In another variation, to reduce complexity, the cost is first calculated on only one component for all candidates considered in the encoding pattern, and the cost of additional components is added only for the best N candidates (e.g., N=3, or N can have different values ​​for each tool / encoding pattern). For example, depending on the metric considered, the best candidate is the test candidate that provides the lowest template-based cost. This variation allows for refining the template-based costs of these N best candidates, and thus, if the template-based cost is used to refine the candidates, then the candidates are improved in the others, or if the template-based cost is used to reorder the candidate list, then the reordering of the candidates is improved.

[0117] In some embodiments, the number of candidates using more than one component in the template-based cost depends on a comparison of the cost of the candidates obtained using only the first component with the cost of the best candidate obtained on that component. For example, if the cost of a candidate is less than t*BestCost (e.g., t = 1.1), the cost of this candidate is refined using another component. The cost of the best candidate is also refined using another component, allowing for further comparison of the two candidates. This variation allows for refinement of the template-based costs of candidates that are too close, and thus further enhances the discriminative power between close candidates.

[0118] As previously described, according to one embodiment, the additional cost may be the cost of using one or more syntax elements of the coded pattern to signal the signal. In ECM, coding tools that use TMs to reorder candidates (e.g., ARMC-TM, GPM-TM, GPM-SPLIT-TM), select candidates (e.g., BCW-TM, TIMD), and / or derive weighted fusions (e.g., TIMD) use difference measurements as a cost function.

[0119] In this embodiment, the estimated cost of signaling (VLC, index number, or log2(index number)) or the actual coding cost (e.g., using CABAC entropy coding) is added to the template-based cost calculation. For example, the signaling may include an index of a candidate in a candidate list sorted according to the template-based cost, a motion vector residual from motion vector prediction values ​​determined using the template-based cost, a flag indicating motion information of block partitions in a block where the geometric block partitions are determined using the template-based cost, and at least one of an intra-prediction mode or a geometric splitting mode index.

[0120] For example, the template-based cost is determined by cost = diff + λ.bits, where diff is the difference measurement determined on the template samples, bits is the estimated coding pattern signaling cost, and λ can be a Lagrangian (which can be a predetermined value as a function of the picture type and / or quantization parameters) or any other equivalent rescaling factor associated with the difference measurement function. In practice, the formula based on the regular Lagrangian uses MSE instead of SAD, and the value of λ can then be adjusted accordingly.

[0121] Figure 10 A method 1000 for encoding / decoding blocks of video according to one embodiment is illustrated, wherein additional costs are used to refine the initially determined template-based costs. At 1010, a template matching cost is determined for all candidates considered in the encoding mode. At 1020, the candidate providing the lowest template matching cost is determined as the best candidate. At 1030, a loop is performed on the other candidates or a subset of the other candidates, and it is verified that the template matching cost of the candidates is close to the template matching cost of the best candidate. For example, it is verified that the difference between the template-based cost of the candidates and the template-based cost of the best candidate is less than a given value.

[0122] If this is the case, additional costs are used to refine the template matching costs of the candidates and the best candidate. This additional cost can be a template-based cost determined for another component or the cost of signaling one or more syntax elements of the coded pattern associated with the candidate. Once the template matching costs have been refined, the parameters of the coded pattern are determined based on the refined template matching costs, and as follows... Figure 8 Or encode / decode the block as in method 9.

[0123] In ECM, the TM shape is fixed (a predetermined set of samples above and / or to the left of the current block to be encoded or decoded) or the TM shape is explicitly sent by a signal (e.g., GPM-TM).

[0124] According to one embodiment, the TM shape is not explicitly sent by signaling, but is implicitly derived from the encoding parameters of the current block.

[0125] In one variant, the TM shape can be derived from the block shape: for example, when W > k*H, the template is composed of reconstructed samples above the block, where {W, H, k} are the width and height of the current block respectively, and k is a predefined value (e.g., k = 4). When H > k*W, the template is composed of reconstructed samples on the left side of the block, using only the left template.

[0126] For example, this variant can be used when the coding mode is intra mode or when the coding mode is inter mode and at least one template-based cost is used to reorder the candidate list.

[0127] In another variant, the TM shape is defined based on candidates considered for template matching cost. That is, the TM shape is different for some candidates. For example, in the case of angular intra prediction, if the intra prediction angle is negative, the TM shape is composed of reconstructed samples on the left side and above the current block. In another example, if the intra prediction angle is zero (or abs(intra prediction angle) < TM threshold) and in the vertical direction, the TM shape is composed only of reconstructed samples above the current block, and if the intra prediction angle is zero (or abs(intra prediction angle) < TM threshold) and in the horizontal direction, the TM shape is composed only of reconstructed samples at the left side of the current block, where the TM threshold is a given threshold. Figure 7 Examples of intra prediction directions for the current block for encoding / decoding are shown.

[0128] In another variant, the template shape is determined based on the position of the candidate relative to the block to be encoded / decoded. Figure 15 Examples of neighboring blocks used to derive spatial candidates are shown.

[0129] For example, if the block is encoded in merge mode, and if the merge candidate is a spatial candidate and the spatial candidate is above the block (or upper right) (e.g., indicated by Figure 15 2, 3 above), then the TM shape is composed only of reconstructed samples above the block ({above}), and if the merge candidate is spatial and the spatial candidate is to the left of the block (or lower left) (e.g., indicated by Figure 15 1, 4 above), then the TM shape is composed only of reconstructed samples on the left side of the block ({left side}). In other cases, the TM shape is composed of reconstructed samples above the block and reconstructed samples on the left side of the block ({above + left side}). When the candidate is a spatial non-neighboring block as shown in Figure 15 blocks 6 - 7, 9 - 12, 14, 17, 19, 22), similar conditions for defining the template shape can also be applied.

[0130] When several candidates use template shapes of different sizes, the differences measured on the template can be renormalized. For example, if the reference TM shape (e.g., {top + left}) contains N... R For each sample point, the following method can be applied using N... C The difference (Diff) in the TM shape measurement composed of individual samples C Renormalization of ) Diff = (Diff C xN C ) / N R , where Diff is the renormalized measurement.

[0131] Figure 11 A method 1100 for encoding / decoding a block of video according to this embodiment is illustrated, wherein a template shape is defined based on the encoding parameters of the block. At 1110, encoding parameters and a template shape are determined for the block based on an encoding mode. The encoding mode uses one or more template-based costs. For each template-based cost used in the encoding mode, the template shape is determined according to any of the variations described above. At 1120, once the parameters of the encoding mode have been determined, a prediction is determined based on the encoding mode, and at 1130, the block is encoded or decoded based on the prediction.

[0132] Figure 12 A block diagram of a system according to another embodiment, in which aspects of the embodiments herein may be implemented, is shown. Figure 12 An embodiment of an apparatus 1200 for encoding or decoding video according to any of the embodiments described herein is illustrated. The apparatus includes a processor 1210 and is interconnected to a memory 1220 via at least one port. Both the processor 1210 and the memory 1220 may also have one or more additional interconnects to external connections.

[0133] The processor 1210 is also configured to: obtain a prediction of at least one block of a video based on an encoding mode, wherein the encoding mode uses at least one template-based cost, the template-based cost being determined using at least one first cost obtained for a first component of the block; and encode or decode the block based on the prediction, wherein the at least one template-based cost is determined using at least one additional cost determined for the block.

[0134] In another embodiment, the processor 1210 is further configured to obtain a prediction of a block based on an encoding mode, wherein the encoding mode uses at least one template-based cost; and to encode or decode the block based on the prediction, wherein the at least one template-based cost is determined using a template shape derived from at least one parameter of the encoding mode.

[0135] For example, processor 1210 uses a computer program product that includes code instructions that implement any of the embodiments described herein.

[0136] exist Figure 13 In one embodiment shown, in the transmission context between two remote devices A and B via a communication network NET, device A includes a processor associated with RAM and ROM, configured to implement methods for encoding video, such as those described above. Figures 1 to 12 As described, device B includes a processor associated with memory RAM and ROM, which is configured to implement methods for decoding video, as per [reference to...]. Figures 1 to 12 As described. Based on one example, the network is a broadcast network, suitable for broadcasting / sending encoded video from device A to decoding devices including device B.

[0137] Figure 14 An example of the syntax for signals transmitted via a packet-based transport protocol is shown. Each transport packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include video data encoded according to any of the embodiments described above. The payload may also include any signaling as described above.

[0138] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process, for example, performing on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also or alternatively includes processes performed by a decoder of the various embodiments described herein, such as entropy decoding of a sequence of binary symbols to reconstruct image or video data.

[0139] As further examples, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding; and in yet another embodiment, "decoding" refers to the entire image reconstruction process including entropy decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific descriptive context and is considered well understood by those skilled in the art.

[0140] Various implementations involve encoding. In a manner similar to the discussion of “decoding” above, “encoding” as used herein can encompass all or part of a process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such a process also or alternatively includes processes performed by an encoder of the various embodiments described herein, such as determining resampling filter coefficients and resampling the decoded image.

[0141] As a further example, in one embodiment, “encoding” refers only to entropy encoding; in another embodiment, “encoding” refers only to differential encoding; and in yet another embodiment, “encoding” refers to a combination of differential and entropy encoding. Whether the phrase “encoding process” is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.

[0142] Note that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.

[0143] This disclosure has described various information fragments that can be transmitted or stored, such as, for example, syntax. This information can be encapsulated or arranged in a variety of ways, including those common in video standards, such as placing the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers, picture headers, or stripe headers), or SEI messages. Other methods are also available, including those common for system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (Session Description Protocol), a format for describing multimedia communication sessions for the purposes of session notification and session invitation, for example, as described in the RFC and used in conjunction with RTP (Real-Time Transport Protocol) transmission.

[0144] b. DASH MPD (Media Presentation Description) descriptors, such as those used in DASH and transmitted via HTTP, are associated with a presentation or a set of presentations to provide additional features for content presentation.

[0145] c. RTP header extensions, such as those used during RTP streaming.

[0146] d. ISO basic media file formats, such as those used in OMAF, and using boxes, which are object-oriented building blocks defined by unique type identifiers and lengths, also referred to as "atoms" in some specifications.

[0147] e. An HLS (HTTP Live Streaming) manifest transmitted over HTTP. The manifest can be associated with, for example, versions or sets of versions of content to provide characteristics of the version or set of versions.

[0148] When the accompanying drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding apparatus. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.

[0149] Some implementations involve rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often with constraints on computational complexity. Rate-distortion optimization is typically expressed as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist for solving rate-distortion optimization problems. For example, methods can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, where their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding are fully evaluated. Faster methods can also be used to save encoding complexity, particularly by calculating approximate distortion based on predicting or predicting the residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, such as by using approximate distortion only for some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.

[0150] The implementations and aspects described herein can be implemented, for example, as methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method), implementations of the discussed features can be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in suitable hardware, software, and firmware. These methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0151] References to “an embodiment” or “an embodiment” or “a implementation” or “an implementation” and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Therefore, the phrases “in an embodiment” or “in an embodiment” or “in an implementation” or “in an implementation” appearing throughout this disclosure and any other variations do not necessarily refer to the same embodiment.

[0152] Additionally, this application may refer to "determining" various information fragments. Determining information may include, for example, one or more of the following: estimation information, calculation information, prediction information, or information retrieved from memory.

[0153] Furthermore, this application may refer to "accessing" various information fragments. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0154] Additionally, this application may refer to "receiving" various pieces of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) one or more. Furthermore, "receiving" is generally referred to in one or more ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0155] It will be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be clear to those skilled in the art or related fields, this can be extended to include the number of items listed.

[0156] Additionally, as used herein, among other things, the phrase “signal” refers to instructing something to the corresponding decoder. Thus, in one embodiment, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicitly signal) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used without transmission (implicitly signal) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It will be understood that signaling can be done in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. While the foregoing refers to the verb form of the word “signal,” the word “signal” can also be used as a noun herein.

[0157] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry information, such as information that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, signals can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

[0158] Several embodiments have been described above. Features of these embodiments may be provided individually or in any combination across various claim classes and types.

Claims

1. A method comprising: Prediction of at least one block of a video is obtained based on an encoding mode, wherein the encoding mode uses at least one template-based cost, the template-based cost being determined using at least one first cost obtained for a first component of the block. The block is encoded or decoded based on the prediction. The cost of the at least one template-based cost is determined using at least one additional cost determined for the block.

2. An apparatus comprising one or more processors, said one or more processors being configured to: Prediction of at least one block of a video is obtained based on an encoding mode, wherein the encoding mode uses at least one template-based cost, the template-based cost being determined using at least one first cost obtained for a first component of the block. The block is encoded or decoded based on the prediction. in, The cost of the at least one template-based cost is determined using at least one additional cost determined for the block.

3. The method according to claim 1 or the apparatus according to claim 2, wherein, The additional cost is a template-based cost obtained for another component of the block.

4. The method or apparatus of claim 3, wherein the first component is a luminance component and the other component is a chromaticity component.

5. The method according to claim 1 or any one of claims 3 to 4, or the apparatus according to any one of claims 2 to 4, wherein the cost of the at least one template-based cost is a weighted sum of the first cost and the at least one additional cost.

6. The method or apparatus of claim 5, wherein the weights depend on the chromaticity format of the chromaticity components.

7. The method of claim 1 or any one of claims 3 to 6, or the apparatus of any one of claims 2 to 6, wherein the at least one additional cost of using the block depends on the encoding mode.

8. The method according to claim 1 or any one of claims 3 to 7, or the apparatus according to any one of claims 2 to 7, wherein, Determining the cost of the at least one template-based template includes: using the first cost to determine the cost of a first template-based template for a given number of candidates, and using the first cost and the additional cost to determine the cost of a second template-based template for a subset of the candidates.

9. The method or apparatus according to claim 8, wherein, Candidates in the subset used to determine the cost of the second template based on the first cost and the additional cost are candidates whose difference between the cost of the first template based on the first candidate and the cost of the first template based on the best candidate is less than a given value.

10. The method of any one of claims 1 to 9 or the apparatus of any one of claims 2 to 9, wherein the additional cost is the cost of signaling one or more syntax elements associated with the encoding pattern.

11. The method or apparatus of claim 10, wherein the one or more syntax elements associated with the encoding mode include at least one of the following: Based on the index of a candidate in a candidate list sorted by cost based on at least one template, The motion vector residuals derived from the motion vector predictions determined using at least one template-based cost. A flag indicating the use of template-based cost to determine the motion information of block partitions within a block of geometric block partitions.

12. A method comprising at least one block for a video: The prediction of the block is obtained based on an encoding pattern, wherein the encoding pattern uses at least one template-based cost. The block is encoded or decoded based on the prediction. The cost of the at least one template-based method is determined using a template shape determined based on at least one parameter of the encoding pattern.

13. An apparatus comprising one or more processors, said one or more processors operable for at least one block of video: The prediction of the block is obtained based on an encoding pattern, wherein the encoding pattern uses at least one template-based cost. The block is encoded or decoded based on the prediction. The cost of at least one template is determined using a template shape derived from at least one parameter of the encoding pattern.

14. The method of claim 12 or the apparatus of claim 13, wherein the encoding mode is an intra-frame mode and the template shape is derived based on the angular intra-frame prediction mode.

15. The method of claim 12 or the apparatus of claim 13, wherein the encoding mode is a merging mode and the template shape depends on the position of the spatial merging candidate relative to the block.

16. The method according to claim 12 or 14, or the apparatus according to claim 13 or 14, wherein, The encoding mode is either intra-frame mode or inter-frame mode, and the at least one template-based cost is used to reorder the candidate list, the template shape being derived from the shape of the block.

17. The method or apparatus of claim 16, wherein if the width of the block is greater than k times the height of the block, then the template is formed by pixels above the block, or if the height of the block is greater than k times the width of the block, then the template is formed by pixels to the left of the block.

18. The method according to any one of claims 1, 3 to 12 or 14 to 17, or the apparatus according to any one of claims 2 to 11 or 13 to 17, wherein the encoding mode is one of the following: The template-based cost is used to derive the intra-prediction mode of the intra-frame mode. The merging pattern that uses template-based cost to reorder the candidate motion list An adaptive motion vector prediction mode that uses template-based cost to reorder the candidate motion list. The template-based cost is used to determine the geometric partitioning pattern of the motion of the geometric partition. Geometric partitioning patterns that use the template-based cost to reorder the list of candidate splitting patterns for blocks The template-based cost is used to determine the merging pattern of weights for block-based weighted bidirectional prediction.

19. A computer program product comprising instructions for causing one or more processors to perform the method of any one of claims 1, 3 to 9, 11 to 12 or 14 to 18.

20. A non-transitory computer-readable medium storing executable program instructions that cause a computer executing the program instructions to perform the method according to any one of claims 1, 3 to 9, 11 to 12 or 14 to 18.

21. An apparatus comprising: The apparatus according to any one of claims 2 to 11 or 13 to 18; as well as At least one of the following: (i) an antenna configured to receive or transmit a signal including data representing video; (ii) a bandwidth limiter configured to limit the signal to a frequency band including data representing video; or (iii) a display configured to display video.

22. The device of claim 21, wherein the device comprises at least one of a television, a cellular phone, a tablet computer, and a set-top box.