Method for encoding or decoding video data, and device for encoding video data

BR112025019990A2Pending Publication Date: 2026-08-11
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025019990
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-08-11

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

1 / 126 Improved fusion and subpel precision mode for template matching tools for video encoding.

[0001] This application claims priority over U.S. Patent Application No. 18 / 617,841, filed March 27, 2024, which claims the benefit of U.S. Provisional Patent Application No. 63 / 492,729, filed March 28, 2023, the entire contents of each of which are incorporated by reference. U.S. Patent Application No. 18 / 617,841, filed March 27, 2024, claims the benefit of U.S. Provisional Patent Application No. 63 / 492,729, filed March 28, 2023. TECHNICAL FIELD

[0002] This disclosure relates to video encoding and video decoding. BACKGROUND

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, so-called smartphones, video conferencing devices, video streaming devices, and the like. Digital video devices implement video encoding techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T Petition 870250084327, dated 09 / 18 / 2025, page 135 / 292 2 / 126 H.264 / MPEG-4, part 10, advanced video coding (AVC), ITU-T H.265 / high-efficiency video coding (HEVC), ITU-T H.266 / versatile video coding (VVC), and extensions of these standards, as well as proprietary video codecs / formats such as AOMedia Video 1 (AV1), which was developed by the Alliance for Open Media. Video devices can transmit, receive, encode, decode, and / or store digital video information more efficiently through the implementation of these video encoding techniques.

[0004] Video coding techniques include spatial (intra-image) and / or temporal (inter-image) prediction to reduce or remove the inherent redundancy in video sequences. For block-based video coding, a video slice (e.g., a video image or a portion of a video image) can be partitioned into video blocks, which may also be called coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intracoded (I) slice of an image are coded using spatial prediction with respect to reference samples in neighboring blocks in the same image. Video blocks in an intercoded (P or B) slice of an image may use spatial prediction with respect to reference samples in neighboring blocks in the same image or temporal prediction with respect to reference samples in other reference images.Images can be called frames, and reference images can be called reference frames. Petition 870250084327, dated 09 / 18 / 2025, p. 136 / 292 3 / 126 SUMMARY

[0005] In general, this disclosure describes techniques for methods (e.g., template type, fusion) and syntax used in template matching (TM) related tools. The disclosed methods can be applied to any of the existing video codecs, such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), Essential Video Coding (EVC), or it can be an efficient coding tool in future video coding standards (e.g., Enhanced Compression Model (ECM)). This disclosure describes techniques in which a video encoder or video decoder can use different template types, use subpel interpolation, store more candidates, and apply fusion to combine these different candidates found by the different methods. These techniques can improve video coding efficiency.

[0006] In one example, a method includes a method for encoding or decoding video data, the method comprising: determining a template pattern from among a plurality of template patterns; identifying, based on the determined template pattern, a prediction block for a current block of video data; and encoding or decoding the current block using the prediction block, wherein a syntax element is signaled to indicate the determined template pattern.

[0007] In another example, this disclosure describes a method of encoding or decoding data. Petition 870250084327, dated 09 / 18 / 2025, page 137 / 292 4 / 126 of video, the method comprising: applying a subpel precision mode to generate a prediction block for a current block of video data, wherein the application of the subpel precision mode comprises: applying an interpolation filter to samples from a reference region to generate a reference region that includes an array of samples with full pel and subpel precision; and using a template pattern to identify, within the array, a prediction block for a current block of video data, wherein a block vector indicating an offset between the current block and the prediction block has subpel accuracy; and encoding or decoding the current block using the prediction block.

[0008] In another example, this disclosure describes a method for encoding or decoding video data, the method comprising: generating template matching (TM) candidates, wherein each of the TM candidates is associated with a different reference block; generating, based on a combination of two or more of the TM candidates, a predictor for a current block of video data; and encoding or decoding the current block using the predictor.

[0009] In another example, this disclosure describes a method for encoding or decoding video data, the method comprising: applying a subpel precision mode to generate a prediction block for a current block of video data, wherein a syntax element indicates that the subpel precision mode is applied to the current block, and applying the subpel precision mode comprises: applying an interpolation filter to samples from a reference region to generate a sample array with full pel precision. Petition 870250084327, dated 09 / 18 / 2025, page 138 / 292 5 / 126 and subpel; and identify, within the array, a prediction block reference template, where the prediction block reference template is the best match for a current block template within the array, where a template pattern defines a format for the prediction block reference template and the current block template; and encode or decode the current block using the prediction block for the current block.

[0010] In another example, the disclosure describes a device for encoding video data, the device comprising: memory for storing the video data; and one or more processors implemented in a circuit assembly, the one or more processors configured to: apply a subpel precision mode to generate a prediction block for a current block of video data, wherein a syntax element is signaled to indicate that the subpel precision mode is applied to the current block and the one or more processors are configured to, by applying the subpel precision mode: apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision;and identify, within the array, a reference template for the prediction block, where the reference template for the prediction block is the best match for a template of the current block within the array, where a template pattern defines a format for the reference template of the prediction block and the template of the current block; and encode or decode the current block using the prediction block for the current block. Petition 870250084327, dated 09 / 18 / 2025, page 139 / 292 6 / 126

[0011] In another example, this disclosure describes devices configured to perform the methods set forth in this disclosure.

[0012] In another example, this disclosure describes a computer-readable storage medium that is encoded with instructions that, when executed, cause a programmable processor to perform the methods set forth in this disclosure.

[0013] The details of one or more examples are set out in the attached drawings and in the description below. Other attributes, objectives and advantages will become apparent from the description, drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a block diagram illustrating an example video encoding and decoding system that can perform the techniques of the present disclosure.

[0015] Figure 2 is a block diagram illustrating an example video encoder that can perform the techniques of the present disclosure.

[0016] Figure 3 is a block diagram illustrating an example video decoder that can perform the techniques of the present disclosure.

[0017] Figure 4 is a conceptual diagram illustrating an example of an intratemplate match search area.

[0018] Figure 5 is a conceptual diagram illustrating the example template matching performed in a search area around an initial motion vector. Petition 870250084327, dated 09 / 18 / 2025, pp. 140 / 292 7 / 126

[0019] Figure 6 is a conceptual diagram illustrating an example of a template and reference samples of the template in reference images.

[0020] Figure 7 is a conceptual diagram illustrating an example template and reference samples of the template for blocks with sub-block movement using the movement information of the sub-blocks for the current block.

[0021] Figure 8 is a conceptual diagram illustrating intrablock copy reference regions depending on the current CU position.

[0022] Figure 9 is a conceptual block diagram illustrating different template types according to the techniques of this disclosure.

[0023] Figure 10 is a conceptual diagram illustrating a method for predicting subpel accuracy according to the techniques of the present disclosure.

[0024] Figure 11 is a flowchart illustrating an example method for coding an actual block in accordance with the techniques of this disclosure.

[0025] Figure 12 is a flowchart illustrating an example method for decoding a current block according to the techniques of the present disclosure.

[0026] Figure 13 is a flowchart illustrating an example operation of a video encoder according to the techniques of this disclosure. DETAILED DESCRIPTION

[0027] Template matching techniques involve a video encoder and a video decoder that searches for a prediction block (i.e., a predictor) for a Petition 870250084327, dated 09 / 18 / 2025, pp. 141 / 292 8 / 126 current block within a search area of ​​a current image or a reference image. The video encoder or video decoder evaluates different reference blocks within the search area by comparing reference samples from the current block with reference samples from the reference blocks. Reference samples from a block can be reconstructed samples above and to the left of the block. Reference samples are confined to specific template patterns, hence the term template matching.

[0028] Typically, when a video encoder (e.g., a video encoder or video decoder) searches the search area, the video encoder analyzes only pixels at full-pel locations. However, if the video encoder were to use an array of interpolated samples from the full-pel locations of the search area to search for a reference template, the video encoder might be able to generate a prediction block that is more similar to the current block than would otherwise be possible. The amount of data used to encode the current block can decrease when the prediction block is more similar to the current block. It is further understood that the use of interpolated samples can unnecessarily increase the complexity of the encoding and decoding process for some blocks. Thus, according to a technique in the present disclosure, a syntax element can indicate whether the subpel precision mode is used applied to a block.If the sub-precision mode is not applied to the block, it may be unnecessary for the video decoder to perform the procedure. Petition 870250084327, dated 09 / 18 / 2025, pp. 142 / 292 9 / 126 interpolation process, which can convert processing resources.

[0029] Thus, according to one or more techniques of the present disclosure, a video encoder can apply a subpel precision mode to generate a prediction block for a current block of video data. A syntax element indicates that the subpel precision mode is applied to the current block. As part of applying the subpel precision mode, the video encoder can apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision. The video encoder can identify, within the array, a reference template for the prediction block. The reference template for the prediction block can be the best match for a template of the current block. A template pattern defines a format for the reference template of the prediction block and the template of the current block. The video encoder can encode or decode the current block using the prediction block.

[0030] Figure 1 is a block diagram illustrating an example 100 video encoding and decoding system that can perform the techniques of this disclosure. The techniques of this disclosure generally refer to the process of encoding (encoding and / or decoding) video data. In general, video data includes any data for processing a video. Thus, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata such as signaling data. Petition 870250084327, dated 09 / 18 / 2025, page 143 / 292 10 / 126

[0031] As shown in Figure 1, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116 in this example. In particular, the source device 102 provides the video data to the destination device 116 via a computer-readable medium 110. The source device 102 and the destination device 116 may be or include any of a wide range of devices, such as desktop computers, notebook computers (i.e., laptops), mobile devices, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, or similar devices. In some cases, the source device 102 and the destination device 116 may be equipped for wireless communication and, in this way, may be called wireless communication devices.

[0032] In the example in Figure 1, source device 102 includes video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply template matching techniques. Thus, source device 102 represents an example of an encoding device of Petition 870250084327, dated 09 / 18 / 2025, page 144 / 292 11 / 126 video, while the destination device 116 represents an example of a video decoding device. In other examples, a source device and a destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source, such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device.

[0033] As shown in Figure 1, system 100 is merely an example. In general, any digital video encoding and / or decoding device can perform template matching techniques. Source device 102 and destination device 116 are merely examples of such encoding devices in which source device 102 generates encoded video data for transmission to destination device 116. This disclosure refers to an encoding device as a device that performs the process of encoding (encoding and / or decoding) data. Thus, video encoder 200 and video decoder 300 represent examples of encoding devices, in particular, a video encoder and a video decoder, respectively.In some examples, source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of the source device 102 and destination device 116 includes video encoding and decoding components. In this way, system 100 can support unidirectional or bidirectional video transmission between the devices. Petition 870250084327, dated 09 / 18 / 2025, page 145 / 292 12 / 126 source 102 and destination device 116, for example, for video streaming, video playback, video broadcasting or video telephony.

[0034] In general, the video source 104 represents a video data source (i.e., raw, unencoded video data) and provides a sequential series of images (also called frames) of video data to the video encoder 200, which encodes data into images. The video source 104 from the source device 102 may include a video capture device, such as a video camera, a video file containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As an additional alternative, the video source 104 may generate computer graphics-based data, such as the source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data.The video encoder 200 can rearrange the images from the received order (sometimes called the display order) into an encoding order for encoding. The video encoder 200 can generate a bitstream that includes encoded video data. The source device 102 can then output the encoded video data via the output interface 108 to the computer-readable medium 110 for reception and / or retrieval, for example, via the input interface 122 of the destination device 116.

[0035] Memory 106 of source device 102 and memory 120 of destination device 116 represent Petition 870250084327, dated 09 / 18 / 2025, pp. 146 / 292 13 / 126 general-purpose memories. In some examples, memories 106, 120 may store raw video data, for example, raw video from video source 104 and raw and decoded video data from video decoder 300. Additionally or alternatively, memories 106, 120 may store executable software instructions, for example, by video encoder 200 and video decoder 300, respectively. Although memory 106 and memory 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memories for functionally similar or equivalent purposes. In addition, memories 106 and 120 can store encoded video data, for example, output from video encoder 200 and input to video decoder 300.In some examples, memory portions 106, 120 can be allocated as one or more video buffers, for example, to store raw, decoded and / or encoded video data.

[0036] The computer-readable medium 110 may represent any type of medium or device with the capability to carry encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium to enable the source device 102 to transmit encoded video data directly to the destination device 116 in real time, for example, via a radio frequency network or computer-based network. The output interface 108 may modulate a transmission signal. Petition 870250084327, dated 09 / 18 / 2025, page 147 / 292 14 / 126 which includes the encoded video data, and the input interface 122 can demodulate the received transmission signal according to a communication standard, such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may be part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication between the source device 102 and the destination device 116.

[0037] In some examples, source device 102 may output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard disk, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other digital storage media suitable for storing encoded video data.

[0038] In some examples, source device 102 may output encoded video data to file server 114 or another intermediate storage device that may store the encoded video data generated by source device 102. The Petition 870250084327, dated 09 / 18 / 2025, pp. 148 / 292 15 / 126 target device 116 can access video data stored from the file server 114 via streaming or download.

[0039] The file server 114 can be any type of server device with the ability to store encoded video data and transmit such encoded video data to the destination device 116.File server 114 can represent a web server (for example, for a website), a server configured to provide a file transfer protocol service (such as File Transfer Protocol (FTP) or File Delivery over Unidirectional Transport (FLUTE)), a content delivery network (CDN) device, a hypertext transfer protocol (HTTP) server, a multimedia broadcast multicast service (MBMS) or enhanced MBMS (eMBMS) server, and / or a network attached storage (NAS) device. File server 114 may additionally or alternatively implement one or more HTTP streaming protocols, such as... Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), Real-Time Streaming Protocol (RTSP), Dynamic HTTP Streaming, or similar.

[0040] Target device 116 can access encoded video data from the file server. Petition 870250084327, dated 09 / 18 / 2025, pp. 149 / 292 16 / 126 114 via any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., digital subscriber line (DSL), cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on the file server 114. The input interface 122 may be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from the file server 114, or other protocols for retrieving media data.

[0041] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired network communication components (e.g., Ethernet cards), wireless communication components operating in accordance with any of a variety of Institute of Engineers standards. Electrical and Electronics Engineers (IEEE - Institute of Electrical and Electronics Engineers) 802.11 or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 can be configured to transfer data, such as encoded video data, according to a cellular communication standard such as 4G, 4G-LTE (long-term evolution), LTE-advanced, 5G, or similar. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to transfer data, such as video data. Petition 870250084327, dated 09 / 18 / 2025, pp. 150 / 292 17 / 126 encoded, according to other wireless standards, such as an IEEE 802.11 specification, an IEEE 802.15 specification (e.g., ZigBee™), a Bluetooth™ standard, or similar. In some examples, the source device 102 and / or the destination device 116 may include the respective system-on-a-chip (SoC) devices. For example, the source device 102 may include an SoC device to perform the functionality assigned to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device to perform the functionality assigned to the video decoder 300 and / or the input interface 122.

[0042] The techniques of this disclosure can be applied to video encoding to support any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television broadcasts, satellite television broadcasts, Internet streaming video transmissions, such as dynamic adaptive streaming over HTTP (DASH), digital video that is encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0043] The input interface 122 of the destination device 116 receives a video bitstream encoded from the computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or similar). The encoded video bitstream may include signaling information defined by the video encoder 200, which is also used by the video decoder 300, such as syntax elements that Petition 870250084327, dated 09 / 18 / 2025, pp. 151 / 292 18 / 126 have values ​​that describe characteristics and / or process video blocks or other encoded units (e.g., slices, images, image groups, sequences, or the like). Display device 118 displays decoded images from the decoded video data to a user. Display device 118 may represent any of several display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0044] Although not shown in Figure 1, in some examples, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or audio decoder and may include suitable MUX-DEMUX units, or other hardware and / or software, to handle multiplexed streams that include audio and video in a common data stream.

[0045] Each of the 200 video encoder and 300 video decoder can be implemented as any of several suitable sets of encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are partially implemented in software, a device can store Petition 870250084327, dated 09 / 18 / 2025, pp. 152 / 292 19 / 126 instructions for the software in a suitable non-transient, computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of the video encoder 200 and video decoder 300 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device. A device that includes video encoder 200 and / or video decoder 300 can implement video encoder 200 and / or video decoder 300 in a processing circuitry, such as an integrated circuit and / or a microprocessor. This device can be a wireless communication device, such as a cell phone or any other type of device described in the present invention.

[0046] The Video Encoder 200 and Video Decoder 300 can operate according to a video encoding standard, such as ITU-T H.265, also called High Efficiency Video Coding (HEVC), or extensions thereof, such as scalable video coding extensions and / or multi-view extensions. Alternatively, the Video Encoder 200 and Video Decoder 300 can operate according to other proprietary or industry standards, such as ITU-T H.266, also called Versatile Video Coding (VVC). In other examples, the Video Encoder 200 and Video Decoder 300 can operate according to a proprietary video codec / format, such as AOMedia Video 1 (AV1), AV1 extensions and / or versions. Petition 870250084327, dated 09 / 18 / 2025, pp. 153 / 292 20 / 126 successors of AV1 (e.g., AV2). In other examples, the 200 video encoder and the 300 video decoder may operate according to other proprietary formats or industry standards. The techniques of this disclosure, however, are not limited to any particular encoding standard or format. In general, the 200 video encoder and the 300 video decoder can be configured to perform the techniques of this disclosure in conjunction with any video encoding techniques that use template matching.

[0047] In general, the Video Encoder 200 and Video Decoder 300 can perform image block-based encoding. The term block generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used in the encoding and / or decoding process). For example, a block might include a two-dimensional array of luminance and / or chrominance data samples. In general, the Video Encoder 200 and Video Decoder 300 can encode video data represented in a YUV format (e.g., Y, Cb, Cr). In other words, instead of encoding red, green, and blue (RGB) data for samples of an image, the 200 video encoder and the 300 video decoder can encode luminance and chrominance components, where the chrominance components can include both red and blue hue chrominance components.In some examples, the 200 video encoder converts received RGB formatted data to a YUV representation before encoding, and the 300 video decoder converts it to... Petition 870250084327, dated 09 / 18 / 2025, pp. 154 / 292 21 / 126 YUV representation in RGB format. Alternatively, pre- and post-processing units (not shown) can perform these conversions.

[0048] This disclosure may generally refer to the encoding (e.g., encoding and decoding) of images to include the process of encoding or decoding image data. Similarly, this disclosure may refer to the encoding of blocks of an image to include the process of encoding or decoding data for the blocks, for example, prediction and / or residual encoding. An encoded video bitstream generally includes a series of values ​​for syntax elements representing encoding decisions (e.g., encoding modes) and image block partitioning. Thus, references to the encoding of an image or a block should generally be understood as encoding values ​​for syntax elements that form the image or block.

[0049] HEVC defines several blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video encoder (such as video encoder 200) partitions a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video encoder partitions CTUs and CUs into four equal, non-overlapping squares, and each node in the quadtree has zero or four child nodes. Nodes without child nodes can be called leaf nodes, and the CUs of these leaf nodes may include one or more PUs and / or one or more TUs. The video encoder may additionally partition Petition 870250084327, dated 09 / 18 / 2025, pages 155 / 292 22 / 126 PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of TUs. In HEVC, PUs represent interprediction data, while TUs represent residual data. Intraprediction CUSs include intraprediction information, such as an indication of intramode.

[0050] As another example, the Video Encoder 200 and Video Decoder 300 can be configured to operate according to VVC. According to VVC, a video encoder (such as the Video Encoder 200) partitions an image into a plurality of CTUs. The Video Encoder 200 can partition a CTU according to a tree structure, such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concepts of multiple partition types, such as the separation between CUs, PUs, and TUs of HEVC. A QTBT structure includes two levels: a first level partitioned according to the quadtree partition and a second level partitioned according to the binary tree partition. A root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary trees correspond to the CUs.

[0051] In an MTT partitioning structure, blocks can be partitioned using a quadtree (QT) partition, a binary tree (BT) partition, and one or more types of triple tree (TT) partitions (also called ternary tree (TT) partitions). A triple or ternary tree partition is a partition where a block is divided into three sub-blocks. In Petition 870250084327, dated 09 / 18 / 2025, pp. 156 / 292 23 / 126 some examples, a triple or ternary tree partition divides a block into three sub-blocks without splitting the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0052] When operating in accordance with the AV1 codec, the Video Encoder 200 and Video Decoder 300 can be configured to encode video data in blocks. In In AV1, the largest coding block that can be processed is called a superblock. In AV1, a superblock can be 128x128 luma samples or 64x64 luma samples. However, in successor video coding formats (e.g., AV2), a superblock can be defined by different (e.g., larger) luma sample sizes. In some examples, a superblock is the top level of a block quadtree. The Video 200 encoder can further partition a superblock into smaller coding blocks. The Video 200 encoder can partition a superblock and other coding blocks into smaller blocks using quadratic or non-quadratic partitioning. Non-quadratic blocks can include blocks N / 2xN, NxN / 2, N / 4xN, and NxN / 4. The 200 video encoder and the 300 video decoder can perform separate prediction and transformation processes on each of the encoding blocks.

[0053] AV1 also defines a video data tile. A tile is a rectangular array of superblocks that can be independently encoded from other tiles. That is, video encoder 200 and video decoder 300 can encode and decode, respectively, blocks Petition 870250084327, dated 09 / 18 / 2025, pp. 157 / 292 24 / 126 encoding on a tile without using video data from other tiles. However, the 200 video encoder and 300 video decoder can perform filtering between tile boundaries. Tiles can be of uniform or non-uniform size. Tile-based encoding can enable parallel and / or multi-threaded processing for encoder and decoder implementations.

[0054] In some examples, the 200 video encoder and the 300 video decoder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the 200 video encoder and the 300 video decoder may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both chrominance components (or two QTBT / MTT structures for the respective chrominance components).

[0055] The 200 video encoder and the 300 video decoder can be configured to use quadtree partition, QTBT partition, MTT partition, superblock partition, or other partition structures.

[0056] In some examples, a CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples from an image that has three sample arrangements, or a CTB of samples from a monochrome image or an image that is encoded using three separate color planes and syntax structures used to encode the samples. A CTB can be an NxN block of samples for some value of N, so that dividing a component into CTBs is a Petition 870250084327, dated 09 / 18 / 2025, pp. 158 / 292 25 / 126 Partitioning. A component is an array or a single sample of one of the three arrays (luma and two chroma) that make up a 4:2:0, 4:2:2, or 4:4:4 color format image, or the array or a single sample of the array that makes up a monochrome format image. In some examples, a coding block is an MxN block of samples for some values ​​of M and N, so a division of a CTB into coding blocks is a partitioning.

[0057] Blocks (e.g., CTUs or CUs) can be grouped in various ways in an image. As an example, a brick can refer to a rectangular region of CTU rows in a particular tile in an image. A tile can be a rectangular region of CTUs in a particular column of tiles and a particular row of tiles in an image. A column of tiles refers to a rectangular region of CTUs that has a height equal to the height of the image and a width specified by syntax elements (e.g., as in an image parameter set). A row of tiles refers to a rectangular region of CTUs that has a height specified by syntax elements (e.g., as in an image parameter set) and a width equal to the width of the image.

[0058] In some examples, a tile can be partitioned into multiple bricks, each of which may include one or more rows of CTUs in the tile. A tile that is not partitioned into multiple bricks may also be called a brick. However, a brick that is a true subset of a tile may not be called a tile. The bricks in an image may also be arranged in a slice. A slice may be an integer number of bricks. Petition 870250084327, dated 09 / 18 / 2025, pp. 159 / 292 26 / 126 of an image that can be contained exclusively within a single network abstraction layer (NAL) unit. In some examples, a slice includes a number of complete tiles or just a consecutive sequence of complete bricks of a tile.

[0059] This disclosure may use NxN and N by N interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in terms of vertical and horizontal dimensions, for example, 16x16 samples or 16 by 16 samples. Generally, a 16x16 CU will have 16 samples in a vertical direction (y = 16) and 16 samples in a horizontal direction (x = 16). Similarly, an NxN CU generally has N samples in a vertical direction and N samples in a horizontal direction, where N represents a non-negative integer value. The samples in a CU may be arranged in rows and columns. Furthermore, CUs do not necessarily have to have the same number of samples in the horizontal direction as in the vertical direction. For example, CUs may include NxM samples, where M is not necessarily equal to N.

[0060] The Video Encoder 200 encodes video data into CUs that represent residual and / or prediction information and other information. Prediction information indicates how the CU should be predicted to form a prediction block for the CU. Residual information generally represents sample-by-sample differences between samples of the CU before encoding and the prediction block.

[0061] To predict a CU, the video encoder 200 can generally form a prediction block for the CU through interprediction or intraprediction. A Petition 870250084327, dated 09 / 18 / 2025, pp. 160 / 292 27 / 126 Interprediction generally refers to predicting the CU from previously encoded image data, while intraprediction generally refers to predicting the CU from previously encoded data of the same image. To perform interprediction, the video encoder Video encoder 200 can generate the prediction block using one or more motion vectors. Video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, for example, in terms of differences between the CU and the reference block. Video encoder 200 can calculate a difference metric using a sum of absolute differences (SAD), a sum of squared differences (SSD), a mean absolute difference (MAD), mean squared differences (MSD), or other such difference calculations to determine whether a reference block closely matches the current CU. In some examples, the Video 200 encoder can predict the current CU using either one-way prediction or two-way prediction.

[0062] Some VVC examples also provide an affine motion compensation mode, which can be considered an interprediction mode. In affine motion compensation mode, the 200 video encoder can determine two or more motion vectors that represent non-translational motion, such as zooming in or out, rotation, perspective motion, or other types of irregular motion. Petition 870250084327, dated 09 / 18 / 2025, pp. 161 / 292 28 / 126

[0063] To perform intraprediction, the Video Encoder 200 can select an intraprediction mode to generate the prediction block. Some VVC examples provide sixty-seven intraprediction modes, including several directional modes, as well as planar mode and DC mode. In general, the Video Encoder 200 selects an intraprediction mode that describes neighboring samples for a current block (e.g., a CU block), from which samples of the current block are predicted. Such samples can generally be above, above and to the left, or to the left of the current block in the same image as the current block, assuming that the Video Encoder 200 encodes CTUs and CUs in raster scan order (left to right, top to bottom).

[0064] The Video Encoder 200 encodes data representing the prediction mode for a current block. For example, for interprediction modes, the Video Encoder 200 can encode data representing which of the various available interprediction modes are used, as well as motion information for the corresponding mode. For one-way or two-way interprediction, for example, the Video Encoder 200 can encode motion vectors using advanced motion vector prediction (AMVP) or fusion mode. The Video Encoder 200 can use similar modes to encode motion vectors for affine motion compensation mode.

[0065] AV1 includes two general techniques for encoding and decoding a block of video data encoding. The two general techniques are intraprediction (by Petition 870250084327, dated 09 / 18 / 2025, pp. 162 / 292 29 / 126 example, intra-frame prediction or spatial prediction) and inter-prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when predicting blocks of a current video data frame using an intraprediction mode, the Video Encoder 200 and Video Decoder 300 do not use video data from other video data frames. For most intraprediction modes, the Video Encoder 200 encodes blocks of a current frame based on the difference between sample values ​​in the current block and predicted values ​​generated from reference samples in the same frame. The Video Encoder 200 determines the predicted values ​​generated from the reference samples based on the intraprediction mode.

[0066] After prediction, such as intraprediction or interprediction of a block, the Video Encoder 200 can calculate residual data for the block. The residual data, such as a residual block, represents the sample by sample differences between the block and a prediction block for the block, formed using the corresponding prediction mode. The Video Encoder 200 can apply one or more transforms to the residual block to produce transformed data in a transform domain, rather than the sample domain. For example, the Video Encoder 200 can apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the Video Encoder 200 can apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform. Petition 870250084327, dated 09 / 18 / 2025, pp. 163 / 292 30 / 126 (MDNSST - mode-dependent non-separable secondary transform), a signal-dependent transform, a Karhunen-Loeve transform (KLT - Karhunen-Loeve transform), or similar. The 200 video encoder produces transform coefficients after the application of one or more transforms.

[0067] As mentioned above, after any transforms to produce transform coefficients, the Video 200 encoder can perform quantization of the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, providing additional compression. By performing the quantization process, the Video 200 encoder can reduce the bit depth associated with some or all of the transform coefficients. For example, the Video 200 encoder can round an n-bit value to an m-bit value during quantization, where n is greater than m. In some instances, to perform quantization, the Video 200 encoder can perform a right shift in the bit direction of the value to be quantized.

[0068] After quantization, the 200 video encoder can sweep the transform coefficients, producing a one-dimensional vector from the two-dimensional array that includes the quantized transform coefficients. The sweep can be designed to place higher-energy (and therefore lower-frequency) transform coefficients at the front of the vector and lower-energy (and therefore higher-frequency) transform coefficients at the back of the vector. In some examples, the Petition 870250084327, dated 09 / 18 / 2025, pp. 164 / 292 Video encoder 200 can use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector and then entropically encode the quantized transform coefficients of the vector. In other examples, video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form the one-dimensional vector, video encoder 200 can entropically encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). Video encoder 200 can also entropically encode the values ​​for syntax elements that describe metadata associated with the encoded video data for use by video decoder 300 in decoding the video data.

[0069] To perform CABAC, the video encoder 200 can assign a context in a context model to a symbol to be transmitted. The context can refer, for example, to the possibility of neighboring values ​​of the symbol being zero or non-zero values. Probability determination can be based on a context assigned to the symbol.

[0070] The 200 video encoder can additionally generate syntax data, such as block-based syntax data, image-based syntax data, and sequence-based syntax data, for the 300 video decoder, for example, in an image header, a block header, a slice header, or other syntax data, such as a sequence parameter set (SPS), a set of parameters of Petition 870250084327, dated 09 / 18 / 2025, pp. 165 / 292 32 / 126 image (PPS - picture parameter set) or a set of video parameters (VPS - video parameter set). The Video Decoder 300 can similarly decode this syntax data to determine how to decode corresponding video data.

[0071] In this way, video encoder 200 can generate a bitstream that includes encoded video data, for example, syntax elements that describe the partitioning of an image into blocks (e.g., CUs) and residual and / or prediction information for the blocks. Finally, video decoder 300 can receive the bitstream and decode the encoded video data.

[0072] In general, the 300 video decoder performs a reciprocal process with respect to that performed by the 200 video encoder to decode the encoded video data from the bitstream. For example, the 300 video decoder can decode values ​​for bitstream syntax elements using CABAC in a substantially similar, though reciprocal, manner to the 200 video encoder's CABAC encoding process. Syntax elements can define partition information for partitioning an image into CTUs, and partitioning each CTU according to a corresponding partition structure, such as a QTBT structure, to define CTU CUs. Syntax elements can additionally define residual and prediction information for blocks (e.g., CUs) of video data.

[0073] Residual information can be represented, for example, by quantized transform coefficients. The 300 video decoder can quantize Petition 870250084327, dated 09 / 18 / 2025, pp. 166 / 292 33 / 126 inversely transforms and inversely transforms the quantized transform coefficients of a block to reproduce a residual block for the block. The Video Decoder 300 uses a signaled prediction mode (intra- or interprediction) and related prediction information (e.g., motion information for interprediction) to form a prediction block for the block. The Video Decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. The Video Decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along block boundaries.

[0074] According to one or more techniques of the present disclosure, a video encoder, such as the Video Encoder 200 and the Video Decoder 300, can apply a subpel precision mode to generate a prediction block for a current block of video data. A syntax element indicates that the subpel precision mode is applied to the current block. The Video Encoder 200 can signal the syntax element in a bitstream. The Video Decoder 300 can obtain the syntax element from the bitstream. As part of applying the subpel precision mode, the video encoder can apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision. The video encoder can identify, within the array, a prediction block reference template, where the prediction block reference template is the best match for a current block template within the Petition 870250084327, dated 09 / 18 / 2025, pp. 167 / 292 34 / 126 arrangement. A template pattern defines a format for the prediction block reference template and the current block template. The video encoder can encode or decode the current block using the prediction block for the current block.

[0075] This disclosure may generally refer to the signaling of certain information, such as syntax elements. The term signaling may generally refer to the communication of values ​​for syntax elements and / or other data used to decode encoded video data. That is, the video encoder 200 may signal values ​​for syntax elements in the bitstream. In general, signaling refers to the generation of a value in the bitstream. As mentioned above, the source device 102 may transport the bitstream to the destination device 116 substantially in real time or in non-real time, as may occur during the storage of syntax elements in the storage device 112 for later retrieval by the destination device 116.

[0076] Figure 2 is a block diagram illustrating an example Video 200 encoder that can perform the techniques of this disclosure. Figure 2 is provided for explanatory purposes and should not be considered limiting of the techniques, as extensively exemplified and described in this disclosure. For explanatory purposes, this disclosure describes the Video 200 encoder according to VVC and HEVC techniques. However, the techniques of this disclosure can be performed by video encoding devices that are configured for other video encoding standards and Petition 870250084327, dated 09 / 18 / 2025, pp. 168 / 292 35 / 126 video encoding formats, such as AV1 and successors to the AV1 video encoding format.

[0077] In the example in Figure 2, the video encoder 200 includes video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, decoded picture buffer (DPB) 218, and entropy encoding unit 220. Any or all of the following are included: video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, the DPB 218 and the entropic encoding unit 220 can be implemented in one or more processors or in a set of processing circuits.For example, the video encoder 200 units can be implemented as one or more circuits or logic elements as part of the hardware circuitry or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 can include additional or alternative processing circuits or sets to perform these and other functions.

[0078] Video data memory 230 can store video data to be encoded by the components of video encoder 200. Video encoder 200 can receive the video data stored in Petition 870250084327, dated 09 / 18 / 2025, pp. 169 / 292 36 / 126 Video data memory 230 originating, for example, from video source 104 (Figure 1). DPB 218 can act as a reference image memory that stores reference video data for use in predicting subsequent video data by video encoder 200. Video data memory 230 and DPB 218 can be formed by any of several memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive random access memory (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or by separate memory devices.In several examples, the video data memory 230 may be inside the chip with other video encoder components 200, as illustrated, or outside the chip relative to the components.

[0079] In this disclosure, the reference to video data memory 230 should not be interpreted as being limited to the internal memory of video encoder 200, unless specifically described as such, or to the external memory of video encoder 200, unless specifically described as such. Instead, the reference to video data memory 230 should be understood as a reference memory that stores video data that video encoder 200 receives for encoding (e.g., video data for a current block to be encoded). Memory 106 in Figure 1 may also provide Petition 870250084327, dated 09 / 18 / 2025, pp. 170 / 292 37 / 126 temporary storage of outputs from the various units of the 200 video encoder.

[0080] The various units in Figure 2 are illustrated to aid in understanding the operations performed by the 200 video encoder. The units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide particular functionality and are predefined in the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and to provide flexible functionality in the operations that can be performed. For example, programmable circuits can execute software or firmware that causes the programmable circuits to operate in the manner defined by the software or firmware instructions.Fixed-function circuits can execute software instructions (for example, to receive parameters or to emit parameters), but the types of operations that fixed-function circuits perform are generally immutable. In some examples, one or more of the units may be distinct circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0081] The Video 200 encoder may include arithmetic logic units (ALUs), elementary function units (EFUs), digital circuits, analog circuits, and / or programmable cores, formed from programmable circuits. In examples where the Video 200 encoder operations are performed using software running Petition 870250084327, dated 09 / 18 / 2025, pp. 171 / 292 38 / 126 by the programmable circuits, memory 106 (Figure 1) can store the instructions (e.g., object code) of the software that the video encoder 200 receives and executes, or another memory in the video encoder 200 (not shown) can store these instructions.

[0082] Video data memory 230 is configured to store received video data. Video encoder 200 can retrieve an image from the video data in video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in video data memory 230 can be raw video data that needs to be encoded.

[0083] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, an intraprediction unit 226, and a template matching (TM) unit 228. The mode selection unit 202 may include additional functional units to perform video prediction according to other prediction modes. As examples, the mode selection unit 202 may include a palette unit, an intrablock copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, or similar units.

[0084] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and resulting rate distortion values ​​for those combinations. Encoding parameters may include partitioning of CTUs into CUs, Petition 870250084327, dated 09 / 18 / 2025, pp. 172 / 292 39 / 126 prediction modes for CUs, transform types for CU residual data, quantization parameters for CU residual data, etc. The mode selection unit 202 can ultimately select the combination of encoding parameters with rate distortion values ​​that are better than the other combinations tested.

[0085] The video encoder 200 can partition an image retrieved from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs in a slice. The mode selection unit 202 can partition a CTU of the image according to a tree structure, such as the MTT structure, the QTBT structure, the superblock structure, or the quadtree structure described above. As described above, the video encoder 200 can form one or more CUs from the partitioning of a CTU according to the tree structure. This CU can also be generally called a video block or a block.

[0086] In general, the mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, intraprediction unit 226, TM unit 228) to generate a prediction block for a current block (e.g., a current CU or, in HEVC, the overlapping portion of a PU and a TU). For interprediction of a current block, the motion estimation unit 222 can perform a motion search to identify one or more closely matching reference blocks in one or more reference images (e.g., one or more previously encoded images stored in DPB 218). In particular, the motion estimation unit 222 can Petition 870250084327, dated 09 / 18 / 2025, pp. 173 / 292 40 / 126 calculate a representative value of how similar a potential reference block is to the current block, for example, according to the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared differences (MSD), or similar methods. The motion estimation unit 222 can generally perform these calculations using sample-by-sample differences between the current block and the reference block being considered. The motion estimation unit 222 can identify a reference block that has a lower value resulting from these calculations, indicating a reference block that corresponds more closely to the current block.

[0087] The motion estimation unit 222 can form one or more motion vectors (MVs) that define the positions of the reference blocks in the reference images relative to the position of the current block in a current image. The motion estimation unit 222 can then provide the motion vectors to the motion compensation unit 224. For example, for one-way interprediction, the motion estimation unit 222 can provide a single motion vector, while for two-way interprediction, the motion estimation unit 222 can provide two motion vectors. The motion compensation unit 224 can then generate a prediction block using the motion vectors. For example, the motion compensation unit 224 can retrieve data from the reference block using the motion vector. As another example, if the motion vector has fractional sample precision, the unit of Petition 870250084327, dated 09 / 18 / 2025, pp. 174 / 292 The 41 / 126 motion compensation 224 can interpolate values ​​for the prediction block according to one or more interpolation filters. Furthermore, for bidirectional interprediction, the 224 motion compensation unit can retrieve data for two reference blocks identified by their respective motion vectors and can combine the retrieved data, for example, by calculating sample-by-sample averaging or weighted averaging.

[0088] When operating in accordance with the AV1 video encoding format, the motion estimation unit 222 and the motion compensation unit 224 can be configured to encode video data encoding blocks (e.g., both luma and chroma encoding blocks) using translational motion compensation, affine motion compensation, overlapped block motion compensation (OBMC), and / or composite inter / intraprediction.

[0089] As another example, for intraprediction or intraprediction coding, the 226 intraprediction unit can generate the prediction block from samples neighboring the current block. For example, for directional modes, the 226 intraprediction unit can generally mathematically combine the values ​​of neighboring samples and fill in these calculated values ​​in the direction defined in the current block to produce the prediction block. As another example, for DC mode, the 226 intraprediction unit can calculate an average of the neighboring samples for the current block and generate the prediction block to include this resulting average for each sample in the prediction block. Petition 870250084327, dated 09 / 18 / 2025, pp. 175 / 292 42 / 126

[0090] When operating in accordance with the AV1 video encoding format, the intraprediction unit 226 can be configured to encode video data encoding blocks (e.g., both luma and chroma encoding blocks) using directional intraprediction, non-directional intraprediction, recursive filter intraprediction, chroma-from-luma (CFL) prediction, intra-block copy (IBC), and / or color palette mode. The mode selection unit 202 can include additional functional units to perform video prediction according to other prediction modes.

[0091] The TM 228 unit can perform template matching to generate a prediction block for a current block. For example, the TM 228 unit can perform intra-template matching or inter-template matching prediction, as described elsewhere in this disclosure. In some examples, the TM 228 unit can use geometric partitioning modes with template matching. In some examples, the TM 228 unit can use template matching with IBC for IBC fusion mode and / or IBC AMVP mode. In some examples, the TM 228 unit is included in the motion estimation unit 222, the motion compensation unit 224, and the intraprediction unit 226.

[0092] According to one or more techniques of the present disclosure, the TM 228 unit can apply a subpel precision mode to generate a prediction block for a current block of video data. A syntax element can indicate that the subpel precision mode is applied to the current block. As part of applying the mode Petition 870250084327, dated 09 / 18 / 2025, pp. 176 / 292 43 / 126 subpel accuracy, the TM 228 unit can apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel accuracy. The TM 228 unit can identify, within the array, a prediction block reference template. The prediction block reference template can be the best match for a current block template within the array. A template pattern defines a format for the prediction block reference template and the current block template. The remaining video encoder 200 units (e.g., residual generation unit 204, transform processing unit 206, quantization unit 208, entropy coding unit 220, etc.) can encode the current block using the prediction block.

[0093] Mode selection unit 202 provides the prediction block to residual generation unit 204. Residual generation unit 204 receives a raw, unencoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates sample-by-sample differences between the current block and the prediction block. The resulting sample-by-sample differences define a residual block for the current block. In some examples, residual generation unit 204 may also determine differences between sample values ​​in the residual block to generate a residual block using residual differential pulse code modulation (RDPCM). In some examples, residual generation unit 204 may be formed using Petition 870250084327, dated 09 / 18 / 2025, pp. 177 / 292 44 / 126 of one or more subtractor circuits that perform binary subtraction.

[0094] In examples where the mode selection unit 202 partitions CUs into PUs, each PU can be associated with a luma prediction unit and corresponding chroma prediction units. The video encoder 200 and video decoder 300 can support PUs that have various sizes. As indicated above, the size of a CU can refer to the luma encoding block size of the CU, and the size of a PU can refer to the size of a luma prediction unit of the PU. Assuming that the size of a particular CU is 2Nx2N, the video encoder 200 can support PU sizes of 2Nx2N or NxN for intraprediction and symmetrical PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar, for interprediction. The 200 video encoder and 300 video decoder can also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for interprediction.

[0095] In instances where the 202 mode selection unit does not further partition a CU into PUs, each CU can be associated with a luma encoding block and corresponding chroma encoding blocks. As above, the size of a CU can refer to the size of the CU's luma encoding block. The 200 video encoder and the 300 video decoder can support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0096] For other video encoding techniques, such as intrablock copy mode encoding, affine mode encoding, and linear model (LM) mode encoding, as some examples, the unit of Petition 870250084327, dated 09 / 18 / 2025, pp. 178 / 292 45 / 126 Mode selection 202, via the respective units associated with the encoding techniques, generates a prediction block for the current block being encoded. In some examples, such as palette mode encoding, the mode selection unit 202 may not generate a prediction block and instead generates syntax elements that indicate how to reconstruct the block based on a selected palette. In these modes, the mode selection unit 202 may provide these syntax elements to the entropic encoding unit 220 to be encoded.

[0097] As described above, residual generation unit 204 receives the video data for the current block and the corresponding prediction block. Residual generation unit 204 then generates a residual block for the current block. To generate the residual block, residual generation unit 204 calculates the sample-by-sample differences between the prediction block and the current block.

[0098] The transform processing unit 206 applies one or more transforms to the residual block to generate a transform coefficient block (referred to in the present invention as a transform coefficient block). The transform processing unit 206 can apply multiple transforms to a residual block to form the transform coefficient block. For example, the transform processing unit 206 can apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to a residual block. In some examples, the transform processing unit 206 can perform multiple transforms on a Petition 870250084327, dated 09 / 18 / 2025, pp. 179 / 292 46 / 126 residual block, for example, a primary transform and a secondary transform, such as a rotational transform. In some examples, the 206 transform processing unit does not apply transforms to a residual block.

[0099] When operating in accordance with AV1, the transform processing unit 206 can apply one or more transforms to the residual block to generate a transform coefficient block (referred to in the present invention as a transform coefficient block). The transform processing unit 206 can apply multiple transforms to a residual block to form the transform coefficient block. For example, the transform processing unit 206 can apply a horizontal / vertical transform combination that may include a discrete cosine transform (DCT), an asymmetric discrete sine transform (ADST), an inverted ADST (e.g., an ADST in reverse order), and an identity transform (IDTX). When an identity transform is used, the transform is skipped in either the vertical or horizontal direction. In some examples, transform processing may be skipped.

[0100] The 208 quantization unit can quantize the transform coefficients in a transform coefficient block to produce a block of quantized transform coefficients. The 208 quantization unit can quantize the transform coefficients of a transform coefficient block according to a quantization parameter (QP) value. Petition 870250084327, dated 09 / 18 / 2025, pages 180 / 292 47 / 126 associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient blocks associated with the current block by adjusting the QP value associated with the CU. Quantization can introduce loss of information and therefore the quantized transform coefficients may have lower precision than the original transform coefficients produced by the transform processing unit 206.

[0101] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transforms to a block of quantized transform coefficients, respectively, to reconstruct a residual block from the block of transform coefficients. The reconstruction unit 214 can produce a reconstructed block corresponding to the current block (although potentially with some degree of distortion) based on the reconstructed residual block and a prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 can add samples from the reconstructed residual block to the corresponding samples from the prediction block generated by the mode selection unit 202 to produce the reconstructed block.

[0102] Filter unit 216 can perform one or more filter operations on reconstructed blocks. For example, filter unit 216 can perform unblocking operations to reduce blocking artifacts along the edges of the CUs. Filter unit 216 operations can be skipped in some instances. Petition 870250084327, dated 09 / 18 / 2025, pages 181 / 292 48 / 126

[0103] When operating in accordance with AV1, filter unit 216 can perform one or more filter operations on reconstructed blocks. For example, filter unit 216 can perform unblocking operations to reduce blocking artifacts along the edges of the CUs. In other examples, filter unit 216 can apply a constrained directional enhancement filter (CDEF), which can be applied after unblocking, and may include the application of non-separable non-linear low-pass directional filters, based on estimated edge directions. Filter unit 216 can also include a loop restoration filter, which is applied after the CDEF, and may include a separable symmetric normalized Wiener filter or a dual self-guided filter.

[0104] Video encoder 200 stores reconstructed blocks in DPB 218. For example, in instances where filter unit 216 operations are not performed, reconstruction unit 214 can store reconstructed blocks in DPB 218. In instances where filter unit 216 operations are performed, filter unit 216 can store the filtered reconstructed blocks in DPB 218. Motion estimation unit 222 and motion compensation unit 224 can retrieve a DPB 218 reference image, formed from the reconstructed (and potentially filtered) blocks, to interpredict blocks from subsequently encoded images. Additionally, intraprediction unit 226 can use reconstructed DPB 218 blocks from a current image to intrapredict other blocks in the current image. Petition 870250084327, dated 09 / 18 / 2025, pages 182 / 292 49 / 126

[0105] In general, the entropy coding unit 220 can entropically encode syntax elements received from other functional components of the video encoder 200. For example, the entropy coding unit 220 can entropically encode quantized transform coefficient blocks from the quantization unit 208. As another example, the entropy coding unit 220 can entropically encode prediction syntax elements (e.g., motion information for interprediction or intramode information for intraprediction) from the mode selection unit 202. The entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, which are another example of video data, in order to generate entropically encoded data.For example, the 220 entropic coding unit can perform a context-adaptive variable length coding (CAVLC) operation, a CABAC operation, a variable-to-variable (V2V) coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a probability interval partitioning entropy (PIPE) coding operation, an Exponential-Golomb coding operation, or another type of entropic coding operation on the data. In some examples, the 220 entropic coding unit... Petition 870250084327, dated 09 / 18 / 2025, pages 183 / 292 50 / 126 can operate in bypass mode, where syntax elements are not entropically encoded.

[0106] Video encoder 200 can output a bitstream that includes the entropically encoded syntax elements needed to reconstruct blocks of a slice or image. In particular, entropic encoding unit 220 can output the bitstream.

[0107] According to AV1, the 220 entropic encoding unit can be configured as an adaptive symbol-to-symbol multisymbol arithmetic encoder. A syntax element in AV1 includes an alphabet of N elements, and a context (e.g., probability model) includes a set of N probabilities. The 220 entropic encoding unit can store the probabilities as n-bit cumulative distribution functions (CDFs) (e.g., 15 bits). The 220 entropic encoding unit can perform recursive scaling, with an update factor based on alphabet size, to update the contexts.

[0108] The operations described above are described in relation to a block. This description should be understood as operations for a luma encoding block and / or for chroma encoding blocks. As described above, in some examples, the luma encoding block and the chroma encoding blocks are luma and chroma components of a CU. In some examples, the luma encoding block and the chroma encoding blocks are luma and chroma components of a PU. Petition 870250084327, dated 09 / 18 / 2025, pages 184 / 292 51 / 126

[0109] In some examples, the operations performed on a luma coding block do not need to be repeated for the chroma coding blocks. As an example, the operations to identify a motion vector (MV) and a reference image for a luma coding block do not need to be repeated to identify an MV and a reference image for the chroma blocks. Instead, the MV for the luma coding block can be scaled to determine the MV for the chroma blocks, and the reference image can be the same. As another example, the intraprediction process can be the same for the luma coding block and for the chroma coding blocks.

[0110] Video encoder 200 represents an example of a device configured to encode video data including a memory configured to store video data and one or more processing units implemented in a circuit set and configured to apply a subpel precision mode to generate a prediction block for a current block of video data, wherein a syntax element indicates that the subpel precision mode is applied to the current block and the application of the subpel precision mode comprises: applying an interpolation filter to samples from a reference region to generate an array of samples with full and subpel precision; and identifying, within the array, a prediction block reference template, wherein the prediction block reference template is the best match for a current block template within the array, wherein a template pattern defines a template format. Petition 870250084327, dated 09 / 18 / 2025, pages 185 / 292 52 / 126 reference the prediction block and the current block template; and encode the current block using the prediction block for the current block.

[0111] In some examples, the 200 video encoder represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units implemented in a circuit set and configured to determine a template pattern from a plurality of template patterns; identify, based on the determined template pattern, a prediction block for a current block of video data; and encode the current block using the prediction block, wherein a syntax element indicates the determined template pattern.

[0112] In some examples, the 200 video encoder represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units implemented in a circuit set and configured to generate template matching (TM) candidates, wherein each of the TM candidates is associated with a different reference block; generate, based on a combination of two or more of the TM candidates, a predictor for a current block of video data; and encode the current block using the predictor.

[0113] In some examples, the 200 video encoder represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units implemented in a circuit assembly and configured Petition 870250084327, dated 09 / 18 / 2025, pages 186 / 292 53 / 126 to generate TM candidates, where each TM candidate is associated with a different reference block; generate, based on a combination of two or more TM candidates, a predictor for a current block of video data; and encode the current block using the predictor.

[0114] Figure 3 is a block diagram illustrating an example 300 video decoder that can perform the techniques of this disclosure. Figure 3 is provided for explanatory purposes and is not limiting to the techniques as extensively exemplified and described in this disclosure. For explanatory purposes, this disclosure describes the 300 video decoder according to VVC and HEVC techniques. However, the techniques of this disclosure can be performed by video encoding devices that are configured for other video encoding standards.

[0115] In the example in Figure 3, the video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropic decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a DPB 314. Any or all of the CPB memory 320, the entropic decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 can be implemented in one or more processors or in a set of circuits. Petition 870250084327, dated 09 / 18 / 2025, pages 187 / 292 54 / 126 processing. For example, the 300 video decoder units can be implemented as one or more circuits or logic elements as part of the hardware circuitry or as part of a processor, ASIC, or FPGA. Furthermore, the 300 video decoder may include additional or alternative processing processors or circuitry to perform these and other functions.

[0116] The prediction processing unit 304 includes a motion compensation unit 316, an intraprediction unit 318, and a TM unit 322. The prediction processing unit 304 may include additional units to perform prediction according to other prediction modes. As examples, the prediction processing unit 304 may include a palette unit, an intrablock copy unit (which may be part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, or similar units. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0117] When operating in accordance with AV1, the motion compensation unit 316 can be configured to decode video data encoding blocks (e.g., both luma and chroma encoding blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite inter-intraprediction, as described above. The intraprediction unit 318 can be configured to decode video data encoding blocks (e.g., both luma and chroma encoding blocks) using Petition 870250084327, dated 09 / 18 / 2025, pages 188 / 292 55 / 126 directional intraprediction, non-directional intraprediction, recursive filter intraprediction, CFL, IBC and / or color palette mode, as described above.

[0118] The CPB 320 memory can store video data, such as an encoded video bitstream, to be decoded by the components of the video decoder 300. The video data stored in the CPB 320 memory can be obtained, for example, from computer-readable media 110 (Figure 1). The CPB 320 memory may include a CPB that stores encoded video data (e.g., syntax elements) of an encoded video bitstream. In addition, the CPB 320 memory may store video data beyond the syntax elements of an encoded image, such as temporary data representing outputs from the various units of the video decoder 300. The DPB 314 generally stores decoded images, which the video decoder 300 can output and / or use as reference video data when decoding subsequent data or images from the encoded video bitstream.The CPB 320 and DPB 314 memory can be comprised of any of several memory devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of memory devices. The CPB 320 and DPB 314 memory can be provided by the same memory device or by separate memory devices. In many instances, the CPB 320 memory may be on-chip with other components of the 300 video decoder or off-chip relative to those components.

[0119] Additionally or alternatively, in some examples, the video decoder 300 can recover encoded video data from memory 120 (Figure 1). That is, the Petition 870250084327, dated 09 / 18 / 2025, pp. 189 / 292 Memory 120, as discussed above, can store data with CPB 320 memory. Similarly, memory 120 can store instructions to be executed by video decoder 300, when some or all of the video decoder 300 functionalities are implemented in software to be executed by the video decoder 300 processing circuitry.

[0120] The various units shown in Figure 3 are illustrated to aid in understanding the operations performed by the 300 video decoder. The units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to Figure 2, fixed-function circuits refer to circuits that provide particular functionality and are predefined in the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and to provide flexible functionality in the operations that can be performed. For example, programmable circuits can execute software or firmware that causes the programmable circuits to operate in the manner defined by the software or firmware instructions.Fixed-function circuits can execute software instructions (for example, to receive parameters or to emit parameters), but the types of operations that fixed-function circuits perform are generally immutable. In some examples, one or more of the units may be distinct circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits. Petition 870250084327, dated 09 / 18 / 2025, pages 190 / 292 57 / 126

[0121] The 300 video decoder may include ALUs, EFUs, digital circuits, analog circuits, and / or programmable cores, formed from programmable circuits. In examples where the operations of the 300 video decoder are performed by software running on the programmable circuits, the on-chip memory or off-chip memory may store instructions (e.g., object code) from the software that the 300 video decoder receives and executes.

[0122] The entropic decoding unit 302 can receive encoded video data from the CPB and entropically decode the video data to reproduce syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0123] In general, the 300 video decoder reconstructs an image on a block-by-block basis. The 300 video decoder can perform a reconstruction operation on each block individually (where the block that is currently being reconstructed, i.e., decoded, can be called the current block).

[0124] The 302 entropic decoding unit can entropically decode syntax elements that define quantized transform coefficients from a block of quantized transform coefficients, as well as transform information such as a quantization parameter (QP) and / or mode indication(s). Petition 870250084327, dated 09 / 18 / 2025, pages 191 / 292 58 / 126 transform. The inverse quantization unit 306 can use the QP associated with the block of quantized transform coefficients to determine a degree of quantization and, similarly, a degree of inverse quantization for the inverse quantization unit 306 to apply. The inverse quantization unit 306 can, for example, perform a left shift operation in the bit direction to inversely quantize the quantized transform coefficients. The inverse quantization unit 306 can thus form a block of transform coefficients that includes transform coefficients.

[0125] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse directional transform, or another inverse transform to the transform coefficient block.

[0126] In addition, prediction processing unit 304 generates a prediction block according to the syntax elements of the prediction information that were entropically decoded by the entropic decoding unit 302. For example, if the syntax elements of the prediction information indicate that the current block is interpredicted, the motion compensation unit 316 Petition 870250084327, dated 09 / 18 / 2025, pages 192 / 292 59 / 126 can generate the prediction block. In this case, the prediction information syntax elements can indicate a reference image in DPB 314 from which a reference block can be retrieved, as well as a motion vector that identifies a location of the reference block in the reference image relative to the location of the current block in the current image. The motion compensation unit 316 can generally perform the interprediction process in a manner that is substantially similar to that described in relation to the motion compensation unit 224 (Figure 2).

[0127] As another example, if the prediction information syntax elements indicate that the current block is intrapredicted, the intraprediction unit 318 can generate the prediction block according to an intraprediction mode indicated by the prediction information syntax elements. Again, the intraprediction unit 318 can generally perform the intraprediction process in a manner that is substantially similar to that described in relation to the intraprediction unit 226 (Figure 2). The intraprediction unit 318 can retrieve data from samples neighboring the current DPB block 314.

[0128] The TM 322 unit can perform template matching to generate a prediction block for a current block. For example, the TM 322 unit can perform intra-template matching or inter-template matching prediction, as described elsewhere in this disclosure. In some examples, the TM 322 unit can use geometric partitioning modes with template matching. In some examples, the unit Petition 870250084327, dated 09 / 18 / 2025, pages 193 / 292 TM 322 60 / 126 can use template matching with IBC for IBC fusion mode and / or IBC AMVP mode. In some examples, the TM 322 unit is included in the motion compensation unit 316 and / or the intraprediction unit 318.

[0129] According to one or more techniques of the present disclosure, the TM 322 unit can apply a subpel precision mode to generate a prediction block for a current block of video data. A syntax element can indicate that the subpel precision mode is applied to the current block. As part of applying the subpel precision mode, the TM 322 unit can apply an interpolation filter to samples from a reference region to generate an array of samples with full and subpel precision. The TM 322 unit can identify, within the array, a prediction block reference template. The prediction block reference template can be the best match for a current block template within the array. A template pattern defines a format for the prediction block reference template and the current block template. The remaining video decoder units 300 (e.g., reconstruction unit 310, filter unit) 312, etc.) can decode the current block using the prediction block.

[0130] Reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block. Petition 870250084327, dated 09 / 18 / 2025, pages 194 / 292 61 / 126

[0131] Filter unit 312 can perform one or more filter operations on reconstructed blocks. For example, filter unit 312 can perform unblocking operations to reduce blocking artifacts along the edges of reconstructed blocks. Filter unit 312 operations are not necessarily performed in all instances.

[0132] Video decoder 300 can store reconstructed blocks in DPB 314. For example, in instances where filter unit 312 operations are not performed, reconstruction unit 310 can store reconstructed blocks in DPB 314. In instances where filter unit 312 operations are performed, filter unit 312 can store filtered reconstructed blocks in DPB 314. As discussed above, DPB 314 can provide reference information, such as samples of a current image for intraprediction and previously decoded images for subsequent motion compensation, to the prediction processing unit 304. In addition, video decoder 300 can output decoded images (e.g., decoded video) from DPB 314 for subsequent presentation on a display device, such as display device 118 in Figure 1.

[0133] Thus, the 300 video decoder represents an example of a video decoding device including a memory configured to store video data and one or more processing units implemented in a circuit set and configured to apply a subpel precision mode to generate a prediction block for a current block of video data, in Petition 870250084327, dated 09 / 18 / 2025, pages 195 / 292 62 / 126 that a syntax element indicates that the subpel precision mode is applied to the current block and the application of the subpel precision mode comprises: applying an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision; and identifying, within the array, a prediction block reference template, wherein the prediction block reference template is the best match for a current block template within the array, wherein a template pattern defines a format of the prediction block reference template and the current block template; and decoding the current block using the prediction block for the current block.

[0134] In this way, the 300 video decoder represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units implemented in a circuit set and configured to determine a template pattern from a plurality of template patterns; identify, based on the determined template pattern, a prediction block for a current block of video data; and decode the current block using the prediction block, where a syntax element indicates the determined template pattern.

[0135] In some examples, the 300 video decoder represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units implemented in a circuit set and configured to generate matching candidates. Petition 870250084327, dated 09 / 18 / 2025, pages 196 / 292 63 / 126 template (TM), in which each of the TM candidates is associated with a different reference block; generate, based on a combination of two or more of the TM candidates, a predictor for a current block of video data; and decode the current block using the predictor.

[0136] In some examples, the 300 video decoder represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units implemented in a circuit set and configured to generate template matching (TM) candidates, wherein each of the TM candidates is associated with a different reference block; generate, based on a combination of two or more of the TM candidates, a predictor for a current block of video data; and decode the current block using the predictor.

[0137] Figure 4 is a conceptual diagram illustrating an example of an intratemplate match search area. Intratemplate match prediction (Intra TMP) is a special intraprediction mode that copies the best prediction block from the reconstructed portion of the current frame whose L-shaped template matches the current template. For a predefined search range, video encoder 200 searches for the template most similar to the current template in a reconstructed portion of the current frame and uses the matching block as a prediction block. Video encoder 200 then signals the use of this mode, and the same prediction operation is performed on the decoder side by video decoder 300. Petition 870250084327, dated 09 / 18 / 2025, pages 197 / 292 64 / 126

[0138] A prediction signal (e.g., matching block 400) is generated by correlating a current template 402 (which is an L-shaped causal neighbor) of a current block 404 with another block (reference template 406) in a predefined search area. In the example in Figure 4, the predefined search area can be one of the following: R1: Current CTU R2: Upper left CTU R3: CTU above R4: Left CTU The sum of absolute differences (SAD) is used as a cost function.

[0139] Within each region, the 300 video decoder searches for the 406 reference template that has less SAD compared to the current 402 template and uses its corresponding block (match block 400) as a prediction block.

[0140] The dimensions of all regions (SearchRange_w, SearchRange_h) are defined proportionally to the block dimension (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * Equation BlkW (1) SearchRange_h = a * Equation BlkH (2) In the equations above, 'a' is a constant that controls the gain / complexity trade-off. In practice, 'a' can be equal to 5.

[0141] The intratemplate matching tool is enabled for CUs with a width and height of 64 or less. This maximum CU size for intratemplate matching is configurable. The mode of Petition 870250084327, dated 09 / 18 / 2025, pp. 198 / 292 65 / 126 intratemplate match prediction is signaled at the CU level via a dedicated flag when decoder-side intra-mode derivation (DIMD) is not used for the current CU.

[0142] Intertemplate matching (InterTM) is a decoder-side MV derivation method that refines the motion information of the current CU by finding the closest match between a template (e.g., neighboring blocks above and / or to the left of the current CU) in the current image and a block (i.e., the same size as the template) in a reference image. The Figure is a conceptual diagram illustrating example template matching performed in a search area around an initial motion vector (MV). In the example in Figure 5, a video encoder is encoding or decoding a current CU 500 into a current frame 502. The current templates 504 of the current CU 500 are above and to the left of the current CU. 500. An initial MV 506 of the current CU 500 indicates a location on a reference frame of 508.

[0143] As illustrated in Figure 5, a video encoder (e.g., video encoder 200 or video decoder 300) can search for a better MV (i.e., a motion vector better than the initial MV 506) around a location indicated by the initial motion vector 506 of the current CU 500 within a search range of [-8, +8]-pel 510. The template matching method in JVET-J0021 can be used with the following modifications: the search step size is determined based on the adaptive motion vector resolution (AMVR) mode, and InterTM can be performed in Petition 870250084327, dated 09 / 18 / 2025, pp. 199 / 292 66 / 126 cascade with bilateral matching process in fusion modes.

[0144] In AMVP mode, a motion vector predictor (MVP) candidate is determined based on template matching error to select the one that achieves the minimum difference between the current block template and the reference block template, and then InterTM is performed only for that particular MVP candidate for MV refinement. InterTM refines this MVP candidate, starting from the motion vector difference (MVD) accuracy of full pel (or 4-pel for 4-pel AMVR mode) in a search range of [-8, +8]-pel using iterative diamond search. The AMVP candidate can be further refined using cross-search with MVD accuracy of full pel (or 4-pel for 4-pel AMVR mode), sequentially followed by half pel and quarter pel, depending on the AMVR mode, as specified in Table 1 below.This search process ensures that the MVP candidate still maintains the same MV accuracy, as indicated by the AMVR mode after the TM process. In the search process, if the difference between the previous minimum cost and the current minimum cost in the current iteration is less than a threshold equal to the block area, the search process will be terminated. Table 1 - AMVR search patterns and fusion mode with AMVR Search pattern AMVR mode Fusion mode 4-pel Full pel Half pel Quarter pel AltIF=0 AltIF=1 Petition 870250084327, dated 09 / 18 / 2025, pages 200 / 292 67 / 126 4-pel diamond v 4-pel cross v Full-pel diamond vvvvv Full-pel cross vvvvv Half-pel cross vvvv Quarter-pel cross vv 1 / 8-pel cross v

[0145] In fusion mode, a similar search method is applied to the fusion candidate indicated by the fusion index. As Table 1 shows, InterTM can perform the entire path up to MVD accuracy of 1 / 8 pel or skip those beyond MVD accuracy of half pel, depending on whether the alternative interpolation filter (which is used when AMVR is in half pel mode) is used according to the fusion motion information. When TM mode is enabled, intratemplate matching can function as a standalone process or an MV refinement process. Petition 870250084327, dated 09 / 18 / 2025, pages 201 / 292 68 / 126 extra between bilateral matching processes (BM bilateral matching) based on blocks and sub-blocks, depending on whether the BM can be enabled or not according to its enabling condition check.

[0146] Adaptive reordering of merge candidates with template matching (ARMC-TM) is now discussed. Merge candidates are adaptively reordered with template matching (TM). The reordering method is applied to regular merge mode, TM merge mode, and affine merge mode (excluding the SbTMVP candidate). For TM merge mode, merge candidates are reordered before the refinement process.

[0147] An initial list of merge candidates is first constructed according to a specific check order, such as spatial, temporal motion vector predictors (TMVPs), non-adjacent, historical motion vector predictors (HMVPs), pairwise, and virtual merge candidates. HMVPs are motion vector predictors of previously encoded blocks that are stored in the first-in, first-out buffer. Then, the candidates in the initial list are divided into several subgroups. For template matching (TM) merge mode, decoder motion vector refinement (DMVR) adaptive merge mode, each merge candidate in the initial list is first refined using multi-pass TM / DMVR. The merge candidates in each subgroup are reordered to generate a reordered merge candidate list, and the Petition 870250084327, dated 09 / 18 / 2025, pages 202 / 292 69 / 126 reordering occurs according to cost values ​​based on template matching. The index of the fusion candidate selected in the reordered fusion candidate list is flagged to video decoder 300. For simplification, fusion candidates in the last, but not the first, subgroup are not reordered. All zero candidates from the ARMC reordering process are excluded during the construction of the fusion motion vector candidate list. The subgroup size is set to 5 for regular fusion mode and for TM fusion mode. The subgroup size is set to 3 for affine fusion mode.

[0148] The template matching cost of a merge candidate during the reordering process is measured by the SAD between samples of a template from the current block and its corresponding reference samples. The template comprises a set of reconstructed samples neighboring the current block. The template's reference samples are located by the merge candidate's movement information.

[0149] Figure 6 is a conceptual diagram illustrating an example of a template and reference samples of the template in reference images. In the example in Figure 6, a video encoder (e.g., video encoder 200 or video decoder 300) is encoding a current block 600 from a current image 602. When a fusion candidate uses bidirectional prediction, the reference samples (e.g., reference block 604A in reference image 606A in reference list 0 and reference block 604B in reference image 606B in reference list 1) Petition 870250084327, dated 09 / 18 / 2025, pages 203 / 292 70 / 126 reference 1) of the 608A, 608B templates of the fusion candidate are also generated by biprediction, as shown in Figure 6. The fusion candidate includes motion vectors 610A indicating positions in the reference image 606A and motion vectors 610B indicating positions in the reference image 606B.

[0150] When multi-pass DMVR is used to derive the refined motion for the initial merge candidate list, only the first pass (i.e., PU level) of multi-pass DMVR is applied in the reordering. When template matching is used to derive the refined motion, the template size is set to 1. Only the template above or to the left is used during TM motion refinement when the block is flat with a block width 2 times greater than the height or narrow with a height 2 times greater than the width. TM is extended to achieve MVD accuracy of 1 / 16-pel. The first four merge candidates are reordered with the refined motion in TM merge mode.

[0151] For subblock-based merge candidates with subblock size equal to Wsub χ Hsub, the template above comprises several subtemplates with size Wsub χ 1 and the template on the left comprises several subtemplates with size 1 χ Hsub. Figure 7 is a conceptual diagram illustrating an example template and reference samples of the template for blocks with subblock movement using the movement information of 700 subblocks for a current block 702 in a current image 704. As shown in Figure 7, the movement information of 700 subblocks in the first row and first column Petition 870250084327, dated 09 / 18 / 2025, pages 204 / 292 71 / 126 of the current block 702 are used to derive the reference samples 708 from each subtemplate in a reference image 706.

[0152] The reordering criteria are now discussed. In the reordering process, a candidate is considered redundant if the cost difference between a candidate and its predecessor is less than a lambda value, for example, |D1-D2| < λ, where D1 and D2 are the costs obtained during the first ARMC ordering and λ is the Lagrange parameter used in the RD criterion on the encoder side. An example algorithm is defined as follows: - Determine the minimum cost difference between a candidate and their predecessor among all candidates on the list. If the minimum cost difference is greater than or equal to λ, the list is considered sufficiently diverse and reordering is stopped. If this minimum cost difference is less than λ, the candidate is considered redundant and is moved to an additional position on the list. This additional position is the first where the candidate is sufficiently diversified compared to its predecessor. The algorithm is interrupted after a finite number of iterations (if the minimum cost difference is not less than λ).

[0153] This algorithm is applied to Regular, TM, BM, and Affine merging modes. A similar algorithm is applied to Merge (merge with motion vector difference (MMVD)) and Sign MVD prediction methods, which also use ARMC for reordering. Petition 870250084327, dated 09 / 18 / 2025, pages 205 / 292 72 / 126

[0154] The value of λ is defined equal to the λ of the rate distortion criterion used to select the best candidate for merging on the encoder side for low-delay configuration and to the λ value corresponding to another QP for Random Access configuration. A set of λ values ​​corresponding to each signaled QP offset is provided in the SPS or in the Slice Header for QP offsets that are not present in the SPS.

[0155] An extension to the AMVP modes is now discussed. The ARMC design is also applicable to the AMVP mode, where AMVP candidates are reordered according to TM cost. For template matching for the advanced motion vector prediction mode (TM-AMVP), an initial list of AMVP candidates is built, followed by a refinement from the TM to build a refined list of AMVP candidates. Additionally, an MVP candidate with a TM cost greater than a threshold, which is equal to five times the cost of the first MVP candidate, is skipped. Note that when engagement around motion compensation is enabled, the MV candidate should be cut with engagement around displacement taken into account.

[0156] Geometric partitioning mode (GPM) with template matching (TM) is now discussed. Template matching is applied to GPM. When GPM mode is enabled for a CU, a CU-level flag is signaled indicating whether TM is applied to both geometric partitions. Motion information for each geometric partition is refined using TM. When TM is chosen, a template is Petition 870250084327, dated 09 / 18 / 2025, pages 206 / 292 73 / 126 constructed using neighboring samples to the left, above, or to the left and above according to the partition angle, as shown in Table 2 below. The motion is then refined by minimizing the difference between the current template and the template in the reference image using the same search pattern as the merge mode with the half-pel interpolation filter disabled. Table 2. Template for the 1st and 2nd geometric partitions, where A represents the use of the samples above, L represents the use of the samples to the left, and L+A represents the use of both samples to the left and above. Partition angle 0 2 3 4 5 8 11 12 13 14 1st partition AAAA L+A L+A L+A L+AAA 2nd partition L+A L+A L+ALLLL L+A L+A L+A Partition angle 16 18 19 20 21 24 27 28 29 30 1a partition AAAA L+A L+A L+A L+AAA 2nd partition L+A L+A L+ALLLL L+A L+A L+A

[0157] A GPM candidate list is constructed as follows: 1. The MV candidates from List-0 and the interleaved MV candidates from List-1 are derived directly from the regular list of merge candidates, where the MV candidates from List-0 have higher priority than the MV candidates from List-1. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates. Petition 870250084327, dated 09 / 18 / 2025, pp. 207 / 292 74 / 126 2. The MV candidates from List-1 and the interleaved MV candidates from List-0 are additionally derived directly from the regular list of merge candidates, where the MV candidates from List-1 have higher priority than the MV candidates from List-0. The same pruning method with the adaptive threshold is also applied to remove redundant MV candidates. 3. MV zero candidates are filled until the GPM candidate list is complete.

[0158] GPM-MMVD and GPM-TM are enabled exclusively for a GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are set to false (that is, GPM-MMVD is disabled for two GPM partitions), the GPMTM flag indicates whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is set to true), the value of the GPM-TM flag is inferred to be false.

[0159] Now IBC with model matching is discussed. Model matching is used in IBC for IBC merge mode and IBC AMVP mode. The IBC-TM merge list is modified compared to that used by regular IBC merge mode, so that candidates are selected according to a pruning method with a movement distance between candidates as in regular TM merge mode. The final zero motion padding is replaced by left (-W, 0), top (0, -H) and top left (-W, -H) motion vectors, where W is the width and H the height of the current CU. Petition 870250084327, dated 09 / 18 / 2025, pp. 208 / 292 75 / 126

[0160] In IBC-TM fusion mode, selected candidates are refined using the template matching method before the rate-distortion optimization process or the decoding process. IBC-TM fusion mode was placed in competition with regular IBC fusion mode and a TM fusion flag is signaled.

[0161] In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM merge list. Each of these 3 selected candidates is refined using the template matching method and ranked according to the resulting template matching cost. Only the top 2 are then considered in the motion estimation process as usual.

[0162] Figure 8 is a conceptual diagram illustrating IBC reference regions depending on the position of a current CU 800. Template matching refinement for both IBC-TM and AMVP fusion modes is quite straightforward since IBC motion vectors are restricted (i) to being integers and (ii) within a reference region, as shown in Figure 10. Thus, in IBC-TM fusion mode, all refinements are performed with integer precision, and in IBC-TM AMVP mode, they are performed with integer precision or 4 pel, depending on the AMVR value. Such refinement accesses only samples without interpolation. In both cases, the refined motion vectors and the template used in each refinement step must respect the reference region restriction.

[0163] To improve the encoding efficiency of template matching, instead of using just one Petition 870250084327, dated 09 / 18 / 2025, pp. 209 / 292 76 / 126 standard and template process in TM, template matching can use different template types, store more candidates and apply merging to combine these different candidates that are found by the different methods.

[0164] For simplicity of description, unless otherwise indicated, the TM mentioned in this disclosure may refer to intratemplate matching, intertemplate matching, adaptive reordering of merge candidates with template matching (ARMC-TM), or IBC template matching. Examples and techniques disclosed may be used individually or in any combination.

[0165] Examples relating to the precision of multiple templates and subpelling in TM are now discussed. In a first example, TM has a set of template patterns comprising multiple template matching patterns, and a syntax element indicates a template pattern used in the template matching process. This syntax element can be signaled in a bitstream by video encoder 200 and obtained from the bitstream by video decoder 300. When template matching is used in the current block and a template matching mode is signaled, the template pattern corresponding to that mode is used in the template matching process. Figure 9 is a conceptual block diagram illustrating different template types according to the techniques of the present disclosure. Each (or a subset of) the template types 900A to 900G (i.e., template patterns) shown in Figure 9 can be used in this example. Petition 870250084327, dated 09 / 18 / 2025, pp. 210 / 292 77 / 126

[0166] Thus, in some examples, a video encoder (for example, video encoder 200 or video decoder 300) can determine a template pattern from among a plurality of template patterns. The video encoder can identify, based on the determined template pattern, a prediction block for a current block of video data. The video encoder can encode or decode the current block using the prediction block. A syntax element can indicate the determined template pattern.As shown in Figure 9, the plurality of templates can include two or more of the following: a set of reference samples above and to the left of the current block, a set of reference samples to the left of the current block, a set of reference samples above the current block, a subset of reference samples to the left of the current block, a subset of reference samples above the current block, a set of reference samples to the left and not adjacent to the current block, and a set of reference samples above and not adjacent to the current block. In some examples, the plurality of template patterns includes a base template pattern. The video encoder can determine the template pattern from among the base template pattern and one or more of the template patterns that are within an area of ​​the base template pattern.

[0167] In a second example, as a simplified method of the first example mentioned above, the set of template patterns is composed of three template matching patterns shown in Figure 9: (Pattern 1) uses neighboring samples both above and to the left, (Pattern 2) uses only neighboring samples above and (Pattern 3) Petition 870250084327, dated 09 / 18 / 2025, pages 211 / 292 78 / 126 uses only left-neighbor samples. The matching cost of template 1 matching pattern is derived from the cost of template 2 matching pattern and the cost of template 3 matching pattern.

[0168] In a third example, as a simplified method of the first example mentioned earlier, the TM has a base template pattern S. The additional template patterns 1...n are all within the region of S. For example, if type 900A in Figure 9 is the base template pattern, types 900B, 900C, 900D, and 900E could be other template patterns, since they are all within region (a). In another example, as a simplified method of the first example mentioned earlier, the base template pattern could be type 900A in Figure 9.

[0169] In another example, TM has a subpel prediction mode that uses an interpolation filter to derive a block vector (BV) of subpel accuracy. Figure 10 is a conceptual diagram illustrating an exemplary subpel accuracy prediction mode according to the techniques of the present disclosure. In the example in Figure 10, circles 1000A to 1000I correspond to full pel samples. Circles 1002A to 1002H correspond to half pel samples.

[0170] In another example, there is a flag to indicate whether the subpel precision mode is applied to the current block or not. In another example, the subpel precision mode flag is inherited from a neighboring block. In another example, the subpel precision mode flag is flagged only when the non-merge TM mode is selected in the current block. Petition 870250084327, dated 09 / 18 / 2025, pages 212 / 292 79 / 126 Otherwise, the subpel precision mode flag is considered to be 0.

[0171] In another example, the subpel precision mode flag is signaled only when the base template pattern is selected in the current block. Otherwise, the subpel precision mode flag is considered to be 0. In other words, a video encoder can determine a template pattern from among a plurality of template patterns, where the plurality of template patterns includes a base template pattern. In this example, the video encoder can determine the template pattern from among the base template pattern and one or more of the template patterns that are within an area of ​​the base template pattern. The subpel precision mode flag is signaled only when the base template pattern is the determined template pattern (e.g., selected).

[0172] In another example, the subpel precision mode flag is signaled only when the TM candidate index is equal to 0 in the current block. Otherwise, the subpel precision mode flag is considered to be 0. In other words, the syntax element that indicates whether subpel precision mode is applied to the current block is signaled based on the determined template pattern being the base template pattern and not any of the other template patterns.

[0173] In another example, the subpel precision mode flag is signaled only when the base template pattern is selected and the TM candidate index is equal to 0 in the current block. Otherwise, the subpel precision mode flag is considered to be 0. Thus, Petition 870250084327, dated 09 / 18 / 2025, pages 213 / 292 In this example, a video encoder can generate a list of template matching candidates. Each of the template matching candidates is associated with a different motion vector. In this example, the video encoder can select a template matching candidate from the list of template matching candidates. Based on the selected template matching candidate having a template matching candidate index equal to 0 and the determined template pattern being the base template pattern and not any of the other template patterns, a syntax element indicates whether subpel precision mode is applied to the current block.

[0174] In another example, the subpel precision mode flag is signaled only when the TM candidate index is less than M, while the number of TM candidates is N and M < N in the current block. Otherwise, the subpel precision mode flag is considered to be 0. In other words, a video encoder can generate a list of template matching candidates, where each of the template matching candidates is associated with a different motion vector. The video encoder can select a template matching candidate from the list of template matching candidates. Based on the selected template matching candidate having a template matching candidate index less than a predetermined value (M), a syntax element indicates whether the subpel precision mode is applied to the current block. The predetermined value is less than a number of template matching candidates in the list (N). Petition 870250084327, dated 09 / 18 / 2025, pp. 214 / 292 81 / 126

[0175] In another example, the subpel precision mode flag is signaled only when the base template pattern is selected and the TM candidate index is less than M, while the number of TM candidates is N and M < N in the current block. Otherwise, the subpel precision mode flag is considered to be 0. In other words, a video encoder can generate a list of template matching candidates. Each of the template matching candidates is associated with a different motion vector. The video encoder can select a template matching candidate from the list of template matching candidates.Based on the selected template matching candidate having a template matching candidate index less than a predetermined value (M) and the determined template pattern being the base template pattern and not any of the other template patterns, a syntax element indicates whether subpel precision mode is applied to the current block. The predetermined value (M) is less than a number of template matching candidates (N) in the list.

[0176] In another example, if subpel precision mode is used, the available signed TM candidates are reduced to M, while the original available signed TM candidates where subpel precision mode is not used are N and M < N in the current block. In other words, the video encoder can generate a list of template matching candidates, where each of the template matching candidates is associated with a different motion vector. The video encoder can select a template matching candidate from Petition 870250084327, dated 09 / 18 / 2025, pages 215 / 292 82 / 126 list of candidate template matchers. Based on the subpel precision mode being used to encode or decode the current block, the list of candidate template matchers is reduced compared to when subpel precision mode is not used.

[0177] In another example, if subpel precision mode is used, the TM candidate index is not signaled and is considered to be 0. In other words, the video encoder can generate a list of template matching candidates. Each of the template matching candidates is associated with a different motion vector. The video encoder can select a template matching candidate from the list of template matching candidates. Based on the subpel precision mode being used to encode or decode the current block, a template matching candidate index of the selected template matching candidate is not signaled.

[0178] In another example, if subpel precision mode is used, the template pattern is not signaled and is considered to be the base template pattern. In other words, based on the subpel precision mode being used to encode or decode the current block, a syntax element indicating the selected template pattern is not signaled.

[0179] In another example, if the subpel precision mode is used, the TM fusion mode flag is not signaled and the TM fusion mode is considered as the non-fusion mode. In other words, based on the subpel precision model being used to encode or decode Petition 870250084327, dated 09 / 18 / 2025, pages 216 / 292 83 / 126 the current block, a template matching merge mode flag is not flagged, where the template matching merge mode flag indicates whether TM merge mode is applied or not.

[0180] In another example, the syntax of the matching pattern is signaled at the CU, PU, ​​CTU, slice, or image level. In other words, an index of the given template pattern is signaled at one of: a CU level, a PU level, a CTU level, a slice level, or an image level.

[0181] In some examples, the subpel precision is half a pixel. In some examples, the subpel precision is a quarter of a pixel. In some examples, another flag signaled in a bitstream or obtained from the bitstream indicates whether the subpel precision is half a pixel or a quarter of a pixel. In other words, a syntax element is signaled (for example, by video encoder 200) that specifies whether the subpel precision is half a pixel or a quarter of a pixel.

[0182] Now, examples related to merge mode reordering and weight selection are discussed. In one example, the prediction block for the current block can be generated from a combination of k TM candidates, where k is greater than 1. The k TM candidates can be arbitrarily selected from the available candidate lists of different template patterns or even the same pattern. In one example, a merge of 2 candidates can combine the lowest and third lowest cost candidates of template pattern A. In another example, a merge of 3 candidates can combine the second lowest candidate of the pattern Petition 870250084327, dated 09 / 18 / 2025, pages 217 / 292 84 / 126 of template A, the second smallest candidate of template B, and the smallest candidate of template C.

[0183] In another example, a video encoder can generate a prediction block (e.g., a prediction block for the current block, a prediction block for a fusion mode, etc.) from a combination of k TM candidates. A prediction block generated from a combination of TM candidates can be more similar to the current block than a prediction block generated from a single TM candidate, which can result in greater encoding efficiency. In one example, a linear combination can be used. In other words, a video encoder can generate a prediction block as a linear combination of matching samples in the reference blocks associated with two or more TM candidates. Matching samples are samples in matching positions within the reference blocks. The 0 ak candidates can be selected from 0 to N candidates from the template matching. The combination can be formulated as follows: P(x,y) = (Wq * Po (X,y) + Wi * Pi (X,y) + · + Wk* Pk(X,y)>) Oo + Wi + ···+ wk) Equation (3) where Po... Pk are the k selected candidates derived from the TM process and wo... wk are the weighting used for each candidate. Thus, in some examples, a video encoder can generate a plurality of template matching (TM) candidates. Each of the TM candidates is associated with a respective prediction block and a respective motion vector predictor. As part of the generation of the plurality of TM candidates, for each of the Petition 870250084327, dated 09 / 18 / 2025, pages 218 / 292 For 85 / 126 motion vector (TM) candidates, the video encoder can apply an interpolation filter to samples from a reference region indicated by the motion vector predictor associated with the TM candidate to generate a respective sample array with full-pel and sub-pel accuracy. The video encoder can identify, within the respective array, a respective reference template for the respective prediction block. The respective reference template for the respective prediction block is the best match for the current block template within the respective array. The video encoder can generate, based on a combination of two or more of the respective prediction blocks associated with two or more of the TM candidates, the prediction block for the current block. The video encoder can generate the prediction block for the current block as a linear combination of matching samples in the prediction blocks associated with two or more TM candidates from the plurality of TM candidates.

[0184] In some examples, the combined weight wk can be derived based on the template matching cost. For example, for each of two or more TM candidates, the video encoder can calculate a template cost for the TM candidate and can generate the prediction block for the current block as a weighted average of the samples across the two or more template candidates. The weights used in the weighted average are based on the template costs for the TM candidates. In one example, the weights are the multiplicative inverse of the template matching cost for that candidate k. The weights are multiplicative inverses of the template matching costs for the TM candidates. Petition 870250084327, dated 09 / 18 / 2025, pages 219 / 292 86 / 126

[0185] In some examples, the weights are derived based on the sum of absolute differences (SAD), mean squared error (MSE), MSE minimization, or block vector (BV). In other words, the video encoder can derive weights based on one or more of an SAD, an MSE, an MSE minimization, or a BV from the two or more template candidates. The video encoder can generate the prediction block for the current block as a weighted average of the corresponding samples in the two or more template candidates. The weights used in the weighted average are based on the template costs for the template candidates.

[0186] In another example, the candidates used in combination can be selected based on the SAD, MSE, or block vector (BV) of the available candidates. In other words, the video encoder can derive an SAD, an MSE, or a BV from the TM candidates. The video encoder can select the two or more TM candidates based on one or more of the SAD, MSE, or BV of the TM candidates.

[0187] In another example, a flag indicates the method for deriving weights from a set of derivation methods. In other words, the video encoder can determine a weight derivation method from a plurality of weight derivation methods. For example, video encoder 200 can perform a rate-distortion optimization process to determine the weight derivation method from a plurality of weight derivation methods. Video decoder 300 can determine the weight derivation method based on a signaled syntax element in a bitstream. The encoder Petition 870250084327, dated 09 / 18 / 2025, pages 220 / 292 87 / 126 video can apply the weight derivation method to derive weights from two or more TM candidates. As part of generating a prediction block (e.g., a prediction block for a current block), the video encoder can generate the prediction block as a weighted average of the corresponding samples in the prediction blocks associated with the two or more TM candidates using the weights of the two or more TM candidates. A syntax element (e.g., flag) indicates the weight derivation method. In another example, a flag indicates whether the weights are derived based on template cost or based on MSE minimization. Video encoder 200 can flag the syntax elements in a bitstream, and video decoder 300 can obtain the syntax elements from the bitstream.Determining the weight derivation method, therefore, can reduce the differences between the prediction block for the actual block and the actual block itself, which can improve coding efficiency.

[0188] In another example, there are multiple combination modes that can be selected and signaled, a merge list is built to indicate the merge candidates for the current block. The merge list includes a list of merge modes, each of which is associated with a different combination of two or more TM candidates. In other words, a video encoder can select and signal multiple combination modes, and the video encoder can build a merge list to indicate the merge candidates for the current block.

[0189] In another example, the number of signaled fusion modes is M and the number of fusion modes Petition 870250084327, dated 09 / 18 / 2025, pages 221 / 292 The number of available 88 / 126 modes is N, where M < N. To select the M modes from the N modes, the combination will be applied to the template of all candidates, and the combined template cost is calculated for each of the N modes. The M modes with the minimum combined template cost will be selected. Using a reduced number of available fusion modes can reduce the amount of work the 300 video decoder may need to perform to determine the fusion mode.

[0190] Thus, in this example, the number of signaled fusion modes is M and the number of available fusion modes is N where M < N. The video encoder can select M fusion modes from N fusion modes. To select the M fusion modes, the video encoder can calculate a combined template cost for each of the N fusion modes. Each of the N fusion modes is a combination of two TM candidates. The video encoder can select the M modes with the minimum combined template cost. The video encoder can select two or more TM candidates from among the M fusion modes, where the two or more TM candidates are used to generate the prediction block for the current block.In other words, each fusion mode from a plurality of available fusion modes is associated with a different combination of two or more TM candidates. In a plurality of TM candidates, a bitstream indicates a subset of the available fusion modes; the number of fusion modes in the subset is M; the number of available fusion modes is N, where M < N; the fusion modes in the subset of available fusion modes have minimum combined template costs among the available fusion modes. As part of the application of... Petition 870250084327, dated 09 / 18 / 2025, pp. 222 / 292 89 / 126 subpel precision mode, the video encoder can, for each fusion mode of the subset of available fusion modes, for each TM candidate of the two or more TM candidates associated with the fusion mode: apply the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective array of samples with full pel and subpel precision. The video encoder can identify, within the respective array, a respective reference template of a respective prediction block for the TM candidate. The respective reference template of the respective prediction block is the best match for the template of the current block within the respective array. The video encoder can generate, based on a combination of the respective prediction blocks for the two or more TM candidates associated with the fusion mode, a prediction block for the fusion mode.The video encoder can determine the prediction block for the current block from among the prediction blocks for the fusion modes.

[0191] In another example, for each merge mode, there are multiple weighting sets that can be selected. To select the weighting set for a merge mode, the combination with each weighting set will be applied to the template and the combined template cost is calculated. The weighting set with the minimum combined template cost will be selected. Thus, in some examples, each merge mode from a plurality of merge modes is associated with a different combination of two or more TM candidates in a plurality of candidates. Petition 870250084327, dated 09 / 18 / 2025, pages 223 / 292 90 / 126 a TM. As part of applying the subpel precision mode, the video encoder can, for each fusion mode of the plurality of fusion modes and for each TM candidate of the two or more TM candidates associated with the fusion mode, apply the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective array of samples with full pel and subpel precision. The video encoder can identify, within the respective array, a respective reference template of a respective prediction block for the TM candidate. The respective reference template of the respective prediction block for the TM candidate can be the best match for the template of the current block within the respective array.For each respective weighting set from a plurality of weighting sets, the video encoder can generate a prediction block for the respective weighting set as a weighted average of samples in the prediction blocks of the two or more TM candidates associated with the fusion mode. The weights used in the weighted average are included in the respective weighting set. The video encoder can calculate a combined template cost of the prediction block for the respective weighting set. The video encoder can select a weighting set for the fusion mode from among the plurality of weighting sets based on a minimum of the combined template cost of the prediction blocks for the weighting sets. The video encoder can determine the prediction block for the current block from among the prediction blocks. Petition 870250084327, dated 09 / 18 / 2025, pages 224 / 292 91 / 126 for the weighting sets selected for the plurality of fusion modes.

[0192] In another example, the weights are derived using MSE minimization. In the training phase, the weights are derived to minimize the MSE between the linear combination of the k candidate templates and the current block template. Thus, in some examples, the video encoder can generate a prediction block as a weighted average of samples from two or more candidate templates. A training process can train the weights used in the weighted average to minimize the mean squared error between linear combinations of training TM candidate templates and training block templates. In some examples, the video encoder or other computing system can derive the weights after determining which TM candidates are merged. After deriving the weights, the video encoder can apply the candidates and generate the prediction block.

[0193] In another example, location information x, y, or both x and y is used as the input parameters in the interpolation filter. In this way, a video encoder can use location information as input parameters in filtering the prediction block for the current block. For example, the filtering function can be Pred(x, y) = w0*x + w1*y + w2*Ref(x, y), Equation (4) where x and y are the relative positions of the current block and the reference block, respectively. The reason for including x and y is to capture the position-related change of the content. Petition 870250084327, dated 09 / 18 / 2025, pages 225 / 292 92 / 126

[0194] In another example, multimodel filter parameters are used in the linear combination of the k candidates. In other words, a video encoder can generate the prediction block as a linear combination of samples in the reference blocks associated with the two or more TM candidates. The video encoder can use a multimodel filter in the linear combination of the k candidates.

[0195] In another example, the difference between the pixel values ​​of the candidates is used to select one of the multi-model filters. In other words, the video encoder can use a difference between the pixel values ​​of the candidates to select the multi-model filter.

[0196] In another example, location information x, y, or both x and y is used to select one of the multi-model filters. In other words, the video encoder can use location information to select the multi-model filters.

[0197] Examples related to spatial filtering in fused mode are now discussed. In some examples, the prediction block for the current block can be generated from a combination of k candidates, where k is greater than 1. The k candidates can be arbitrarily selected from the available candidate lists of different template patterns or even the same pattern. The combination weight wk can be derived based on the template matching cost, SAD, MSE, or block vector (BV). A spatial filter is additionally applied to the fused block.

[0198] In another example, the spatial filter is a six-point sampling filter, including the value of Petition 870250084327, dated 09 / 18 / 2025, pages 226 / 292 93 / 126 current pixel, the pixel value above, the pixel value to the left, the pixel value to the right, the pixel value below, and a bias term.

[0199] In another example, the spatial filter includes location information, such as x, y position or both x and y.

[0200] In another example, the multi-model spatial filter is used. In other words, the video encoder can use a difference between a pixel value of a current pixel and a pixel value of a neighboring pixel to select one of the spatial filters.

[0201] In another example, the difference between the pixel value of the current pixel and the pixel value of a neighboring pixel is used to select one of the multi-model filters. In other words, the video encoder can use a pixel value of a current pixel and an average of pixel values ​​to select the spatial filter.

[0202] In another example, the pixel value of the current pixel and the average of the pixel values ​​are used to select one of the multimodel filters.

[0203] In another example, a flag indicates whether the spatial filter is applied in fused mode. Video encoder 200 can signal the flag in a bitstream and video decoder 300 can obtain the flag from the bitstream.

[0204] In another example, a flag used to decide whether the spatial filter is applied is inherited from a neighboring block. Video encoder 200 can signal the flag in a bitstream and video decoder 300 can obtain the flag from the bitstream. Petition 870250084327, dated 09 / 18 / 2025, pages 227 / 292 94 / 126

[0205] In another example, a flag indicates whether the spatial filter is applied only when the fused mode uses MSE minimization to derive the weights. Otherwise, the spatial filter is considered not applied. Video encoder 200 can signal the flag in a bitstream and video decoder 300 can obtain the flag from the bitstream. In this way, a video encoder can determine samples from a prediction block based on weighted averages of corresponding samples from the two or more TM candidates, where the weighted averages use the derived weights.

[0206] Figure 11 is a flowchart illustrating an example method for encoding a current block according to the techniques of the present disclosure. The current block may be or include a current CU. Although described in relation to the Video Encoder 200 (Figures 1 and 2), it should be understood that other devices may be configured to perform a method similar to that of Figure 11.

[0207] In this example, video encoder 200 initially predicts the current block (1100). For example, video encoder 200 can form a prediction block for the current block. According to one or more techniques of the present disclosure, video encoder 200 can apply a subpel precision mode to generate the prediction block. Video encoder 200 can signal a syntax element to indicate that the subpel precision mode is applied to the current block. As part of applying the subpel precision mode, video encoder 200 can apply an interpolation filter to samples from a reference region to generate a sample array with precision. Petition 870250084327, dated 09 / 18 / 2025, pages 228 / 292 95 / 126 total and sub-layer. Video encoder 200 can identify, within the array, a prediction block reference template. The prediction block reference template may be the best match for a current block template within the array. A template pattern defines a format for the prediction block reference template and the current block template.

[0208] Video encoder 200 can then calculate a residual block for the current block (1102). To calculate the residual block, video encoder 200 can calculate a difference between the original, unencoded block and the prediction block for the current block. Video encoder 200 can then transform the residual block and quantize the transform coefficients of the residual block (1104). Next, video encoder 200 can scan the quantized transform coefficients of the residual block (1106). During or after scanning, video encoder 200 can entropically encode the transform coefficients (1108). For example, video encoder 200 can encode the transform coefficients using CAVLC or CABAC. Video encoder 200 can then output the entropically encoded block data (1110).

[0209] Figure 12 is a flowchart illustrating an example method for decoding a current block of video data according to the techniques of the present disclosure. The current block may be or include a current CU. Although described in relation to the 300 video decoder (Figures 1 and 3), it should be understood that other devices may Petition 870250084327, dated 09 / 18 / 2025, pages 229 / 292 96 / 126 can be configured to perform a method similar to that in Figure 12.

[0210] Video decoder 300 can receive entropically encoded data for the current block, such as entropically encoded prediction information and entropically encoded data for transform coefficients of a residual block corresponding to the current block (1200). Video decoder 300 can entropically decode the entropically encoded data to determine prediction information for the current block and to reproduce transform coefficients of the residual block (1202).

[0211] The Video Decoder 300 can predict the current block (1204), for example using an intra- or interprediction mode, as indicated by the prediction information for the current block, in order to calculate a prediction block for the current block. According to one or more techniques of the present disclosure, the Video Decoder 300 can apply a subpel precision mode to generate the prediction block. A syntax element indicates that the subpel precision mode is applied to the current block. The Video Decoder 300 can obtain the flag from the bitstream. As part of applying the subpel precision mode, the Video Decoder 300 can apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision. The Video Decoder 300 can identify, within the array, a reference template of the prediction block.The reference template for the prediction block may be the best match for a template of the current block within the array. A template pattern defines a format for... Petition 870250084327, dated 09 / 18 / 2025, pages 230 / 292 97 / 126 reference template for the prediction block and the current block template.

[0212] The video decoder 300 can then perform an inverse scan on the reproduced transform coefficients (1206) to create a block of quantized transform coefficients. The video decoder 300 can then inversely quantize the transform coefficients and apply an inverse transform to the transform coefficients to produce a residual block (1208). The video decoder 300 can finally decode the current block by combining the prediction block and the residual block (1210).

[0213] Figure 13 is a flowchart illustrating an example operation 1300 of a video encoder according to the techniques of the present disclosure. Operation 1300 can be performed by video encoder 200 or video decoder 300. In the example in Figure 13, the video encoder can apply a subpel precision mode to generate a prediction block for a current block of video data (1302). In some examples, a syntax element indicates whether the subpel precision mode is applied to the current block. Video encoder 200 can signal the flag on a bitstream, and video decoder 300 can obtain the flag from the bitstream. The syntax element can be signaled based on a non-merge template matching mode being used for the current block. In some examples, based on the subpel precision mode being used to generate the prediction block for the current block, a template matching merge mode flag is not signaled.In the examples in Figure 2 and Figure 3, the unit. Petition 870250084327, dated 09 / 18 / 2025, pages 231 / 292 TM 228 98 / 126 or TM 322 unit can apply the subpel precision mode to generate the prediction block for the current block.

[0214] As part of applying subpel precision mode, the video encoder can apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision (1304). In some examples, the subpel precision is quarter-pixel precision. In other examples, the subpel precision may be different, such as half pel, 1 / 8pel, and so on.

[0215] The video encoder can identify, within the array, a prediction block reference template, where the prediction block reference template is the best match for a current block template within the array (1306). A template pattern defines a format for the prediction block reference template and the current block template. prediction block for a current block of video data. In this way, the video encoder can search the array for a reference template that has the closest match to a current template. The reference template and the current template have a format defined by the template pattern, and the current template includes samples neighboring the current block.

[0216] In some examples, the video encoder generates a list of template matching candidates. Each of the template matching candidates is associated with a different motion vector. The video encoder can select a template matching candidate from the list of matching candidates. Petition 870250084327, dated 09 / 18 / 2025, pages 232 / 292 99 / 126 template. Based on the subpel precision mode being used to generate the prediction block for the current block, the list of template matching candidates may be reduced compared to when subpel precision mode is not used.

[0217] In some examples, the video encoder uses a fusion of two or more reference sample blocks to generate the prediction block for the current block. In such examples, the video encoder may generate a plurality of TM candidates. Each of the TM candidates is associated with a respective prediction block and a respective motion vector predictor. As part of generating the plurality of TM candidates, the video encoder may, for each of the TM candidates, apply the interpolation filter to samples from a reference region indicated by the motion vector predictor associated with the TM candidate to generate a respective sample array with full-pel and sub-pel accuracy. The video encoder may identify, within the respective array, a respective reference template for the respective prediction block.The respective reference template of the respective prediction block may be the best match for the template of the current block within the respective array. The video encoder may generate, based on a combination of two or more of the respective prediction blocks associated with the TM candidates, the prediction block for the current block.

[0218] The video encoder can encode or decode the current block using the prediction block for the current block (1308). For example, in an instance where the video encoder is video encoder 200, video encoder 200 can generate residual data based Petition 870250084327, dated 09 / 18 / 2025, pages 233 / 292 100 / 126 in the prediction block for the current block, apply a transform to the residual data to generate transform coefficients, quantize the transform coefficients, and entropically encode syntax elements representing the quantized transform coefficients. In an example where the video encoder is video decoder 300, video decoder 300 can reconstruct samples from the current block based on the prediction block and the residual data.

[0219] The following numbered clauses illustrate one or more aspects of the techniques and devices described in this disclosure.

[0220] Clause 1A. A method for encoding or decoding video data, the method comprising: determining a template pattern from among a plurality of template patterns; identifying, based on the determined template pattern, a prediction block for a current block of video data; and encoding or decoding the current block using the prediction block, wherein a syntax element is signaled to indicate the determined template pattern.

[0221] Clause 2A. The method of clause 1A, wherein the plurality of templates may include two or more of the following: a set of reference samples above and to the left of the current block, a set of reference samples to the left of the current block, a set of reference samples above the current block, a subset of reference samples to the left of the current block, a subset of reference samples above the current block, a set of reference samples to the left and not Petition 870250084327, dated 09 / 18 / 2025, pages 234 / 292 101 / 126 adjacent to the current block, a set of reference samples above and not adjacent to the current block.

[0222] Clause 3A. The method of any of clauses 1A to 2A, wherein: the plurality of template patterns includes a base template pattern and the determination of the template pattern comprises determining the template pattern from among the base template pattern and one or more of the template patterns that are within an area of ​​the base template pattern.

[0223] Clause 1B. A method for encoding or decoding video data, the method comprising: applying a subpel precision mode to generate a prediction block for a current block of video data, wherein the application of the subpel precision mode comprises: applying an interpolation filter to samples from a reference region to generate a reference region that includes an array of samples with full pel and subpel precision; and using a template pattern to identify, within the array, a prediction block for a current block of video data, wherein a block vector indicating an offset between the current block and the prediction block has subpel accuracy; and encoding or decoding the current block using the prediction block.

[0224] Clause 2B. The method of clause 1B, in which a syntax element is signaled to indicate whether subpel precision mode is applied to the current block.

[0225] Clause 3B. The method of clause 2B, where the syntax element is flagged based on a non-merge template matching mode being used for the current block. Petition 870250084327, dated 09 / 18 / 2025, pages 235 / 292 102 / 126

[0226] Clause 4B. The method of clause 1B, which additionally comprises inheriting a subpel precision model flag from a block that is adjacent to the current block.

[0227] Clause 5B. The method of any of clauses 1B to 4B, wherein the method further comprises determining the template pattern from among a plurality of template patterns.

[0228] Clause 6B. The method of clause 5B, wherein: the plurality of template patterns includes a base template pattern and the determination of the template pattern comprises determining the template pattern from among the base template pattern and one or more of the template patterns that are within an area of ​​the base template pattern.

[0229] Clause 7B. The method of clause 6B, where a syntax element is flagged to indicate whether the subpel precision mode is applied to the current block based on the determined template pattern being the base template pattern and not any of the other template patterns.

[0230] Clause 8B. The method of clause 6B, which further comprises: generating a list of template matching candidates, where each of the template matching candidates is associated with a different motion vector; selecting a template matching candidate from the list of template matching candidates, wherein, based on the selected template matching candidate having a template matching candidate index equal to 0 and the determined template pattern being the base template pattern and not any of the other template patterns, an element Petition 870250084327, dated 09 / 18 / 2025, pp. 236 / 292 The syntax flag (103 / 126) indicates whether subpel precision mode is applied to the current block.

[0231] Clause 9B. The method of clause 6B, in which, based on the subpel precision mode being used to encode the current block, a syntax element indicating the selected template pattern is not signaled.

[0232] Clause 10B. The method of any of clauses 1B to 9B, which further comprises: generating a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; and selecting a template matching candidate from the list of template matching candidates, wherein, based on the selected template matching candidate having a template matching candidate index equal to a predetermined value less than a number of template matching candidates in the list and the determined template pattern being the base template pattern and not any of the other template patterns, a syntax element is signaled to indicate whether subpel precision mode is applied to the current block.

[0233] Clause 11B. The method of any of clauses 1B to 10B, which further comprises: generating a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; and selecting a template matching candidate from the list of template matching candidates, wherein, based on the selected template matching candidate having a template matching candidate index equal to Petition 870250084327, dated 09 / 18 / 2025, pp. 237 / 292 104 / 126 to 0, a syntax element is flagged to indicate whether subpel precision mode is applied to the current block.

[0234] Clause 12B. The method of any of clauses 1B to 11B, which further comprises: generating a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; and selecting a template matching candidate from the list of template matching candidates, wherein, based on the selected template matching candidate having a template matching candidate index equal to a predetermined value less than a number of template matching candidates in the list, a syntax element is signaled to indicate whether subpel precision mode is applied to the current block.

[0235] Clause 13B. The method of any of clauses 1B to 12B, which further comprises: generating a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; selecting a template matching candidate from the list of template matching candidates; and wherein, based on the subpel precision mode being used to encode the current block, the list of template matching candidates is reduced compared to when the subpel precision mode is not used.

[0236] Clause 14B. The method of any of clauses 1B to 13B, which further comprises: generating a list of candidate template matches, wherein each candidate template match is Petition 870250084327, dated 09 / 18 / 2025, pp. 238 / 292 105 / 126 associated with a different motion vector; and select a template match candidate from the list of template match candidates, wherein, based on the subpel precision mode being used to encode the current block, a template match candidate index of the selected template match candidate is not flagged.

[0237] Clause 15B. The method of any of clauses 1B to 14B, wherein, based on the subpel precision model being used to encode the current block, a template matching merge mode flag is not signaled.

[0238] Clause 16B. The method of any of clauses 1A to 15B, wherein an index of the given template pattern is signaled at one of: a coding unit (CU) level, a prediction unit (PU) level, a coding tree unit (CTU) level, a slice level, or an image level.

[0239] Clause 17B. The method of any of clauses 1B to 16B, wherein the subpel precision is one of half a pixel or one-quarter of a pixel.

[0240] Clause 18B. The method of clause 17B, which additionally comprises signaling a syntax element that specifies whether the subpel precision is half a pixel or a quarter of a pixel.

[0241] Clause 1C. A method for encoding or decoding video data, the method comprising: generating template matching (TM) candidates, wherein each TM candidate is associated with a different reference block; generating, based on a combination of Petition 870250084327, dated 09 / 18 / 2025, pages 239 / 292 106 / 126 two or more of the TM candidates, a predictor for a current block of video data; and encode or decode the current block using the predictor.

[0242] Clause 2C. The method of clause 1C, which further comprises: evaluating the costs of the reference blocks associated with the MT candidates based on two or more model standards; and selecting the two or more MT candidates based on costs.

[0243] Clause 3C. The method of any of clauses 1C to 2C, wherein predictor generation comprises generating the predictor as a linear combination of samples in the reference blocks associated with the two or more TM candidates.

[0244] Clause 4C. The method of clause 3C, wherein the generation of the predictor as the linear combination of the samples comprises: for each of the two or more TM candidates, calculate a template cost for the template pattern candidate and generate the predictor as a weighted average of the samples in the two or more template candidates, wherein the weights used in the weighted average are based on the template matching costs for the template pattern candidates.

[0245] Clause 5C. The method of clause 3C, wherein the generation of the predictor as the linear combination of the samples comprises: deriving weights based on one or more of a sum of absolute differences (SAD), a mean squared error (MSE), an MSE minimization, or a block vector (BV) of the two or more template candidates and generating the predictor as a weighted average of the samples in the two or more template candidates, wherein the weights used in the average Petition 870250084327, dated 09 / 18 / 2025, pages 240 / 292 The weighted 107 / 126 is based on template matching costs for TM applicants.

[0246] Clause 6C. The method of any of clauses 3C to 5C, wherein the generation of the predictor as the linear combination of the samples comprises: determining a weight derivation method from a plurality of weight derivation methods; applying the weight derivation method to derive weights from the two or more TM candidates; and generating the predictor as a weighted average of the samples in the two or more TM candidates, wherein the weights used in the weighted average are based on the template matching costs for the template pattern candidates, wherein a syntax element is signaled to indicate the determined weight derivation method.

[0247] Clause 7C. The method of any of clauses 1C to 6C, which further comprises: deriving a sum of absolute differences (SAD), a mean squared error (MSE), an MSE minimization or a block vector (BV) of the template pattern candidates and selecting the two or more template pattern candidates based on one or more of the SAD, MSE, MSE minimization or BV of the template pattern candidates.

[0248] Clause 8C. The method of any of clauses 1C to 7C, where a flag is set to indicate whether the weights are derived based on template cost or based on MSE minimization.

[0249] Clause 9C. The method of any of the clauses 1C to 8C, in which multiple combination modes can be selected and signaled, the method comprising Petition 870250084327, dated 09 / 18 / 2025, pp. 241 / 292 108 / 126 additionally the construction of a merger list to indicate merger candidates for the current bloc.

[0250] Clause 10C. The method of any of clauses 1C to 9C, where a number of signed fusion modes is M and a number of available fusion modes is N, where M < N, comprises selecting the M fusion modes from N fusion modes; applying the combination to the template of all TM candidates; calculating a combined template cost for each of the N fusion modes; selecting the M modes with the minimum combined template cost.

[0251] Clause 11C. The method of any of clauses 1C to 10C, wherein: the method comprises, for each fusion mode of a plurality of fusion modes, selecting multiple weighting sets, wherein selecting the weighting set comprises: applying a combination with each weighting set to the model and calculating a combined model cost; and selecting a weighting set with a minimum combined model cost; generating the predictor as a sample weighted average across the two or more template candidates, wherein the weights used in the weighted average are included in the weighting set.

[0252] Clause 12C. The method of any of clauses 1C to 11C, which additionally comprises: training weights to minimize MSE between linear combinations of the k candidate templates and the current block template; generating the predictor as a weighted average of samples across the two or more candidate templates, wherein the weights used in the weighted average are included in the weighting set. Petition 870250084327, dated 09 / 18 / 2025, pages 242 / 292 109 / 126

[0253] Clause 13C. The method of any of clauses 1C to 12C, which additionally includes the use of location information as input parameters in predictor filtering.

[0254] Clause 14C. The method of any of clauses 1C to 13C, wherein: generate the predictor as a linear combination of samples in the reference blocks associated with the two or more TM candidates, and the method additionally comprising the use of a multimodel filter in the linear combination of the k candidates.

[0255] Clause 15C. The method of clause 14C, which additionally comprises the use of a difference between the pixel values ​​of the candidates to select the multimodel filter.

[0256] Clause 16C. The method of clause 14C, which additionally includes the use of location information to select multi-model filters.

[0257] Clause 17C. The method of any of clauses 1C to 16C, which additionally includes the application of a spatial filter to the predictor.

[0258] Clause 18C. The method of clause 17C, where the spatial filter is a 6-point sampling filter.

[0259] Clause 19C. The method of any of clauses 17C to 18C, wherein the spatial filter includes location information.

[0260] Clause 20C. The method of any of clauses 17C to 19C, which additionally comprises using a difference between a pixel value of a current pixel and a pixel value of a neighboring pixel to select one of the spatial filters. Petition 870250084327, dated 09 / 18 / 2025, pages 243 / 292 110 / 126

[0261] Clause 21C. The method of any of clauses 17C to 20C, which additionally comprises the use of a pixel value of a current pixel and an average of pixel values ​​to select the spatial filter.

[0262] Clause 22C. The method of any of clauses 17C to 21C, where a flag is set to indicate whether the spatial filter is applied to the predictor.

[0263] Clause 23C. The method of any of clauses 17C to 22C, where a flag used to decide whether the spatial filter is applied to the predictor is inherited from a neighboring block.

[0264] Clause 24C. The method of any of clauses 17C to 23C, wherein: a flag is signaled to indicate whether the spatial filter is applied to the predictor only when a fused mode uses mean squared error minimization (MSE) to derive weights, and wherein predictor generation comprises determining predictor samples based on weighted averages of corresponding samples from the two or more TM candidates, wherein the weighted averages use the derived weights.

[0265] Clause 25C. A combination of any of the methods in any of clauses 1A to 24C.

[0266] Clause 1D. A device for encoding video data, the device comprising one or more means for carrying out the method of any of the method clauses listed above.

[0267] Clause 2D. The device of clause 1D, wherein the one or more means comprise one or more processors implemented in a circuit assembly. Petition 870250084327, dated 09 / 18 / 2025, pp. 244 / 292 111 / 126

[0268] 3D Clause. The device of either of the 1D and 2D clauses, which additionally comprises a memory for storing video data.

[0269] Clause 4D. The device of any of the clauses 1D to 3D, which additionally comprises a display configured to display decoded video data.

[0270] Clause 5D. The device of any of clauses 1D to 4D, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver or a set-top box.

[0271] Clause 6D. The device of any of clauses 1D to 5D, wherein the device comprises a video decoder.

[0272] Clause 7D. The device of any of clauses 1D to 6D, wherein the device comprises a video encoder.

[0273] Clause 8D. A computer-readable storage medium that has stored within it instructions which, when executed, cause one or more processors to perform the method of any of the method clauses listed above.

[0274] Clause 1E. A method for encoding or decoding video data, the method comprising: applying a subpel precision mode to generate a prediction block for a current block of video data, wherein a syntax element indicates that the subpel precision mode is applied to the current block, and applying the subpel precision mode comprises: applying an interpolation filter to samples from a reference region to generate an array of samples with full and subpel precision; and Petition 870250084327, dated 09 / 18 / 2025, pages 245 / 292 112 / 126 identify, within the array, a reference template for the prediction block, where the reference template for the prediction block is the best match for a template of the current block within the array, where a template pattern defines a format for the reference template of the prediction block and the template of the current block; and encode or decode the current block using the prediction block for the current block.

[0275] Clause 2E. The method of clause 1E, where the syntax element is flagged based on a non-merge template matching mode being used for the current block.

[0276] Clause 3E. The method of any of clauses 1E to 2E, which further comprises: generating a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; and selecting a template matching candidate from the list of template matching candidates, wherein, based on the subpel precision mode being used to encode or decode the current block, the list of template matching candidates is reduced compared to when the subpel precision mode is not used.

[0277] Clause 4E. The method of any of clauses 1E to 3E, wherein, based on the subpel precision mode being used to encode or decode the current block, a template match merge mode flag is not signaled. Petition 870250084327, dated 09 / 18 / 2025, pp. 246 / 292 113 / 126

[0278] Clause 5E. The method of any of the clauses 1E to 3E, where the subpel precision is quarter-pixel precision.

[0279] Clause 6E. The method of any of clauses 1E to 5E, wherein: the method further comprises generating a plurality of template matching (TM) candidates, wherein each TM candidate is associated with a respective prediction block and a respective motion vector predictor, and generating the plurality of TM candidates comprises, for each TM candidate: applying the interpolation filter to samples from a reference region indicated by the motion vector predictor associated with the TM candidate to generate a respective array of samples with full and sub-pel accuracy; and identifying, within the respective array, a respective reference template of the respective prediction block, wherein the respective reference template of the respective prediction block is the best match for the current block template within the respective array;Generating the prediction block for the current block involves generating, based on a combination of the respective prediction blocks associated with two or more TM candidates from the plurality of TM candidates, the prediction block for the current block.

[0280] Clause 7E. The method of clause 6E, in which the generation of the prediction block for the current block comprises generating the prediction block for the current block as a linear combination of corresponding samples in the prediction blocks associated with the two or more TM candidates.

[0281] Clause 8E. The method of clause 7E, in which the generation of the prediction block for the current block as Petition 870250084327, dated 09 / 18 / 2025, pages 247 / 292 114 / 126 The linear combination of corresponding samples comprises: for each of the two or more TM candidates, calculating a template cost for the TM candidate and generating the prediction block for the current block as a weighted average of the corresponding samples in the prediction blocks associated with the two or more TM candidates, where the weights used in the weighted average are based on the template cost for the TM candidates.

[0282] Clause 9E. The method of clause 8E, where the weights are multiplicative inverses of the template matching costs for TM candidates.

[0283] Clause 10E. The method of any of clauses 8E to 9E, wherein a flag indicates whether the weights are derived based on template cost or based on mean squared error minimization (MSE).

[0284] Clause 11E. The method of any of clauses 7E to 10E, wherein the syntax element is a first syntax element and generate the prediction block for the current block as the linear combination of the corresponding samples comprises: determining a weight derivation method from a plurality of weight derivation methods; applying the weight derivation method to derive weights from the two or more TM candidates; and generating the prediction block for the current block as a weighted average of the corresponding samples in the prediction blocks associated with the two or more TM candidates using the weights of the two or more TM candidates, wherein a second syntax element indicates the weight derivation method.

[0285] Clause 12E. The method of any of the clauses 1E to 11E, where: each merging mode of a Petition 870250084327, dated 09 / 18 / 2025, pp. 248 / 292 115 / 126 plurality of available fusion modes is associated with a different combination of two or more TM candidates in a plurality of TM candidates, a bitstream indicates a subset of the available fusion modes, a quantity of fusion modes in the subset is M, a quantity of available fusion modes is N, where M < N, the fusion modes in the subset of available fusion modes have minimum combined template costs among the available fusion modes, the application of the subpel precision mode comprises: for each fusion mode of the subset of available fusion modes: for each TM candidate of the two or more TM candidates associated with the fusion mode: apply the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective sample array with full pel and subpel precision;and identify, within the respective array, a respective reference template for a respective prediction block for the TM candidate, where the respective reference template of the respective prediction block is the best match for the template of the current block within the respective array; and generate, based on a combination of the respective prediction blocks for the two or more TM candidates associated with the fusion mode, a prediction block for the fusion mode; and determine the prediction block for the current block from among the prediction blocks for the fusion modes.

[0286] Clause 13E. The method of any of clauses 1E to 12E, wherein each fusion mode of a plurality of fusion modes is associated with a different combination of two or more TM candidates in a plurality Petition 870250084327, dated 09 / 18 / 2025, pp. 249 / 292 116 / 126 of TM candidates, the application of the subpel precision mode comprises: for each fusion mode of the plurality of fusion modes: for each TM candidate of the two or more TM candidates associated with the fusion mode: applying the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective array of samples with full and subpel precision; and identifying, within the respective array, a respective reference template of a respective prediction block for the TM candidate, wherein the respective reference template of the respective prediction block for the TM candidate is the best match for the template of the current block within the respective array;For each respective weighting set from a plurality of weighting sets: generate a prediction block for the respective weighting set as a weighted average of corresponding samples in the prediction blocks of the two or more TM candidates associated with the fusion mode, where the weights used in the weighted average are included in the respective weighting set; and calculate a combined template cost of the prediction block for the respective weighting set; and select a weighting set for the fusion mode from among the plurality of weighting sets based on a minimum of the combined template cost of the prediction blocks for the weighting sets; and determine the prediction block for the current block from among the prediction blocks for the weighting sets selected for the plurality of fusion modes.

[0287] Clause 14E. A device for encoding video data, the device comprising: Petition 870250084327, dated 09 / 18 / 2025, pp. 250 / 292 117 / 126 a memory for storing video data; and one or more processors implemented in a circuit set, the one or more processors configured to: apply a subpel precision mode to generate a prediction block for a current block of video data, wherein a syntax element is signaled to indicate that the subpel precision mode is applied to the current block and the one or more processors are configured to, applying the subpel precision mode: apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision; and identify, within the array, a prediction block reference template, wherein the prediction block reference template is the best match for a current block template within the array, wherein a template pattern defines a format of the prediction block reference template and the current block template;and encode or decode the current block using the prediction block for the current block.

[0288] Clause 15E. The device of clause 14E, in which the syntax element is flagged based on a non-merge template matching mode being used for the current block.

[0289] Clause 16E. The device of any of clauses 14E to 15E, wherein one or more processors are additionally configured to: generate a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; and select a template matching candidate from the list of candidates. Petition 870250084327, dated 09 / 18 / 2025, pages 251 / 292 118 / 126 template matching, where based on the subpel precision mode being used to encode or decode the current block, the list of candidates for template matching is reduced compared to when subpel precision mode is not used.

[0290] Clause 17E. The device of any of clauses 14E to 16E, in which, based on the subpel precision mode being used to encode or decode the current block, a template match merge mode flag is not signaled.

[0291] Clause 18E. The device of any of clauses 14E to 17E, wherein the subpel precision is quarter-pixel precision.

[0292] Clause 19E. The device of any of clauses 14E to 18E, wherein: one or more processors are additionally configured to generate a plurality of template matching (TM) candidates, wherein each TM candidate is associated with a respective prediction block and a respective motion vector predictor, and one or more processors are configured to, as part of generating the plurality of TM candidates, for each TM candidate: apply the interpolation filter to samples from a reference region indicated by the motion vector predictor associated with the TM candidate to generate a respective array of samples with full and sub-pel accuracy; and identify, within the respective array, a respective reference template of the respective prediction block, wherein the respective reference template of the respective prediction block is the best match for the current block template within the Petition 870250084327, dated 09 / 18 / 2025, pages 252 / 292 119 / 126 respective arrangement; and one or more processors are configured to, as part of the generation of the prediction block for the current block, generate, based on a combination of the respective prediction blocks associated with two or more TM candidates from the plurality of TM candidates, the prediction block for the current block.

[0293] Clause 20E. The device of clause 19E, wherein one or more processors are configured to, as part of the generation of the prediction block for the current block, generate the prediction block for the current block as a linear combination of matching samples in the prediction blocks associated with the two or more TM candidates.

[0294] Clause 21E. The device of clause 20E, wherein the generation of the prediction block for the current block as the linear combination of the corresponding samples comprises: for each of the two or more TM candidates, calculating a template cost for the TM candidate and generating the prediction block for the current block as a weighted average of the corresponding samples in the two or more TM candidates, wherein the weights used in the weighted average are based on the template cost for the TM candidate.

[0295] Clause 22E. The provision of clause 21E, wherein the weights are multiplicative inverses of the template matching costs for TM applicants.

[0296] Clause 23E. The device of any of clauses 21E to 22E, wherein a flag is signaled to indicate whether the weights are derived based on template cost or based on mean squared error minimization (MSE). Petition 870250084327, dated 09 / 18 / 2025, pages 253 / 292 120 / 126

[0297] Clause 24E. The device of any of clauses 20E to 23E, wherein the syntax element is a first syntax element and one or more processors are configured to, as part of generating the prediction block for the current block as the linear combination of the corresponding samples: determine a weight derivation method from a plurality of weight derivation methods; apply the weight derivation method to derive weights from the two or more TM candidates; and generate the prediction block for the current block as the weighted average of the corresponding samples in the prediction blocks associated with the two or more TM candidates using the weights of the two or more TM candidates, wherein a second syntax element indicates the weight derivation method.

[0298] Clause 25E. The device of any of clauses 14E to 24E, wherein: each merge mode from a plurality of available merge modes is associated with a different combination of two or more TM candidates from a plurality of TM candidates, a bitstream indicates a subset of the available merge modes, a quantity of merge modes in the subset is M, a quantity of available merge modes is N, where M < N, the merge modes in the subset of available merge modes have minimum combined template costs among the available merge modes, as part of applying the subpel precision mode, the one or more processors are configured to: for each merge mode from the subset of available merge modes: for each TM candidate from the two or more TM candidates associated with the merge mode: apply the interpolation filter to samples from a reference region Petition 870250084327, dated 09 / 18 / 2025, pp. 254 / 292 121 / 126 indicated by a motion vector predictor associated with the TM candidate to generate a respective sample array with full and sub-pel accuracy; and identify, within the respective array, a respective reference template of a respective prediction block for the TM candidate, wherein the respective reference template of the respective prediction block is the best match for the template of the current block within the respective array; and generate, based on a combination of the respective prediction blocks for the two or more TM candidates associated with the fusion mode, a prediction block for the fusion mode; and determine the prediction block for the current block from among the prediction blocks for the fusion modes.

[0299] Clause 26E. The device of any of clauses 14E to 25E, wherein each fusion mode of a plurality of fusion modes is associated with a different combination of two or more TM candidates in a plurality of TM candidates, and the one or more processors are configured to, as part of the application of the subpel precision mode: for each fusion mode of the plurality of fusion modes: for each TM candidate of the two or more TM candidates associated with the fusion mode: apply the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective array of samples with full and subpel precision; and identify, within the respective array, a respective reference template of a respective prediction block for the TM candidate, wherein the respective reference template of the respective prediction block is the best match for the template of Petition 870250084327, dated 09 / 18 / 2025, pages 255 / 292 122 / 126 current block within the respective array; for each respective weighting set from a plurality of weighting sets: generate a prediction block for the respective weighting set as a weighted average of corresponding samples in the prediction blocks of the two or more TM candidates associated with the fusion mode, where the weights used in the weighted average are included in the respective weighting set; and calculate a combined template cost of the prediction block for the respective weighting set; and select a weighting set for the fusion mode from among the plurality of weighting sets based on a minimum of the combined template cost of the prediction blocks for the weighting sets; and determine the prediction block for the current block from among the prediction blocks for the weighting sets selected for the plurality of fusion modes.

[0300] Clause 27E. The device of any of clauses 14E to 26E, which additionally comprises a display configured to display decoded video data.

[0301] Clause 28E. The device of any of clauses 14E to 27E, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver or a set-top box.

[0302] Clause 29E. The device of any of clauses 14E to 28E, wherein the device comprises a video decoder. Petition 870250084327, dated 09 / 18 / 2025, pp. 256 / 292 123 / 126

[0303] Clause 30E. The device of any of clauses 14E to 29E, wherein the device comprises a video encoder.

[0304] It should be recognized that, depending on the example, certain actions or events of any of the techniques described in the present invention may be performed in a different sequence, may be added, combined or completely omitted (for example, not all actions or events described are necessary for the practice of the techniques). Furthermore, in certain examples, the actions or events may be performed simultaneously, for example, through multi-threaded processing, interrupt processing or multiple processors, instead of sequentially.

[0305] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored or transmitted as one or more instructions or code in a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to a tangible medium, such as data storage media, or communication media, including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transient or (2) a communication medium, such as a Petition 870250084327, dated 09 / 18 / 2025, pages 257 / 292 124 / 126 signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0306] By way of example, and not limitation, such computer-readable storage media may include one or more of the following: RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory or any other media that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is properly termed a computer-readable medium.For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave will be included in the definition of medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transient storage media. Petition 870250084327, dated 09 / 18 / 2025, pages 258 / 292 125 / 126 tangible. As used in the present invention, discs (disk and disc) include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where discs typically reproduce data magnetically, while discs reproduce data optically by means of lasers. Combinations of the above items should also be included in the scope of computer-readable media.

[0307] The instructions may be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent discrete or integrated logic circuitry. Consequently, the terms processor and processing circuitry, as used in the present invention, may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described in the present invention. Furthermore, in some respects, the functionality described in the present invention may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Additionally, the techniques may be fully implemented in one or more logic circuits or elements.

[0308] The techniques in this disclosure can be implemented in a wide variety of devices or appliances, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured for Petition 870250084327, dated 09 / 18 / 2025, pages 259 / 292 126 / 126 performs the disclosed techniques, but does not necessarily require implementation by different hardware units. Instead, as described above, multiple units can be combined into a single codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors, as described above, in conjunction with appropriate software and / or firmware.

[0309] Several examples have been described. These and other examples are within the scope of the following claims. Petition 870250084327, dated 09 / 18 / 2025, pp. 260 / 292

Claims

1 / 14 CLAIMS 1. A method for encoding or decoding video data characterized in that it comprises: applying a subpel precision mode to generate a prediction block for a current block of video data, wherein a syntax element indicates that the subpel precision mode is applied to the current block, and applying the subpel precision mode comprises: applying an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision; and identifying, within the array, a prediction block reference template, wherein the prediction block reference template is the best match for a current block template within the array, wherein a template pattern defines a format of the prediction block reference template and the current block template; and encoding or decoding the current block using the prediction block for the current block.

2. A method according to claim 1, characterized in that the syntax element is signaled based on a non-merge template matching mode being used for the current block.

3. Method according to claim 1, characterized in that it further comprises: generating a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; and selecting a template matching candidate from the list of template matching candidates, wherein, based on the subpel precision mode being used to encode or decode the current block, the list of template matching candidates is reduced relative to when the subpel precision mode is not used.

4. A method according to claim 1, characterized in that, based on the subpel precision mode being used to encode or decode the current block, a template matching merge mode flag is not signaled.

5. Method, according to claim 1, characterized in that the subpel precision is quarter-pixel precision.

6. Method, according to claim 1, characterized in that: the method further comprises generating a plurality of template matching (TM) candidates, wherein each TM candidate is associated with a respective prediction block and a respective motion vector predictor, and generating the plurality of TM candidates comprises, for each TM candidate: applying the interpolation filter to samples from a reference region indicated by the motion vector predictor associated with the TM candidate to generate a respective array of samples with full and sub-pel accuracy; and identifying, within the respective array, a respective reference template of the respective block of Petition 870250084327, dated 09 / 18 / 2025, p.262 / 292 3 / 14 prediction, where the respective reference template of the respective prediction block is the best match for the template of the current block within the respective array; and generating the prediction block for the current block comprises generating, based on a combination of the respective prediction blocks associated with two or more TM candidates from the plurality of TM candidates, the prediction block for the current block.

7. Method, according to claim 6, characterized in that the generation of the prediction block for the current block comprises generating the prediction block for the current block as a linear combination of corresponding samples in the prediction blocks associated with the two or more TM candidates.

8. Method, according to claim 7, characterized in that the generation of the prediction block for the current block as the linear combination of the corresponding samples comprises: for each of the two or more TM candidates, calculating a template cost for the TM candidate, and generating the prediction block for the current block as a weighted average of the corresponding samples in the prediction blocks associated with the two or more TM candidates, wherein the weights used in the weighted average are based on the template cost for the TM candidates.

9. Method according to claim 8, characterized in that the weights are multiplicative inverses of the template matching costs for TM candidates. Petition 870250084327, dated 09 / 18 / 2025, pp. 263 / 292 4 / 14 10. Method according to claim 8, characterized in that a flag indicates whether the weights are derived based on template cost or based on mean squared error minimization (MSE).

11. A method according to claim 7, characterized in that the syntax element is a first syntax element and the generation of the prediction block for the current block as the linear combination of the corresponding samples comprises: determining a weight derivation method from among a plurality of weight derivation methods; applying the weight derivation method to derive weights from the two or more TM candidates; and generating the prediction block for the current block as a weighted average of the corresponding samples in the prediction blocks associated with the two or more TM candidates using the weights of the two or more TM candidates, wherein a second syntax element indicates the weight derivation method.

12. Method according to claim 1, characterized in that: each fusion mode of a plurality of available fusion modes is associated with a different combination of two or more TM candidates in a plurality of TM candidates, a bitstream indicates a subset of the available fusion modes, a quantity of fusion modes in the subset is M, a quantity of available fusion modes is N, where M < N, the fusion modes in the subset of available fusion modes have minimum combined template costs among the available fusion modes, Petition 870250084327, dated 09 / 18 / 2025,Page 264 / 292 5 / 14 The application of the subpel precision mode comprises: for each fusion mode of the subset of available fusion modes; for each TM candidate of the two or more TM candidates associated with the fusion mode: applying the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective array of samples with full and subpel precision; and identifying, within the respective array, a respective reference template of a respective prediction block for the TM candidate, where the respective reference template of the respective prediction block is the best match for the template of the current block within the respective array; and generating, based on a combination of the respective prediction blocks for the two or more TM candidates associated with the fusion mode,a prediction block for the merge mode; and determine the prediction block for the current block from among the prediction blocks for the merge modes.

13. Method, according to claim 1, characterized in that each fusion mode of a plurality of fusion modes is associated with a different combination of two or more TM candidates in a plurality of TM candidates, wherein the application of the subpel precision mode comprises: Petition 870250084327, dated 09 / 18 / 2025, pp. 265 / 292 6 / 14 for each fusion mode of the plurality of fusion modes: for each TM candidate of the two or more TM candidates associated with the fusion mode: applying the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective sample array with full pel and subpel precision;and identify, within the respective array, a respective reference template of a respective prediction block for the TM candidate, wherein the respective reference template of the respective prediction block for the TM candidate is the best match for the current block template within the respective array; for each respective weighting set from a plurality of weighting sets: generate a prediction block for the respective weighting set as a weighted average of matching samples in the prediction blocks of the two or more TM candidates associated with the merge mode, wherein the weights used in the weighted average are included in the respective weighting set; and calculate a combined template cost of the prediction block for the respective weighting set;and select a weighting set for the merge mode from among the plurality of weighting sets based on a minimum of the combined template cost of the prediction blocks for the weighting sets; and Petition 870250084327, dated 09 / 18 / 2025, pp. 266 / 292 7 / 14 determine the prediction block for the current block from among the prediction blocks for the weighting sets selected for the plurality of merge modes.

14. Device for encoding video data characterized in that it comprises: a memory for storing video data; and one or more processors implemented in a circuit assembly, the one or more processors configured to: apply a subpel precision mode to generate a prediction block for a current block of video data, wherein a syntax element is signaled to indicate that the subpel precision mode is applied to the current block and the one or more processors are configured, when applying the subpel precision mode, to: apply an interpolation filter to samples from a reference region to generate an array of samples with full pel and subpel precision;and identify, within the array, a reference template for the prediction block, where the reference template for the prediction block is the best match for a template of the current block within the array, where a template pattern defines a format for the reference template of the prediction block and the template of the current block; and encode or decode the current block using the prediction block for the current block.

15. Device according to claim 14, characterized in that the syntax element is signaled based on a non-merge template matching mode being used for the current block. Petition 870250084327, dated 09 / 18 / 2025, pp. 267 / 292 8 / 14 16. Device according to claim 14, characterized in that one or more processors are additionally configured to: generate a list of template matching candidates, wherein each of the template matching candidates is associated with a different motion vector; and select a template matching candidate from the list of template matching candidates, wherein, based on the subpel precision mode being used to encode or decode the current block, the list of template matching candidates is reduced compared to when the subpel precision mode is not used.

17. Device according to claim 14, characterized in that, based on the subpel precision mode being used to encode or decode the current block, a template matching merge mode flag is not signaled.

18. Device according to claim 14, characterized in that the subpel accuracy is quarter-pixel accuracy.

19. Device according to claim 14, characterized in that: one or more processors are additionally configured to generate a plurality of template matching (TM) candidates, wherein each of the TM candidates is associated with a respective prediction block and a respective motion vector predictor, and the one or more processors are configured to, as part of Petition 870250084327, dated 09 / 18 / 2025, p.268 / 292 9 / 14 of the generation of the plurality of TM candidates, for each of the TM candidates: apply the interpolation filter to samples from a reference region indicated by the motion vector predictor associated with the TM candidate to generate a respective array of samples with full and sub-pel accuracy; and identify, within the respective array, a respective reference template of the respective prediction block, where the respective reference template of the respective prediction block is the best match for the template of the current block within the respective array; and one or more processors are configured to, as part of the generation of the prediction block for the current block, generate, based on a combination of the respective prediction blocks associated with two or more TM candidates from the plurality of TM candidates, the prediction block for the current block.

20. Device according to claim 19, characterized in that one or more processors are configured to, as part of the generation of the prediction block for the current block, generate the prediction block for the current block as a linear combination of corresponding samples in the prediction blocks associated with the two or more TM candidates.

21. Device according to claim 20, characterized in that the generation of the prediction block for the current block as the linear combination of the corresponding samples comprises: Petition 870250084327, dated 09 / 18 / 2025, pp. 269 / 292 10 / 14 for each of the two or more TM candidates, calculating a template cost for the TM candidate, and generating the prediction block for the current block as a weighted average of the corresponding samples in the two or more TM candidates, wherein the weights used in the weighted average are based on the template cost for the TM candidate.

22. Device according to claim 21, characterized in that the weights are multiplicative inverses of the template matching costs for TM candidates.

23. Device according to claim 21, characterized in that a flag is displayed to indicate whether the weights are derived based on template cost or based on mean squared error minimization (MSE).

24. Device according to claim 20, characterized in that the syntax element is a first syntax element and the one or more processors are configured to, as part of generating the prediction block for the current block as the linear combination of the corresponding samples: determine a weight derivation method from among a plurality of weight derivation methods; apply the weight derivation method to derive weights from the two or more TM candidates; and generate the prediction block for the current block as the weighted average of the corresponding samples in the prediction blocks associated with the two or more TM candidates using the weights of the two or more TM candidates, Petition 870250084327, dated 09 / 18 / 2025, pp. 270 / 292 11 / 14 wherein a second syntax element indicates the weight derivation method.

25. Device according to claim 14, characterized in that: each fusion mode from a plurality of available fusion modes is associated with a different combination of two or more TM candidates in a plurality of TM candidates, a bitstream indicates a subset of the available fusion modes, the number of fusion modes in the subset is M, the number of available fusion modes is N, where M < N, the fusion modes in the subset of available fusion modes have minimum combined template costs among the available fusion modes, as part of the application of the subpel precision mode,One or more processors are configured to: for each fusion mode of the subset of available fusion modes; for each TM candidate of the two or more TM candidates associated with the fusion mode: apply the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective array of samples with full and sub-pel accuracy; and identify, within the respective array, a respective reference template of a respective prediction block for the TM candidate, where the respective reference template of the respective prediction block is the best match for the template of the current block within the respective array; and Petition 870250084327, dated 09 / 18 / 2025, pp. 271 / 292 12 / 14 generate, based on a combination of the respective prediction blocks for the two or more TM candidates associated with the fusion mode,a prediction block for the merge mode; and determine the prediction block for the current block from among the prediction blocks for the merge modes.

26. Device according to claim 14, characterized in that each fusion mode of a plurality of fusion modes is associated with a different combination of two or more TM candidates in a plurality of TM candidates, and the one or more processors are configured to, as part of the application of the subpel precision mode: for each fusion mode of the plurality of fusion modes: for each TM candidate of the two or more TM candidates associated with the fusion mode: apply the interpolation filter to samples from a reference region indicated by a motion vector predictor associated with the TM candidate to generate a respective sample array with full pel and subpel precision;and identify, within the respective array, a respective reference template of a respective prediction block for the TM candidate, wherein the respective reference template of the respective prediction block is the best match for the current block template within the respective array; for each respective weighting set from a plurality of weighting sets, Petition 870250084327, dated 09 / 18 / 2025, pp. 272 / 292 13 / 14 generate a prediction block for the respective weighting set as a weighted average of matching samples in the prediction blocks of the two or more TM candidates associated with the merge mode, wherein the weights used in the weighted average are included in the respective weighting set; and calculate a combined template cost of the prediction block for the respective weighting set;and select a weighting set for the merge mode from among the plurality of weighting sets based on a minimum of the combined template cost of the prediction blocks for the weighting sets; and determine the prediction block for the current block from among the prediction blocks for the weighting sets selected for the plurality of merge modes.

27. Device according to claim 14, characterized in that it further comprises a display configured to show decoded video data.

28. Device according to claim 14, characterized in that it comprises one or more of a camera, a computer, a mobile device, a broadcast receiver or a set-top box.

29. Device according to claim 14, characterized in that it comprises a video decoder. Petition 870250084327, dated 09 / 18 / 2025, pp. 273 / 292 14 / 14 30. Device according to claim 14, characterized in that it comprises a video encoder. Petition 870250084327, dated 09 / 18 / 2025, pp. 274 / 292