Multi-hypothesis prediction with template matching in video coding

By optimizing weight determination in the video decoder using template matching technology, the problems of signaling overhead and bandwidth utilization in the multi-hypothesis prediction process are solved, thereby improving video decoding efficiency.

CN120917733APending Publication Date: 2025-11-07QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024642.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-18
Filing Date
2024-04-19
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing video decoding technologies require significant signaling overhead and bandwidth utilization during multi-hypothesis prediction, resulting in low efficiency.

Method used

Template matching (TM) technology is used to determine the weights in the video decoder. The cost ranking of template matching reduces the amount of information sent for signaling and optimizes the multiple hypothesis prediction (MHP) process.

Benefits of technology

By reducing signaling overhead and bandwidth utilization, the efficiency and bandwidth utilization of video decoding are improved, thus enhancing the overall video decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917733A_ABST
    Figure CN120917733A_ABST
Patent Text Reader

Abstract

A method of encoding or decoding video data includes: for a multi-hypothesis prediction (MHP) process, determining a plurality of prediction templates based on a plurality of weights; comparing the plurality of prediction templates with a current template of the current block; determining a weight from the plurality of weights based on the comparison of the plurality of prediction templates and the current template of the current block; determining one or more prediction hypotheses; determining a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encoding or decoding the current block based on the prediction signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Patent Application No. 18 / 639,557, filed April 18, 2024, and U.S. Provisional Application No. 63 / 497,384, filed April 20, 2023, the entire contents of which are hereby incorporated by reference. U.S. Patent Application No. 18 / 639,557, filed April 18, 2024, claims the benefit of U.S. Provisional Application No. 63 / 497,384, filed April 20, 2023. TECHNICAL FIELD

[0002] The present disclosure relates to video encoding and video decoding. BACKGROUND

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio telephones, so-called “smart phones,” video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC), and extensions of such standards, as well as proprietary video codecs / formats such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. By implementing such video coding techniques, video devices can more efficiently send, receive, encode, decode, and / or store digital video information.

[0004] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) can be partitioned into video blocks, which can also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to other video blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture can be encoded using either spatial prediction relative to other video blocks in the same picture or temporal prediction relative to other video blocks in other reference pictures. Pictures can be referred to as frames, and reference pictures can be referred to as reference frames. SUMMARY

[0005] In general, this disclosure describes techniques for multi-hypothesis prediction (MHP) with template matching (TM). In an MHP process, to encode or decode a current block, a video coder (e.g., a video encoder and a video decoder) determines a prediction block (e.g., based on motion or a block vector or based on an intra mode) and determines one or more prediction hypotheses (e.g., prediction signals from other blocks). The video coder determines a prediction signal for the current block based on the prediction block and the one or more prediction hypotheses based on weights that can be applied to the prediction block and / or the one or more prediction hypotheses.

[0006] In one or more examples, a video coder can determine weights for an MHP process based on a template matching technique. For example, the video coder can determine a plurality of weights. In some examples, the video coder can select a weight of the plurality of weights that corresponds to a minimum template matching cost, which reduces the amount of information signaled to indicate the weights applied to the MHP process. In some examples, the video coder can rank a plurality of weights and corresponding hypotheses in a list based on corresponding template matching costs. A weight with a higher probability of application can be associated with a lower index in the list of weights. Signaling lower index values tends to require less bandwidth than higher index values. By ranking the weights as described, the amount of information signaled can be reduced.

[0007] In this way, example techniques can facilitate improvements to MHP processes for video coding. For example, example techniques provide practical applications of video coding techniques that can result in reduced bandwidth and signaling overhead.

[0008] In one example, this disclosure describes a method of encoding or decoding video data, the method comprising: determining, for a multi-hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights; comparing the plurality of prediction templates to a current template of a current block; determining a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determining one or more prediction hypotheses; determining a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encoding or decoding the current block based on the prediction signal.

[0009] In one example, the disclosure describes a device for encoding or decoding video data, the device comprising: one or more memories configured to store the video data; and processing circuitry coupled to the one or more memories, wherein the processing circuitry is configured to: determine, for a multi-hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights; compare the plurality of prediction templates to a current template of a current block; determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determine one or more prediction hypotheses; determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encode or decode the current block based on the prediction signal.

[0010] In one example, the disclosure describes a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine, for a multi-hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights; compare the plurality of prediction templates to a current template of a current block; determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determine one or more prediction hypotheses; determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encode or decode the current block based on the prediction signal.

[0011] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a block diagram illustrating an example video encoding and decoding system that can perform the techniques of this disclosure.

[0013] Figure 2 is a block diagram illustrating an example video encoder that can perform the techniques of this disclosure.

[0014] Figure 3 is a block diagram illustrating an example video decoder that can perform the techniques of this disclosure.

[0015] Figure 4 is a flowchart illustrating an example method for encoding a current block according to the techniques of this disclosure.

[0016] Figure 5 is a flowchart illustrating an example method for decoding a current block according to the techniques of this disclosure.

[0017] Figure 6 is a conceptual diagram illustrating an example of template matching on a search region around an initial motion vector (MV).

[0018] Figure 7 is a conceptual diagram illustrating an example of a template and reference samples of the template in a reference picture.

[0019] Figure 8 is a conceptual diagram illustrating an example of a template and reference samples of the template for a block with subblock motion using motion information of subblocks of the current block.

[0020] Figure 9A and Figure 9B are conceptual diagrams illustrating examples of merge mode with motion vector difference (MMVD) search points for a first reference picture list and a second reference picture list, respectively.

[0021] Figure 10 is a flowchart illustrating an example method of operation. DETAILED DESCRIPTION

[0022] In video coding, a video coder determines a prediction block for a current block. The video coder can determine the prediction block using a motion vector or block vector of the current block or based on an intra mode of the current block. From the prediction block, the video coder can determine a prediction signal used to encode or decode the current block.

[0023] A multi-hypothesis prediction (MHP) process is a video coding tool to determine a prediction signal for a current block based on a prediction block that provides a base prediction signal and one or more additional prediction signals, referred to as one or more prediction hypotheses. In the MHP process, the video coder applies weights to the base prediction signal and / or the one or more prediction hypotheses to determine the prediction signal for the current block.

[0024] A video encoder and a video decoder can each use the same technique to determine the prediction signal for the current block. The video encoder can determine residual values based on a difference between the prediction signal and the current block, and signal information indicative of the residual values. The video decoder can receive the information indicative of the residual values, and reconstruct the current block based on the residual values and the prediction signal (e.g., add the residual values to the prediction signal).

[0025] In some techniques, the video encoder can signal information indicative of weights applied to the MHP process to the video decoder. Such signaling can require additional signaling overhead and bandwidth utilization.

[0026] The disclosure describes example techniques that utilize template matching (TM) in MHP processes. By using TM, in some examples, a video decoder can determine weights for MHP processes without a video encoder signaling information indicating which weight to use. In some examples, by using TM, a video encoder and a video decoder can construct a list such that weights with higher application probabilities are at lower index values than weights with lower application probabilities. The amount of information needed to signal lower index values is less than the amount of information needed to signal higher index values. By using TM to construct the list, there is a higher probability of a reduction in the amount of information needed to signal to indicate applied weights than if TM is not used to construct the list for MHP processes.

[0027] In this way, example techniques can facilitate efficient bandwidth utilization and improve overall video coding processes, such as video coding using MHP processes. Moreover, example techniques can also apply MHP processes where one or more prediction hypotheses are derived from a merge list or an advanced motion vector predictor (AMVP) list.

[0028] For example, a video coder (e.g., a video encoder or a video decoder) can determine, for a multiple hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights. As one example, the video coder can determine a reference template (e.g., a template identified from an initial vector for a current block) and can determine a hypothesis template (e.g., a template identified from a vector used to identify a prediction hypothesis). The reference template can be referred to as T ref , and the hypothesis template can be referred to as T add .

[0029] The video coder can determine a first prediction template by applying a weight to the reference template and the hypothesis template based on a first weight of the plurality of weights, and determine a second prediction template by applying a weight to the reference template and the hypothesis template based on a second weight of the plurality of weights. For example, a first prediction template (referred to as T pred1 ) can be equal to (1 - a1)T ref + a1T add , where a1 is the first weight. A second prediction template (referred to as T pred2 ) can be equal to (1 - a2)T ref + a2T add , where a2 is the second weight.

[0030] The video coder can repeat such techniques for each weight of the plurality of weights (e.g., a3, a4, etc.). In some examples, a video encoder can signal to identify a particular hypothesis template (e.g., T addThe video decoder can receive this information, and the video decoder can repeat this technique for each of the multiple weights of a particular hypothetical template.

[0031] However, in some examples, the video decoder can determine the identity for multiple hypothetical templates (e.g., T). add1 T add2 T for each hypothetical template in (etc.) add Instead of the existence of a T add The video decoder can repeat the example techniques described above for each of a plurality of hypothesis templates to determine a plurality of prediction templates. That is, the video decoder can iterate through all weights in a plurality of weights for a first hypothesis template. The video decoder can repeat these techniques for all hypothesis templates.

[0032] For example, a video decoder can be based on applications such as T add1 Each of the multiple weights determines T. pred The first set. Video decoders can be based on applications such as T. add2 Each of the multiple weights determines T. pred The second set, and so on.

[0033] In the example above, the video decoder may have determined multiple prediction templates. The video decoder can compare the multiple prediction templates with the current template of the current block. For example, for comparison, the video decoder can determine the template matching cost of each prediction template in the prediction templates (e.g., the sum of absolute differences (SAD) value or a value obtained using some other technique between the current template and each prediction template in the prediction templates).

[0034] In one or more examples, the video decoder may determine the weights from a plurality of weights that result in the lowest template matching cost, and use these weights to fuse predictive hypotheses to generate a predictive signal. For example, the video encoder signals a specific hypothetical template (e.g., T) to... add In an example where the video decoder receives information about a specific hypothetical template, the video decoder can use a sampling technique to determine which weights result in the lowest template matching cost. In an example where the video encoder does not signal information about a particular hypothetical template, the video decoder can use the sampling technique across multiple hypothetical templates to determine the weights and which predictive hypotheses should be used to generate the predictive signal.

[0035] In some examples, the video coder can order the list based on the template matching cost. A prediction template or prediction hypothesis associated therewith having a lower temporal matching cost is identified at a lower index in the list than a prediction template or prediction hypothesis associated therewith having a higher temporal matching cost. The video encoder can signal the index into the list and the video decoder can receive the index and the video decoder can determine the weight or possible weights applied to generate the prediction signal or prediction hypotheses.

[0036] The example techniques described in this disclosure can be applied as extensions to any of the existing video codecs, such as HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), Essential Video Coding (EVC), or can be a high efficiency coding tool in future video coding standards, e.g., ECM (Enhanced Compression Model).

[0037] Figure 1 is a block diagram illustrating an example video encoding and decoding system 100 that can perform the techniques of this disclosure. The techniques of this disclosure generally relate to coding (encoding and / or decoding) video data. In general, video data includes any data for processing video. Thus, video data can include uncoded raw video, coded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0038] As Figure 1 shown in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. In particular, source device 102 provides the video data to destination device 116 via a computer- readable medium 110. Source device 102 and destination device 116 can be or include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, mobile devices, tablet computers, set-top boxes, handheld phones such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, broadcast receiver devices, and the like. In some cases, source device 102 and destination device 116 can be equipped to communicate wirelessly and thus can be referred to as wireless communication devices.

[0039] In Figure 1In the example of FIG. 1, source device 102 includes video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes input interface 122, video decoder 300, memory 120, and display device 118. In accordance with this disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 can be configured to apply techniques for utilizing template matching (TM) in a multi-hypothesis prediction (MHP) process. Thus, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, a source device and a destination device can include other components or arrangements. For example, source device 102 can receive video data from an external video source, such as an external camera. Likewise, destination device 116 can interface with an external display device, rather than include an integrated display device.

[0040] As shown in system 100, video source 104 can provide the video data to be encoded by video encoder 200. In general, video source 104 represents a source of video data to be encoded. For example, video source 104 can include a video camera that captures a video image, which is then stored in memory 106 as an uncompressed video file. As another example, video source 104 can include a modem or other communication interface that receives a video file from another device. In some cases, video source 104 can include a video archive that stores previously captured or received video files. In general, video source 104 can represent a source of video data that is to be encoded by video encoder 200. Figure 1 System 100 as shown in FIG. 1 is merely one example. In general, any digital video encoding and / or decoding device can perform techniques for utilizing TM in a MHP process. Source device 102 and destination device 116 are merely examples of such coding devices in which source device 102 generates coded video data for transmission to destination device 116. This disclosure refers to a "coding" device as a device that performs coding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of coding devices, specifically, a video encoder and a video decoder, respectively. In some examples, source device 102 and destination device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Hence, system 100 can support one-way or two-way video transmission between source device 102 and destination device 116, e.g., for video streaming, video playback, video broadcasting, or video telephony.

[0041] In general, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes data for the pictures. Video source 104 of source device 102 can include a video capture device, such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface to receive video from a video content provider. As a further alternative, video source 104 can generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 can rearrange the pictures from the received order (sometimes referred to as “display order”) into the coding order for coding. Video encoder 200 can generate a bitstream including encoded video data. Source device 102 can then output the encoded video data via output interface 108 onto computer-readable medium 110 for reception and / or retrieval by, for example, input interface 122 of destination device 116.

[0042] Memory 106 of source device 102 and memory 120 of destination device 116 represent general purpose memories. In some examples, memories 106, 120 can store raw video data, e.g., raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106, 120 can store software instructions capable of execution by, e.g., video encoder 200 and video decoder 300, respectively. While memories 106 and 120 are shown as separate from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 can also include internal memories for similar or equivalent purposes. Furthermore, memories 106, 120 can store encoded video data, e.g., from an output of video encoder 200 and an input of video decoder 300. In some examples, portions of memories 106, 120 can be allocated as one or more video buffers, e.g., to store raw, decoded, and / or encoded video data.

[0043] Computer-readable medium 110 can represent any type of medium or device capable of storing coding video data for transmission from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium to enable source device 102 to transmit encoded video data directly to destination device 116 in real-time, e.g., via a radio frequency network or computer-based network. Output interface 108 can modulate a transmission signal including the encoded video data, and input interface 122 can demodulate received transmission signals, according to a communication standard, such as a wireless communication protocol. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from source device 102 to destination device 116.

[0044] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.

[0045] In some examples, source device 102 can output encoded video data to file server 114, or another intermediate storage device, which can store the encoded video data generated by source device 102. Destination device 116 can access stored video data from file server 114 via streaming or download.

[0046] The file server 114 can be any type of server device capable of storing encoded video data and transmitting that encoded video data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a server configured to provide file delivery protocol services such as File Delivery Protocol (FTP) or the File Delivery over Unidirectional Transport (FLUTE) protocol, a content delivery network (CDN) device, a hypertext transfer protocol (HTTP) server, a multimedia broadcast multicast service (MBMS) or enhanced MBMS (eMBMS) server, and / or a network attached storage (NAS) device. The file server 114 can additionally or alternatively implement one or more HTTP streaming protocols, such as Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), Real Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, and the like.

[0047] The destination device 116 can access the encoded video data from the file server 114 through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on the file server 114. The input interface 122 can be configured to operate according to any one or more of various protocols discussed above for retrieving or receiving media data from the file server 114, or other such protocols for retrieving media data.

[0048] The output interface 108 and the input interface 122 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 comprise wireless components, the output interface 108 and the input interface 122 can be configured to transfer data, such as encoded video data, according to a cellular communication standard, such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, or the like. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 can be configured to transfer data, such as encoded video data, according to other wireless standards ™ ™ ​Standard, etc. In some examples, source device 102 and / or destination device 116 can include respective system on a chip (SoC) devices. For example, source device 102 can include SoC devices to perform the functionality attributable to video encoder 200 and / or output interface 108, and destination device 116 can include SoC devices to perform the functionality attributable to video decoder 300 and / or input interface 122.

[0049] The techniques of this disclosure can be applied to video coding in support of any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions, such as dynamic adaptive streaming over HTTP (DASH), digital video that is encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0050] Input interface 122 of destination device 116 receives an encoded video bitstream from computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or the like). The encoded video bitstream can include signaling information defined by video encoder 200, and also used by video decoder 300, such as syntax elements having values

[0051] Although in Figure 1Although not shown, in some examples, video encoder 200 and video decoder 300 can each be integrated with an audio encoder and / or audio decoder (e.g., an audio codec), and can include appropriate MUX-DEMUX units or other hardware and / or software to handle multiplexed streams including both audio and video in a common data stream. Example audio codecs can include AAC, AC-3, AC-4, ALAC, ALS, AMBE, AMR, AMR-WB (G.722.2), AMR-WB+, aptx (various versions), ATRAC, BroadVoice (BV16, BV32), CELT, Enhanced AC-3 (E-AC-3), EVS, FLAC, G.711, G.722, G.722.1, G.722.2 (AMR-WB), G.723.1, G.726, G.728, G.729, G.729.1, GSM-FR, HE-AAC, iLBC, iSAC, LA Lyra, Monkey's Audio, MP1, MP2 (MPEG-1, 2 Audio Layer II), MP3, Musepack, Nellymoser Asao, OptimFROG, Opus, Sac, Satin, SBC, SILK, Siren 7, Speex, SVOPC, True Audio (TTA), TwinVQ, USAC, Vorbis (Ogg), WavPack, and Windows Media Audio.

[0052] Video encoder 200 and video decoder 300 each can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof. When the techniques are implemented partially in software, a device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of video encoder 200 and video decoder 300 can be included in one or more encoders or decoders, any of which alone can be a part of a combined encoder / decoder (CODEC). An apparatus including video encoder 200 and / or video decoder 300 can implement video encoder 200 and / or video decoder 300 in a processing circuitry, such as an integrated circuit and / or a microprocessor. Such an apparatus can be a wireless communication device, such as a cellular phone or any of the other types of apparatus described herein.

[0053] Video encoder 200 and video decoder 300 can operate according to video coding standards, such as ITU-T H.265, also referred to as High Efficiency Video Coding (HEVC), or extensions thereof, such as the multi-view and / or scalable video coding extensions. Alternatively, video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards, such as ITU-T H.266, also referred to as Versatile Video Coding (VVC). In other examples, video encoder 200 and video decoder 300 can operate according to proprietary video codecs / formats, such as AOMedia Video 1 (AV1), extensions of AV1, and / or subsequent versions of AV1 (e.g., AV2). In other examples, video encoder 200 and video decoder 300 can operate according to other proprietary formats or industry standards. The techniques of this disclosure, however, are not limited to any particular coding standard or format. Generally, video encoder 200 and video decoder 300 can be configured to perform the techniques of this disclosure in connection with any video coding technology that uses MHP processes. For example, this disclosure describes examples that use TM as part of MHP processes, which can improve MHP processes.

[0054] In general, video encoder 200 and video decoder 300 can perform block-based coding of pictures. The term “block” generally refers to a structure containing data to be processed (e.g., encoded, decoded, or otherwise used) during the encoding and / or decoding process. For example, a block can include a two-dimensional matrix of samples of luma and / or chroma data. In general, video encoder 200 and video decoder 300 can code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data for samples of a picture, video encoder 200 and video decoder 300 can code luminance and chrominance components, where the chrominance components can include both red hue chrominance components and blue hue chrominance components. In some examples, video encoder 200 converts received RGB format data to a YUV representation prior to encoding, and video decoder 300 converts the YUV representation to the RGB format. Alternatively, pre- and post-processing units (not shown) can perform these conversions.

[0055] The disclosure can generally relate to coding (e.g., encoding and decoding) of pictures to include processes that encode or decode data of pictures. Similarly, the disclosure can relate to coding of blocks of pictures to include processes that encode or decode (e.g., prediction and / or residual coding) data for blocks. An encoded video bitstream generally includes a series of values for syntax elements that represent coding decisions (e.g., coding modes) and partitioning of pictures into blocks. Thus, a reference to coding of a picture or block should generally be understood to be a reference to coding values of syntax elements that form the picture or block.

[0056] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) partitions a coding tree unit (CTU) into CUs according to a quad tree structure. That is, the video coder partitions a CTU and a CU into four equal, non overlapping squares, and each node of the quad tree has either zero or four child nodes. Nodes with zero child nodes can be referred to as“leaf nodes,” and CUs of such leaf nodes can include one or more PUs and / or one or more TUs. The video coder can further partition PUs and TUs. For example, in HEVC, a residual quad tree (RQT) represents partitioning of TUs. In HEVC, PUs represent inter prediction data, while TUs represent residual data. Intra predicted CUs include intra prediction information, such as an intra mode indication.

[0057] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, a video coder (such as video encoder 200) partitions a picture into CTUs. Video encoder 200 can partition a CTU according to a tree structure such as a quad-tree binary tree (QTBT) structure or Multi-Type Tree (MTT) structure. The QTBT structure removes the concepts of multiple partition types, such as the separation between CUs, PUs, and TUs of HEVC. The QTBT structure includes two levels: a first level of partitioning according to quad tree partitioning, and a second level of partitioning according to binary tree partitioning. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to CUs.

[0058] In the MTT partitioning structure, blocks can be partitioned using quad tree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) (also referred to as tri-tree (TT)) partitioning. A ternary tree or tri-tree partitioning is a partitioning in which a block is split into three sub-blocks. In some examples, a ternary tree or tri-tree partitioning divides a block into three sub-blocks without dividing the original block through a center. The partitioning types (e.g., QT, BT, and TT) in the MTT can be symmetric or asymmetric.

[0059] When operating according to the AV1 codec, video encoder 200 and video decoder 300 can be configured to code video data in units of blocks. In AV1, the largest coding block that can be processed is referred to as a superblock. In AV1, a superblock can be 128x128 luma samples or 64x64 luma samples. However, in subsequent video coding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luma sample sizes. In some examples, a superblock is the top level of a block quad tree. Video encoder 200 can further partition a superblock into smaller coding blocks. Video encoder 200 can partition superblocks and other coding blocks into smaller blocks using square or non-square partitions. Non-square blocks can include N / 2xN blocks, NxN / 2 blocks, N / 4xN blocks, and NxN / 4 blocks. Video encoder 200 and video decoder 300 can perform separate prediction and transform processing for each coding block.

[0060] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that can be coded independently of other tiles. That is, video encoder 200 and video decoder 300 can encode and decode coding blocks within a tile without using video data from other tiles. However, video encoder 200 and video decoder 300 can perform filtering across tile boundaries. The size of a tile can be uniform or non-uniform. Tile-based coding can enable parallel processing and / or multi-threading of encoder and decoder implementations.

[0061] In some examples, video encoder 200 and video decoder 300 can use a single QTBT or MTT structure to represent each of luma and chroma components, while in other examples, video encoder 200 and video decoder 300 can use two or more QTBT or MTT structures, such as one QTBT / MTT structure for luma components and another QTBT / MTT structure for two chroma components (or two QTBT / MTT structures for respective chroma components).

[0062] Video encoder 200 and video decoder 300 can be configured to use quad tree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partition structures.

[0063] In some examples, a CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a monochrome picture or a CTB of samples of a picture coded using three separate color planes and syntax structures for coding samples. A CTB can be an NxN block of samples of some N value, such that one partitioning is to divide components into CTBs. A component is an array or a single sample from one of the three arrays (luma and two chroma) that make up a 4:2:0, 4:2:2, or 4:4:4 color format picture, or an array or a single sample of an array that makes up a monochrome format picture. In some examples, a coding block is an MxN block of samples of some M value and N value, such that one partitioning is to divide a CTB into coding blocks.

[0064] Blocks (e.g., CTUs or CUs) can be grouped in pictures in various ways. As one example, a brick can refer to a rectangular region of CTU rows within a particular tile in a picture. A tile can be a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular region of CTUs having a height equal to a height of a picture and a width specified by a syntax element (e.g., such as in a picture parameter set). A tile row refers to a rectangular region of CTUs having a height specified by a syntax element (e.g., such as in a picture parameter set) and a width equal to a width of a picture.

[0065] In some examples, a tile can be divided into multiple bricks, each of which can include one or more CTU rows within the tile. A tile that is not divided into multiple bricks can also be referred to as a brick. However, a brick that is a true subset of a tile cannot be referred to as a tile. Bricks in a picture can also be arranged in slices. A slice can be an integer number of bricks of a picture, which can be uniquely contained in a single network abstraction layer (NAL) unit. In some examples, a slice includes multiple complete tiles or only a contiguous sequence of complete bricks of one tile.

[0066] The disclosure can use“NxN” and“N by N” interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions, e.g., 16x16 samples or 16 by 16 samples. In general, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Likewise, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a nonnegative integer value. The samples in a CU can be arranged in rows and columns. Moreover, a CU need not necessarily have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can comprise NxM samples, where M need not necessarily equal N.

[0067] Video encoder 200 encodes video data representing prediction and / or residual information for CUs, among other information. Prediction information indicates how to predict a CU in order to form a prediction block for the CU. Residual information generally represents sample-by-sample differences between the CU prior to encoding and the prediction block.

[0068] To predict a CU, video encoder 200 generally forms a prediction block for the CU through inter prediction or intra prediction. Inter prediction generally refers to predicting the CU from data of a previously coded picture, whereas intra prediction generally refers to predicting the CU from previously coded data of the same picture. To perform inter prediction, video encoder 200 can use one or more motion vectors to generate the prediction block. Video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, e.g., according to a difference between the CU and the reference block. Video encoder 200 can calculate the difference metric using a sum of absolute difference (SAD), sum of squared difference (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether a reference block closely matches a current CU. In some examples, video encoder 200 can use uni -prediction or bi-prediction to predict a current CU.

[0069] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter prediction mode. Under the affine motion compensation mode, video encoder 200 can determine two or more motion vectors that represent non-translational motion, such as scaling or zooming, rotation, perspective motion, or other irregular types of motion.

[0070] To perform intra prediction, video encoder 200 can select an intra prediction mode to generate the prediction block. Some examples of VVC provide sixty-seven intra prediction modes, including various directional modes, as well as a planar mode and a DC mode. Generally, video encoder 200 selects an intra prediction mode that describes neighboring samples of a current block (e.g., a block of a CU) from which to predict samples of the current block. Such samples can generally be located above, above and to the left, or to the left of the current block in the same picture as the current block, assuming video encoder 200 is coding CTUs and CUs in a raster scan order (left-to-right, top-to-bottom).

[0071] Video encoder 200 encodes data representing the prediction mode of a current block. For example, for inter prediction modes, video encoder 200 can encode data indicating which of various available inter prediction modes to use, as well as motion information for the corresponding mode. For example, for uni - or bi-prediction, video encoder 200 can encode motion vectors using advanced motion vector prediction (AMVP) or merge mode. Video encoder 200 can use similar modes to encode motion vectors for the affine motion compensation mode.

[0072] AV1 includes two general techniques for encoding and decoding blocks of video data. The two general techniques are intra prediction (e.g., intra prediction or spatial prediction) and inter prediction (e.g., inter prediction or temporal prediction). In the context of AV1, when using an intra prediction mode to predict a block of a current frame of video data, video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra prediction modes, video encoder 200 encodes the block of the current frame based on differences between sample values in the current block and prediction values generated from reference samples in the same frame. Video encoder 200 determines the prediction values generated from the reference samples based on the intra prediction mode.

[0073] After prediction, such as intra prediction or inter prediction of a block, video encoder 200 can calculate residual data for the block. The residual data, such as a residual block, represents sample-by-sample differences between the block and a prediction block formed using the corresponding prediction mode. Video encoder 200 can apply one or more transforms to the residual block to produce transform data in a transform domain instead of in the sample domain. For example, video encoder 200 can apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform. In addition, video encoder 200 can apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), or the like. Video encoder 200 produces transform coefficients after applying the one or more transforms.

[0074] As noted above, after any transforms that produce transform coefficients, video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, providing further compression. By performing the quantization process, video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, video encoder 200 can round the values of the transform coefficients down to lower precision n bit values to m bit values, where n is greater than m In some examples, to perform quantization, video encoder 200 can perform a bitwise right-shift of the values to be quantized.

[0075] After quantization, video encoder 200 can scan the transform coefficients, producing a one-dimensional vector from the two-dimensional matrix comprising the quantized transform coefficients. The scan can be designed to place higher energy (and hence less frequent) transform coefficients earlier in the vector and lower energy (and hence more frequent) transform coefficients later in the vector. In some examples, video encoder 200 can utilize a pre-defined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 can entropy encode the one-dimensional vector, e.g., according to context adaptive binary arithmetic coding (CABAC). Video encoder 200 can also entropy encode values for syntax elements that describe metadata associated with encoded video data, which is used by video decoder 300 when decoding the video data.

[0076] To perform CABAC, video encoder 200 can assign a context within a context model to a symbol to be transmitted. The context can relate to, for example, whether neighboring values of the symbol are zero-valued or not. The probability determination can be based on the context assigned to the symbol.

[0077] Video encoder 200 can further generate syntax data, such as block-based, picture-based, and sequence-based syntax data, to video decoder 300, e.g., in picture headers, block headers, slice headers, or generate other syntax data such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS), for example. Video decoder 300 can likewise decode such syntax data to determine how to decode corresponding video data.

[0078] In this way, video encoder 200 can generate a bitstream that includes encoded video data, e.g., syntax elements that describe partitioning of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Ultimately, video decoder 300 can receive the bitstream and decode the encoded video data.

[0079] In general, video decoder 300 performs a reciprocal process to that performed by video encoder 200 to decode the encoded video data of the bitstream. For example, video decoder 300 can use CABAC in substantially a reciprocal manner, but reversed, to video encoder 200’s CABAC encoding process to decode values for syntax elements of the bitstream. The syntax elements can define partitioning information for partitioning a picture into CTUs, and partitioning each CTU according to a corresponding partition structure such as a QTBT structure to define CUs of the CTU. The syntax elements can further define prediction and residual information for blocks (e.g., CUs) of video data.

[0080] The residual information can be represented by, for example, quantized transform coefficients. Video decoder 300 can inverse quantize and inverse transform the quantized transform coefficients of the block to reproduce a residual block for the block. Video decoder 300 uses the signaled prediction mode (intra prediction or inter prediction) and related prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. Video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. Video decoder 300 can perform additional processing such as performing a deblocking process to reduce visual artifacts along boundaries of the blocks.

[0081] This disclosure can generally relate to “signaling” certain information, such as syntax elements. The term “signaling” can generally refer to the communication of values for syntax elements and / or other data used to decode encoded video data. That is, video encoder 200 can signal values for syntax elements in a bitstream. In general, signaling refers to generating values in a bitstream. As noted above, source device 102 can transmit the bitstream to destination device 116 in real time or not in real time, such as can occur when syntax elements are stored to storage device 112 for later retrieval by destination device 116.

[0082] According to techniques of this disclosure, video encoder 200 and video decoder 300 can be configured to perform MHP processes with TM. MHP is described in greater detail below.

[0083] Multiple hypothesis prediction (MHP)

[0084] In example video coding standards, temporally predicted blocks can employ uni-prediction or bi-prediction. Multi-hypothesis prediction (MHP) was proposed in Winken et al., “Multi-Hypothesis Inter Prediction,” JVET-J0041, and later in Winken et al., “CE 10: Multi-hypothesis inter prediction (Tests 1.5-1.8),” JVET-K0269, Winken et al., “CE 10: Multi-hypothesis inter prediction (Tests 1.2.a-1.2.c),” JVET-L0148, and Winken et al., “CE 10: Multi-hypothesis inter prediction (Test 10.1.2),” JVET-M0425. In MHP, the inter prediction method allows for a weighted superposition of more than two motion-compensated prediction signals. The resulting overall prediction signal is obtained by a sample-wise weighted superposition. With the uni-prediction / bi-prediction signal (referred to as base prediction signal) and a first additional inter prediction signal / hypothesis , the resulting prediction signal is obtained as follows: .

[0085] The weighting factor is specified by the syntax element add_hyp_weight_idx according to the following mapping:

[0086] In some examples, more than one additional prediction signal can be used. The resulting overall prediction signal is accumulated iteratively with each additional prediction signal.

[0087] .

[0088] The resulting overall prediction signal can be obtained as the last (i.e., having the largest index ). For inter prediction blocks using MERGE mode (but in some cases non-SKIP mode), an additional inter prediction signal can also be specified. For the additional prediction signal, one of the two AMVP candidate lists is used: a. If the POC of the reference picture of the additional prediction signal is equal to the POC of the used List 1 reference picture (e.g., the second reference picture list), the List 1 AMVP candidate list is used; b. Otherwise, the List 0 (e.g., the first reference picture list) AMVP candidate list is used.

[0089] The motion parameters of each additional prediction hypothesis can be explicitly signaled by specifying the reference index, the motion vector predictor index, and the motion vector difference, or implicitly signaled by specifying the merge index. A separate multi-hypothesis merge flag distinguishes the two signaling modes. In ECM (enhanced compression model), the number of additional hypothesis prediction signals is set to 2.

[0090] In other words, video encoder 200 and video decoder 300 can determine a base prediction signal, where the base prediction signal is based on a motion vector, a block vector, or an intra prediction mode. Video encoder 200 and video decoder 300 can also determine one or more prediction hypotheses. The prediction hypotheses can be additional prediction signals. As one example, video encoder 200 and video decoder 300 can determine the motion vector of each of the entries in the merge list or the AMVP list. Video encoder 200 and video decoder 300 can determine which blocks the motion vectors in the merge list or the AMVP list point to, and determine the prediction hypotheses (e.g., additional prediction signals) based on the blocks that the motion vectors in the merge list or the AMVP list point to. The above is one example way of determining one or more prediction hypotheses, and other techniques can also exist.

[0091] As an example, with the base prediction signal and one or more additional prediction signals (e.g., from the prediction hypotheses), video encoder 200 and video decoder 300 can generate a prediction signal for the current block using the above equation for p n+1 In some examples, video encoder 200 can signal and video decoder 300 can receive the add_hyp_weight_idx syntax element to determine the value of a (e.g., the weight parameter). In some examples, video encoder 200 can signal and video decoder 300 can receive information identifying which prediction hypotheses to use. For example, video encoder 200 can signal and video decoder 300 can receive the indices into the merge list or the AMVP list, and use that information to determine which prediction hypotheses to use.

[0092] To have the video encoder 200 signal the add_hyp_weight_idx syntax element and have the video decoder 300 parse the syntax element, signaling bandwidth and overhead is required. As described in greater detail, the present disclosure describes example techniques of using template matching to determine weights applied to one or more prediction hypotheses (e.g., additional prediction signals), which can reduce the signaling bandwidth and overhead. Moreover, in some examples, using template matching, the video decoder 300 can determine which prediction hypotheses to use without signaling from the video encoder 200 or signaling in a manner that utilizes less bandwidth, which further improves bandwidth utilization.

[0093] Template matching prediction

[0094] Template matching (TM) prediction is a special merge mode based on the frame rate up conversion (FRUC) technique. With this mode, the motion information of a block is not signaled but derived at the decoder side by the video decoder 300. TM prediction is applied to the AMVP (advanced motion vector predictor) mode and the regular merge mode. In the AMVP mode, the MVP (motion vector predictor) candidate selection is determined based on template matching to pick the candidate that achieves the smallest difference between the current block template and the reference block template. In the regular merge mode, a TM mode flag is signaled to indicate the use of TM, and then TM is applied to the merge candidate indicated by the merge index for MV refinement.

[0095] As shown in FIG. 6, Figure 6 Template matching is used to derive the motion information of a current CU by finding the closest match between a template (top and / or left neighboring blocks of the current CU) in the current picture and a block (same size as the template) in the reference picture. For example, Figure 6 A current frame 600 (e.g., current picture) with a current CU 602 is illustrated, which has a current template 604 including an upper template 604A and a left template 604B. A reference frame 606 (e.g., reference picture) includes a reference template as illustrated.

[0096] In the case that an AMVP candidate is selected based on the initial matching error, the MVP of the AMVP candidate is refined by template matching. In the case that a merge candidate is indicated by the signaled merge index, the merge MVs of the merge candidate corresponding to L0 (first reference picture list) and L1 (second reference picture list) are independently refined by template matching, and then the less accurate MV is further refined again with the better MV as a priori value.

[0097] For the cost function, motion compensation interpolation is needed when the motion vector points to fractional sample positions. To reduce complexity, instead of regular 8-tap DCT-IF interpolation, bilinear interpolation is used for template matching to generate the template on the reference picture. The matching cost of template matching is calculated as follows: .

[0098] In the above, is a weighting factor set to 4 empirically, and indicate the current test MV and the initial MV (i.e., the MVP candidate in AMVP mode or the merge motion in merge mode, respectively. SAD is used as the matching cost of template matching.

[0099] When TM is used, the motion can be refined by using only the luma samples. The derived motion can be used for both luma and chroma for MC inter prediction. After the MV is decided, the final MC is performed using 8-tap interpolation filter for luma and 4-tap interpolation filter for chroma.

[0100] For the search method, MV refinement is a pattern-based MV search with criteria of template matching cost and hierarchical structure. Two search patterns can be supported - diamond search and cross search for MV refinement. The hierarchical structure specifies the iterative process to refine the MV, starting with coarse MVD precision (e.g., quarter-pixel) and ending with fine MVD precision (e.g., 1 / 8-pixel). The MV is searched with quarter-luma-sample MVD precision in diamond pattern directly, followed by quarter-luma-sample MVD precision in cross pattern, and then eighth-luma-sample MVD refinement in cross pattern. The search range of MV refinement is set to equal to (-8, +8) luma samples around the initial MV. When the current block has bi-prediction, the two MVs are refined independently, and then the best one (in terms of matching cost) of the two MVs is set as the prior value to further refine the other MV with BCW weight value.

[0101] Merge candidate adaptive reordering (ARMC)

[0102] In ECM, the merge candidates are adaptively reordered with template matching (TM). This reordering method is applied to regular merge candidate list, TM merge candidate list, and affine merge candidate list (subblock merge candidate list except for SbTMVP candidate). For TM merge mode, the merge candidates are reordered before the TM refinement process.

[0103] After constructing the merge candidate list, the merge candidates are divided into several subgroups. For regular merge mode and TM merge mode, the subgroup size is set to 5. For affine merge mode, the subgroup size is set to 3. The merge candidates in each subgroup are reordered incrementally according to the TM-based cost value. For simplicity, the merge candidates in the last subgroup, not the first subgroup, are not reordered.

[0104] The TM cost of a merge candidate is measured by the sum of absolute difference (SAD) between the samples of the template of the current block and their corresponding reference samples. The template includes a set of reconstructed samples neighboring the current block. The reference samples of the template are located by the motion information of the merge candidate.

[0105] When a merge candidate utilizes bi-prediction, the reference samples of the template of the merge candidate are also generated by bi-prediction, as shown in Figure 7 For example, Figure 7 A current picture 700 including a current block 702 with a template 703 is shown. A reference block 704 is in a reference picture in the reference picture list 1 with a reference template 708. A reference block 706 is in a reference picture in the reference picture list 0 with a reference template 710.

[0106] For a subblock-based merge candidate with a subblock size equal to Wsubx Hsub, the top template includes several sub-templates with a size of Wsubx 1, and the left template includes several sub-templates with a size of 1 x Hsub. As shown in Figure 8 The motion information of the subblocks in the first row and the first column of the current block is used to derive the reference samples of each sub-template. For convenience, Figure 8 The same reference numerals are used. Figure 7 The same reference numerals are used.

[0107] Merge mode with MVD (MMVD)

[0108] In VVC and ECM, in addition to the merge mode (in which the implicitly derived motion information is directly used for the prediction sample generation of the current CU), a merge mode with motion vector difference (MMVD) is also introduced. The MMVD flag is signaled immediately after the regular merge flag to specify whether the MMVD mode is used for the CU.

[0109] In MMVD, after the selection of the merge candidate, the merge candidate can be further refined by the signaled MVD information. The additional information includes a merge candidate flag, an index for specifying the motion magnitude value, and an index for the indication of the motion direction. In MMVD mode, one of the first two candidates in the merge list is selected to be used as the MV basis. The MMVD candidate flag can be signaled to specify which one is used between the first merge candidate and the second merge candidate.

[0110] The distance index specifies the motion value information and indicates a predefined offset from the starting point. As shown in Figure 9A and Figure 9B The offset is added to the horizontal component or vertical component of the starting MV. The relationship of the distance index and the predefined offset is specified in Table 1.

[0111] Table 1 - Relationship of distance index and predefined offset

[0112] The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions as shown in Table 2. The meaning of the sign of the MVD can vary depending on the information of the starting MV. When the starting MV is a uni-predicted MV or a bi-predicted MV and both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture or both are less than the POC of the current picture), the sign in Table 2 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bi-predicted MV and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), and the difference of the POC in list 0 (first reference picture list) is greater than the difference in list 1, the sign in Table 2 specifies the sign of the MV offset added to the list 0 MV component of the starting MV and the sign of the list 1 MV has the opposite value. Otherwise, if the difference of the POC in list 1 (second reference picture list) is greater than list 0, the sign in Table 2 specifies the sign of the MV offset added to the list 1 MV component of the starting MV and the sign of the list 0 MV has the opposite value.

[0113] The MVD is scaled according to the POC difference in each direction. If the difference of the POC in both lists is the same, no scaling is needed. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, the MVD of list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb. If the POC difference of L1 is greater than L0, the MVD of list 0 is scaled in the same way. If the starting MV is uni-predicted, the MVD is added to the available MV.

[0114] Table 2 - Sign of MV offset specified by direction index

[0115] Currently in ECM, the weight index for multi-hypothesis prediction (MHP) is signaled regardless of whether the MHP predictor is derived from the merge list or the AMVP list. According to one or more examples described in this disclosure, template matching techniques can be used with MHP. That is, video encoder 200 and video decoder 300 can apply template matching techniques to MHP to provide additional information to further improve the coding performance of the MHP coding tool.

[0116] For example, the template matching cost can be an effective measure that describes how similar a predictor (e.g., a prediction block) is compared to a current block. The templates of the current block and the reference block are defined as described above with respect to template matching prediction and Figure 6 In some examples, a sum of absolute difference (SAD) value is computed and used as the template matching cost between the current template and the reference template.

[0117] For template matching based MHP weight selection for hypotheses derived from the merge list (as described above for multi-hypothesis prediction (MHP)), the additional hypothesis in MHP can be derived from either the merge list or the AMVP list. When the additional hypothesis is derived from the merge list, the MHP weight can be selected based on the template matching cost, and thus the signaling overhead for MHP weight index is saved. Given the template of the current block, the template of the reference block, and the template of the additional hypothesis predictor defined as , and the template of the total prediction is defined as .

[0118] In the above, is the current weight being examined. The template matching cost between is generated by computing the SAD of the two templates. For a total of n possible weights, the n TM costs can be derived, denoted as In one example, the MHP weight is selected as the weight corresponding to the smallest TM cost among all In some examples, only the first m weight indices are kept, where m n are then signaled. In this way, the signaling overhead of the MHP weight index can be reduced.

[0119] For example, video encoder 200 and video decoder 300 can determine a reference block (e.g., based on the motion vector of the current block). Video encoder 200 and video decoder 300 can determine a reference template for the reference block. The reference template is referred to as T​ref The reference template can include samples neighboring the reference block.

[0120] Video encoder 200 and video decoder 300 can also determine one or more prediction hypotheses. As one example, video encoder 200 can signal information identifying the prediction hypotheses and video decoder 300 can receive the information (such as based on an index in a merge list or an AMVP list). As another example, video encoder 200 and video decoder 300 can select the top N entries in a merge list or an AMVP list. There can be other ways of determining the prediction hypotheses, such as with MMVD extensions, ARMC, and reordering, as will be described in more detail below.

[0121] For each of the prediction hypotheses, video encoder 200 and video decoder 300 can determine a respective hypothesis template. For example, for a prediction hypothesis, the hypothesis template for the prediction hypothesis can be samples neighboring the prediction hypothesis. The hypothesis template is referred to above as T add There can be multiple hypothesis templates (e.g., multiple T add ), one for each of the one or more prediction hypotheses.

[0122] Using the above equations, video encoder 200 and video decoder 300 can determine a prediction template based on the current weight being evaluated, which is referred to above as T pred Video encoder 200 and video decoder 300 can loop through each of the multiple weights and determine a respective prediction template. That is, video encoder 200 and video decoder 300 can determine multiple prediction templates based on the multiple weights for a multiple hypothesis prediction (MHP) process.

[0123] For example, video encoder 200 and video decoder 300 can determine a first prediction template based on the reference template, the hypothesis template, and a first weight of the multiple weights, and determine a second prediction template based on the reference template, the hypothesis template, and a second weight of the multiple weights. As an example, to determine the first prediction template, video encoder 200 and video decoder 300 can determine T pred1 = (1 - a1)T ref + a1T add , where T pred1 is the first prediction template, a1 is the first weight (e.g., 1 / 4 or -1 / 8), T ref is the reference template, and T add is the hypothesis template. To determine the second prediction template, video encoder 200 and video decoder 300 can determine T pred2 = (1 - a2)T ref + a2Tadd where T pred2 is a second prediction template, a2 is a second weight (e.g., ¼ or another of -1 / 8), T ref is a reference template, and T add is a hypothetical template.

[0124] Video encoder 200 and video decoder 300 can repeat this process for N weights. For example, in the example above, there were two weights, a1 and a2 (e.g., one was ¼ and the other was -1 / 8). However, there can be more than two weights (e.g., N weights), and video encoder 200 and video decoder 300 can repeat the example techniques above to generate N pred (i.e., T pred1 through T predN ).

[0125] In the example above, video encoder 200 can have signaled information identifying one or more prediction hypotheses. However, in some examples, video encoder 200 can not signal such information. In such examples, video encoder 200 and video decoder 300 can test different prediction hypotheses, including different hypothetical templates.

[0126] For example, assume there are M prediction hypothesis candidates. Video encoder 200 and video decoder 300 can determine T pred1 through T predN for a first prediction hypothesis candidate. For example, video encoder 200 and video decoder 300 can determine a first hypothetical template (T add1 ) for a first prediction hypothesis candidate (e.g., a first entry in a merge or AMVP list). Video encoder 200 and video decoder 300 can use T add1 to determine T pred1 through T predN . Video encoder 200 and video decoder 300 can determine a second hypothetical template (T add2 ) for a second prediction hypothesis candidate (e.g., a second entry in a merge or AMVP list). Video encoder 200 and video decoder 300 can use T add2 to determine T pred1 through T predN . Video encoder 200 and video decoder 300 can repeat these operations for all M prediction hypothesis candidates. In this example, video encoder 200 and video decoder can generate N M T pred .

[0127] Thus, in some examples, a hypothetical template (e.g., T add ) is a first hypothetical template (e.g., Tadd1 ). Video encoder 200 and video decoder 300 can determine a third prediction template (e.g., T ref ) based on the reference template (e.g., T add2 ), the second hypothetical template (e.g., T pred3 ), and a first weight (e.g., a1) of the plurality of weights. Video encoder 200 and video decoder 300 can determine a fourth prediction template (e.g., T ref ) based on the reference template (e.g., T add2 ), the second hypothetical template (e.g., T pred4 ), and a second weight (e.g., a2) of the plurality of weights. Video encoder 200 and video decoder 300 can repeat these steps for all M options of T add .

[0128] Further, the above example equations for T pred include only one T add . However, in other examples, there can be more than one T add , such as in examples where a prediction signal is generated using multiple additional prediction signals (e.g., using multiple prediction hypotheses). For ease of description only, examples are described with respect to having one prediction hypothesis for generating a prediction signal.

[0129] Video encoder 200 and video decoder 300 can compare the plurality of prediction templates (e.g., T pred1 through T predN , and possibly for all T add1 through T addM ) to a current template of the current block. The current template of the current block can include samples neighboring the current block. As one example, video encoder 200 and video decoder 300 can determine a template matching cost for each T pred . One example way of determining the template matching cost is based on a sum of absolute differences (SAD) between each of T pred and the current template to compare each of the plurality of prediction templates to the current template.

[0130] In one or more examples, video encoder 200 and video decoder 300 can determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block. As one example, video encoder 200 and video decoder 300 can determine which of the plurality of prediction templates has a smallest difference to the current template, and determine the weight based on the prediction template that has the smallest difference to the current template. For example, it can be assumed that the SAD between T pred2 and the current template is the smallest difference, and T pred2is determined based on a2. In this example, video encoder 200 and video decoder 300 can determine that the weight applied to the MHP process is a2. In this way, video encoder 200 can not need to signal a syntax element indicating the weight, and video decoder 300 can not need to receive the syntax element.

[0131] As another example, video encoder 200 and video decoder 300 can construct the list based on a comparison of a plurality of prediction templates to a current template of the current block. For example, video encoder 200 and video decoder 300 can order the weights and / or prediction hypotheses based on template matching costs. Video encoder 200 and video decoder 300 can include the weights and / or prediction hypotheses with lower template matching costs at lower index entries in the list, and the weights and / or prediction hypotheses with lower template matching costs at higher index entries in the list.

[0132] Video decoder 300 can determine the weight and / or prediction hypothesis to use based on the entries in the list. For example, video encoder 200 can signal an index into the list, and video decoder 300 can use the index to determine the weight and / or prediction hypothesis to use. In this example, there can be signaling for the index into the list. However, the list is ordered based on which weight and / or prediction hypothesis is most likely to be selected. In general, signaling a smaller value as compared to a larger value requires less bandwidth. Thus, by ordering the list such that the weights and / or prediction hypotheses are identified earlier in the list, bandwidth utilization can be reduced as a lower index value will be signaled.

[0133] Video encoder 200 and video decoder 300 can determine one or more prediction hypotheses. As an example, video encoder 200 can signal video decoder 300 information to receive to determine the one or more prediction hypotheses (e.g., an index into the merge list). In this example, video encoder 200 and video decoder 300 can have determined the hypothesis templates (e.g., T add ).

[0134] As another example, video encoder 200 and video decoder 300 can determine one or more prediction hypotheses based on a comparison of a plurality of prediction templates to a current template of the current block. For example, video encoder 200 and video decoder 300 can loop through the weights and prediction hypotheses to determine N M T predwhere N is the number of weights and M is the number of prediction hypotheses. Video encoder 200 and video decoder 300 can determine which combination of weights and hypothesis templates results in the lowest template matching cost, and select that weight and one or more prediction hypotheses as the weight and one or more prediction hypotheses used for the MHP process. In this example, video encoder 200 can not need to signal information indicating the weight and / or prediction hypotheses, and video decoder 300 can not need to receive this information, which can further reduce bandwidth utilization.

[0135] As yet another example, video encoder 200 can signal an index into the list and video decoder 300 can receive the index, as described above. Based on the index, video decoder 300 can determine both the weight and one or more prediction hypotheses used for the MHP process. In this example, signaling can be utilized, but bandwidth utilization can be reduced because video encoder 200 is more likely to signal lower index values, as lower index values will be signaled.

[0136] Video encoder 200 and video decoder 300 can determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight. To determine the prediction signal for the current block based on the one or more prediction hypotheses and the weight, video encoder 200 and video decoder 300 can apply the weight to at least one of the one or more prediction hypotheses, and combine the base prediction signal with the at least one prediction hypothesis to which the weight is applied to determine the prediction signal for the current block.

[0137] For example, video encoder 200 and video decoder 300 can determine In this example, the determined weight (e.g., a) is applied to h3, which is a prediction hypothesis. For example, as described above, p3 is a prediction signal, a is a weight determined using the template matching techniques described above (e.g., comparing a plurality of prediction templates to a current template of the current block), and h3 is a prediction hypothesis (e.g., determined by video encoder 200 and signaled to video decoder 300 or determined using the template matching described above).

[0138] If there is more than one prediction hypothesis (e.g., where ), video encoder 200 and video decoder 300 can perform similar techniques. For example, video encoder 200 and video decoder 300 can perform the example techniques described above to determine a respective weight for each of a plurality of hypotheses.

[0139] Video encoder 200 and video decoder 300 can encode or decode the current block based on the prediction signal. For example, to decode the current block, video decoder 300 can determine residual values that indicate a difference between the prediction signal and the current block, and add the residual values to the prediction signal to reconstruct the current block. To encode the current block, video encoder 200 can determine residual values that indicate a difference between the prediction signal and the current block, and signal information indicative of the residual values.

[0140] In the above examples, it is assumed that the template can be based on prediction hypotheses determined from the merge or AMVP list. However, the examples are not limited thereto, and can be extended to other techniques for determining prediction hypotheses. For example, in MMVD, there are motion vectors and motion vector differences (MVDs). In one or more examples, video encoder 200 and video decoder 300 can add different MVDs (e.g., signaled or not signaled) to the motion vectors to determine additional prediction hypotheses, and thus determine additional hypothesis templates.

[0141] For example, for template matching based MHP weight selection with MMVD extension on the merge list, the source of the additional hypothesis prediction factors can originate from the merge list. For the merge candidate, MMVD techniques can be used. As described above for merge mode with MVD (MMVD), a motion vector offset is added to the original merge candidate, and thus different additional hypothesis prediction factors can be generated. For each prediction factor, a corresponding template can be generated, and based on the examples described for template matching based MHP weight selection for hypotheses derived from the merge list, the template matching cost for each of these MMVD candidates can be calculated. Video encoder 200 and video decoder 300 can derive the TM cost for all possible MMVD offsets and all possible MHP weights, and then sort the TM costs in ascending order, not only can the best MHP weight be decided, but the motion vector of the merge candidate can be refined as well.

[0142] In some examples, only the candidate with the smallest TM cost is selected, thus no signaling is needed for both MHP weight and MMVD information. In some examples, the top m candidates are kept, and the corresponding indices are signaled. In some examples, the sorting is applied in subgroups with each MMVD offset as a group. In this way, it can only be necessary to signal the MMVD information, and still decide the MHP weight based on the TM cost. In some examples, the subgroups are divided based on the MHP weight, in this way, signaling can only be necessary for the MHP weight, but the MMVD information can be omitted from the signaling.

[0143] For template matching based MHP merge list with ARMC, in the above example, TM cost based ordering is performed. However, the ordering can be limited within each of these merge candidates. In some examples, when deriving MHP from the merge list, not only the MHP weight needs to be signaled, but also the merge index. To reduce the signaling overhead, ARMC can be applied to the merge list. Similar to the example techniques described above, the best MHP weight can be decided based on the TM cost. Thus, for each of these merge candidates, this TM cost can be used as the best TM cost for the particular merge candidate. Looping over all the merge candidates in the merge list, the best TM cost for each of these merge candidates can be determined. Then, the ARMC technique can be applied to shuffle the merge index randomly. In this way, a good merge candidate with a large merge index can end up with a much smaller merge index after ARMC.

[0144] In some examples, only the best merge candidate is selected and the merge index signaling is skipped. In some examples, only the top m candidates are kept, where m n are signaled. In some examples, all candidates are kept and signaled.

[0145] For template matching based MHP weight reordering for hypotheses derived from the AMVP list, similar techniques as described above can be used when additional hypothesis predictors are generated from the AMVP list, where an ARMC process is applied to reorder the MHP weights. For MHP from the AMVP list, the motion estimation process is applied considering the MHP weight index. Thus, the motion vector is optimized for each weight. Thus, each weight in the MHP derived from the AMVP list is counted compared to the MHP derived from the merge list. Thus, after signaling the ARMC, only the ARMC process can be applied based on the TM cost of each weight and the weight index instead of the original index to save signaling overhead.

[0146] Figure 2 is a block diagram illustrating an example video encoder 200 that can perform the techniques of this disclosure. Figure 2 is provided for purposes of explanation and should not be considered to be limiting of the techniques broadly illustrated and described in this disclosure. For purposes of explanation, this disclosure describes video decoder 200 according to the techniques of VVC and HEVC. However, the techniques of this disclosure can be performed by video coding devices configured to other video coding standards and video coding formats, such as AV1 and subsequent formats of the AV1 video coding format.

[0147] In Figure 2 ​In the example of FIG. 2, video encoder 200 includes video data memory 230, mode select unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, decoded picture buffer (DPB) 218, and entropy encoding unit 220. Any or all of video data memory 230, mode select unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy encoding unit 220 can be implemented in one or more processors or in processing circuitry. For instance, the units of video encoder 200 can be implemented as part of a hardware circuit, as part of a software program, or as a combination of a hardware circuit and a software program. Additionally, video encoder 200 can include additional or alternative processors or processing circuitry to perform these and other functions.

[0148] Video data memory 230 can store video data to be encoded by the components of video encoder 200. Video encoder 200 can receive the video data stored in video data memory 230 from, for example, video source 104 Figure 1 DPB 218 can function as a reference picture memory that stores reference video data for use in encoding future video data by video encoder 200. Video data memory 230 and DPB 218 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, video data memory 230 can be on-chip with other components of video encoder 200, as illustrated, or off-chip relative to those components.

[0149] In this disclosure, reference to video data memory 230 should not be interpreted as being limited to memory internal to video encoder 200 (unless specifically described as such) or memory external to video encoder 200 (unless specifically described as such). Rather, reference to video data memory 230 should be understood as a reference memory that stores video data that video encoder 200 receives for encoding (e.g., video data for a current block to be encoded). Figure 1 Memory 106 of source device 102 can also provide temporary storage of the outputs from the various units of video encoder 200.

[0150] FIG. 2 illustrates an example of a video encoder 200 that can implement techniques for encoding video data. Video encoder 200 can be included in a device that can be configured to perform the techniques described in this disclosure. In the example of FIG. 2, video encoder 200 includes video data memory 230, mode select unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, decoded picture buffer (DPB) 218, and entropy encoding unit 220. Any or all of video data memory 230, mode select unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy encoding unit 220 can be implemented in one or more processors or in processing circuitry. For instance, the units of video encoder 200 can be implemented as part of a hardware circuit, as part of a software program, or as a combination of a hardware circuit and a software program. Additionally, video encoder 200 can include additional or alternative processors or processing circuitry to perform these and other functions. Figure 2The various units of video encoder 200 are shown to help understand the operations performed by video encoder 200. The units can be implemented as fixed- function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide particular functionality, and are preset on the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality over the operations that can be performed. For instance, programmable circuits can execute software or firmware that causes the programmable circuits to operate in the manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations that the fixed-function circuits perform are generally immutable. In some examples, one or more of the units can be distinct circuit blocks (fixed-function or programmable), and in some examples, one or more of the units can be integrated circuits.

[0151] Video encoder 200 can include arithmetic logic units (ALUs), elementary function units (EFUs), digital circuits, analog circuits, and / or programmable cores formed from programmable circuits. In examples where the operations of video encoder 200 are performed using software executed by the programmable circuits, memory 106 Figure 1 ) can store the instructions (e.g., object code) of the software that video encoder 200 receives and executes, or another memory within video encoder 200 (not shown) can store such instructions.

[0152] Video data memory 230 is configured to store video data that is received for encoding. Video encoder 200 can retrieve pictures of the video data from video data memory 230 and provide the video data to residual generation unit 204 and mode selection unit 202. Video data in video data memory 230 can be raw video data that is to be encoded.

[0153] Mode selection unit 202 includes motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226. Mode selection unit 202 can include additional functional units that perform video prediction according to other prediction modes. As examples, mode selection unit 202 can include a palette unit, an intra block copy unit (which can be part of motion estimation unit 222 and / or motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0154] Mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and resulting rate-distortion values for such combinations. The encoding parameters can include partitioning of CTUs into CUs, prediction modes for CUs, transform types for residual data of CUs, quantization parameters for residual data of CUs, etc. Mode selection unit 202 can ultimately select the combination of encoding parameters that has the best rate-distortion value compared to other tested combinations.

[0155] Video encoder 200 can partition a picture retrieved from video data memory 230 into a series of CTUs, and encapsulate one or more CTUs within a slice. Mode select unit 202 can partition CTUs of a picture according to a tree structure such as the MTT structure, the QTBT structure, the superblock structure, or the quadtree structure described above. As described above, video encoder 200 can form one or more CUs by partitioning a CTU according to the tree structure. Such CUs can also be referred to as “video blocks” or “blocks” generally.

[0156] In general, mode select unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate a prediction block for a current block (e.g., a current CU, or in HEVC, an overlapping portion of a PU and a TU). To perform inter-prediction for the current block, motion estimation unit 222 can perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously coded pictures stored in DPB 218). Specifically, motion estimation unit 222 can calculate a value that represents how similar a potential reference block would be to the current block, e.g., according to a sum of absolute difference (SAD), a sum of squared difference (SSD), a mean absolute difference (MAD), a mean squared difference (MSD), etc. Motion estimation unit 222 can generally perform these calculations using a sample-by-sample difference between the current block and the reference block under consideration. Motion estimation unit 222 can identify the reference block with the lowest value resulting from these calculations, to indicate the reference block that is most matching to the current block.

[0157] Motion estimation unit 222 can form one or more motion vectors (MVs) that define the location of a reference block in a reference picture relative to the location of the current block in the current picture. Motion estimation unit 222 can then provide the motion vector(s) to motion compensation unit 224. For example, motion estimation unit 222 can provide a single motion vector for single -prediction inter-prediction, and two motion vectors for bi-prediction inter-prediction. Motion compensation unit 224 can then generate the prediction block using the motion vector(s). For example, motion compensation unit 224 can use the motion vector(s) to retrieve data for the reference block. As another example, where the motion vector has fractional sample precision, motion compensation unit 224 can interpolate values for the prediction block according to one or more interpolation filters. Further, for bi-prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by the respective motion vectors and combine the retrieved data, e.g., by sample-wise averaging or weighted averaging.

[0158] When operating according to the AV1 video coding format, the motion estimation unit 222 and the motion compensation unit 224 can be configured to use translational motion compensation, affine motion compensation, overlapped block motion compensation (OBMC), and / or compound inter-intra prediction to encode coding blocks (e.g., both luma coding blocks and chroma coding blocks) of video data.

[0159] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 can generate a prediction block from neighboring samples of the current block. For example, for directional modes, the intra prediction unit 226 can generally mathematically combine values of the neighboring samples and fill these computed values across the current block in a defined direction to produce the prediction block. As another example, for a DC mode, the intra prediction unit 226 can compute an average of the neighboring samples of the current block and generate the prediction block to include the resulting average for each sample of the prediction block.

[0160] When operating according to the AV1 video coding format, the intra prediction unit 226 can be configured to use directional intra prediction, non-directional intra prediction, recursive filter intra prediction, luma-chroma (CFL) prediction, intra block copy (IBC), and / or palette mode to encode coding blocks (e.g., both luma coding blocks and chroma coding blocks) of video data. The mode selection unit 202 can include additional functional units that perform video prediction according to other prediction modes.

[0161] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the unencoded original version of the current block from the video data memory 230 and the prediction block from the mode selection unit 202. The residual generation unit 204 computes the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines a residual block for the current block. In some examples, the residual generation unit 204 can also determine differences between sample values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 can be formed using one or more subtractor circuits that perform binary subtraction.

[0162] In examples in which mode selection unit 202 partitions a CU into PUs, each PU can be associated with a luma prediction unit and corresponding chroma prediction units. Video encoder 200 and video decoder 300 can support PUs having various sizes. As noted above, the size of a CU can refer to the size of the CU's luma coding block, while the size of a PU can refer to the size of the PU's luma prediction unit. Assuming that a particular CU has a size of 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-prediction, and 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetric PU sizes for inter-prediction. Video encoder 200 and video decoder 300 can also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-prediction.

[0163] In examples in which mode selection unit 202 does not partition a CU into PUs, each CU can be associated with a luma coding block and corresponding chroma coding blocks. As above, the size of a CU can refer to the size of the CU's luma coding block. Video encoder 200 and video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0164] For other video coding techniques, such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, as some examples, mode selection unit 202 generates a prediction block for the current block being encoded via the respective unit associated with the coding technique. In some examples, such as palette mode coding, mode selection unit 202 can not generate a prediction block, but instead generate syntax elements indicative of the manner in which the block is to be reconstructed based on a selected palette. In such modes, mode selection unit 202 can provide these syntax elements to entropy encoding unit 220 for encoding.

[0165] As described above, residual generation unit 204 receives video data for a current block and a corresponding prediction block. Residual generation unit 204 then generates a residual block for the current block. To generate the residual block, residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0166] Transform processing unit 206 applies one or more transforms to the residual block to produce a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 can apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 can apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform. In some examples, transform processing unit 206 can perform multiple transforms on the residual block, e.g., a primary transform and a secondary transform such as a rotation transform. In some examples, transform processing unit 206 does not apply a transform to the residual block.

[0167] When operating according to AV1, transform processing unit 206 can apply one or more transforms to the residual block to produce a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 can apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 can apply a horizontal / vertical transform combination, which can include a discrete cosine transform (DCT), an asymmetric discrete sine transform (ADST), a flipped ADST (e.g., ADST in reverse order), and an identity transform (IDTX). When the identity transform is used, the transform is skipped in one of the vertical or horizontal directions. In some examples, transform processing can be skipped.

[0168] Quantization unit 208 can quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode select unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization can introduce loss of information, and thus, quantized transform coefficients can have lower precision than the original transform coefficients produced by transform processing unit 206.

[0169] Inverse quantization unit 210 and inverse transform processing unit 212 can apply inverse quantization and inverse transforms, respectively, to the quantized transform coefficient block to reconstruct the residual block from the transform coefficient block. Reconstruction unit 214 can produce a reconstructed block corresponding to the current block (although potentially with some degree of distortion) based on the reconstructed residual block and the prediction block generated by mode select unit 202. For example, reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples from the prediction block generated by mode select unit 202 to produce the reconstructed block.

[0170] Filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In some examples, the operations of filter unit 216 can be skipped.

[0171] When operating according to AV1, filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In other examples, filter unit 216 can apply a constrained direction enhancement filter (CDEF), which can be applied after deblocking and can include the application of a non-separable, non-linear, low-pass directional filter based on an estimated edge direction. Filter unit 216 can also include a loop restoration filter applied after CDEF and can include a separable, symmetric, normalized Wiener filter or a double- self-guided filter.

[0172] Video encoder 200 stores the reconstructed block in DPB 218. For example, in examples in which the operations of filter unit 216 are not performed, reconstructed unit 214 can store the reconstructed block into DPB 218. In examples in which the operations of filter unit 216 are performed, filter unit 216 can store the filtered reconstructed block into DPB 218. Motion estimation unit 222 and motion compensation unit 224 can retrieve reference pictures formed from reconstructed (and potentially filtered) blocks from DPB 218 to inter-predict blocks of subsequent encoded pictures. In addition, intra-prediction unit 226 can use reconstructed blocks of a current picture in DPB 218 to intra-predict other blocks within the current picture.

[0173] In general, entropy encoding unit 220 can entropy encode syntax elements received from other functional components of video encoder 200. For example, entropy encoding unit 220 can entropy encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy encoding unit 220 can entropy encode prediction syntax elements (e.g., motion information for inter-prediction or intra-mode information for intra-prediction) from mode select unit 202. Entropy encoding unit 220 can perform one or more entropy encoding operations on the syntax elements, which are another example of video data, to generate entropy encoded data. For example, entropy encoding unit 220 can perform a context- adaptive variable length coding (CAVLC) operation, a CABAC operation, a variable- to-variable (V2V) length coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a Probability Interval Partitioning Entropy (PIPE) coding operation, an Exponential-Golomb coding operation, or another type of entropy encoding operation. In some examples, entropy encoding unit 220 can operate in a bypass mode in which syntax elements are not entropy encoded.

[0174] Video encoder 200 can output a bitstream that includes the entropy encoded syntax elements needed to reconstruct blocks of a slice or picture. In particular, entropy encoding unit 220 can output the bitstream.

[0175] According to AV1, entropy encoding unit 220 can be configured as a symbol-to-symbol adaptive multi-symbol arithmetic coder. Syntax elements in AV1 include an alphabet of N elements, and a context (e.g., a probability model) includes a set of N probabilities. Entropy encoding unit 220 can store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). Entropy encoding unit 220 can perform recursive scaling using an update factor based on the alphabet size to update the context.

[0176] The operations described above are described with respect to a block. Such description should be understood as operations for a luma coding block and / or a chroma coding block. As described above, in some examples, the luma coding block and the chroma coding block are luma and chroma components of a CU. In some examples, the luma coding block and the chroma coding block are luma and chroma components of a PU.

[0177] In some examples, operations performed with respect to a luma coding block do not need to be repeated for a chroma coding block. As one example, operations to identify a motion vector (MV) and a reference picture for a luma coding block do not need to be repeated for identifying an MV and a reference picture for a chroma block. Instead, the MV for the luma coding block can be scaled to determine the MV for the chroma block, and the reference picture can be the same. As another example, an intra prediction process can be the same for a luma coding block and a chroma coding block.

[0178] Video encoder 200 represents an example of a device configured to encode video data, the device comprising: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: determine a plurality of weights based on a current template for a current block and a reference template for a reference; determine, from the plurality of weights, a weight to apply to a multi-hypothesis prediction (MHP) process. The MHP process can comprise determining a prediction signal for the current block based on a prediction block, one or more prediction hypotheses, and the weight. Video encoder 200 can be configured to encode the current block based on the prediction signal.

[0179] As one example, video encoder 200 can determine, for a multiple hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights, and compare the plurality of prediction templates to a current template of a current block. Video encoder 200 can determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block. Video encoder 200 can determine one or more prediction hypotheses, and determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight. Video encoder 200 can encode the current block based on the prediction signal.

[0180] For example, video encoder 200 can be configured to determine a plurality of template matching costs for each weight of the plurality of weights, and determine a weight corresponding to a plurality of the template matching costs having a smallest template matching cost. As another example, video encoder 200 can construct a list of weights based on the plurality of weights, and signal an index into the list of weights. Video encoder 200 can order the list of weights (e.g., weights of the plurality of weights) based on the plurality of template matching costs, and signal an index into the ordered list.

[0181] Video encoder 200 can be configured to determine residual values that indicate a difference between the current block and the prediction signal. Video encoder 200 can signal information indicative of the residual values.

[0182] Figure 3 FIG. 3 is a block diagram illustrating an example video decoder 300 that can perform the techniques of this disclosure. Figure 3 FIG. 3 is provided for purposes of explanation and is not limiting on the techniques of this disclosure as broadly exemplified and described. For purposes of explanation, this disclosure is described in the context of a video decoder 300 according to the techniques of VVC and HEVC. However, the techniques of this disclosure can be performed by video coding devices configured to other video coding standards.

[0183] In Figure 3In the example of FIG. 3, video decoder 300 includes coded picture buffer (CPB) memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314. Any or all of CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or in processing circuitry. For instance, the units of video decoder 300 can be implemented as one or more circuits or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Also, video decoder 300 can include additional or alternative processors or processing circuitry to perform these and other functions.

[0184] Prediction processing unit 304 includes motion compensation unit 316 and intra-prediction unit 318. Prediction processing unit 304 can include additional units to perform prediction according to other prediction modes. As examples, prediction processing unit 304 can include a palette unit, an intra-block copy unit (which can form a part of motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, video decoder 300 can include more, less, or different functional components.

[0185] When operating according to AV1, motion compensation unit 316 can be configured to decode coding blocks (e.g., both luma coding blocks and chroma coding blocks) of video data using translational motion compensation, affine motion compensation, OBMC, and / or compound inter-intra prediction, as described above. Intra-prediction unit 318 can be configured to decode coding blocks (e.g., both luma coding blocks and chroma coding blocks) of video data using directional intra-prediction, non-directional intra-prediction, recursive filter intra-prediction, CFL, IBC, and / or palette mode, as described above.

[0186] CPB memory 320 can store video data, such as an encoded video bitstream, to be decoded by the components of video decoder 300. For instance, the encoded video bitstream can be obtained (e.g., from computer- readable medium 110 Figure 1The video data stored in the CPB memory 320 is obtained. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Furthermore, the CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores a decoded picture, which the video decoder 300 may output, and / or uses as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed from any of a variety of memory devices, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. The CPB memory 320 and the DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.

[0187] Alternatively or concurrently, in some examples, the video decoder 300 may be located from the memory 120 ( Figure 1 Decoded video data can be retrieved from the memory. That is, memory 120 can utilize CPB memory 320 to store data, as discussed above. Similarly, when some or all of the functionality of video decoder 300 is implemented in software to be executed by the processing circuitry of video decoder 300, memory 120 can store instructions to be executed by video decoder 300.

[0188] Examples Figure 3 The various units shown help to understand the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 2 Fixed-function circuits refer to circuits that provide specific functionality and are pre-defined for the operations they can perform. Programmable circuits, on the other hand, are circuits that can be programmed to perform various tasks and provide flexible functionality for the operations they can perform. For example, a programmable circuit can execute software or firmware that causes it to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more units in a unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in a unit may be integrated circuits.

[0189] Video decoder 300 can include ALUs, EFUs, digital circuits, analog circuits, and / or programmable cores formed from programmable circuitry. In examples where the operations of video decoder 300 are performed by software executing on the programmable circuitry, on-chip or off-chip memory can store instructions (e.g., object code) of the software that video decoder 300 receives and executes.

[0190] Entropy decoding unit 302 can receive encoded video data from the CPB and entropy decode the video data to reconstruct syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0191] In general, video decoder 300 reconstructs a picture on a block-by-block basis. Video decoder 300 can perform the reconstruction operations individually for each block (where the block that is currently being reconstructed (i.e., decoded) can be referred to as the “current block”).

[0192] Entropy decoding unit 302 can entropy decode syntax elements defining quantized transform coefficients of a quantized transform coefficient block, as well as transform information such as a quantization parameter (QP) and / or an indication of a transform mode. Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine a degree of quantization, and likewise a degree of inverse quantization for inverse quantization unit 306 to apply. Inverse quantization unit 306 may, for example, perform a bitwise left-shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 may

[0193] After inverse quantization unit 306 forms the transform coefficient block, inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve Transform (KLT), an inverse rotational transform, an inverse directional transform, or another inverse transform.

[0194] Further, prediction processing unit 304 generates the prediction block from the prediction information syntax elements entropy decoded by entropy decoding unit 302. For instance, in a case where the prediction information syntax elements indicate that the current block is inter-predicted, motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements can indicate a reference picture, in the DPB 314, to retrieve a reference block from, and a motion vector identifying a location of the reference block in the reference picture relative to a location of the current block in the current picture. Motion compensation unit 316 generally can perform the inter-prediction process in a manner substantially similar to that described with respect to motion compensation unit 224 Figure 2 ).

[0195] As another example, in a case where the prediction information syntax element indicates that the current block is intra-predicted, intra-prediction unit 318 can generate the prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, intra-prediction unit 318 can generally perform the intra-prediction process in a manner substantially similar to that described with respect to intra-prediction unit 226 Figure 2 ). Intra-prediction unit 318 can retrieve data for neighboring samples of the current block from DPB 314.

[0196] Reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, reconstruction unit 310 can add the samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.

[0197] Filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce blocking artifact pseudomorphs along edges of the reconstructed block. The operations of filter unit 312 are not necessarily performed in all examples.

[0198] Video decoder 300 can store the reconstructed block in DPB 314. For example, in examples in which the operations of filter unit 312 are not performed, reconstruction unit 310 can store the reconstructed block to DPB 314. In examples in which the operations of filter unit 312 are performed, filter unit 312 can store the filtered reconstructed block to DPB 314. As discussed above, DPB 314 can provide reference information, such as samples of the current picture for intra-prediction and previously decoded pictures for subsequent motion compensation, to prediction processing unit 304. In addition, video decoder 300 can output decoded pictures (e.g., a decoded video) from DPB 314 for subsequent presentation on a display device, such as display device 118 of FIG. 1. Figure 1

[0199] In this way, video decoder 300 represents an example of a video decoding device including a memory configured to store video data and one or more processing units implemented in circuitry and configured to determine a plurality of weights based on a current template for a current block and a reference template for a reference, determine a weight from the plurality of weights to apply to a multiple hypothesis prediction (MHP) process. The MHP process can include determining a prediction signal for the current block based on a prediction block, one or more prediction hypotheses, and the weight. Video decoder 300 can be configured to decode the current block based on the prediction signal.

[0200] ​As one example, video decoder 300 can determine a plurality of prediction templates based on a plurality of weights for a multi-hypothesis prediction (MHP) process and compare the plurality of prediction templates to a current template of a current block. Video decoder 300 can determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block. Video decoder 300 can determine one or more prediction hypotheses based on the one or more prediction hypotheses and the weight and determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight. Video decoder 300 can encode the current block based on the prediction signal.

[0201] For example, video decoder 300 can be configured to determine a plurality of template matching costs for each weight of the plurality of weights and determine the weight corresponding to the plurality of template matching costs having a minimum template matching cost. As another example, video decoder 300 can construct a list of weights based on the plurality of weights and receive an index into the list of weights. Video decoder 300 can order the list of weights (e.g., weights of the plurality of weights) based on the plurality of template matching costs and receive an index into the ordered list.

[0202] Video decoder 300 can be configured to receive information indicating residual values that indicate a difference between the current block and the prediction signal. Video decoder 300 can reconstruct the current block based on the prediction signal and the residual values.

[0203] Figure 4 is a flowchart illustrating an example method for encoding a current block in accordance with the techniques of this disclosure. The current block can be or can include a current CU. While described with respect to video encoder 200 Figure 1 and Figure 2 ) it should be understood that other devices can be configured to perform methods similar to those of Figure 4 .

[0204] In this example, video encoder 200 initially predicts a current block (400). For example, video encoder 200 can use the example techniques described in this disclosure to form a prediction block for the current block and determine a prediction signal. Video encoder 200 can then calculate a residual block for the current block (402). To calculate the residual block, video encoder 200 can calculate a difference between an unencoded original block for the current block and the prediction block. Video encoder 200 can then transform the residual block and quantize transform coefficients of the residual block (404). Next, video encoder 200 can scan the quantized transform coefficients of the residual block (406). During or after the scan, video encoder 200 can entropy encode the transform coefficients (408). For example, video encoder 200 can use CAVLC or CABAC to encode the transform coefficients. Video encoder 200 can then output the entropy encoded data for the block (410).

[0205] Figure 5 is a flowchart illustrating an example method for decoding a current block of video data in accordance with the techniques of this disclosure. The current block can be or can include a current CU. Although described with respect to video decoder 300 Figure 1 and Figure 3 ), it should be understood that other devices can be configured to perform similar methods as those of Figure 5 .

[0206] Video decoder 300 can receive entropy encoded data for a current block, such as entropy encoded prediction information and entropy encoded transform coefficients for a residual block corresponding to the current block (500). Video decoder 300 can entropy decode the entropy encoded data to determine prediction information for the current block and to reproduce transform coefficients of the residual block (502). Video decoder 300 can predict the current block, e.g., using an intra-prediction mode or an inter-prediction mode as indicated by the prediction information for the current block, to calculate a prediction block for the current block using the techniques described in this disclosure (including a prediction signal for the current block) (504). Video decoder 300 can then inverse scan the reproduced transform coefficients to create a block of quantized transform coefficients (506). Video decoder 300 can then inverse quantize the transform coefficients and apply an inverse transform to the transform coefficients to produce a residual block (508). Video decoder 300 can finally decode the current block by combining the prediction block and the residual block (510).

[0207] Figure 10 is a flowchart illustrating an example method of operation. In general, Figure 10 example techniques of Figure 10 example techniques of Figure 10 are described with respect to processing circuitry. Examples of processing circuitry include fixed function and / or programmable circuitry of video encoder 200 and video decoder 300. For example, one or more memories can be configured to store video data. Examples of the one or more memories include memory 106, memory 120, video data memory 230, decoded picture buffer 218, CPB memory 320, DPB 314, or some other memory including distributed memory. Processing circuitry of video encoder 200 or video decoder 300 can be coupled to the one or more memories and configured to perform example techniques of Figure 10 .

[0208] The processing circuitry of the video encoder 200 or the video decoder 300 can determine a plurality of prediction templates based on a plurality of weights for a multi-hypothesis prediction (MHP) process (1000). For example, to determine the plurality of prediction templates based on the plurality of weights, the processing circuitry of the video encoder 200 and the video decoder 300 can determine a first prediction template based on a reference template, a hypothesis template, and a first weight of the plurality of weights, and determine a second prediction template based on the reference template, the hypothesis template, and a second weight of the plurality of weights.

[0209] In one or more examples, the reference template can be samples adjacent to a reference block pointed to by a motion or a block vector. The hypothesis template can be samples adjacent to a prediction hypothesis. The prediction hypothesis can be a previously encoded or decoded block in the same or a different picture as the current block. The first weight and the second weight can be example weights of the plurality of weights.

[0210] As an example, to determine the first prediction template, the processing circuitry of the video encoder 200 and the video decoder 300 can determine: pred1 = (1 - a1)T ref + a1T add where T pred1 is the first prediction template, a1 is the first weight, T ref is the reference template, and T add is the hypothesis template. To determine the second prediction template, the processing circuitry of the video encoder 200 and the video decoder 300 can determine: pred2 = (1 - a2)T ref + a2T add where T pred2 is the second prediction template, a2 is the second weight, T ref is the reference template, and T add is the hypothesis template. The processing circuitry can repeat these operations for each weight of the plurality of weights.

[0211] In some examples, there can be multiple hypothesis templates or multiple reference templates, and the processing circuitry of the video encoder 200 and the video decoder 300 can repeat such techniques for each weight of the plurality of weights, each reference template of the multiple reference templates, and each hypothesis template of the multiple hypothesis templates. For example, the hypothesis template can be a first hypothesis template, and the processing circuitry of the video encoder 200 or the video decoder 300 can determine a third prediction template based on the reference template, a second hypothesis template, and a first weight of the plurality of weights, and determine a fourth prediction template based on the reference template, the second hypothesis template, and a second weight of the plurality of weights.

[0212] With these various example techniques, the processing circuitry of video encoder 200 or video decoder 300 can determine a plurality of prediction templates based on a plurality of weights. The processing circuitry of video encoder 200 or video decoder 300 can compare the plurality of prediction templates to a current template of a current block (1002). The current template can be samples neighboring the current block. As one example, to compare, the processing circuitry of video encoder 200 or video decoder 300 can determine a template matching cost for each of the plurality of prediction templates, such as based on a SAD determination between each of the plurality of prediction templates and the current template of the current block.

[0213] The processing circuitry of video encoder 200 and video decoder 300 can determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block (1004). As one example, the processing circuitry of video encoder 200 and video decoder 300 can determine which of the prediction templates is associated with a lowest temporal matching cost (e.g., has a smallest difference from the current template). The processing circuitry of video encoder 200 and video decoder 300 can determine the weight used to compute the prediction template with the lowest temporal matching cost as the weight to be used for the MHP process.

[0214] As another example, the processing circuitry of video encoder 200 or video decoder 300 can construct a list based on the comparison of the plurality of prediction templates to the current template of the current block. For example, the processing circuitry of video encoder 200 or video decoder 300 can identify a weight and / or prediction hypothesis associated with the prediction template with the lowest temporal matching cost in a first entry of the list, identify a weight and / or prediction hypothesis associated with the prediction template with the next lowest temporal matching cost in a second entry of the list, and so on. The processing circuitry of video encoder 200 can signal an index for the entries in the list and the processing circuitry of video decoder 300 can receive the index, which the processing circuitry of video decoder 300 uses to determine the weight to be used for the MHP process, and in some examples, the prediction hypothesis to be used.

[0215] The processing circuitry of video encoder 200 or video decoder 300 can determine one or more prediction hypotheses (1006). As one example, the processing circuitry of video encoder 200 can signal information identifying the one or more prediction hypotheses and the processing circuitry of video decoder 300 can receive the information.

[0216] In some examples, the processing circuitry of video encoder 200 or video decoder 300 can determine the one or more prediction hypotheses based on a comparison of the plurality of prediction templates to a current template of the current block. For example, such as where the processing circuitry of video encoder 200 and video decoder 300 loop through the plurality of prediction hypotheses, the processing circuitry of video encoder 200 and video decoder 300 can determine the prediction hypotheses associated with the prediction templates having the lowest temporal matching cost and select these prediction hypotheses as the one or more prediction hypotheses for the MHP process. As another example, the processing circuitry of video encoder 200 can signal an index for the entries in the list and the processing circuitry of video decoder 300 can receive the index (as described above), and the processing circuitry of video decoder 300 can determine the one or more prediction hypotheses based on the index for the entries in the list or a plurality of indices for the entries in the list.

[0217] The processing circuitry of video encoder 200 or video decoder 300 can determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight (1008). For example, the processing circuitry of video encoder 200 or video decoder 300 can apply the weight to at least one of the one or more prediction hypotheses and combine the base prediction signal with the at least one prediction hypothesis to which the weight is applied to determine the prediction signal for the current block. For example, the processing circuitry of video encoder 200 or video decoder 300 can determine In this example, the determined weight (e.g., a, which is determined using the example techniques described above) is applied to h3, which is a prediction hypothesis. For example, as described above, p3 is a prediction signal, a is a weight determined using the template matching techniques described above (e.g., comparing a plurality of prediction templates to a current template of a current block), and h3 is a prediction hypothesis (e.g., determined by video encoder 200 and signaled to video decoder 300 or determined using the template matching described above).

[0218] If there is more than one prediction hypothesis (e.g., where ), the processing circuitry of video encoder 200 or video decoder 300 can perform similar techniques. For example, the processing circuitry of video encoder 200 or video decoder 300 can perform the example techniques described above to determine a respective weight for each of the plurality of hypotheses.

[0219] The processing circuitry of video encoder 200 or video decoder 300 can encode or decode the current block based on the prediction signal (1010). For example, to decode the current block, the processing circuitry of video decoder 300 can determine residual values that indicate a difference between the prediction signal and the current block, and add the residual values to the prediction signal to reconstruct the current block. To encode the current block, the processing circuitry of video encoder 200 can determine residual values that indicate a difference between the prediction signal and the current block, and signal information indicative of the residual values.

[0220] The following numbered clauses exemplify one or more aspects of the devices and techniques described in this disclosure.

[0221] Clause 1. A method of encoding or decoding video data, the method comprising: determining a plurality of weights based on a current template for a current block and a reference template for a reference; determining a weight to apply to a multi-hypothesis prediction (MHP) process from the plurality of weights, wherein the MHP process comprises: determining a prediction signal for the current block based on a prediction block, one or more prediction hypotheses, and the weight; and encoding or decoding the current block based on the prediction signal.

[0222] Clause 2. The method of clause 1, further comprising: determining a plurality of template matching costs for each weight of the plurality of weights, wherein determining the weight comprises determining the weight corresponding to the plurality of template matching costs having a minimum template matching cost.

[0223] Clause 3. The method of clause 1, further comprising: constructing a weight list based on the plurality of weights, wherein determining the weight comprises receiving an index into the weight list.

[0224] Clause 4. The method of clause 3, wherein the weight list comprises weights for less than all of the plurality of weights.

[0225] Clause 5. The method of any of clauses 3 and 4, further comprising: determining a plurality of template matching costs for each weight of the plurality of weights, wherein constructing the weight list comprises ordering the weights of the plurality of weights based on the plurality of template matching costs.

[0226] Clause 6. The method of any of clauses 1 to 5, further comprising constructing a merge list.

[0227] Clause 7. The method of clause 6, wherein the one or more prediction hypotheses are based on a motion vector offset of a merge candidate added to the merge list.

[0228] Clause 8. The method of any of clauses 6 and 7, the method further comprising performing merge candidate adaptive reordering (ARMC) on the merge list, wherein the one or more prediction hypotheses are based on the merge list after performing ARMC on the merge list.

[0229] Clause 9. The method of any of clauses 1 to 5, the method further comprising constructing an advanced motion vector predictor (AMVP) list; and performing merge candidate adaptive reordering (ARMC) on the AMVP list, wherein the one or more prediction hypotheses are based on the AMVP list after performing ARMC on the AMVP list.

[0230] Clause 10. The method of any of clauses 1 to 9, wherein encoding or decoding the current block based on the prediction signal comprises decoding the current block, and wherein decoding the current block comprises receiving information indicative of residual values that are indicative of a difference between the current block and the prediction signal; and reconstructing the current block based on the prediction signal and the residual values.

[0231] Clause 11. The method of any of clauses 1 to 9, wherein encoding or decoding the current block based on the prediction signal comprises encoding the current block, and wherein encoding the current block comprises determining residual values that are indicative of a difference between the current block and the prediction signal; and signaling information indicative of the residual values.

[0232] Clause 12. A device for encoding or decoding video data, the device comprising a memory configured to store the video data; and one or more processors implemented in circuitry, coupled to the memory, and configured to perform a method recited in any of clauses 1 to 11.

[0233] Clause 13. The device of clause 12, the device further comprising a display configured to display decoded video data.

[0234] Clause 14. The device of any of clauses 12 and 13, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0235] Clause 15. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to perform a method recited in any of clauses 1 to 11.

[0236] Clause 16. A device for encoding or decoding video data, the device comprising means for performing the method of any of clauses 1-11.

[0237] Clause 1A. A method of encoding or decoding video data, the method comprising: for a multi-hypothesis prediction (MHP) process, determining a plurality of prediction templates based on a plurality of weights; comparing the plurality of prediction templates to a current template of a current block; determining a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determining one or more prediction hypotheses; determining a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encoding or decoding the current block based on the prediction signal.

[0238] Clause 2A. The method of clause 1A, wherein determining the plurality of prediction templates based on the plurality of weights comprises: determining a first prediction template based on a reference template, a hypothesis template, and a first weight of the plurality of weights; and determining a second prediction template based on the reference template, the hypothesis template, and a second weight of the plurality of weights.

[0239] Clause 3A. The method of clause 2A, wherein determining the first prediction template comprises: Tpred1 = (1 - a1)Tref + a1Tadd, where Tpred1 is the first prediction template, a1 is the first weight, Tref is the reference template, and Tadd is the hypothesis template.

[0240] Clause 4A. The method of clause 3A, wherein determining the second prediction template comprises: Tpred2 = (1 - a2)Tref + a2Tadd, where Tpred2 is the second prediction template, a2 is the second weight, Tref is the reference template, and Tadd is the hypothesis template.

[0241] Clause 5A. The method of any of clauses 2A-4A, wherein the hypothesis template is a first hypothesis template, the method further comprising: determining a third prediction template based on the reference template, a second hypothesis template, and the first weight of the plurality of weights; and determining a fourth prediction template based on the reference template, the second hypothesis template, and the second weight of the plurality of weights.

[0242] Clause 6A. The method of any of clauses 1A-5A, the method further comprising: constructing a list based on the comparison of the plurality of prediction templates to the current template of the current block, and wherein determining the weight comprises determining the weight based on an entry in the list.

[0243] Clause 7A. The method of any of clauses 1A-6A, wherein determining the one or more prediction hypotheses comprises determining the one or more prediction hypotheses based on the comparison of the plurality of prediction templates to the current template of the current block.

[0244] Clause 8A. The method of any of clauses 1A-7A, wherein determining the prediction signal for the current block based on the one or more prediction hypotheses and the weight comprises: applying the weight to at least one prediction hypothesis of the one or more prediction hypotheses; and combining a base prediction signal with the at least one prediction hypothesis to which the weight is applied to determine the prediction signal for the current block.

[0245] Clause 9A. The method of any of clauses 1A-8A, wherein encoding or decoding the current block based on the prediction signal comprises decoding the current block, wherein decoding the current block comprises: determining residual values indicative of a difference between the prediction signal and the current block; and adding the residual values to the prediction signal to reconstruct the current block.

[0246] Clause 10A. The method of any of clauses 1A-8A, wherein encoding or decoding the current block based on the prediction signal comprises encoding the current block, wherein encoding the current block comprises: determining residual values indicative of a difference between the prediction signal and the current block; and signaling information indicative of the residual values.

[0247] Clause 11A. A device for encoding or decoding video data, the device comprising: one or more memories configured to store the video data; and processing circuitry coupled to the one or more memories, wherein the processing circuitry is configured to: for a multiple-hypothesis prediction (MHP) process, determine a plurality of prediction templates based on a plurality of weights; compare the plurality of prediction templates to a current template of a current block; determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determine one or more prediction hypotheses; determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encode or decode the current block based on the prediction signal.

[0248] Clause 12A. The device of clause 11A, wherein to determine the plurality of prediction templates based on the plurality of weights, the processing circuitry is configured to: determine a first prediction template based on a reference template, a hypothesis template, and a first weight of the plurality of weights; and determine a second prediction template based on the reference template, the hypothesis template, and a second weight of the plurality of weights.

[0249] Clause 13A. The device of clause 12A, wherein to determine the first prediction template, the processing circuitry is configured to determine: Tpred1 = (1 - ai)Tref + aiTadd, where Tpred1 is the first prediction template, ai is the first weight, Tref is the reference template, and Tadd is the hypothesis template.

[0250] Clause 14A. The device of clause 13A, wherein to determine the second prediction template, the processing circuitry is configured to determine: Tpred2 = (1 - a2)Tref + a2Tadd, where Tpred2 is the second prediction template, a2 is the second weight, Tref is the reference template, and Tadd is the hypothesis template.

[0251] Clause 15A. The device of any of clauses 12A to 14A, wherein the hypothesis template is a first hypothesis template, and wherein the processing circuitry is configured to: determine a third prediction template based on the reference template, a second hypothesis template, and the first weight of the plurality of weights; and determine a fourth prediction template based on the reference template, the second hypothesis template, and the second weight of the plurality of weights.

[0252] Clause 16A. The device of any of clauses 11A to 15A, wherein the processing circuitry is configured to: construct a list based on the comparison of the plurality of prediction templates to the current template of the current block; and wherein to determine the weights, the processing circuitry is configured to determine the weights based on entries in the list.

[0253] Clause 17A. The device of any of clauses 11A to 16A, wherein to determine the one or more prediction hypotheses, the processing circuitry is configured to determine the one or more prediction hypotheses based on the comparison of the plurality of prediction templates to the current template of the current block.

[0254] Clause 18A. The device of any of clauses 11A-17A, wherein to determine the prediction signal for the current block based on the one or more prediction hypotheses and the weight, the processing circuitry is configured to: apply the weight to at least one prediction hypothesis of the one or more prediction hypotheses; and combine a base prediction signal with the at least one prediction hypothesis to which the weight is applied to determine the prediction signal for the current block.

[0255] Clause 19. The device of any of clauses 11A-18A, wherein to encode or decode the current block based on the prediction signal, the processing circuitry is configured to decode the current block, wherein to decode the current block, the processing circuitry is configured to: determine residual values that indicate a difference between the prediction signal and the current block; and add the residual values to the prediction signal to reconstruct the current block.

[0256] Clause 20A. A computer-readable storage medium storing instructions thereon that, when executed, cause one or more processors to: determine, for a multiple-hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights; compare the plurality of prediction templates to a current template of a current block; determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determine one or more prediction hypotheses; determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encode or decode the current block based on the prediction signal.

[0257] It is recognized that, in accordance with examples, certain acts or events of any of the techniques described herein can be performed in a different sequence, can be added, merged, or omitted altogether (for example, not all described acts or events are necessary to practice the techniques). Moreover, in certain examples, acts or events can be performed concurrently, for example, through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.

[0258] In one or more examples, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer- readable media generally can correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product can include a computer-readable medium.

[0259] By way of example, and not limitation, such computer-readable storage media can include one or more of RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any

[0260] Instructions can be executed by one or more processors, such as one or more DSPs, general purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein can refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.

[0261] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described herein to emphasize functionality that can be provided by devices configured to perform the techniques described herein. But, it should be understood that the various components, modules, or units can be combined in a codec hardware unit, or provided by a collection of interoperating hardware units, including one or more processors as described above, in conjunction with appropriate software and / or firmware.

[0262] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method of encoding or decoding video data, the method comprising: determining, for a multi-hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights; comparing the plurality of prediction templates to a current template of a current block; determining a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determining one or more prediction hypotheses; determining a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encoding or decoding the current block based on the prediction signal.

2. The method of claim 1, wherein determining the plurality of prediction templates based on the plurality of weights comprises: determining a first prediction template based on a reference template, a hypothesis template, and a first weight of the plurality of weights; and determining a second prediction template based on the reference template, the hypothesis template, and a second weight of the plurality of weights.

3. The method of claim 2, wherein determining the first prediction template comprises:

4. The method of claim 3, wherein determining the second prediction template comprises:

5. The method of claim 2, wherein the hypothesis template is a first hypothesis template, the method further comprising: T pred1 = (1 – α1)T ref + α1T add , where T pred1 is the first prediction template, a1 is the first weight, T ref is the reference template, and T add is the hypothesis template. determining a third prediction template based on the reference template, a second hypothesis template, and the first weight of the plurality of weights; and T pred2 = (1 – α2)T ref + α2T add , where T pred2 is the second prediction template, α2 is the second weight, T ref is the reference template, and T add is the hypothesis template. determining a fourth prediction template based on the reference template, the second hypothesis template, and the second weight of the plurality of weights.

6. The method of claim 1, the method further comprising: constructing a list based on the comparison of the plurality of prediction templates to the current template of the current block, and wherein determining the weight comprises determining the weight based on an entry in the list.

7. The method of claim 1, wherein determining the one or more prediction hypotheses comprises determining the one or more prediction hypotheses based on the comparison of the plurality of prediction templates to the current template of the current block.

8. The method of claim 1, wherein determining the prediction signal for the current block based on the one or more prediction hypotheses and the weight comprises: applying the weight to at least one prediction hypothesis of the one or more prediction hypotheses; and combining a base prediction signal with the at least one prediction hypothesis to which the weight is applied to determine the prediction signal for the current block.

9. The method of claim 1, wherein encoding or decoding the current block based on the prediction signal comprises decoding the current block, wherein decoding the current block comprises: determining residual values that indicate a difference between the prediction signal and the current block; and adding the residual values to the prediction signal to reconstruct the current block.

10. The method of claim 1, wherein encoding or decoding the current block based on the prediction signal comprises encoding the current block, wherein encoding the current block comprises: ​ ​ ​ ​ ​ determining residual values indicative of a difference between the prediction signal and the current block; and signaling information indicative of the residual values.

11. A device for encoding or decoding video data, the device comprising: one or more memories configured to store the video data; and processing circuitry coupled to the one or more memories, wherein the processing circuitry is configured to: determine, for a multiple hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights; compare the plurality of prediction templates to a current template of a current block; determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determine one or more prediction hypotheses; determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encode or decode the current block based on the prediction signal.

12. The device of claim 11, wherein to determine the plurality of prediction templates based on the plurality of weights, the processing circuitry is configured to: determine a first prediction template based on a reference template, a hypothesis template, and a first weight of the plurality of weights; and determine a second prediction template based on the reference template, the hypothesis template, and a second weight of the plurality of weights.

13. The device of claim 12, wherein to determine the first prediction template, the processing circuitry is configured to determine:

14. The device of claim 13, wherein to determine the second prediction template, the processing circuitry is configured to determine: T pred1 = (1 – α1)T ref + α1T add , where T pred1 is the first prediction template, α1 is the first weight, T ref is the reference template, and T add is the hypothesis template.

15. The device of claim 12, wherein the hypothesis template is a first hypothesis template, and wherein the processing circuitry is configured to: T pred2 = (1 – α2)T ref + α2T add , where T pred2 is the second prediction template, α2 is the second weight, T ref is the reference template, and T add is the hypothesis template. determine a third prediction template based on the reference template, a second hypothesis template, and the first weight of the plurality of weights; and determine a fourth prediction template based on the reference template, the second hypothesis template, and the second weight of the plurality of weights.

16. The device of claim 11, wherein the processing circuitry is configured to: construct a list based on the comparison of the plurality of prediction templates to the current template of the current block; and wherein to determine the weight, the processing circuitry is configured to determine the weight based on an entry in the list.

17. The device of claim 11, wherein to determine the one or more prediction hypotheses, the processing circuitry is configured to determine the one or more prediction hypotheses based on the comparison of the plurality of prediction templates to the current template of the current block.

18. The device of claim 11, wherein to determine the prediction signal for the current block based on the one or more prediction hypotheses and the weight, the processing circuitry is configured to: apply the weight to at least one prediction hypothesis of the one or more prediction hypotheses; and ​ combining the base prediction signal with the at least one prediction hypothesis to which the weights are applied to determine the prediction signal for the current block.

19. The device of claim 11, wherein, to encode or decode the current block based on the prediction signal, the processing circuitry is configured to decode the current block, wherein, to decode the current block, the processing circuitry is configured to: determine residual values that indicate a difference between the prediction signal and the current block; and add the residual values to the prediction signal to reconstruct the current block.

20. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: determine, for a multiple-hypothesis prediction (MHP) process, a plurality of prediction templates based on a plurality of weights; compare the plurality of prediction templates to a current template of a current block; determine a weight from the plurality of weights based on the comparison of the plurality of prediction templates to the current template of the current block; determine one or more prediction hypotheses; determine a prediction signal for the current block based on the one or more prediction hypotheses and the weight; and encode or decode the current block based on the prediction signal.