Derivation of Dual Prediction Using Coding Unit-Level Weight Index of Merge Candidates
By deriving coding unit-level weight indices for merge candidates using template matching cost, the method addresses the issue of suboptimal weighting in VVC's dual prediction mode, improving video coding efficiency.
Patent Information
- Application Number
- JP2024577157
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-28
- Filing Date
- 2023-07-03
- Publication Date
- 2025-07-10
AI Technical Summary
The VVC standard's dual prediction mode inherits BCW indices from adjacent blocks, which may not be suitable for merge-coded blocks, leading to suboptimal weighting in bi-prediction.
Derive a dual prediction merge candidate's coding unit-level weight index using template matching cost to determine the most suitable weighting for bi-prediction.
Improves the accuracy of weight indexing in bi-prediction, enhancing video coding efficiency by optimizing the weighting of prediction signals.
Smart Images

Figure 2025521804000001_ABST
Abstract
Description
Technical Field
[0001] Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 358,215, filed on July 4, 2022, entitled "DERIVING BI-PREDICTION WITH CODING UNIT-LEVEL WEIGHT INDICES FOR MERGE CANDIDATES"; U.S. Provisional Application No. 63 / 403,199, filed on September 1, 2022, entitled "DERIVING BI-PREDICTION WITH CODING UNIT-LEVEL WEIGHT INDICES FOR MERGE CANDIDATES"; and U.S. Patent Application No. 18 / 215,753, filed on June 28, 2023, entitled "DERIVING BI-PREDICTION WITH CODING UNIT-LEVEL WEIGHT INDICES FOR MERGE CANDIDATES". All of the above applications are hereby expressly incorporated by reference in their entirety.
[0002] The present disclosure generally relates to video processing, and more particularly, to methods and systems for deriving bi-prediction using coding unit-level weight indices for merge candidates.
Background Art
[0003] In 2020, the Joint Video Experts Team (「JVET」) of the ITU-T Video Coding Expert Group (「ITU-T VCEG」) and the ISO / IEC Moving Picture Expert Group (「ISO / IEC MPEG」) published the final draft of Versatile Video Coding (「VVC」), the next-generation video codec specification. This specification further improves the video coding performance compared to conventional standards such as H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High Efficiency Video Coding). Furthermore, at the time of writing this article, the VVC standard has been extended by the latest draft of the Enhanced Compression Model (「ECM」), which was published as 「Algorithm description of Enhanced Compression Model 8 (ECM 8)」 at the 29th meeting of JVET in January 2023.
[0004] Inter prediction can use single prediction or dual prediction. In single prediction, only one motion vector pointing to one reference picture is used to generate the predictor of the current block. In dual prediction, two motion vectors each pointing to its own reference picture are used to generate the predictor of the current block. According to the previous HEVC standard, the dual prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or by using two different motion vectors. According to the VVC standard, the dual prediction mode is extended to enable weighted averaging of the two prediction signals beyond simple averaging, and this technique is called dual prediction using Coding Unit (「CU」)-level weights (「BCW」) for brevity.
[0005] According to decoder-side motion vector refinement ("DMVR"), dual prediction can be performed on the current coding unit (CU) such that the motion information of the current CU includes a weighted average of two prediction signals, and the weight index is inferred from adjacent blocks based on the merge candidate index. For merge candidates that meet the conditions of DMVR, the adaptive decoder-side motion vector refinement technique further extends multi-path DMVR and refines the motion vector in only one of the two dual prediction directions.
[0006] According to the current VVC and ECM specifications, the BCW index of the merge-coded block is inherited from adjacent blocks according to the signaled merge candidate index. However, the inherited BCW index may not be suitable for the merge-coded block.
SUMMARY OF THE INVENTION
MEANS FOR SOLVING THE PROBLEM
[0007] Embodiments of the present disclosure are directed to methods and systems for deriving dual prediction using a merge candidate's coding unit level weight index.
[0008] In a first aspect, an embodiment of the present disclosure provides a computer system, the computer system comprising one or more processors and a computer-readable storage medium communicatively coupled to the one or more processors, which when executed by the one or more processors derives a dual prediction (BCW) index using the CU-level weight of the dual prediction merge candidate according to the template matching cost for the dual prediction merge candidates of the merge-coded coding unit (CU) list and stores computer-readable instructions executable by the one or more processors to perform related operations including thereof.
[0009] In a second aspect, one embodiment of the present disclosure provides a non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method including steps of deriving a bi-prediction (BCW) index using CU-level weights of bi-prediction merge candidates according to template matching costs for the bi-prediction merge candidates in a merge candidate list of a merged coded coding unit (CU).
[0010] In a third aspect, one embodiment of the present disclosure provides a computer program product including computer program instructions, which cause a computer to be capable of executing a method including steps of deriving a bi-prediction (BCW) index using CU-level weights of bi-prediction merge candidates according to template matching costs for the bi-prediction merge candidates in a merge candidate list of a merged coded coding unit (CU).
[0011] In a fourth aspect, one embodiment of the present disclosure provides a computer program, which causes a computer to be capable of executing a method including steps of deriving a bi-prediction (BCW) index using CU-level weights of bi-prediction merge candidates according to template matching costs for the bi-prediction merge candidates in a merge candidate list of a merged coded coding unit (CU).
[0012] The detailed description is described with reference to the accompanying drawings. In the figures, the leftmost digit of the reference number identifies the figure in which the reference number first appears. When the same reference number is used in different figures, it indicates similar or identical items or features.
Brief Description of the Drawings
[0013]
Fig. 1A
Fig. 1B
Fig. 2
Fig. 3
Fig. 4
DETAILED DESCRIPTION OF THE INVENTION
[0014] According to the VVC video coding standard (the "VVC standard") and the motion prediction described therein, a computing system includes at least one or a plurality of processors and a computer-readable storage medium communicatively coupled to the one or more processors. The computer-readable storage medium is a non-transitory or non-temporary computer-readable storage medium that stores computer-readable instructions. At least some of the computer-readable instructions stored in the computer-readable storage medium are executable by one or more processors of the computing system to perform related operations of the computer-readable instructions, including at least the operations of the encoder described by the VVC standard and the operations of the decoder described by the VVC standard. Although some of these encoder operations and decoder operations according to the VVC standard are described in more detail below, these subsequent descriptions should not be understood to cover all encoder operations and decoder operations according to the VVC standard. Hereinafter, "VVC standard encoder" and "VVC standard decoder" shall describe the respective computer-readable instructions stored in a computer-readable storage medium that configure one or more processors to perform their respective operations (this can be referred to, for example, as the "reference implementation" of the encoder or decoder).
[0015] Furthermore, according to an exemplary embodiment of the present disclosure, a VVC standard encoder and a VVC standard decoder include computer-readable instructions stored in a computer-readable storage medium that are executable by one or more processors of a computing system to configure the one or more processors to perform operations not specified by the VVC standard. The VVC standard encoder should not be understood as being limited to the operations of the reference implementation of the encoder, but is understood to include additional computer-readable instructions that configure one or more processors of a computing system to perform additional operations described herein. The VVC standard decoder should not be understood as being limited to the operations of the reference implementation of the decoder, but is understood to include additional computer-readable instructions that configure one or more processors of a computing system to perform additional operations described herein.
[0016] FIG. 1A and FIG. 1B respectively show exemplary block diagrams of an encoding process 100 and a decoding process 150 according to an exemplary embodiment of the present disclosure.
[0017] In the symbolization process 100, the VVC standard encoder configures one or more processors of the computing system to receive one or more input pictures from the image source 102 as input. The input picture includes several pixels sampled by an image capture device such as a photosensor array, and includes an uncompressed stream of multiple color channels (such as RGB color channels) that store color data at the original resolution of the picture. Each channel uses several bits to store the color data of each pixel of the picture. The VVC standard encoder configures one or more processors of the computing system to store this uncompressed color data in a compressed format, and the color data is stored at a resolution lower than the original resolution of the picture and is encoded as one luminance ("Y") channel and two chrominance ("U" and "V") channels at a resolution lower than the luminance channel.
[0018] The VVC standard encoder encodes a picture (the picture to be encoded is referred to as the "current picture" so as to be distinguished from any other picture received from the image source 102) by configuring one or more processors of the computing system to divide the original picture into units and sub-units according to a partitioning structure. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into macroblocks ("MBs") each having dimensions of 16x16 pixels, and to further subdivide these into partitions. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into coding tree units ("CTUs"), the luma component and chroma component of which can be further subdivided into coding tree blocks ("CTBs"), which are further subdivided into coding units ("CUs"). Alternatively, the VVC standard encoder configures one or more processors of the computing system to subdivide the picture into units of NxN pixels, which can be further subdivided into sub-units. Each of these maximum subdivided units of the picture can generally be referred to as a "block" for the purposes of the present disclosure.
[0019] The CU is encoded using one block of luma samples and two corresponding blocks of chroma samples, and the picture is encoded using one coding tree and is not monochrome.
[0020] The VVC standard encoder configures one or more processors of the computing system to subdivide the block into partitions having dimensions that are multiples of 4x4 pixels. For example, the partitions of the block can have dimensions of 8x4 pixels, 4x8 pixels, 8x8 pixels, 16x8 pixels, or 8x16 pixels.
[0021] Rather than the pixel color information of the original picture at full resolution, by encoding the color information of the picture blocks and their subdivisions, the VVC standard encoder configures one or more processors of the computing system to encode the color information of the picture at a lower resolution than the input picture and store the color information with fewer bits than the input picture.
[0022] Furthermore, the VVC standard encoder encodes the picture by configuring one or more processors of the computing system to perform motion prediction on the blocks of the current picture. Motion prediction coding refers to using motion information and prediction units ("PUs") rather than pixel data to store the image data of the blocks of the current picture (the original blocks of the picture before coding are called "input blocks") according to intra prediction 104 or inter prediction 106.
[0023] Motion information refers to data that describes the motion of the block structure of the picture or its units or subunits, such as motion vectors and references to blocks in the current picture or reference pictures. A PU can refer to a unit or multiple subunits corresponding to one block structure among multiple block structures of the picture, such as an MB or a CTU. The block is divided based on the picture data and encoded according to the VVC standard. The motion information corresponding to the PU can describe the motion prediction encoded by the VVC standard encoder as described herein.
[0024] The VVC standard encoder configures one or more processors of the computing system to encode the motion prediction information for each block of the picture in the coding order within the block, such as raster scan order where the first decoded block is the topmost and leftmost block of the picture. The block to be encoded is called the "current block" to distinguish it from any other block of the same picture.
[0025] According to intra prediction 104, one or more processors of a computing system are configured to encode a block by referring to motion information and PUs of one or more other blocks of the same picture. According to intra prediction coding, one or more processors of a computing system execute intra prediction 104 (also called spatial prediction) calculation by encoding motion information of a current block based on spatially adjacent samples from spatially adjacent blocks of the current block.
[0026] According to inter prediction 106, one or more processors of a computing system are configured to encode a block by referring to motion information and PUs of one or more other pictures. One or more processors of the computing system are configured to store one or more previously encoded and decoded pictures in a reference picture buffer for the purpose of inter prediction coding, and these stored pictures are called reference pictures.
[0027] One or more processors are configured to execute inter prediction 106 (also called temporal prediction or motion compensation prediction) calculation by encoding motion information of a current block based on samples from one or more reference pictures. Inter prediction can be further calculated according to single prediction or dual prediction. In single prediction, a prediction signal of the current block is generated using only one motion vector pointing to one reference picture. In dual prediction, a prediction signal of the current block is generated using two motion vectors each pointing to a respective reference picture.
[0028] According to bi-prediction implemented by the previous HEVC standard, the bi-prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or by using two different motion vectors. In contrast, the bi-prediction mode implemented by the VVC standard is extended to enable weighted averaging of two prediction signals beyond simple averaging, and this technique is called bi-prediction using a coding unit (「CU」)-level weight (「BCW」) for simplicity. The weighted average of two prediction signals P0 and P1 is calculated as P bi-pred as follows, where >> is a bitwise right shift operator. P bi-pred =((8 - w)*P0+w*P1+4)≫3
[0029] Equation 1 is mathematically equivalent to the following equation.
[0030]
Equation
[0031] Therefore, depending on the value of the weight w applied to Equation 1, P0 and P1 can be weighted equally or unequally. When the weight w = 4, P0 and P1 are weighted equally. Therefore, the weight w = 4 is the "equal weight" when applied to Equation 1 (however, this is not necessarily the case when applied to other weighted average bi-prediction equations in this specification). Furthermore, a weight w that gives positive weighting to both P0 and P1 when applied to Equation 1 is a "positive weight" with respect to Equation 1 (where equal weight is also a positive weight), and a weight w that gives negative weighting to one of P0 and P1 when applied to Equation 1 is a "negative weight" with respect to Equation 1.
[0032] According to the bi-prediction mode implemented by the VVC standard, in weighted average bi-prediction, a set of five possible weight values \(w\in\{-2, 3, 4, 5, 10\}\) is allowed. For each bi-predicted CU, the weight \(w\) is determined in one of two ways. 1) For a non-merge-coded CU, the weight index is signaled after the motion vector difference. 2) For a merge-coded CU, the weight index is inferred from adjacent blocks based on the merge candidate index.
[0033] According to the bi-prediction mode implemented by the VVC standard, BCW is applied only to CUs having 256 or more luma samples (i.e., CU width × CU height is 256 or more). In low-delay pictures, all five weights are used. In non-low-delay pictures, only a set of three possible weight values (\(w\in\{3, 4, 5\}\)) is used.
[0034] The VVC standard encoder configures one or more processors of the computing system to code the CUs to include a reference index for identifying the prediction signal of the current block for reference by the VVC standard decoder. The motion vector and the reference index are sent to the decoder to identify where the prediction signal of the current block comes from. One or more processors of the computing system can code the CUs to include an inter-prediction indicator. The inter-prediction indicator indicates list 0 prediction that refers to a first reference picture list called list 0, list 1 prediction that refers to a second reference picture list called list 1, or bi-prediction that refers to both reference picture lists called list 0 and list 1, respectively.
[0035] When the inter prediction indicator indicates a list 0 prediction or a list 1 prediction, one or more processors of the computing system are each configured to encode a CU that includes a reference index for referencing the reference picture of the reference picture buffer referenced by list 0 or list 1. In the case of an inter prediction indicator indicating dual prediction, one or more processors of the computing system are configured to encode a CU that includes a first reference index for referencing the first reference picture of the reference picture buffer referenced by list 0 and a second reference index for referencing the second reference picture of the reference picture referenced by list 1.
[0036] The VVC standard encoder configures one or more processors of the computing system to individually encode each current block of a picture and output a prediction block for each. According to the VVC standard, a CTU can be of a size of up to 128x128 luma samples (and corresponding chroma samples depending on the chroma format). A CTU can be further divided into CUs according to a quadtree, binary tree, or ternary tree. One or more processors of the computing system are configured to finally record a set of coding parameters such as a coding mode (intra mode or inter mode), motion information (reference index, motion vector, etc.) of an inter-coded block, and quantized residual coefficients in the syntax structure of a leaf node of the partitioning structure.
[0037] After the prediction block is output, the VVC standard encoder configures one or more processors of the computing system to send a set of coding parameters such as a coding mode (i.e., intra prediction or inter prediction), an intra prediction mode or an inter prediction mode, and motion information to an entropy encoder 124 (described later).
[0038] The VVC standard provides semantics for recording the coding parameter set of a CU. For example, regarding the above-mentioned coding parameter set, the pred_mode_flag of a CU is set to 0 for an inter-coded block and 1 for an intra-coded block. The general_merge_flag of a CU is set to indicate whether the merge mode is used in the inter-prediction of the CU. The inter_affine_flag and cu_affine_type_flag of a CU are set to indicate whether affine motion compensation is used in the inter-prediction of the CU. The mvp_l0_flag and mvp_l1_flag are set to indicate the motion vector index of list 0 or list 1 respectively. The ref_idx_l0 and ref_idx_l1 are set to indicate the reference picture index of list 0 or list 1 respectively. It should be understood that the VVC standard includes semantics for recording various other information, flags, and options beyond the scope of this disclosure.
[0039] The VVC standard encoder further implements one or more mode decisions and encoder control settings 108, including rate control settings. One or more processors of the computing system are configured to perform mode decision by selecting an optimized prediction mode for the current block based on a rate-distortion optimization method after intra-prediction or inter-prediction.
[0040] The rate control settings configure one or more processors of the computing system to assign different quantization parameters ( "QP") to different pictures. The magnitude of the QP determines the scale at which picture information is quantized by one or more processors during encoding, and thus determines the range in which the encoding process 100 discards picture information (due to information between scale steps) from the MBs of the sequence during coding.
[0041] The VVC standard encoder further implements an adder 110. One or more processors of the computing system are configured to perform an addition operation by calculating the difference between the input block and the prediction block. Based on the optimized prediction mode, the prediction block is subtracted from the input block. The difference between the input block and the prediction block is the prediction residual, or simply called the "residual" for short.
[0042] Based on the prediction residual, the VVC standard encoder further implements a transform 112. One or more processors of the computing system are configured to perform a transform operation on the residual by matrix arithmetic operations to calculate an array of coefficients (which can be called "residual coefficients", "transformation coefficients", etc.), thereby encoding the current block as a transform block ("TB"). The transform coefficients may refer to coefficients representing one of several spatial transforms that can be applied to sub-blocks, such as diagonal inversion, vertical inversion, or rotation.
[0043] As will be described in more detail later, it should be understood that the coefficients can be stored as two components: an absolute value and a sign.
[0044] Sub-blocks of CUs such as PUs and TBs can be arranged in any combination of sub-block dimensions as described above. The VVC standard encoder configures one or more processors of the computing system to subdivide the CU into a residual quad-tree ("RQT") that is a hierarchy of TBs. The RQT provides the order of motion prediction and residual coding at each level of sub-blocks and descends recursively through each level of the RQT.
[0045] The VVC standard encoder further implements quantization 114. One or more processors of the computing system are configured to perform a quantization operation on the residual coefficients by matrix arithmetic operations based on the quantization matrix and QP assigned above. Residual coefficients within the interval are retained, and residual coefficients outside the interval step are discarded.
[0046] The VVC standard encoder further implements inverse quantization 116 and inverse transformation 118. One or more processors of a computing system are configured to perform an inverse quantization operation and an inverse transformation operation on the quantized residual coefficients by matrix arithmetic operations that are the inverse of the quantization operation and the transformation operation described above. The inverse quantization operation and the inverse transformation operation result in a reconstructed residual.
[0047] The VVC standard encoder further implements an adder 120. One or more processors of a computing system are configured to perform an addition operation by adding a prediction block and a reconstructed residual to output a reconstructed block.
[0048] The VVC standard encoder further implements a loop filter 122. One or more processors of a computing system are configured to apply loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF) to the reconstructed block to output a filtered reconstructed block.
[0049] The VVC standard encoder further configures one or more processors of a computing system to output the filtered reconstructed block to a decoded picture buffer (DPB) 200. The DPB 200 stores reconstructed pictures that are used by one or more processors of a computing system as reference pictures when encoding pictures other than the current picture, as described above for inter prediction.
[0050] The VVC standard encoder further implements an entropy encoder 124. One or more processors of a computing system are configured to perform entropy coding, and symbols constituting quantized residual coefficients are encoded by mapping to a binary string (hereinafter referred to as "bin") according to a context-dependent binary arithmetic coder ("CABAC") and can be transmitted in an output bitstream at a compressed bitrate. The symbols of the quantized residual coefficients to be encoded include the absolute values of the residual coefficients (these absolute values are hereinafter referred to as "residual coefficient levels").
[0051] Accordingly, the entropy encoder encodes the residual coefficient levels of a block, bypasses the encoding of the residual coefficient signs, records the residual coefficient signs together with the encoded block, and configures one or more processors of the computing system to record a set of coding parameters such as a coding mode, an intra prediction mode, or an inter prediction mode, and motion information encoded in the syntax structure of the encoded block (a picture parameter set ("PPS") found in a picture header, and a sequence parameter set ("SPS") found in a sequence of multiple pictures, etc.) and output the encoded block.
[0052] The VVC standard encoder configures one or more processors of a computing system to output an encoded picture composed of encoded blocks from the entropy encoder 124. The encoded picture is output to a transmission buffer and finally packed into a bitstream output from the VVC standard encoder. The bitstream is written by one or more processors of the computing system to a non-transitory or non-volatile computer-readable storage medium of the computing system for transmission.
[0053] In the decoding process 150, the VVC standard decoder configures one or more processors of the computing system to receive one or more coded pictures from the bitstream as input.
[0054] The VVC standard decoder implements an entropy decoder 152. One or more processors of the computing system are configured to perform entropy decoding. According to CABAC, bins are decoded by reversing the mapping of symbols to bins, thereby restoring the entropy-coded quantized residual coefficients. The entropy decoder 152 outputs the quantized residual coefficients, outputs the coding bypassed residual coefficient signs, and also outputs syntax structures such as PPS and SPS.
[0055] The VVC standard decoder further implements inverse quantization 154 and inverse transform 156. One or more processors of the computing system are configured to perform inverse quantization and inverse transform operations on the decoded quantized residual coefficients by matrix arithmetic operations that are the inverse of the quantization and transform operations described above. The inverse quantization and inverse transform operations result in a reconstructed residual.
[0056] Furthermore, based on the coding parameter set recorded in syntax structures such as PPS and SPS by the entropy encoder 124 (or received by out-of-band transmission or coded by the decoder) and the coding mode included in the coding parameter set, the VVC standard decoder determines whether to apply intra prediction 156 (i.e., spatial prediction) or motion compensation prediction 158 (i.e., temporal prediction) to the reconstructed residual.
[0057] When the coding parameter set specifies intra prediction, the VVC standard decoder configures one or more processors of the computing system to perform intra prediction 158 using the prediction information specified in the coding parameter set. Thereby, intra prediction 158 generates a prediction signal.
[0058] When the coding parameter set specifies inter prediction, the VVC standard decoder configures one or more processors of the computing system to perform motion compensation prediction 160 using reference pictures from the DPB 200. Thereby, motion compensation prediction 160 generates a prediction signal.
[0059] The VVC standard decoder further implements an adder 162. The adder 162 configures one or more processors of the computing system to perform an addition operation on the reconstructed residual and the prediction signal, thereby outputting the reconstructed block.
[0060] The VVC standard decoder further implements a loop filter 164. One or more processors of the computing system are configured to apply loop filters such as a deblocking filter, an SAO filter, and an ALF to the reconstructed block to output a filtered reconstructed block.
[0061] The VVC standard decoder further configures one or more processors of the computing system to output the filtered reconstructed block to the DPB 200. As described above, the DPB 200 stores reconstructed pictures used by one or more processors of the computing system as reference pictures when coding pictures other than the current picture, as described above for motion compensation prediction.
[0062] The VVC standard decoder further configures one or more processors of the computing system to output the reconstructed picture from the DPB to a display that can be viewed by a user of the computing system, such as a television display, a personal computing monitor, a smartphone display, or a tablet display.
[0063] Therefore, as shown by the above-described encoding process 100 and decoding process 150, the VVC standard encoder and the VVC standard decoder each implement motion prediction coding according to the VVC specification. The VVC standard encoder and the VVC standard decoder each configure one or more processors of the computing system to generate a reconstructed picture based on the reconstructed picture before the DPB according to motion compensation prediction described by the VVC standard, and the previous reconstructed picture functions as a reference picture in the motion compensation prediction described herein.
[0064] The VVC standard encoder and the VVC standard decoder each configure one or more processors of the computing system to predict the motion information of the CU of the reconstructed picture by various merge modes, such as a normal merge mode, a template matching merge mode, a decoder-side motion vector refinement ("DMVR") mode, a motion vector difference ("MMVD") mode, an affine merge mode, an affine + MMVD merge mode, a composite inter prediction and intra prediction ("CIIP") mode. Such a CU having motion information predicted by the merge mode is hereinafter referred to as a "merge-coded CU". The motion information may include a plurality of motion vectors. The derivation of the motion vector is known to those skilled in the art and need not be described again herein.
[0065] The motion information of the CU of the reconstructed picture may include a motion candidate list. The motion candidate list may be a data structure including references to a plurality of motion candidates. A motion candidate may be a block structure or its sub-unit, such as a pixel or any other suitable subdivision of the block structure of the current picture, or a reference to a motion candidate of another picture. A motion candidate may be a spatial motion candidate or a temporal motion candidate. By applying motion vector compensation (“MVC”), the VVC standard decoder configures one or more processors of the computing system to select a motion candidate from the motion candidate list and derive the motion vector of the motion candidate as the motion vector of the CU of the reconstructed picture.
[0066] The motion candidate list may be a merge candidate list and may include up to five types of merge candidates or six types of merge candidates according to ECM. The merge candidate list can be signaled to one or more processors of the computing system during the coding of the CU by a merge candidate index. The spatial merge candidates of the list can be derived, for example, by searching for adjacent blocks of the current CU and adding the motion information of those adjacent blocks to the merge candidate list.
[0067] The temporal merge candidates of the list can be derived, for example, by deriving the motion information of the co-located CU belonging to the co-located reference picture.
[0068] The merge candidate list of the current CU coded according to the merge mode may include the following merge candidates, in order, from the spatial motion vector prediction (“MVP”) candidates from the CU spatially adjacent to the current CU, the temporal MVP candidates from the co-located CU of the current CU, the history-based MVP (“HMVP”) candidates from the FIFO table, the pairwise average MVP candidates, and the zero motion vector.
[0069] Furthermore, after the merge candidate list is constructed, the merge candidates are sorted according to an adaptive rearrangement of the merge candidates by template matching, hereinafter referred to as "ARMC-TM". The merge candidates are sorted in ascending order of cost values based on template matching. For simplicity, the merge candidates of the last subgroup instead of the first subgroup are not sorted. The template matching cost of the merge candidates is measured by the sum of absolute differences ("SAD") calculated between the samples of the template of the current block and the corresponding reference samples. The template of the current block includes a set of reconstructed samples adjacent to the current block. The reference samples of the template of the current block are located by the motion information of the merge candidates.
[0070] ARMC-TM can be applied to motion prediction by merge modes including the normal merge mode, CIIP mode, adaptive decoder-side motion vector refinement mode, template matching ("TM") merge mode, and affine merge mode, excluding sub-block-based temporal motion vector prediction ("SbTMVP").
[0071] Any merge candidate in the merge candidate list can include motion information predicted by bi-prediction. In such a case, the reference samples of the template of the merge candidate are also generated by bi-prediction, and thus, two-directional template matching can be performed based on such merge candidates, hereinafter referred to as "two-directional matching (BM) candidates".
[0072] Figure 2 shows the motion prediction performed on the current picture 202 according to dual prediction. The current picture 202 includes a current block 202A that includes a template 202B. The reference samples of the template refer to two co-located reference pictures 204 and 206, one from the first temporal reference list 0 and one from the second temporal reference list 1, according to dual prediction. The motion information of the current block 202A refers to the co-located reference block 204A of the co-located reference picture 204 and the co-located reference block 206A of the co-located reference picture 206. The template 202B of the current block 202A refers to the reference sample 204B of the co-located reference picture 204 and the reference sample 206B of the co-located reference picture 206.
[0073] According to the adaptive decoder-side motion vector refinement technique, for merge candidates that satisfy the DMVR conditions, two additional merge modes are provided. One is a merge mode in which the motion vector of the current CU is refined by multi-pass DMVR in the first temporal direction but not in the second temporal direction, and the other is a merge mode in which the motion vector of the current CU is refined by multi-pass DMVR in the second temporal direction but not in the first temporal direction.
[0074] For both of the additional merge modes, a common merge candidate list is constructed only from merge candidates that satisfy the DMVR conditions. The merge candidates in the common merge candidate list are derived from spatially adjacent coded blocks, TMVP, non-adjacent blocks, HMVP, and pairwise candidates in a manner similar to the above-described merge candidate list.
[0075] If the merge candidate list includes BM candidates as described above in Figure 2, such dual-prediction merge candidates inherit the weight index inferred from adjacent blocks based on the signaled merge candidate index for weighted averaging of the prediction signals.
[0076] Given this merge candidate list, a multi-pass DMVR process is applied to the merge candidates to refine the motion vectors. The DMVR process is modified by measuring distortion by calculating the motion vector difference ("MVD") between the motion vector and the motion vector prediction ("MVP"). The MVD can be calculated by the mean removed sum of absolute differences ("MRSAD"), or, if the weights are unequal and bi-prediction is weighted by BCW weights, by the mean removed sum of absolute transform differences ("MRSATD"). During the first pass of the multi-pass DMVR process (i.e., at the PU level), either MVD0 (the first temporal MVD) or MVD1 (the second temporal MVD) is set to zero. The merge index is coded in the same way as in the normal merge mode.
[0077] According to an exemplary embodiment of the present disclosure, additional BCW weights can be derived based on the picture order count ("POC") difference.
[0078] As an example, if both reference pictures are in the same temporal direction with respect to the current picture (i.e., both from the past or both from the future) and the current picture is a low-delay picture, the weight pair (-3, 11) can be added.
[0079] As another example, if both reference pictures are in the same temporal direction with respect to the current picture (i.e., both from the past or both from the future) and the current picture is a non-low-delay picture, the weight pair (-2, 10) can be added.
[0080] As yet another example, in any other situation, the weight pair (2, 6) can be added.
[0081] According to any of the above examples, when the POC distances are the same, the larger value from the weight pair can be assigned to the nearest POC reference picture or the list 0 reference picture.
[0082] According to any of the above examples, for the weighted average double prediction equation, instead of equal weights, additional BCW weights can be assigned to pairwise candidates and zero merge candidates.
[0083] As described above, according to the VVC standard and ECM specification, the BCW index of the merge-coded block is inherited from adjacent blocks according to the signaled merge candidate index. However, the inherited BCW index may not be suitable for the merge-coded block. Therefore, exemplary embodiments of the present disclosure perform the derivation of the BCW index according to a cost value based on template matching.
[0084] In one or more aspects, exemplary embodiments of the present disclosure provide for the derivation of a merge candidate BCW index according to a template matching cost.
[0085] In one or more aspects, exemplary embodiments of the present disclosure provide for the derivation of a merge candidate BCW index according to a template matching cost by reducing a set of possible weight values.
[0086] In one or more aspects, exemplary embodiments of the present disclosure provide for the derivation of a merge candidate BCW index according to a template matching cost with an extended set of possible weight values.
[0087] In one or more aspects, exemplary embodiments of the present disclosure provide for the derivation of a merge candidate BCW index according to a template matching cost with an extended set of possible weight values, improving the weighting accuracy.
[0088] In one or more aspects, exemplary embodiments of the present disclosure provide for the derivation of a merge candidate BCW index according to a template matching cost while adjusting the template matching cost of the inherited BCW weight.
[0089] In one or more aspects, exemplary embodiments of the present disclosure provide the application of bidirectional optical flow to bidirectional prediction candidates with equal weights with respect to the weighted average bi-prediction equation.
[0090] In one or more aspects, exemplary embodiments of the present disclosure provide the derivation of a merge candidate BCW index according to a template matching cost before applying ARMC-TM to the merge candidate.
[0091] In one or more aspects, exemplary embodiments of the present disclosure provide the derivation of a merge candidate BCW index according to a template matching cost after applying ARMC-TM to the merge candidate.
[0092] In one or more aspects, exemplary embodiments of the present disclosure provide the derivation of a merge candidate BCW index according to a template matching cost for a subset of possible merge modes.
[0093] In one or more aspects, exemplary embodiments of the present disclosure provide the derivation of a merge candidate BCW index while inheriting the BCW index from non-adjacent spatial merge candidates.
[0094] Each of the above aspects of the exemplary embodiments of the present disclosure will be described in more detail below.
[0095] FIG. 3 shows a flowchart of a method 300 of configuring one or more processors of a computing system such that a VVC standard encoder or VVC standard decoder derives a merge candidate BCW index according to a template matching cost while utilizing a set of possible weight values according to the VVC standard and ECM specification.
[0096] As used herein, "on the decoder side" does not mean that this method is only implemented by a decoder. Rather, it should be understood that the steps of this method can be implemented similarly or identically by an encoder and a decoder. References to a "VVC standard encoder or VVC standard decoder" mean that the steps are executed in a similar or identical manner by one or more processors of a computing system, regardless of whether the processor is configured by a VVC standard encoder or a VVC standard decoder.
[0097] As described above, in step 302, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to construct a merge candidate list for a merged coded CU.
[0098] In step 304, for each merge candidate in the merge candidate list, if the merge candidate is a bi-prediction candidate, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to derive a BCW index of the bi-prediction merge candidate according to a template matching cost. The bi-prediction weight of the merge candidate is determined by a weight index of the bi-prediction merge candidate (hereinafter referred to as "BCW index"). A VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to derive the BCW index in one or more sub-steps described later with reference to sub-steps 304A, 304B, 304C, 304D, and 304E.
[0099] In step 304A, for each dual-prediction weight (i.e., BCW index), the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to calculate the template matching cost using the sum of absolute differences (SAD) between the samples of the template of the merged-coded CU and their respective corresponding reference samples. (According to the dual prediction based on the VVC standard, it should be understood that the template includes the reconstructed samples adjacent to the left side and / or the upper side of the merged-coded CU, such as those shown in template 202B of FIG. 2).
[0100] In step 304B, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to generate the reference samples of the template by dual prediction using the corresponding dual-prediction weight values.
[0101] In step 304C, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to calculate the template matching cost for each of the possible sets of BCW weight values and select the BCW weight value that results in the lowest template matching cost among the calculated template matching costs as the BCW index of the dual-prediction merge candidate. Various examples of performing step 304C will be described subsequently according to the exemplary embodiments of the present disclosure.
[0102] For example, in the case of a dual-prediction merge candidate in a non-low-delay picture, for each of the three possible sets of BCW weight values, three template matching cost values are calculated. The three possible sets of BCW weight values each represent three different weight values (w ∈ {3, 4, 5}, where w is as defined above with reference to Equation 1).
[0103] Therefore, among the three template matching costs calculated from these three possible weight values, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to determine the weight value that results in the lowest template matching cost as the BCW index of the bi-prediction merge candidate.
[0104] According to the example of step 304C, the set of possible weight values corresponds to the set of possible weight values in the weighted average bi-prediction provided by the VVC standard and / or ECM specification. That is, for low-delay pictures, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to calculate the template matching cost from a set of five possible weight values (w ∈ {-2, 3, 4, 5, 10}), and for non-low-delay pictures, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to calculate the template matching cost from a set of three possible weight values (w ∈ {3, 4, 5}). Further, the weighted average of the two prediction signals P0 and P1 is P bi-pred calculated as follows.
[0105] According to a further example of step 304C, the set of possible weight values is a reduced set of possible weight values compared to the set of possible weight values in the weighted average bi-prediction provided by the VVC standard and / or ECM specification. For both low-delay pictures and non-low-delay pictures, the VVC standard encoder or VVC standard decoder calculates the template matching cost from a set of three possible weight values (w ∈ {3, 4, 5}), that is, configures one or more processors of the computing system so that the negative weights for Equation 1 are excluded. Further, the weighted average of the two prediction signals P0 and P1 is P bi-pred calculated as follows.
[0106] According to a further example of step 304C, the set of possible weight values is an extended set of possible weight values compared to the possible weight values in the weighted average bi-prediction provided by the VVC standard and / or the ECM specification. As an example, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to calculate a template matching cost from a set of seven possible weight values (w ∈ {1, 2, 3, 4, 5, 6, 7}) for both low-delay pictures and non-low-delay pictures. This can be implemented by removing bit overhead in the coding of the CU for signaling the BCW index. Further, the weighted average of two prediction signals P0 and P1 is P bi-pred calculated as follows.
[0107] According to a further example of step 304C, the set of possible weight values is an extended set of possible weight values compared to the possible weight values in the weighted average bi-prediction provided by the VVC standard and / or the ECM specification. As an example, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to calculate a template matching cost from a set of five possible weight values (w ∈ {6, 7, 8, 9, 10}) for both low-delay pictures and non-low-delay pictures. Further, the weighted average of two prediction signals P0 and P1 is calculated as P bi-pred where >> is a bitwise right shift operator. P bi-pred =((16 - w)*P0 + w*P1 + 8) >> 4
[0108] Equation 2 is mathematically equivalent to the following equation.
[0109]
Equation
[0110] Depending on the value of the weight w applied to Equation 2, P0 and P1 can be weighted equally or unequally. When the weight w = 8, P0 and P1 are weighted equally. Therefore, the weight w = 8 is the "equal weight" when applied to Equation 2 (however, this is not necessarily the case when applied to other weighted average double prediction equations in this specification). Further, a weight w that assigns positive weights to both P0 and P1 when applied to Equation 2 is a "positive weight" with respect to Equation 2 (where equal weights are also positive weights), and a weight w that assigns a negative weight to one of P0 and P1 when applied to Equation 2 is a "negative weight" with respect to Equation 2.
[0111] According to a further exemplary embodiment of the present disclosure, in step 304D, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to calculate a template matching cost for each of a subset of possible BCW weight values based on the inherited BCW weights and select the BCW weight value that results in the lowest template matching cost among the calculated template matching costs as the BCW index of the double prediction merge candidate. Various examples of performing step 304D will be subsequently described according to exemplary embodiments of the present disclosure.
[0112] According to an example of step 304D, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to calculate a template matching cost for each of a subset of possible BCW weight values based on the fact that the inherited BCW weight is a positive weight with respect to the weighted average double prediction equations (i.e., w in Equations 1 and 2). When the inherited weight is positive, the subset of possible weight values for which the template matching cost is calculated includes only positive weights with respect to the weighted average double prediction equations (positive weights are determined differently with respect to Equations 1 and 2 above). Note that when the weight w is positive, both the weight for P0 (i.e., 8 - w with respect to Equation 1 and 16 - w with respect to Equation 2) and the weight of P1 (i.e., w) are greater than 0.
[0113] As an example, the set of possible weight values for Equation 1 is determined to be {-2, 3, 4, 5, 10}. If the inherited weight is positive (i.e., for Equation 1, the inherited weight is any of {3, 4, 5}), then according to step 304D, the template matching costs for {3, 4, 5} are calculated, and according to step 304D, the template matching costs for {-2, 10} are not calculated. If the inherited weight is not positive, then the template matching costs for {-2, 3, 4, 5, 10} are calculated as in step 304C.
[0114] According to an example of step 304D, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to calculate the template matching cost for each of the subsets of possible BCW weight values based on the fact that the inherited BCW weight is a negative weight for the weighted average bi-prediction equation. If the inherited weight is negative, the subset of possible weight values for which the template matching cost is calculated includes only negative weights for the weighted average bi-prediction equation (the negative weights are determined to be different for Equations 1 and 2 above). If the weight w is negative, at least one of the weights of P0 and P1 will be less than 0.
[0115] As an example, the set of possible weight values according to one or more examples of step 304C above is determined to be {-2, 3, 4, 5, 10}. If the inherited weight is negative (i.e., either {-2, 10} for Equation 1), then according to step 304D, the template matching costs for {-2, 10} are calculated.
[0116] According to a further embodiment that extends the foregoing example, even if the inherited weight is negative, the subset of possible BCW weight values for which the template matching cost is calculated includes equal weights for the weighted average bi-prediction equation.
[0117] As an example, a set of possible weight values according to one or more examples of step 304C above is determined to be {-2, 3, 4, 5, 10}. When the inherited weight is negative (i.e., either of {-2, 10} with respect to Equation 1), according to step 304D, a template matching cost of {-2, 4, 10} is calculated (4 is not a negative weight with respect to Equation 1 but is an equal weight with respect to Equation 1).
[0118] According to an example of step 304D, one or more processors of a computing system are configured such that a VVC standard encoder or a VVC standard decoder calculates a template matching cost for each in a subset of possible BCW weight values similar to the inherited BCW weight. When the absolute difference between the weight w and the inherited weight is less than a predefined threshold, the weight w is similar to the inherited weight. The predefined threshold is a positive integer other than zero.
[0119] As an example, a set of possible weight values according to one or more examples of step 304C above is determined to be {-2, 3, 4, 5, 10}, and the predefined threshold is set to 2. As a result, when the inherited weight is 4, a template matching cost of {3, 4, 5} is calculated according to step 304D. When the inherited weight is 3, a template matching cost of {3, 4} is calculated according to step 304D. When the inherited weight is -2, a template matching cost of {-2} is calculated according to step 304D.
[0120] As another example, a set of possible weight values is determined to be {-4, -3, -2, -1, 1, 2, 3, 4, 5, 6, 7, 9, 10, 11, 12}, and a predefined threshold is set to 2. As a result, when the inherited weight is 4, the template matching cost of {3, 4, 5} is calculated according to step 304D. When the inherited weight is 3, the template matching cost of {2, 3, 4} is calculated according to step 304D. When the inherited weight is -2, the template matching cost of {-3, -2, -1} is calculated according to step 304D.
[0121] As yet another example, a set of possible weight values according to one or more examples of step 304C above is determined to be {1, 2, 3, 4, 5, 6, 7}, and a predefined threshold is set to 2.
[0122] As yet another example, a set of possible weight values according to one or more examples of step 304C above is determined to be {0, 1, 2, 3, 4, 5, 6, 7, 8}, and a predefined threshold is set to 2.
[0123] According to an example of step 304D, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to calculate a template matching cost for each of a subset of possible BCW weight values based on an inherited BCW weight, and the subset is further based on a POC distance between a current picture and a reference picture.
[0124] As an example, when using a template matching method to derive a BCW index of a dual-prediction merge candidate, a POC distance (hereinafter referred to as Poc_diff_L0) between the current picture and a reference picture from reference picture list 0, and a POC distance (hereinafter referred to as Poc_diff_L1) between the current picture and a reference picture from reference picture list 1 are calculated according to step 304D.
[0125] When Poc_diff_L0 and Poc_diff_L1 are equal, the template matching cost for all weights within the set of possible weight values is calculated as in step 304C. When Poc_diff_L0 is smaller than Poc_diff_L1, the template matching cost for weights that are less than or equal to the equal weights with respect to the weighted average dual prediction equation is calculated according to step 304D. That is, the weight applied to P1 (i.e., w) is less than or equal to the weight applied to P0 (8 - w for equation 1, or 16 - w for equation 2), P1 is a prediction signal using a reference picture from reference picture list 1, and P0 is a prediction signal using a reference picture from reference picture list 0. On the other hand, when Poc_diff_L0 is larger than Poc_diff_L1, the template matching cost for weights that are greater than or equal to the equal weights with respect to the weighted average dual prediction equation is calculated according to step 304D. That is, the weight applied to P1 is greater than or equal to the weight applied to P0.
[0126] As an example, the set of possible weight values according to one or more examples of step 304C above is determined to be {-2, 3, 4, 5, 10}. When Poc_diff_L0 and Poc_diff_L1 are equal, the template matching cost of {-2, 3, 4, 5, 10} is calculated as in step 403C. When Poc_diff_L0 is smaller than Poc_diff_L1, the template matching cost of {-2, 3, 4} is calculated according to step 304D. When Poc_diff_L0 is larger than Poc_diff_L1, the template matching cost of {4, 5, 10} is calculated according to step 304D.
[0127] According to an example of step 304D, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to calculate the template matching cost for each in a subset of possible BCW weight values similar to the inherited BCW weights, and the subset is further based on the POC distance between the current picture and the reference picture.
[0128] When deriving the BCW index of the dual prediction merge candidate using the template matching method, in addition to Poc_diff_L0 and Poc_diff_L1, the value diff_w is calculated by subtracting the weights inherited from each of the possible sets of BCW weight values. When Poc_diff_L0 and Poc_diff_L1 are equal, the template matching cost of the weight whose absolute value of diff_w is smaller than the first predefined threshold is calculated. When Poc_diff_L0 is smaller than Poc_diff_L1, diff_w is 0 or less, and the template matching cost of the weight whose absolute value of diff_w is smaller than the second predefined threshold is calculated. When Poc_diff_L0 is larger than Poc_diff_L1, diff_w is 0 or more, and the template matching cost of the weight whose absolute value of diff_w is smaller than the third predefined threshold is calculated. The first, second, and third predefined thresholds may be the same as or different from each other.
[0129] As an example, the set of possible weight values according to one or more examples of step 304C above is determined to be {1, 2, 3, 4, 5, 6, 7}. The first, second, and third predefined thresholds are set to 2, 3, and 3, respectively. Assume that the inherited weight is 4.
[0130] As a result, when Poc_diff_L0 and Poc_diff_L1 are equal, the template matching cost of {3, 4, 5} is calculated according to step 304D. When Poc_diff_L0 is smaller than Poc_diff_L1, the template matching cost of {2, 3, 4} is calculated according to step 304D. When Poc_diff_L0 is larger than Poc_diff_L1, the template matching cost of {4, 5, 6} is calculated according to step 304D.
[0131] According to a further embodiment, when deriving the BCW index of the dual-prediction merge candidate using the template matching method, the absolute difference between each of Poc_diff_L0, Poc_diff_L1, and the set of possible BCW weight values and the inherited weight is calculated. If Poc_diff_L0 and Poc_diff_L1 are equal, the template matching cost of the weight whose absolute difference is smaller than the first predefined threshold is calculated. If Poc_diff_L0 and Poc_diff_L1 are not equal, the template matching cost of the weight whose absolute difference is smaller than the second predefined threshold is calculated. The first predefined threshold is different from the second predefined threshold.
[0132] As an example, the set of possible weight values according to one or more examples of step 304C above is determined to be {1, 2, 3, 4, 5, 6, 7}. The first and second predefined thresholds are set to 3 and 2, respectively. Assume the inherited weight is 3.
[0133] As a result, if Poc_diff_L0 and Poc_diff_L1 are equal, the template matching cost of {1, 2, 3, 4, 5} is calculated according to step 304D. If Poc_diff_L0 and Poc_diff_L1 are not equal, the template matching cost of {2, 3, 4} is calculated according to step 304D.
[0134] Furthermore, in step 304E, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to derive a merge candidate BCW index according to the template matching cost while adjusting the template matching cost of the BCW weight inherited from the value calculated according to the VVC standard and the ECM specification. As described above, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to calculate the template matching cost using the sum of absolute differences (SAD) between the samples of the template of the merge-coded CU and their respective corresponding reference samples for each dual-prediction weight (i.e., BCW index). However, since the BCW index inherited from the adjacent block is likely to be more accurate than another BCW index, the VVC standard encoder or VVC standard decoder further configures one or more processors of the computing system to adjust so that the template matching cost of the inherited BCW index is preferentially selected over the template matching cost of other BCW indexes.
[0135] As an example, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to multiply the template matching cost (TMcost) of the inherited BCW index by a weight less than 1. Thus, the inherited BCW index is likely to be selected. To reduce the computational complexity, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to perform bitwise shift operations and subtraction operations to avoid multiplication. For example, the weight applied to the template matching cost of the inherited BCW index is set to 0.90625, which can be replaced by TMcost-(TMcost>>4)-(TMcost>>5).
[0136] As another example, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to subtract an offset value from the template matching cost of an inherited BCW index. The offset value is a positive number other than zero and is determined according to a quantization parameter (“QP”). For example, the offset value can be set equal to the Lagrange multiplier λ.
[0137] Considering the inherited weights with higher accuracy, weights that are not similar to the inherited weights may not be highly expected. According to a further embodiment, a value greater than 1 is multiplied by the template matching cost of a weight that is not similar to the inherited weight. Based on the absolute difference between the weight and the inherited weight being greater than a predefined threshold, the weight is determined to not be similar to the inherited weight. As described above, weights that are not similar to the inherited weights can be excluded from the calculation of the template matching cost.
[0138] As an example, the predefined threshold is set to 1, and the set of possible weight values according to one or more examples of step 304C above is determined to be {−2, 3, 4, 5, 10}. Assume the inherited weight is 3. The template matching cost TMcost of {−2, 5, 10} is multiplied by 1.09375, which can be replaced with TMcost+(TMcost>>4)+(TMcost>>5).
[0139] According to an exemplary embodiment of the present disclosure, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to apply bidirectional optical flow (“BDOF”) to a bi-prediction block with equal weights. By the BDOF technique, the prediction samples of the bi-prediction block can be refined.
[0140] Furthermore, in step 304F, as an option, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to adjust the template matching cost of the BCW weights with equal weights from the values calculated according to the VVC standard and the ECM specification. Since the application of BDOF is likely to be beneficial for the prediction samples, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to adjust the template matching cost of the BCW index with equal weights for the weighted average bi-prediction equation so that it is preferentially selected over the template matching costs of other BCW indexes. As an example, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to multiply the template matching cost of the BCW index with equal weights for the weighted average bi-prediction equation by a weight less than 1, or to subtract a positive value from the template matching cost of the BCW index with equal weights for the weighted average bi-prediction equation.
[0141] Alternatively, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to apply BDOF to all bi-prediction blocks regardless of the respective merge candidate BCW indexes.
[0142] Furthermore, according to an exemplary embodiment of the present disclosure, step 304 in which the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to derive a merge candidate BCW index according to the template matching cost can be executed before applying ARMC-TM to the merge candidate, or can be executed after applying ARMC-TM to the merge candidate.
[0143] According to some exemplary embodiments, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to derive a merge candidate BCW index according to a template matching cost before applying ARMC-TM to the merge candidates. After the merge candidate list is constructed, the VVC standard encoder or the VVC standard decoder first configures one or more processors of the computing system to derive the BCW index of each bi-prediction merge candidate in the merge candidate list according to the template matching cost, and then configures one or more processors of the computing system to apply ARMC-TM to reorder the merge candidates. The VVC standard encoder or the VVC standard decoder further configures one or more processors of the computing system to use the derived BCW index to calculate the template matching cost value of the bi-prediction merge candidates during ARMC-TM.
[0144] According to some exemplary embodiments, after applying ARMC-TM to merge candidates, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to derive a merge candidate BCW index according to the template matching cost, reducing the decoding complexity. The VVC standard encoder or VVC standard decoder first sorts the merge candidates in the merge candidate list by applying ARMC-TM, and then configures one or more processors of the computing system to calculate the template matching cost value of the bi-predicted merge candidates using the BCW index inherited from adjacent blocks during ARMC-TM. The VVC standard encoder or VVC standard decoder then further configures one or more processors of the computing system to select a merge candidate according to the signaled merge index, and for the selected merge candidate that is bi-predicted, the VVC standard encoder or VVC standard decoder configures one or more processors of the computing system to derive its BCW index according to the template matching cost. In this way, compared with the derivation of the BCW index before applying ARMC-TM, the derivation of the BCW index after applying ARMC-TM may be less.
[0145] The merge candidate BCW index technique described herein can be applied to any subset of possible merge modes that configure one or more processors of the computing system for the VVC standard encoder or VVC standard decoder to apply when coding a CU according to the VVC standard and ECM specification. As an example, a subset of possible merge modes includes the standard merge mode, the template matching merge mode, the decoder-side motion vector refinement mode, the MMVD mode, the affine merge mode, the affine + MMVD merge mode, and the CIIP mode, excluding other merge modes.
[0146] As another example, the merge candidate BCW index technique described herein is applied to a first subset of possible merge modes for non-low-delay pictures and a second subset of possible merge modes for low-delay pictures. For example, the first subset of possible merge modes includes the standard merge mode, the template matching merge mode, the decoder-side motion vector refinement mode, and the affine merge mode, excluding other merge modes, and the second subset of possible merge modes includes the normal merge mode, the template matching merge mode, the decoder-side motion vector refinement mode, the MMVD mode, the affine merge mode, and the affine + MMVD merge mode, excluding other merge modes.
[0147] As another example, the merge candidate BCW index technique described herein is applied to a subset of possible merge modes for both non-low-delay pictures and low-delay pictures, including the standard merge mode, the template matching merge mode, the decoder-side motion vector refinement mode, the CIIP mode, and the affine merge mode, excluding other merge modes.
[0148] According to some exemplary embodiments, a VVC standard encoder or a VVC standard decoder configures one or more processors of a computing system to derive a merge candidate BCW index according to a template matching cost while inheriting the BCW index from non-adjacent spatial merge candidates. For a current CU encoded using the decoder-side motion vector refinement mode, if the BCW index of the non-adjacent spatial merge candidate is not inherited by the current CU, when the current CU inherits motion information from the non-adjacent spatial merge candidate, the BCW index is always set to equal weights. However, the BCW index is inherited from spatially adjacent blocks. Thus, according to the exemplary embodiments of the present disclosure, the BCW index from non-adjacent spatial merge candidates is also inherited for the current CU for the decoder-side motion vector refinement mode, thereby aligning both modes.
[0149] When additional BCW weights are derived based on the POC difference (as described above), the BCW index may or may not be derived using the template matching cost in merge mode, as described below.
[0150] According to one embodiment, when the BCW index derived using the template matching cost is valid for merge mode, the default weights are always assigned to pairwise and zero merge candidates regardless of whether the additional BCW weights are valid.
[0151] According to another embodiment, when the inherited weight is the additional BCW weight, the method of deriving the BCW index using the template matching cost in merge mode is not applied.
[0152] According to yet another embodiment, when the method of deriving the BCW index using the template matching cost is valid for merge mode, the additional BCW weight is applied only to the low temporal layers to further reduce the bit overhead.
[0153] Those skilled in the art will understand that all of the above aspects of the present disclosure may be implemented simultaneously in any combination, and all aspects of the present disclosure may be implemented in combination as yet another embodiment of the present disclosure.
[0154] FIG. 4 shows an exemplary system 400 for implementing the processes and methods described above for performing the derivation of the merge candidate BCW index.
[0155] The techniques and mechanisms described herein can be implemented not only by multiple instances of system 400, but also by any other computing device, system, and / or environment. System 400 shown in FIG. 4 is merely an example of a system and is not intended to imply any limitation as to the scope of use or functionality of any computing device utilized to execute the processes and / or procedures described above. Other well-known computing devices, systems, environments, and / or configurations that may be suitable for use with the embodiments include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, game consoles, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, implementations using field programmable gate arrays (“FPGAs”) and application specific integrated circuits (“ASICs”).
[0156] System 400 may include one or more processors 402 and a system memory 404 communicatively coupled to the processors 402. The processors 402 can execute one or more modules and / or processes to cause the processors 402 to perform various functions. In some embodiments, the processors 402 may include a central processing unit (“CPU”), a graphics processing unit (“GPU”), both a CPU and a GPU, or other processing devices or components known in the art. Additionally, each of the processors 402 can have its own local memory that can also store program modules, program data, and / or one or more operating systems.
[0157] Depending on the exact configuration and type of system 400, system memory 404 may be volatile such as RAM, ROM, flash memory, non-volatile memory such as a small hard drive, memory card, or some combination thereof. System memory 404 can include one or more computer-executable modules 406 executable by processor 402.
[0158] Module 406 can include, but is not limited to, encoder 408 and decoder 410.
[0159] Encoder 408 can be a VVC standard encoder executable by processor 402 to implement any, some, or all aspects of the exemplary embodiments of the present disclosure described above and configure processor 402 to perform operations as described above.
[0160] Decoder 410 can be a VVC standard encoder executable by processor 402 to implement any, some, or all aspects of the exemplary embodiments of the present disclosure described above and configure processor 402 to perform operations as described above.
[0161] In some embodiments, a computer program product is provided, the program product includes computer program instructions, and the computer program instructions enable a computer to execute the steps of the method described in any of the embodiments of the present disclosure.
[0162] In some embodiments, a computer program is provided, and the computer program enables a computer to execute the steps of the method described in any of the embodiments of the present disclosure.
[0163] System 400 may further include an input / output (I / O) interface 440 for receiving video source data and bitstream data and outputting the decoded pictures to a reference picture buffer and / or a display buffer. System 400 may also include a communication module 450 that enables System 400 to communicate with other devices (not shown) via a network (not shown). The network may include wired media such as the Internet, a wired network, or a direct wired connection, as well as wireless media such as acoustic, radio frequency (“RF”), infrared, and other wireless media.
[0164] Some or all of the operations of the above-described method can be performed by the execution of computer-readable instructions stored on a computer-readable storage medium, as defined below. As used herein and in the claims, the term “computer-readable instructions” includes routines, applications, application modules, program modules, programs, components, data structures, algorithms, and the like. Computer-readable instructions can be implemented in various system configurations, including single-processor or multiprocessor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like.
[0165] The computer-readable storage medium can include volatile memory (such as random access memory (“RAM”)) and / or non-volatile memory (such as read-only memory (“ROM”), flash memory, etc.). The computer-readable storage medium can also include additional removable and / or non-removable storage that can provide non-volatile storage of computer-readable instructions, data structures, program modules, etc., including, but not limited to, flash memory, magnetic storage, optical storage, and / or tape storage.
[0166] A non-transitory or non-volatile computer-readable storage medium is an example of a computer-readable medium. Computer-readable media include at least two types of computer-readable media, namely computer-readable storage media and communication media. Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any process or technology for storing information such as computer-readable instructions, data structures, program modules, and other data. Computer-readable storage media include, but are not limited to, phase change memory ("PRAM"), static random access memory ("SRAM"), dynamic random access memory ("DRAM"), other types of random access memory ("RAM"), read-only memory ("ROM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory or other memory technologies, compact disc read-only memory ("CD-ROM"), digital versatile disc ("DVD") or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information for access by a computing device. In contrast, communication media may incorporate computer-readable instructions, data structures, program modules, or other data into a modulated data signal such as a carrier wave or other transmission mechanism. As used herein, a computer-readable storage medium is not to be construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal propagating through a wire.
[0167] Computer-readable instructions stored on one or more non-transitory or non-volatile computer-readable storage media, when executed by one or more processors, can perform the operations described above with reference to FIGS. 1A-3. In general, program-readable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular abstract data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0168] The subject matter is described in terms of language specific to structural features and / or methodological acts, but it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms for carrying out the claims.
Explanation of Signs
[0169] 100 Video encoding process 102 Image source 104 Intra prediction 106 Inter prediction 108 Encoder control settings 110 Subtractor 114 Quantization 116 Inverse quantization 118 Inverse transform 120 Adder 122 Loop filter 124 Entropy coder 150 Video decoding process 152 Entropy decoder 154 Inverse quantization 156 Inverse transform 158 Intra prediction 160 Motion compensation prediction 164 Loop filter 200 DPB 202 Bidirectional prediction 204 Reference Picture 206 Reference Picture 300 Method 400 System 402 Processor 404 System Memory 406 Module 408 Encoder 410 Decoder 450 Communication Module
Claims
1. One or more processors, A computer-readable storage medium communicatively coupled to the one or more processors, which, when executed by the one or more processors, Deriving a binary prediction (BCW) index using the CU-level weight of the binary prediction merge candidate for the binary prediction merge candidates in the merge candidate list of the merged coded coding unit (CU) according to the template matching cost A computer-readable storage medium storing computer-readable instructions executable by the one or more processors to perform related operations including A computing system comprising.
2. The computing system according to claim 1, wherein the template matching cost includes the sum of the absolute differences between the samples of the template of the merged coded CU and each corresponding reference sample.
3. Deriving the BCW index of the binary prediction merge candidate according to the template matching cost includes Calculating the template matching cost for each BCW weight value in a set of possible BCW weight values, and Selecting the BCW weight value that results in the lowest template matching cost among the template matching costs calculated as the BCW index of the binary prediction merge candidate The computing system according to claim 1.
4. The related operations further include P bi-pred =((16-w)*P 0 +w*P 1 +8)≫4 to calculate the weighted average bi-prediction according to the formula The computing system according to claim 3.
5. The computing system according to claim 4, wherein the set of possible BCW weight values includes {6, 7, 8, 9, 10}.
6. Deriving the BCW index of the binary prediction merge candidate according to the template matching cost includes Calculating the template matching cost for each BCW weight value in a subset of possible BCW weight values based on an inherited BCW weight, and Selecting the BCW weight value that results in the lowest template matching cost among the template matching costs calculated as the BCW index of the binary prediction merge candidate The computing system according to claim 1.
7. The computing system according to claim 6, wherein the template matching cost is calculated for each BCW weight value in a subset of possible BCW weight values based on the fact that the inherited BCW weight is a negative weight with respect to the weighted average dual prediction equation.
8. wherein the weighted average double prediction equation is P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3, the computing system according to claim 7.
9. Deriving the BCW index of the dual prediction merge candidate according to the template matching cost, Calculating a template matching cost for each BCW weight value in a subset of possible BCW weight values based on an inherited BCW weight, wherein the subset is further based on a respective picture order count (POC) distance between a current picture and a plurality of reference pictures, and selecting a BCW weight value that results in the lowest template matching cost among the template matching costs calculated as the BCW index of the dual prediction merge candidate The computing system according to claim 1, comprising:
10. The computing system according to claim 9, wherein each of the POC distances includes a first POC distance between the current picture and a reference picture in a first reference picture list and a second POC distance between the current picture and a reference picture in a second reference picture list.
11. Calculating a template matching cost for each BCW weight value in a subset of possible BCW weight values, when the first POC distance is less than the second POC distance, calculating a template matching cost for each BCW weight value equal to or less than a weight with respect to the weighted average dual prediction equation; and when the first POC distance is greater than the second POC distance, calculating a template matching cost for each BCW weight value equal to or greater than the weight The computing system according to claim 10, comprising:
12. The weighted average double prediction equation is P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3, the computing system according to claim 11.
13. Deriving the BCW index of the dual prediction merge candidate according to the template matching cost, calculating a template matching cost for each BCW weight value in a subset of possible BCW weight values similar to the inherited BCW weight, selecting a BCW weight value that results in the lowest template matching cost among each of the template matching costs calculated as the BCW index of the double prediction merge candidate The computing system according to claim 1, comprising the above.
14. calculating a template matching cost for each BCW weight value among a subset of possible BCW weight values, calculating each absolute difference between each BCW weight value and the inherited BCW weight, calculating a template matching cost for each BCW weight value whose respective absolute difference is less than a predefined threshold The computing system according to claim 13, comprising the above.
15. The related operation is adjusting so that the template matching cost of the inherited BCW weight is preferentially selected over the template matching costs of other BCW weights The computing system according to any one of claims 1 to 14, further comprising the above.
16. Adjusting the template matching cost of the inherited BCW weight includes multiplying the template matching cost of the inherited BCW weight by a weight less than 1. The computing system according to claim 15.
17. Adjusting the template matching cost of the inherited BCW weight includes applying bitwise shift operations and subtraction operations to scale the template matching cost of the inherited BCW weight by a factor less than 1. The computing system according to claim 15.
18. The related operation is adjusting so that the template matching cost of a BCW weight including equal weights for the weighted average double prediction equation is preferentially selected over the template matching costs of other BCW weights The computing system according to any one of claims 1 to 14, further comprising the above.
19. Adjusting the template matching cost of the BCW weight includes multiplying the template matching cost of the equal BCW weight by a weight less than 1. The computing system according to claim 18.
20. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium is For a dual-prediction merge candidate in the merge candidate list of a merged-coded CU, a step of deriving a BCW index of the dual-prediction merge candidate according to a template matching cost A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of the apparatus to cause the apparatus to initiate a method including the above step. **Claim 21** A computer program product comprising computer program instructions, where the computer program instructions For a dual-prediction merge candidate in the merge candidate list of a merged-coded CU, a step of deriving a BCW index of the dual-prediction merge candidate according to a template matching cost A computer program product that enables a computer to execute a method including the above step. **Claim 22** A computer program, where the computer program For a dual-prediction merge candidate in the merge candidate list of a merged-coded CU, a step of deriving a BCW index of the dual-prediction merge candidate according to a template matching cost A computer program that enables a computer to execute a method including the above step.