Decoder-side motion vector refinement with prediction sample offset
Patent Information
- Application Number
- PCT/CN2026/086134
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-10-01
- Filing Date
- 2026-03-26
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026086134_01102026_PF_FP_ABST
Abstract
Description
DECODER-SIDE MOTION VECTOR REFINEMENT WITH PREDICTION SAMPLE OFFSETCROSS REFERENCE TO RELATED PATENT APPLICATION (S)
[0001] The present disclosure is part of a non-provisional application that claims the priority benefit of U.S. Provisional Patent Application Nos. 63 / 777,795, 63 / 858,428, and 63 / 891,396, filed on 26 March 2025, 6 August 2025, and 1 October 2025, respectively. Contents of above-listed applications are herein incorporated by reference.TECHNICAL FIELD
[0002] The present disclosure relates generally to video coding. In particular, the present disclosure relates to methods of coding pixel blocks by decoder-side motion vector refinement (DMVR) .BACKGROUND
[0003] Unless otherwise indicated herein, approaches described in this section are not prior art to the claims listed below and are not admitted as prior art by inclusion in this section.
[0004] High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) . HEVC is based on the hybrid block-based motion-compensated DCT-like transform coding architecture. The basic unit for compression, termed coding unit (CU) , is a 2Nx2N square block of pixels, and each CU can be recursively split into four smaller CUs until the predefined minimum size is reached. Each CU contains one or multiple prediction units (PUs) .
[0005] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Expert Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11. The input video signal is predicted from the reconstructed signal, which is derived from the coded picture regions. The prediction residual signal is processed by a block transform. The transform coefficients are quantized and entropy coded together with other side information in the bitstream. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transform on the de-quantized transform coefficients. The reconstructed signal is further processed by in-loop filtering for removing coding artifacts. The decoded pictures are stored in the frame buffer for predicting the future pictures in the input video signal.
[0006] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . The leaf nodes of a coding tree correspond to the coding units (CUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs.
[0007] Each CU contains one or more prediction units (PUs) . The prediction unit, together with the associated CU syntax, works as a basic unit for signaling the predictor information. The specified prediction process is employed to predict the values of the associated pixel samples inside the PU.
[0008] For each inter-predicted CU, motion parameters consisting of motion vectors, reference picture indices and reference picture list usage index, and additional information are used for inter-predicted sample generation. The motion parameter can be signalled in an explicit or implicit manner. When a CU is coded with skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. A merge mode is specified whereby the motion parameters for the current CU are obtained from neighbouring CUs, including spatial and temporal candidates, and additional schedules introduced in VVC. The merge mode can be applied to any inter-predicted CU. The alternative to merge mode is the explicit transmission of motion parameters, where motion vector, corresponding reference picture index for each reference picture list and reference picture list usage flag and other needed information are signalled explicitly per each CU.SUMMARY
[0009] The following summary is illustrative only and is not intended to be limiting in any way. That is, the following summary is provided to introduce concepts, highlights, benefits and advantages of the novel and non-obvious techniques described herein. Select and not all implementations are further described below in the detailed description. Thus, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended for use in determining the scope of the claimed subject matter.
[0010] Some embodiments of the disclosure provide a method for using sample-based prediction offset (SPO) to enhance decoder side motion vector refinement (DMVR) when coding pixel blocks. A video coder receives data to be encoded or decoded as a current block of pixels of a current picture of a video. The video coder refines a motion vector based on a bilateral matching cost between first and second predictors that are derived based on first and second reference blocks that are located by the motion vector. Sample-based prediction offset (SPO) values are determined and applied to the first and second predictors. The SPO values may be derived from position-related weighting (PRW) , similarity-check weighting (SCW) and the difference between reconstructed template and reference template (DRR) . The video coder encodes or decodes the current block by using the refined motion vector to generate a final predictor. In some embodiments, the SPO values are derived by right-shifting the difference values (by 1 or 2 to divide by 2 or 4) to weaken the effect of the SPO on the bilateral matching process.
[0011] In some embodiments, the current block comprises multiple subblocks, such that the first and second predictors may be derived for a subblock of the current block and with the subblock reconstructed by using the refined motion vector. In some embodiments, the reconstructed samples of a first subblock neighboring a second subblock of the current block are used to derive SPO values for the second subblock. In some embodiments, the reconstructed samples of the template region neighboring the current block are shared by different subblocks for deriving the SPO values. In some embodiments, the first and second predictors may be for a subblock of the current block and the PRW scheme implements weight decay over the entire the current block.
[0012] In some embodiments, the SPO values are derived and applied to the first and second predictors only certain conditions are met. For example, in some embodiments, the SPO values are applied only if the first predictor is for a subblock of the current block. In some embodiments, the SPO values are applied only if a size of the current block meets a particular condition. In some embodiments, the SPO values are applied to the first predictor only if the motion vector is from a merge candidate list that meets a particular condition. In some embodiments, the SPO values are derived for only a sub-partition of the current block. In some embodiments, the SPO values are applied only if the sub-partition is a boundary partition of the current block. In some embodiments, the SPO values are applied to the first and second predictors only when a particular coding tool (e.g., local illumination compensation or LIC) is not applied to the current block.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of the present disclosure. The drawings illustrate implementations of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It is appreciable that the drawings are not necessarily in scale as some components may be shown to be out of proportion than the size in actual implementation in order to clearly illustrate the concept of the present disclosure.
[0014] FIG. 1 conceptually illustrates refinement of a prediction candidate by bilateral matching for decoder-side motion vector refinement (DMVR) .
[0015] FIG. 2 conceptually illustrates using sample-based prediction offset (SPO) to generate a predictor with offset compensation.
[0016] FIG. 3 conceptually illustrates DMVR refinement in which the L0 and L1 predictors are refined by SPO offset when computing the bilateral matching cost.
[0017] FIG. 4 illustrates a CU that is partitioned into sub-CUs A, B, C, and D during the DMVR process.
[0018] FIGS. 5A-5D illustrate reusing or sharing templates for determining SPO among different sub-CUs.
[0019] FIG. 6 illustrates an example video encoder that may implement SPO and DMVR.
[0020] FIG. 7 illustrates portions of the video encoder that implement SPO and DMVR.
[0021] FIG. 8 conceptually illustrates a video encoding process that encodes a block using DMVR with SPO.
[0022] FIG. 9 illustrates an example video decoder that may implement SPO and DMVR.
[0023] FIG. 10 illustrates portions of the video decoder that implement SPO and DMVR.
[0024] FIG. 11 conceptually illustrates a video decoding process that decodes a block using DMVR with SPO.
[0025] FIG. 12 conceptually illustrates an electronic system with which some embodiments of the present disclosure are implemented.DETAILED DESCRIPTION
[0026] In the following detailed description, numerous specific details are set forth by way of examples in order to provide a thorough understanding of the relevant teachings. Any variations, derivatives and / or extensions based on teachings described herein are within the protective scope of the present disclosure. In some instances, well-known methods, procedures, components, and / or circuitry pertaining to one or more example implementations disclosed herein may be described at a relatively high level without detail, in order to avoid unnecessarily obscuring aspects of teachings of the present disclosure. I. Decoder-side motion vector refinement (DMVR)
[0027] A. Bilateral Matching
[0028] In order to increase the accuracy of the MVs of the merge mode, a bilateral-matching (BM) based decoder side motion vector refinement (DMVR) may be applied to refine MVs. In bi-prediction operation, candidate MVs around the initial MVs in the reference picture list L0 and reference picture list L1 are examined. For each candidate MV, the BM method calculates the distortion between the two reference blocks in the reference picture list L0 and list L1 that are located by the candidate MV. The candidate MV with the lowest SAD becomes the refined MV and used to generate the bi-predicted signal.
[0029] FIG. 1 conceptually illustrates refinement of a prediction candidate (e.g., merge candidate) by bilateral matching for DMVR. MV0 is an initial motion vector and MV1 is the mirror of MV0. MV0 references an initial reference block 120 in L0 reference picture 110. MV1 references an initial reference block 121 in a L1 reference picture 111. The figure shows MV0 and MV1 being refined to form MV0’ and MV1’ , which locate updated reference blocks 130 and 131, respectively. The refinement is performed according to bilateral matching (BM) , such that the refined motion vector pair MV0’ and MV1’ has better bilateral matching cost than the initial motion vector pair MV0 and MV1. MV0’ -MV0 (i.e., MVD0) and MV1’ -MV1 (i.e., MVD1) are constrained to be equal in magnitude but opposite in direction. In some embodiments, the bilateral matching cost of a pair of mirrored motion vectors (e.g., MV0’ and MV1’ ) is calculated based on the difference (i.e., the distortion) between the two reference blocks (e.g., reference blocks 130 and 131) referred by the mirrored motion vectors.
[0030] B. Multi-pass DMVR
[0031] In some embodiments, a multi-pass decoder-side motion vector refinement (DMVR) method is applied in regular merge mode if the selected merge candidate meets the DMVR conditions. In the first pass, bilateral matching (BM) is applied to the coding block. In the second pass, BM is applied to each 16x16 subblock within the coding block. In the third pass, MV in each 8x8 subblock is refined by applying bi-directional optical flow (BDOF) .
[0032] C. Adaptive Multi-Pass DMVR
[0033] Adaptive decoder side motion vector refinement method includes two new merge modes to refine MV only in one direction, either L0 or L1, of the bi prediction for the merge candidates that meet the DMVR conditions. The multi-pass DMVR process is applied for the selected merge candidate to refine the motion vectors, however either MVD0 or MVD1 is set to zero in the first pass (i.e. PU level) DMVR.
[0034] Like the regular merge mode, merge candidates for the proposed merge modes are derived from the spatial neighboring coded blocks, TMVPs, non-adjacent blocks, HMVPs, and pair-wise candidate. The difference is that only those meet DMVR conditions are added into the candidate list. The same merge candidate list is used by the two new merge modes and merge index is coded as in regular merge mode. There are two syntax elements to indicate this mode, including bmMergeFlag and bmDirFlag. The syntax element bmMergeFlag is used to indicate the on-off of this kind of prediction (refine MV only on one direction) . The syntax element bmDirFlag is used to indicate the refined MV direction. For example, when bmDirFlag is equal to 0, the refined MV is from List0; when bmDirFlag is equal to 1, the refined MV is from List1. As shown in the following table.
[0035] After decoding bm_merge_flag and bm_dir_flag, bmDir can be decided. For example, if bm_merge_flag is equal to 1, bm_dir_flag is equal to 0, bmDir will be set as 1. And it is used to represent the adaptive MP-DMVR only refine the MV in List0. For another example, if bm_merge_flag is equal to 1, bm_dir_flag is equal to 1, bmDir will be set as 2. And it is used to represent the adaptive MP-DMVR only refine the MV in List1.
[0036] D. Scaling Template Differences for Bilateral Matching
[0037] In some embodiments, bilateral matching (BM) technology is used to find a deltaMV_P0 and a deltaMV_P1 to make P0 and P1 more similar. Absolute values of delta_P0 and deltaMV_P1 are the same. And signs of delta_P0 and delta_P1 are opposite. In some embodiments, to improve the performance of bilateral matching, not only the SAD / SSD of predictor P0 and predictor P1 are considered during deltaMV_PX derivation, but template differences information are also used (X: list0 or list1) .
[0038] In some embodiments, the sum of template differences of P0 and P1 is considered during bilateral matching. In that, if two candidates (two deltaMVs) introduce the same SAD of predictor P0 and P1, the candidate that results in lower sum of template differences of P0 and P1 will be given preference. For example, the distortion calculated for each candidate (deltaMV) includes SAD of predictor P0 and P1, and the sum of template differences of P0 and P1. The distortion function is shown as following.
[0039] For example, a ratio is used on the sum of template differences of P0 and P1 during distortion calculation. The ratio is determined based on CU size, prediction mode, MAD value, or QP. MAD is the variance of current block calculated by the following function:
[0040] The distortion function is shown as following.
[0041] In some embodiments, the sum of scaled template differences of P0 and P1 is considered during bilateral matching. In that, for each candidate of deltaMV, a distortion_final is determined as the summation of distortion of each sample in current block region. The distortion of each sample is not only the difference between predictor P0 and P1, but also the sum of corresponding scaled template differences.
[0042] For example, in some embodiments, the corresponding template differences are calculated by the top template sample (top_t) of P0 and P1, and the left template sample (left_t) of P0 and P1. For another example, in some embodiments, the corresponding template differences are calculated by top_t and its neighboring samples, and left_t and its neighboring samples. For example, in some embodiments, the selection of corresponding template differences is made based on the current block’s direction. For example, DIMD can be used to determine the direction of current block. After that, the corresponding samples can be selected to calculate template differences.
[0043] In some embodiments, the scaling factors used to generate scaled template differences of P0 and P1 are position dependent. That is, for generating scaled template differences, the scaling factors for the samples of the current CU closer to the CU boundary are larger than those for samples far away from CU boundary.
[0044] In some embodiments, a region of a CU which uses template differences to calculate distortion is constrained, and the samples outside a predefined region of a CU use only SAD / SSD to calculate distortion. For example, the predefined region may include samples in the first 1 / 4 rows closest to CU top boundary or left boundary. The predefined region may include only 64 rows, or 64 columns closet to CU top boundary or left boundary.
[0045] In some embodiments, if the scaled template differences are considered during bilateral matching, LIC cannot be applied.
[0046] In some embodiments, the scaled template differences are considered only on true-bi-prediction CUs. In that, one reference picture of current block is coded before current picture, and the other reference picture of current block is coded after current picture. In some embodiments, the scaled template differences are considered only on non-true-bi-prediction CUs. In that, two reference pictures of current block are coded before current picture, or coded after current picture. In some embodiments, the scaled template differences are considered only on CUs with equal distance reference pictures. In that, the POC distance of L0 and current picture is the same of POC distance of L1 and current picture.
[0047] In some embodiment, the scaled template differences are considered during bilateral matching only when motions in both predictors can be changed. That is, adaptive MP-DMVR cannot be applied this scheme. (Adaptive MP-DMVR only modifies one side of motion for a bi-prediction MV. ) In some embodiments, a region used to calculate bilateral matching cost can be larger than current block region. For example, the region used to calculate bilateral matching cost may include the current block as well as top N lines or left M lines neighboring samples. N and M can be any integer larger than 0.
[0048] In some embodiments, the distortions calculated by using samples outside current block region will be multiplied with scaling factors before adding with other distortions calculated by using samples inside current block region. The scaling factors are smaller than 1 and larger than 0 values.
[0049] In some embodiments, the scaling factors are determined according to the distance of samples and the CU boundary. Larger scaling factors (closer to 1 value) will be used for samples closer to CU boundary. In that, the impact of the samples closer to CU boundary will be higher than the samples far away from CU boundary.
[0050] In some embodiments, the above-mentioned technology can be applied to any bilateral matching application. For example, DMVR, adaptive DMVR, and MV refinement.
[0051] One skilled in the art would understand that any of the methods described above or their combination can be implemented individually or jointly in a video coding system.
[0052] E. DMVR with predictors refined by SPO
[0053] In some embodiments, a sample-based prediction offset (SPO) is used to refine the L0 predictor and the L1 predictor. In some embodiments, the SPO is derived from position-related weighting (PRW) , similarity-check weighting (SCW) , and the difference between reconstructed template and reference template (DRR) .
[0054] FIG. 2 conceptually illustrates using sample-based prediction offset (SPO) to generate a predictor with offset compensation. As illustrated, for a current block 210, a reference block 220 is used to generate predictor 250 for the current block. A neighboring template area of the reference block 220 is used as the reference template 225, and a reconstructed neighboring template area of the current block 210 is used as the reconstructed template 215. Difference values 235 (DiffTemp, including DiffAbove 231 and DiffLeft 232) between the reference template 225 and the reconstructed template 215 are used to generate estimated offset values 230 (OffsetSPO) over the region of the current block 210 based on DiffAbove 231 and DiffLeft 232. Specifically, SPO values at positions within the current block 210 are obtained by applying position-related weighting look-up tables 240 (LUTPRW) to DiffAbove 231 and DiffLeft 232 at positions in the above and left templates. In other words, SPO_Offseti, j= DiffAbove, i* LUTh, j + DiffLeft, j* LUTw, i where SPO_Offseti, j is the estimated offset value 230 at position (i, j) of the current block 210, DiffAbove, i is the difference values in a top template area 231 and Diffleft, j is the difference values in left template area 232. LUTh, j and LUTw, i are lookup tables for mapping values in top and left template areas into values at position (i, j) of the current block. The estimated offset values (or SPO offset) 230 are then applied to the reference block predictor 220 to generate a predictor with offset compensation 250 for the current block.
[0055] When generating the estimated offset values 230 using the difference values 231 and the difference values 232, LUTh, j and LUTw, i may implement position-related weighting (PRW) scheme, such that positions further away from the above and left templates are given less weight when generating the SPO offsets. Such a position-related weighting scheme for SPO over a region (e.g., a CU or a sub-CU) may be referred to as a SPO weight decay scheme over the region.
[0056] The template areas used for determining the SPO offsets (e.g., the reference template 225 and the reconstructed template 215) can be 1 line or multiple line. In some embodiments, the template areas may be refined by LIC firstly, and then used to derive the SPO offset. If the L0 predictor and the L1 predictor used to derive DMVR MV refinement is refined by LIC, the template may also be refined by LIC. After that, the refined template may be used to derive the SPO.
[0057] In some embodiments, the SPO process is used to refine predictor of L0 and predictor of L1, and the refined predictors are used to derive refined MV under DMVR by minimizing bilateral matching cost. FIG. 3 conceptually illustrates DMVR refinement in which the L0 and L1 predictors are refined by SPO offset when computing the bilateral matching cost.
[0058] As illustrated, DMVR is performed for the current block 100 in the current picture 101, with BM cost being determined for a candidate motion vector MV0” and its mirror MV1” . MV0” is referencing a reference block 140 in the L0 reference picture 110 and MV1” is referencing a reference block 141 in the L1 reference picture 111.
[0059] To compute the BM cost for the mirrored MV pair MV0” and MV1” , the SPO offsets are determined and applied to their respective L0 and L1 predictors 140 and 141. Specifically, the SPO offsets for the L0 predictor 140 is computed using the reference template 340 neighboring the L0 reference block 140, and the SPO offsets for the L1 predictor 141 is computed using the reference template 341 neighboring the L1 reference block 141 (using the SPO process described by reference to FIG. 2 above) . The computed SPO offsets are then respectively applied to the L0 and L1 predictors to generate L0 and L1 SPO compensated predictors 350 and 351. The difference between the L0 and L1 SPO compensated predictors 350 and 351 is then used as the BM cost 390 of the mirrored pair MV0” and MV1” under a DMVR process.
[0060] The process of SPO derivation for DMVR may be aligned with (same or similar to) the process of SPO derivation after motion compensation (i.e., for producing the final predictor of the current block. ) The parameters used for SPO derivation for DMVR may be aligned with the parameters used for SPO derivation after motion compensation. in some embodiments, the parameters used for SPO derivation for DMVR can be derived by adding an offset to the parameters used for SPO derivation after motion compensation. In some embodiments, the process of SPO derivation for DMVR is aligned with the process of SPO derivation after motion compensation. In some embodiments, the parameters used for SPO derivation for DMVR is aligned with the parameters used for SPO derivation for skip mode predictor. In some embodiments, the parameters used for SPO derivation for DMVR is aligned with the parameters used for SPO derivation for non-skip mode predictor.
[0061] In some embodiments, before applying derived SPO offset values to the L0 and L1 predictors in DMVR, the derived SPO offset may be modified by a right shift technology (i.e., right shift by 1 or right shift by 2) to weaken the effect of SPO offsets on the final predictor.
[0062] In some embodiments, the parameters used for SPO derivation for DMVR is adaptively selected by the merge index of the current merge candidate. For example, if merge index modulo 2 is equal to 0, the first parameter set is used to derive SPO; and if merge index modulo 2 is equal to 1, the second parameter set is used to derive SPO. In some embodiments, a SPO is used to refine predictors of L0 and L1 for DMVR is conditional applied. In some embodiments, the SPO is only used for block-based DMVR to refine L0 and L1 predictors. In some embodiments, the SPO is only used for subblock-based DMVR to refine L0 and L1 predictors. In some embodiments, the SPO is only used for DMVR on subblock modes to refine L0 and L1 predictors. In some embodiments, the SPO is only used for DMVR on non-subblock modes to refine L0 and L1 predictors. In some embodiments, SPO for DMVR is only enabled when LIC is disabled. In some embodiments, the SPO is only used for the first N candidates in merge list on DMVR to refine L0 and L1 predictor. In some embodiments, the SPO for DMVR is only used for CU larger than M and smaller than N (i.e., M is equal to 8x8, N is equal to 128x128. )
[0063] In some embodiments, the SPO is only used for the first group of candidates in merge list on DMVR to refine L0 and L1 predictors. That is, the first group of candidates contains merge candidates with index modulo 2 equal to zero. In some embodiments, the SPO is only used for the first group of candidates in merge list on DMVR to refine L0 and L1 predictors. That is, the first group of candidates contains merge candidates with index modulo 2 equal to one.
[0064] In some embodiments, if the SPO for DMVR is only applied on the best N candidates during searching process, then the SPO is used on DMVR to find the final best candidate by minimizing predictor L0 and predictor L1 differences.
[0065] In some embodiments, the SPO for DMVR is applied on true bi-prediction CU only. That is, one reference picture is coded before the current picture, and the other reference picture is coded after the current picture. In some embodiments, the SPO for DMVR is applied on block with equal distance reference pictures only. In some embodiments, the SPO for DMVR will be disabled if the difference between L0 and L1 predictors is smaller than a predefined threshold. In some embodiments, the SPO for DMVR will be disabled if the difference between L0 and L1 predictors is larger than a predefined threshold. In the above-mentioned predefined threshold, it can be designed based on current CU size, the size of DMVR processing unit or quantization parameters.
[0066] A CU may be partitioned into sub-CUs during DMVR. FIG. 4 illustrates a CU 400 that is partitioned into sub-CUs A, B, C, and D during the DMVR process. In some embodiments, DMVR with SPO (as described by reference to FIG. 3 above) can be applied on a sub-CU only if the sub-CU is on the CU boundary. For example, if a 32x32 CU is divided into 4 16x16 sub-CU, DMVR may be performed for each 16x16 sub-CU independently, but only the sub-CU in the first sub-CU row (e.g. sub-CUs ‘A’ and ‘B’ ) or the first sub-CU column (e.g., sub-CUs ‘A’ and ‘C’ ) would have DMVR performed with SPO. For another example, if a 32x32 CU is divided into 4 16x16 sub-CU, each 16x16 sub-CU will have DMVR performed independently, but only the sub-CU in top-left position (e.g., sub-CU ‘A’ ) may have DMVR performed with SPO.
[0067] In some embodiments, if a CU is partitioned into sub-CU to apply DMVR, only if a sub-CU is near the CU boundary, the SPO for DMVR can be applied. For example, the sub-CU in the first N sub-CU row or the first M sub-CU column may have SPO with DMVR performed. In some embodiments, the SPO for DMVR is only applied on the first N rows or the first M columns closet to CU boundary.
[0068] In some embodiments, SPO may be applied to each sub-CU individually during DMVR, and the templates used for SPO of each sub-CU can be shared. For example, the very top template may be used / shared for sub-CUs in the same sub-CU column, or the very left template may be used / shared for sub-CUs in the same sub-CU row. FIGS. 5A-5D illustrate reusing or sharing templates for determining SPO among different sub-CUs. In the figure, DMVR with SPO is applied for each sub-CU of the current block 400.
[0069] FIG. 5A shows that for SPO during DMVR of sub-CU A, template areas 410 and 430 neighboring the current block 400 are used as reconstructed templates and template areas 510A and 530A neighboring a reference block 500A are used as reference templates. FIG. 5B shows that for SPO during DMVR of sub-CU B, template areas 420 and 430 neighboring the current block 400 are used as reconstructed templates and template areas 520B and 530B neighboring a reference block 500B are used as reference templates. FIG. 5C shows that for SPO during DMVR of sub-CU C, template areas 410 and 440 neighboring the current block 400 are used as reconstructed templates and template areas 510C and 540C neighboring a reference block 500C are used as reference templates. FIG. 5D shows that for SPO during DMVR of sub-CU D, template areas 420 and 440 neighboring the current block 400 are used as reconstructed templates and template areas 520D and 540D neighboring a reference block 500D are used as reference templates. Thus, for SPO during DMVR of the sub-CUs, the reconstructed template 410 above sub-CU A is shared with sub-CU C, the reconstructed template 420 above sub-CU B is shared with sub-CU D, the reconstructed template 430 left of sub-CU A is shared with sub-CU B, and the reconstructed template 440 left of sub-CU C is shared with sub-CU D.
[0070] In some embodiments, if a CU is partitioned into sub-CU to apply DMVR, the SPO of DMVR for not first sub-CU row (second sub-CU row and below) is only derived from left template, and the SPO of DMVR for not first sub-CU column (second sub-CU column and after) is only derived from top template.
[0071] In some embodiments, a CU-level weight decay scheme of SPO (or PRW scheme over the CU for SPO) is used for DMVR, regardless of whether the CU is partitioned into N sub-CUs for performing DMVR or not. Thus, if a 32x32 CU uses a 32x32 SPO weight decay scheme for DMVR, and even if the 32x32 CU is partitioned into 4 16x16 sub-CUs for performing DMVR, the 32x32 weight decay scheme may be used for all 4 sub-CUs. For example, the top-left 16x16 may use the top-left region of the 32x32 weight decay scheme, the top-right 16x16 may use the top-right region of the 32x32 weight decay scheme, the bottom-left 16x16 may use the bottom-left region of the 32x32 weight decay scheme, the bottom-right 16x16 may use the bottom-right region of the 32x32 weight decay scheme.
[0072] In some embodiments, the SPO on DMVR is only applied on merge candidate with odd or even candidate number. In some embodiments, the SPO on DMVR is only applied on the first N merge candidate with odd or even candidate number. In some embodiments, the SPO on DMVR is derived based on the reference template of current block. That is, two different motion shifts of current motion vector can use the same reference template to derive SPO. In some embodiments, the SPO on DMVR is only applied when Hadamard transform is enabled. In some embodiments, the SPO on DMVR is only applied when Hadamard transform isn’ t enabled. In some embodiments, the SPO on DMVR is only applied when BCW is disabled.
[0073] In some embodiments, when a CU is partitioned into more than one sub-CU to perform DMVR, the template used for SPO derivation can be generated by using the refined sub-CU motion. For example, a 32x32 CU is partitioned into 4 16x16 sub-CUs to perform DMVR. After the top-left sub-CU motion is refined by bilateral matching, the very right line of the corresponding predictor of top-left sub-CU may be used to derive SPO for top-right sub-CU, and the very bottom line of the corresponding of predictor of top-left sub-CU may be used to derive SPO for bottom-left sub-CU.
[0074] In some embodiments, if M reordering stage is performed before determining the final candidate list, the SPO on DMVR may only be performed after N reordering stage. In that, a merge index is signaled to indicate a merge candidate from the final candidate list to predict current CU. N may be an integer smaller than or equal to M.
[0075] In some embodiments, to mimic the behavior of SPO after motion compensation stage, if DMVR is performed with M sub-CU (M is an integer larger than or equal to 1) , the template used for some sub-CUs to derive SPO can be shared. For example, as shown in FIG. 4, the sub-CUs within the same sub-CU row (sub-CU A and B) may use the same left template to derive SPO. In that, the left template of best bilateral matching results of the very left sub-CU (sub-CU A) will be shared for each sub-CU in the same sub-CU row (sub-CU A and B) to derive SPO. For another example, the sub-CUs within the same sub-CU column (sub-CU A and C) may use the same top template to derive SPO. In that, the top template of best bilateral matching results of the very top sub-CU (sub-CU A) may be shared for each sub-CU in the same sub-CU column (sub-CU A and C) to derive SPO.
[0076] In some embodiments, considering the motion differences between two sub-CUs may be large after bilateral matching, the predictors of a sub-CU after bilateral matching may be used to derive SPO for other sub-CUs. In the example of FIG. 4, after the motion of sub-CU A is refined by bilateral matching, the very right line of the predictor of sub-CU A may be used to derive SPO for sub-CU B. After the motion of sub-CU A is refined by bilateral matching, the very bottom line of predictor of sub-CU A can be used to derive SPO for sub-CU C.
[0077] In some embodiments, a SPO is used to refine predictor of L0 and / or predictor of L1. After that, the refined predictors is used to derive DMVR MV refinement by minimizing bilateral matching cost. For example, in some embodiments, for neighboring search points of DMVR, the template used for SPO derivation can be shared. In some embodiments, if two search points distance is within 2 integer samples, the same template is used for two search points to perform SPO derivation. In some embodiments, search points in one region will use the same template to perform SPO derivation.
[0078] In some embodiments, the enabling of proposed method is explicitly indicated by one or more flags. The flags are signaled / parsed at CTU / slice / tile / sub-picture / picture / sequence level, or at APS / PPS / VPS / SPS level.
[0079] In some embodiments, the proposed method in the above is enabled or disabled, according to one or the combination of the selected reference pictures indices, temporal distance between reference picture and current picture, quantization parameter, the coded information of current CU, prediction mode, motion vectors, motion vector resolution, residual of current CU, and reference samples.
[0080] Any of the foregoing proposed methods or combination thereof can be implemented in encoders and / or decoders. For example, any of the proposed methods or combination thereof can be implemented in a inter module of a encoder and / or decoder. Alternatively, any of the proposed methods or combination thereof can be implemented as a circuit coupled to a inter module of the encoders and / or decoder, so as to provide the information needed by the inter module used in encoders and / or decoder. The proposed aspects, methods, related embodiments, and combination thereof can be implemented individually or jointly in a video coding system. II. Example Video Encoder
[0081] FIG. 6 illustrates an example video encoder 600 that may implement SPO and DMVR. As illustrated, the video encoder 600 receives input video signal from a video source 605 and encodes the signal into bitstream 695. The video encoder 600 has several components or modules for encoding the signal from the video source 605, at least including some components selected from a transform module 610, a quantization module 611, an inverse quantization module 614, an inverse transform module 615, an intra estimation module 624, an intra prediction module 625, a motion compensation module 630, a motion estimation module 635, an in-loop filter 645, a reconstructed picture buffer 650, a MV buffer 665, and a MV prediction module 675, and an entropy encoder 690. The motion compensation module 630 and the motion estimation module 635 are part of an inter prediction module 640. The intra prediction module 625 and the intra estimation module 624 are part of a current picture prediction module 620, which uses current picture reconstructed samples as reference samples for prediction of the current block.
[0082] In some embodiments, the modules 610 –690 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device or electronic apparatus. In some embodiments, the modules 610 –690 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Though the modules 610 –690 are illustrated as being separate modules, some of the modules can be combined into a single module.
[0083] The video source 605 provides a raw video signal that presents pixel data of each video frame without compression. A subtractor 608 computes the difference between the raw pixel data 602 provided by the video source 605 and the predicted pixel data 613 from the inter prediction module 640 or the current picture prediction module 620 as prediction residual 609. The transform module 610 converts the difference (or the residual pixel data or residual signal 609) into transform coefficients 616 (e.g., by performing Discrete Cosine Transform or DCT, or Discrete Sine Transform or DST) . The quantization module 611 quantizes the transform coefficients 616 into quantized data (or quantized coefficients) 612, which is encoded into the bitstream 695 by the entropy encoder 690.
[0084] The inverse quantization module 614 de-quantizes the quantized data (or quantized coefficients) 612 to obtain transform coefficients 618, and the inverse transform module 615 performs inverse transform on the transform coefficients 618 to produce reconstructed residual 619. The reconstructed residual 619 is added with the predicted pixel data 613 to produce reconstructed pixel data 617. In some embodiments, the reconstructed pixel data 617 is temporarily stored in a line buffer 627 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction. The reconstructed pixels are filtered by the in-loop filter 645 and stored in the reconstructed picture buffer 650. In some embodiments, the reconstructed picture buffer 650 is a storage external to the video encoder 600. In some embodiments, the reconstructed picture buffer 650 is a storage internal to the video encoder 600.
[0085] The intra estimation module 624 derives intra prediction data (e.g., intra prediction modes) based on the reconstructed pixel data 617 (stored in the line buffer 627) . The intra prediction data is provided to the entropy encoder 690 to be encoded into bitstream 695. The intra prediction data is also used by the intra prediction module 625 to produce the predicted pixel data 613.
[0086] The motion estimation module 635 performs inter prediction by producing MVs to reference pixel data of previously decoded frames stored in the reconstructed picture buffer 650. These MVs are provided to the motion compensation module 630 to produce predicted pixel data.
[0087] Instead of encoding the complete actual MVs in the bitstream, the video encoder 600 uses MV prediction to generate predicted MVs, and the difference between the MVs used for motion compensation and the predicted MVs is encoded as residual motion data and stored in the bitstream 695.
[0088] The MV prediction module 675 generates the predicted MVs based on reference MVs that were generated for encoding previously video frames, i.e., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 675 retrieves reference MVs from previous video frames from the MV buffer 665. The video encoder 600 stores the MVs generated for the current video frame in the MV buffer 665 as reference MVs for generating predicted MVs.
[0089] The MV prediction module 675 uses the reference MVs to create the predicted MVs. The predicted MVs can be computed by spatial MV prediction or temporal MV prediction. The difference between the predicted MVs and the motion compensation MVs (MC MVs) of the current frame (residual motion data) are encoded into the bitstream 695 by the entropy encoder 690.
[0090] The entropy encoder 690 encodes various parameters and data into the bitstream 695 by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding. The entropy encoder 690 encodes various header elements, flags, along with the quantized transform coefficients 612, and the residual motion data as syntax elements into the bitstream 695. The bitstream 695 is in turn stored in a storage device or transmitted to a decoder over a communications medium such as a network.
[0091] The in-loop filter 645 performs filtering or smoothing operations on the reconstructed pixel data 617 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 645 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.
[0092] FIG. 7 illustrates portions of the video encoder 600 that implement SPO and DMVR. Specifically, the figure illustrates the components of the inter prediction module 640 of the video encoder 600. As illustrated, a DMVR module 700 receives an initial motion vector 705 that may be a prediction candidate selected from a merge candidate list 765 constructed with candidates derived from previously used motion vectors. The selection of the merge candidate is made by the motion estimation module 635. The selection (e.g., a merge index) is also provided to the entropy encoder 690 to be signaled in the bitstream 695.
[0093] The DMVR module 700 generates a refined motion vector 795 by refining the initial motion vector 705. A MV refinement module 720 starts the bilateral matching process by using the initial motion vector 705 as the MV being refined 750. The MV being refined is provided to a bilateral matching module 710, which calculates a bilateral matching cost (BM cost) 760 of the motion vector being refined 750. The motion vector 750 may then be further refined by the MV refinement module 720 based on the calculated BM cost 760, so and so forth, until the motion vector 750 is sufficiently refined to have an acceptable BM cost (e.g., within a certain threshold. ) This finalized refined motion vector 795 is then provided to the motion compensation module 630 to produce a final predictor 790 of the current block as the predicted pixel data 613.
[0094] The bilateral matching module 710 performs bilateral matching based on the motion vector being refined 750, as described by reference to FIG. 1 above. In some embodiments, the bilateral matching module 710 may determine and apply SPO offset values to the L0 and L1 predictors when computing the bilateral matching cost, as described in Section I. E by reference to FIGS. 2 and 3 above. The SPO offset values are generated based on content of the reconstructed picture buffer 650, which provides the samples of the reference and reconstructed templates used for generating the SPO offset values. In some embodiments, the parameters used for deriving the SPO are aligned with parameters used for SPO derivation after motion compensation for use on the final predictor, or are aligned with the parameters used for SPO derivation for skip mode predictor.
[0095] The operations of the bilateral matching module 710 are further described in Section I.E above. For example, in some embodiments, the bilateral matching module applies SPO offset values to the L0 and L1 predictors only when certain conditions are met. In some embodiments, DMVR may be performed on subblock basis, i.e., the motion vector is refined for each individual subblock. In some of these embodiments, the templates used for determining the SPO offset values may be shared by multiple different subblocks. In some embodiments, the reconstructed pixel samples of one subblock may be used as templates for generating the SPO offset values of a neighboring subblock. In some embodiments, the PRW scheme implements weight decay over the entire the current block may be used even if DMVR is performed on a subblock basis (i.e., the L0 and L1 predictors used for bilateral matching upon which SPO is applied are for a subblock of the current block. ) In some embodiments, the SPO offset values are right-shifted before being applied to the L0 and L1 predictors to weaken the effect of SPO offsets on the final predictor 790.
[0096] FIG. 8 conceptually illustrates a video encoding process 800 that encodes a block using DMVR with SPO. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the encoder 600 performs the process 800 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the encoder 600 performs the process 800.
[0097] The encoder receives (at block 810) data to be encoded as a current block of pixels in a current picture. The encoder receives (at block 820) a motion vector to be refined. The motion vector may be provided by a merge candidate selected by the encoder.
[0098] The encoder derives (at block 830) first and second predictors based on L0 and L1 reference blocks located by the motion vector. The L0 and L1 reference blocks are located by the motion vector and its mirror vector in L0 and L1 reference pictures.
[0099] The encoder modifies (at block 840) the first and second predictors by deriving and applying SPO values. In some embodiments, the parameters used for deriving the SPO are aligned with parameters used for SPO derivation after motion compensation for use on the final predictor, or are aligned with the parameters used for SPO derivation for skip mode predictor. In some embodiments, the SPO values are derived by right-shifting the difference values (by 1 or 2 to divide by 2 or 4) to weaken the effect of the SPO on the bilateral matching process.
[0100] In some embodiments, the SPO values are derived from position-related weighting (PRW) , similarity-check weighting (SCW) and the difference between reconstructed template and reference template (DRR) , which are difference values between reference sample values in a template region neighboring the first reference block and reconstructed sample values in a template region neighboring the current block. In some embodiments, the first and second predictors are for a subblock of the current block even when the PRW scheme implements weight decay over the entire the current block.
[0101] In some embodiments, the current block comprises multiple subblocks, such that the first and second predictors may be derived for a subblock of the current block and with the subblock reconstructed by using the refined motion vector. In some embodiments, the reconstructed samples of a first subblock neighboring a second subblock of the current block are used to derive SPO values for the second subblock. In some embodiments, the reconstructed samples of the template region neighboring the current block are shared by different subblocks for deriving the SPO values as described by reference to FIGS. 5A-5D above.
[0102] In some embodiments, the SPO values are derived and applied to the first and second predictors only certain conditions are met. For example, in some embodiments, the SPO values are applied only if the first predictor is for a subblock of the current block. In some embodiments, the SPO values are applied only if a size of the current block meets a particular condition. In some embodiments, the SPO values are applied to the first predictor only if the motion vector is from a merge candidate list that meets a particular condition. In some embodiments, the SPO values are derived for only a sub-partition of the current block. In some embodiments, the SPO values are applied only if the sub-partition is a boundary partition of the current block. In some embodiments, the SPO values are applied to the first and second predictors only when a particular coding tool (e.g., local illumination compensation or LIC) is not applied to the current block.
[0103] The encoder computes (at block 850) bilateral matching cost based on the first and second predictors. The bilateral matching cost is determined based on a distortion between the first and second predictors.
[0104] The encoder determines (at block 860) whether the motion vector is sufficiently refined (e.g., whether the BM cost is less than a threshold. ) If the motion vector is sufficiently refined, the process proceeds to block 870. If the motion vector is not sufficiently refined, the process proceeds to block 865 to refine the motion vector and returns to block 830 to re-derive the first and second predictors based on the refined motion vector.
[0105] The encoder encodes (at block 870) the current block by using the refined motion vector to generate a final predictor. III. Example Video Decoder
[0106] In some embodiments, an encoder may signal (or generate) one or more syntax element in a bitstream, such that a decoder may parse said one or more syntax element from the bitstream.
[0107] FIG. 9 illustrates an example video decoder 900 that may implement SPO and DMVR. As illustrated, the video decoder 900 is an image-decoding or video-decoding circuit that receives a bitstream 995 and decodes the content of the bitstream into pixel data of video frames for display. The video decoder 900 has several components or modules for decoding the bitstream 995, including some components selected from an inverse quantization module 914, an inverse transform module 915, an intra prediction module 925, a motion compensation module 930, an in-loop filter 945, a decoded picture buffer 950, a MV buffer 965, a MV prediction module 975, and a parser 990. The motion compensation module 930 is part of an inter prediction module 940. The intra prediction module 925 is part of a current picture prediction module 920, which uses current picture reconstructed samples as reference samples for prediction of the current block.
[0108] In some embodiments, the modules 914 –990 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device. In some embodiments, the modules 914 –990 are modules of hardware circuits implemented by one or more ICs of an electronic apparatus. Though the modules 914 –990 are illustrated as being separate modules, some of the modules can be combined into a single module.
[0109] The parser 990 (or entropy decoder) receives the bitstream 995 and performs initial parsing according to the syntax defined by a video-coding or image-coding standard. The parsed syntax element includes various header elements, flags, as well as quantized data (or quantized coefficients) 912. The parser 990 parses out the various syntax elements by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding.
[0110] The inverse quantization module 914 de-quantizes the quantized data (or quantized coefficients) 912 to obtain transform coefficients, and the inverse transform module 915 performs inverse transform on the transform coefficients 918 to produce reconstructed residual signal 919. The reconstructed residual signal 919 is added with predicted pixel data 913 from the intra prediction module 925 or the motion compensation module 930 to produce decoded pixel data 917. The decoded pixels data are filtered by the in-loop filter 945 and stored in the decoded picture buffer 950. In some embodiments, the decoded picture buffer 950 is a storage external to the video decoder 900. In some embodiments, the decoded picture buffer 950 is a storage internal to the video decoder 900.
[0111] The intra prediction module 925 receives intra prediction data from bitstream 995 and according to which, produces the predicted pixel data 913 from the decoded pixel data 917 stored in the decoded picture buffer 950. In some embodiments, the decoded pixel data 917 is also stored in a line buffer 927 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction.
[0112] In some embodiments, the content of the decoded picture buffer 950 is used for display. A display device 905 either retrieves the content of the decoded picture buffer 950 for display directly, or retrieves the content of the decoded picture buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 950 through a pixel transport.
[0113] The motion compensation module 930 produces predicted pixel data 913 from the decoded pixel data 917 stored in the decoded picture buffer 950 according to motion compensation MVs (MC MVs) . These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 995 with predicted MVs received from the MV prediction module 975.
[0114] The MV prediction module 975 generates the predicted MVs based on reference MVs that were generated for decoding previous video frames, e.g., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 975 retrieves the reference MVs of previous video frames from the MV buffer 965. The video decoder 900 stores the motion compensation MVs generated for decoding the current video frame in the MV buffer 965 as reference MVs for producing predicted MVs.
[0115] The in-loop filter 945 performs filtering or smoothing operations on the decoded pixel data 917 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 945 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.
[0116] FIG. 10 illustrates portions of the video decoder 900 that implement SPO and DMVR. Specifically, the figure illustrates the components of the inter prediction module 940 of the video decoder 900. As illustrated, a DMVR module 1000 receives an initial motion vector 1005 that may be a prediction candidate selected from a merge candidate list 1065 that is constructed with candidates derived from previously used motion vectors. The selection may be signaled in the bitstream 995 (as a merge index) and parsed out by the entropy decoder 990.
[0117] The DMVR module 1000 generates a refined motion vector 1095 by refining the initial motion vector 1005. A MV refinement module 1020 starts the bilateral matching process by using the initial motion vector 1005 as the MV being refined 1050. The MV being refined is provided to a bilateral matching module 1010, which calculates a bilateral matching cost (BM cost) 1060 of the motion vector being refined 1050. The motion vector 1050 may then be further refined by the MV refinement module 1020 based on the calculated BM cost 1060, so and so forth, until the motion vector 1050 is sufficiently refined to have an acceptable BM cost (e.g., within a certain threshold. ) This finalized refined motion vector 1095 is then provided to the motion compensation module 930 to produce a final predictor 1090 of the current block as the predicted pixel data 913.
[0118] The bilateral matching module 1010 performs bilateral matching based on the motion vector being refined 1050, as described by reference to FIG. 1 above. In some embodiments, the bilateral matching module 1010 may determine and apply SPO offset values to the L0 and L1 predictors when computing the bilateral matching cost, as described in Section I. E by reference to FIGS. 2 and 3 above. The SPO offset values are generated based on content of the decoded picture buffer 950, which provides the samples of the reference and reconstructed templates used for generating the SPO offset values. In some embodiments, the parameters used for deriving the SPO are aligned with parameters used for SPO derivation after motion compensation for use on the final predictor, or are aligned with the parameters used for SPO derivation for skip mode predictor.
[0119] The operations of the bilateral matching module 1010 are further described in Section I.E above. For example, in some embodiments, the bilateral matching module applies SPO offset values to the L0 and L1 predictors only when certain conditions are met. In some embodiments, DMVR may be performed on subblock basis, i.e., the motion vector is refined for each individual subblock. In some of these embodiments, the templates used for determining the SPO offset values may be shared by multiple different subblocks. In some embodiments, the reconstructed pixel samples of one subblock may be used as templates for generating the SPO offset values of a neighboring subblock. In some embodiments, the PRW scheme implements weight decay over the entire the current block may be used even if DMVR is performed on a subblock basis (i.e., the L0 and L1 predictors used for bilateral matching upon which SPO is applied are for a subblock of the current block. ) In some embodiments, the SPO offset values are right-shifted before being applied to the L0 and L1 predictors to weaken the effect of SPO offsets on the final predictor 1090.
[0120] FIG. 11 conceptually illustrates a video decoding process 1100 that decodes a block using DMVR with SPO. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the decoder 900 performs the process 1100 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the decoder 900 performs the process 1100.
[0121] The decoder receives (at block 1110) data to be decoded as a current block of pixels in a current picture. The decoder receives (at block 1120) a motion vector to be refined. The motion vector may be provided by a merge candidate selected by the decoder.
[0122] The decoder derives (at block 1130) first and second predictors based on L0 and L1 reference blocks located by the motion vector. The L0 and L1 reference blocks are located by the motion vector and its mirror vector in L0 and L1 reference pictures.
[0123] The decoder modifies (at block 1140) the first and second predictors by deriving and applying SPO values. In some embodiments, the parameters used for deriving the SPO are aligned with parameters used for SPO derivation after motion compensation for use on the final predictor, or are aligned with the parameters used for SPO derivation for skip mode predictor. In some embodiments, the SPO values are derived by right-shifting the difference values (by 1 or 2 to divide by 2 or 4) to weaken the effect of the SPO on the bilateral matching process.
[0124] In some embodiments, the SPO values are derived from position-related weighting (PRW) , similarity-check weighting (SCW) and the difference between reconstructed template and reference template (DRR) , which are difference values between reference sample values in a template region neighboring the first reference block and reconstructed sample values in a template region neighboring the current block. In some embodiments, the first and second predictors are for a subblock of the current block even when the PRW scheme implements weight decay over the entire the current block.
[0125] In some embodiments, the current block comprises multiple subblocks, such that the first and second predictors may be derived for a subblock of the current block and with the subblock reconstructed by using the refined motion vector. In some embodiments, the reconstructed samples of a first subblock neighboring a second subblock of the current block are used to derive SPO values for the second subblock. In some embodiments, the reconstructed samples of the template region neighboring the current block are shared by different subblocks for deriving the SPO values, as described by reference to FIGS. 5A-5D above.
[0126] In some embodiments, the SPO values are derived and applied to the first and second predictors only certain conditions are met. For example, in some embodiments, the SPO values are applied only if the first predictor is for a subblock of the current block. In some embodiments, the SPO values are applied only if a size of the current block meets a particular condition. In some embodiments, the SPO values are applied to the first predictor only if the motion vector is from a merge candidate list that meets a particular condition. In some embodiments, the SPO values are derived for only a sub-partition of the current block. In some embodiments, the SPO values are applied only if the sub-partition is a boundary partition of the current block. In some embodiments, the SPO values are applied to the first and second predictors only when a particular coding tool (e.g., local illumination compensation or LIC) is not applied to the current block.
[0127] The decoder computes (at block 1150) bilateral matching cost based on the first and second predictors. The bilateral matching cost is determined based on a distortion between the first and second predictors.
[0128] The decoder determines (at block 1160) whether the motion vector is sufficiently refined (e.g., whether the BM cost is less than a threshold. ) If the motion vector is sufficiently refined, the process proceeds to block 1170. If the motion vector is not sufficiently refined, the process proceeds to block 1165 to refine the motion vector and returns to block 1130 to re-derive the first and second predictors based on the refined motion vector.
[0129] The decoder reconstructs (at block 1170) the current block by using the refined motion vector to generate a final predictor. The decoder may then provide the reconstructed current block for display or output as part of the reconstructed current picture. IV. Example Electronic System
[0130] Many of the above-described features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer readable storage medium (also referred to as computer readable medium) . When these instructions are executed by one or more computational or processing unit (s) (e.g., one or more processors, cores of processors, or other processing units) , they cause the processing unit (s) to perform the actions indicated in the instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, random-access memory (RAM) chips, hard drives, erasable programmable read only memories (EPROMs) , electrically erasable programmable read-only memories (EEPROMs) , etc. The computer readable media does not include carrier waves and electronic signals passing wirelessly or over wired connections.
[0131] In this specification, the term “software” is meant to include firmware residing in read-only memory or applications stored in magnetic storage which can be read into memory for processing by a processor. Also, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while remaining distinct software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement a software invention described here is within the scope of the present disclosure. In some embodiments, the software programs, when installed to operate on one or more electronic systems, define one or more specific machine implementations that execute and perform the operations of the software programs.
[0132] FIG. 12 conceptually illustrates an electronic system 1200 with which some embodiments of the present disclosure are implemented. The electronic system 1200 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc. ) , phone, PDA, or any other sort of electronic device. Such an electronic system includes various types of computer readable media and interfaces for various other types of computer readable media. Electronic system 1200 includes a bus 1205, processing unit (s) 1210, a graphics-processing unit (GPU) 1215, a system memory 1220, a network 1225, a read-only memory 1230, a permanent storage device 1235, input devices 1240, and output devices 1245.
[0133] The bus 1205 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of the electronic system 1200. For instance, the bus 1205 communicatively connects the processing unit (s) 1210 with the GPU 1215, the read-only memory 1230, the system memory 1220, and the permanent storage device 1235.
[0134] From these various memory units, the processing unit (s) 1210 retrieves instructions to execute and data to process in order to execute the processes of the present disclosure. The processing unit (s) may be a single processor or a multi-core processor in different embodiments. Some instructions are passed to and executed by the GPU 1215. The GPU 1215 can offload various computations or complement the image processing provided by the processing unit (s) 1210.
[0135] The read-only-memory (ROM) 1230 stores static data and instructions that are used by the processing unit (s) 1210 and other modules of the electronic system. The permanent storage device 1235, on the other hand, is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 1200 is off. Some embodiments of the present disclosure use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as the permanent storage device 1235.
[0136] Other embodiments use a removable storage device (such as a floppy disk, flash memory device, etc., and its corresponding disk drive) as the permanent storage device. Like the permanent storage device 1235, the system memory 1220 is a read-and-write memory device. However, unlike storage device 1235, the system memory 1220 is a volatile read-and-write memory, such a random access memory. The system memory 1220 stores some of the instructions and data that the processor uses at runtime. In some embodiments, processes in accordance with the present disclosure are stored in the system memory 1220, the permanent storage device 1235, and / or the read-only memory 1230. For example, the various memory units include instructions for processing multimedia clips in accordance with some embodiments. From these various memory units, the processing unit (s) 1210 retrieves instructions to execute and data to process in order to execute the processes of some embodiments.
[0137] The bus 1205 also connects to the input and output devices 1240 and 1245. The input devices 1240 enable the user to communicate information and select commands to the electronic system. The input devices 1240 include alphanumeric keyboards and pointing devices (also called “cursor control devices” ) , cameras (e.g., webcams) , microphones or similar devices for receiving voice commands, etc. The output devices 1245 display images generated by the electronic system or otherwise output data. The output devices 1245 include printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD) , as well as speakers or similar audio output devices. Some embodiments include devices such as a touchscreen that function as both input and output devices.
[0138] Finally, as shown in FIG. 12, bus 1205 also couples electronic system 1200 to a network 1225 through a network adapter (not shown) . In this manner, the computer can be a part of a network of computers (such as a local area network ( “LAN” ) , a wide area network ( “WAN” ) , or an Intranet, or a network of networks, such as the Internet. Any or all components of electronic system 1200 may be used in conjunction with the present disclosure.
[0139] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as computer-readable storage media, machine-readable media, or machine-readable storage media) . Some examples of such computer-readable media include RAM, ROM, read-only compact discs (CD-ROM) , recordable compact discs (CD-R) , rewritable compact discs (CD-RW) , read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM) , a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc. ) , flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc. ) , magnetic and / or solid state hard drives, read-only and recordable discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media may store a computer program that is executable by at least one processing unit and includes sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter.
[0140] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, many of the above-described features and applications are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) . In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In addition, some embodiments execute software stored in programmable logic devices (PLDs) , ROM, or RAM devices.
[0141] As used in this specification and any claims of this application, the terms “computer” , “server” , “processor” , and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. For the purposes of the specification, the terms display or displaying means displaying on an electronic device. As used in this specification and any claims of this application, the terms “computer readable medium, ” “computer readable media, ” and “machine readable medium” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.
[0142] While the present disclosure has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the present disclosure can be embodied in other specific forms without departing from the spirit of the present disclosure. In addition, a number of the figures (including FIG. 8 and FIG. 11) conceptually illustrate processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in one continuous series of operations, and different specific operations may be performed in different embodiments. Furthermore, the process could be implemented using several sub-processes, or as part of a larger macro process. Thus, one of ordinary skill in the art would understand that the present disclosure is not to be limited by the foregoing illustrative details, but rather is to be defined by the appended claims. Additional Notes
[0143] The herein-described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are merely examples, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively "associated" such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as "associated with" each other such that the desired functionality is achieved, irrespective of architectures or intermediate components. Likewise, any two components so associated can also be viewed as being "operably connected" , or "operably coupled" , to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being "operably couplable" , to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable and / or physically interacting components and / or wirelessly interactable and / or wirelessly interacting components and / or logically interacting and / or logically interactable components.
[0144] Further, with respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.
[0145] Moreover, it will be understood by those skilled in the art that, in general, terms used herein, and especially in the appended claims, e.g., bodies of the appended claims, are generally intended as “open” terms, e.g., the term “including” should be interpreted as “including but not limited to, ” the term “having” should be interpreted as “having at least, ” the term “includes” should be interpreted as “includes but is not limited to, ” etc. It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim recitation to implementations containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an, " e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more; ” the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number, e.g., the bare recitation of "two recitations, " without other modifiers, means at least two recitations, or two or more recitations. Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. In those instances where a convention analogous to “at least one of A, B, or C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B. ”
[0146] From the foregoing, it will be appreciated that various implementations of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various implementations disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
1.A video coding method comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;refining a motion vector based on a bilateral matching cost between first and second predictors that are derived based on first and second reference blocks that are located by the motion vector, wherein sample-based prediction offset (SPO) values are determined and applied to the first and second predictors; andencoding or decoding the current block by using the refined motion vector to generate a final predictor.2.The video coding method of claim 1, wherein the SPO values applied to the first predictor are derived based on difference values between reference sample values in a template region neighboring the first reference block and reconstructed sample values in a template region neighboring the current block, wherein the SPO values applied to the second predictor are derived based on difference values between reference sample values in a template region neighboring the second reference block and reconstructed sample values in a template region neighboring the current block.3.The video coding method of claim 2, wherein the SPO values are derived by applying a position-related weighting (PRW) scheme to the difference values.4.The video coding method of claim 3, wherein the first and second predictors are for a subblock of the current block and the PRW scheme implements weight decay over the entire the current block.5.The video coding method of claim 2, wherein the SPO values are derived by right-shifting the difference values.6.The video coding method of claim 2, wherein the reconstructed samples of the template region neighboring the current block are shared by different subblocks for deriving SPO values.7.The video coding method of claim 1, wherein the current block comprises multiple subblocks, wherein the first and second predictors are derived for a first subblock of the current block, wherein the first subblock is reconstructed by using the refined motion vector, wherein reconstructed samples of the first subblock neighboring a second subblock of the current block are used to derive SPO values for the second subblock.8.The video coding method of claim 1, wherein the SPO values are applied to the first and second predictors only if the first predictor is for a subblock of the current block.9.The video coding method of claim 1, wherein the SPO values are applied to the first and second predictors only if a size of the current block meets a particular condition.10.The video coding method of claim 1, wherein the SPO values are applied to the first and second predictors only if the motion vector is provided by a merge candidate having a merge index that meets a particular condition.11.The video coding method of claim 1, wherein the SPO values are derived for only a sub-partition of the current block.12.The video coding method of claim 11, wherein the SPO values are applied only if the sub-partition is a boundary partition of the current block.13.The video coding method of claim 1, wherein the SPO values are applied to the first and second predictors only when a particular coding tool is not applied to the current block.14.An electronic apparatus comprising:a video coder circuit configured to perform operations comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;refining a motion vector based on a bilateral matching cost between first and second predictors that are derived based on first and second reference blocks that are located by the motion vector, wherein sample-based prediction offset (SPO) values are determined and applied to the first and second predictors; andencoding or decoding the current block by using the refined motion vector to generate a final predictor.15.A video decoding method comprising:receiving data to be decoded as a current block of pixels of a current picture of a video;refining a motion vector based on a bilateral matching cost between first and second predictors that are derived based on first and second reference blocks that are located by the motion vector, wherein sample-based prediction offset (SPO) values are determined and applied to the first and second predictors; andreconstructing the current block by using the refined motion vector to generate a final predictor.