Methods and apparatus of neighbouring SKIP mode and regression derived weighting in overlapped blocks motion compensation for video coding
Adaptive OBMC settings for video coding improve efficiency by tailoring OBMC processes to neighboring block modes, reducing complexity and enhancing visual quality and coding efficiency.
Patent Information
- Application Number
- PCT/CN2025/083978
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-28
- Filing Date
- 2025-03-21
- Publication Date
- 2025-10-16
AI Technical Summary
Existing video coding technologies face inefficiencies in Overlapped Block Motion Compensation (OBMC) processes, particularly when neighboring blocks are coded in skip mode, leading to increased computation complexity and memory bandwidth without optimizing visual quality and coding efficiency.
Adaptive OBMC settings are introduced to differentiate between neighboring blocks coded in skip mode and non-skip mode, employing specific OBMC blending rules, weightings, and predictor generation settings tailored for each mode to optimize OBMC processes.
This approach reduces computation complexity and memory bandwidth while enhancing visual quality and coding efficiency by optimizing OBMC operations based on neighboring block coding modes.
Smart Images

Figure CN2025083978_16102025_PF_FP_ABST
Abstract
Description
METHODS AND APPARATUS OF NEIGHBOURING SKIP MODE AND REGRESSION DERIVED WEIGHTING IN OVERLAPPED BLOCKS MOTION COMPENSATION FOR VIDEO CODINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 632,565, filed on April 11, 2024, and U.S. Provisional Patent Application No. 63 / 665,519, filed on June 28, 2024. The U.S. Provisional Patent Applications are hereby incorporated by reference in their entireties.FIELD OF THE INVENTION
[0002] The present invention relates to video coding system using Overlapped Block Motion Compensation (OBMC) . In particular, the present invention relates to different OBMC settings for blocks with neighbouring blocks coded in skip mode and blocks with neighbouring blocks coded in non-skip mode. BACKGROUND AND RELATED ART
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0006] The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
[0007] Overlapped Block Motion Compensation (OBMC)
[0008] Overlapped Block Motion Compensation (OBMC) is to find a Linear Minimum Mean Squared Error (LMMSE) estimate of a pixel intensity value based on motion-compensated signals derived from its nearby block motion vectors (MVs) . From estimation-theoretic perspective, these MVs are regarded as different plausible hypotheses for its true motion, and to maximize coding efficiency, their weights should minimize the mean squared prediction error subject to the unit-gain constraint.
[0009] When High Efficient Video Coding (HEVC) was developed, several proposals were made using OBMC to provide coding gain. Some of them are described as follows.
[0010] In JCTVC-C251 (Peisong Chen, et. al., “Overlapped block motion compensation in TMuC” , Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, 3rd Meeting: Guangzhou, CN, 7-15 October, 2010, Document: JCTVC-C251) , OBMC was applied to geometry partition. In geometry partition, it is very likely that a transform block contains pixels belonging to different partitions. In geometry partition, since two different motion vectors are used for motion compensation, the pixels at the partition boundary may have large discontinuities that can produce visual artefacts similar to blockiness. This in turn decreases the transform efficiency. Let the two regions created by a geometry partition be denoted by region 1 and region 2. A pixel from region 1 (2) is defined to be a boundary pixel if any of its four connected neighbours (left, top, right, and bottom) belongs to region 2 (1) . Fig. 2 shows an example where grey-dotted pixels belong to the boundary of region 1 (grey region) and white-dotted pixels belong to the boundary of region 2 (white region) . If a pixel is a boundary pixel, the motion compensation is performed using a weighted sum of the motion predictions from the two motion vectors. The weights are 3 / 4 for the prediction using the motion vector of the region containing the boundary pixel and 1 / 4 for the prediction using the motion vector of the other region. The overlapping boundaries improve the visual quality of the reconstructed video while also providing BD-rate gain.
[0011] In JCTVC-F299 (Liwei Guo, et. al., “CE2: Overlapped Block Motion Compensation for 2NxN and Nx2N Motion Partitions” , Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, 6th Meeting: Torino, 14-22 July, 2011, Document: JCTVC-F299) , OBMC was applied to symmetrical motion partitions. If a coding unit (CU) is partitioned into 2 2NxN or Nx2N prediction units (PUs) , OBMC is applied to the horizontal boundary of the two 2NxN prediction blocks, and the vertical boundary of the two Nx2N prediction blocks. Since those partitions may have different motion vectors, the pixels at partition boundaries may have large discontinuities, which may generate visual artefacts and also reduce the transform / coding efficiency. In JCTVC-F299, OBMC is introduced to smooth the boundaries of motion partition.
[0012] Figs. 3A-B illustrate an example of OBMC for 2NxN (Fig. 3A) and Nx2N blocks (Fig. 3B) . The grey pixels are pixels belonging to Partition 0 and white pixels are pixels belonging to Partition 1. The overlapped region in the luma component is defined as 2 rows (columns) of pixels on each side of the horizontal (vertical) boundary. For pixels which are 1 row (column) apart from the partition boundary, i.e., pixels labelled as A in Figs. 3A-B, OBMC weighting factors are (3 / 4, 1 / 4) . For pixels which are 2 rows (columns) apart from the partition boundary, i.e., pixels labelled as B in Figs. 3A-B, OBMC weighting factors are (7 / 8, 1 / 8) . For chroma components, the overlapped region is defined as 1 row (column) of pixels on each side of the horizontal (vertical) boundary, and the weighting factors are (3 / 4, 1 / 4) .
[0013] Currently, the OBMC is performed after normal MC, and BIO is also applied in these two MC processes, separately. That is, the MC results for the overlapped region between two CUs or PUs is generated by another process not in the normal MC process. BIO (Bi-Directional Optical Flow) is then applied to refine these two MC results. This can help to skip the redundant OBMC and BIO processes, when two neighbouring MVs are the same. However, the required bandwidth and MC operations for the overlapped region is increased compared to integrating OBMC process into the normal MC process. For example, the current PU size is 16x8, the overlapped region is 16x2, and the interpolation filter in MC is 8-tap. If the OBMC is performed after normal MC, then we need (16+7) x (8+7) + (16+7) x (2+7) = 552 reference pixels per reference list for the current PU and the related OBMC. If the OBMC operations are combined with normal MC into one stage, then only (16+7) x (8+2+7) = 391 reference pixels per reference list for the current PU and the related OBMC. Therefore, in the following, in order to reduce the computation complexity or memory bandwidth of BIO, several methods are proposed, when BIO and OBMC are enabled simultaneously.
[0014] In the JEM (Joint Exploration Model) , the OBMC is also applied. In the JEM, unlike in H. 263, OBMC can be switched on and off using syntax at the CU level. When OBMC is used in the JEM, the OBMC is performed for all motion compensation (MC) block boundaries except for the right and bottom boundaries of a CU. Moreover, it is applied to both the luma and chroma components. In the JEM, a MC block corresponds to a coding block. When a CU is coded with sub-CU mode (includes sub-CU merge, affine and FRUC mode) , each sub-block of the CU is a MC block. To process CU boundaries in a uniform fashion, OBMC is performed at sub-block level for all MC block boundaries, where sub-block size is set equal to 4×4, as illustrated in Figs. 4A-B.
[0015] When OBMC is applied to the current sub-block, besides current motion vectors, motion vectors of four connected neighbouring sub-blocks, if available and are not identical to the current motion vector, are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal of the current sub-block. Prediction block based on motion vectors of a neighbouring sub-block is denoted as PNn, with n indicating an index for the neighbouring above, below, left and right sub-blocks and prediction block based on motion vectors of the current sub-block is denoted as PC. Fig. 4A illustrates an example of OBMC for sub-blocks of the current CU 410 using a neighbouring above sub-block (i.e., PN1) , left neighbouring sub-block (i.e., PN2) , left and above sub-blocks i.e., PN3) . Fig. 4B illustrates an example of OBMC for the ATMVP mode, where block PN of the current CU 420 uses MVs from four neighbouring sub-blocks for OBMC. When PN is based on the motion information of a neighbouring sub-block that contains the same motion information as the current sub-block, the OBMC is not performed from PN. Otherwise, every sample of PN is added to the same sample in PC, i.e., four rows / columns of PN are added to PC. The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for PN and the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for PC. The exception are small MC blocks (i.e., when height or width of the coding block is equal to 4 or a CU is coded with sub-CU mode) , for which only two rows / columns of PN are added to PC. In this case, weighting factors {1 / 4, 1 / 8} are used for PN and weighting factors {3 / 4, 7 / 8} are used for PC. For PN generated based on motion vectors of vertically (horizontally) neighbouring sub-block, samples in the same row (column) of PN are added to PC with a same weighting factor.
[0016] In the JEM, for a CU with size less than or equal to 256 luma samples, a CU level flag is signalled to indicate whether OBMC is applied or not for the current CU. For the CUs with size larger than 256 luma samples or not coded with the AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied for a CU, its impact is taken into account during the motion estimation stage. The prediction signal formed by OBMC using motion information of the top neighbouring block and the left neighbouring block is used to compensate the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.
[0017] In JEM (Joint Exploration Model for VVC development) , the OBMC is applied. For example, as shown in Fig. 5, for a current block 510, if the above block and the left block are coded in an inter mode, it takes the MV of the above block to generate an OBMC block A and takes the MV of the left block to generate an OBMC block L. The predictors of OBMC block A and OBMC block L are blended with the current predictors. To reduce the memory bandwidth of OBMC, it is proposed to do the above 4-row MC and left 4-column MC with the neighbouring blocks. For example, when doing the above block MC, 4 additional rows are fetched to generate a block of (above block + OBMC block A) . The predictors of OBMC block A are stored in a buffer for coding the current block. When doing the left block MC, 4 additional columns are fetched to generate a block of (left block + OBMC block L) . The predictors of OBMC block L are stored in a buffer for coding the current block. Therefore, when doing the MC of the current block, four additional rows and four additional columns of reference pixels are fetched to generate the predictors of the current block, the OBMC block B, and the OBMC block R as shown in Fig. 6A (may also generate the OBMC block BR as shown in Fig. 6B) . The OBMC block B and the OBMC block R are stored in buffers for the OBMC process of the bottom neighbouring blocks and the right neighbouring blocks.
[0018] For an MxN block, if the MV is not integer and a 8-tap interpolation filter is applied, a reference block with size of (M+7) x (N+7) is used for motion compensation. However, if the BIO and OBMC is applied, additional reference pixels are required, which increases the worst case memory bandwidth.
[0019] Template Matching Based OBMC
[0020] Recently, a template matching-based OBMC scheme has been proposed (JVET-Y0076) to the emerging international coding standard. As shown in Fig. 7, for each top block with a size of 4×4 at the top CU boundary, the above template size equals to 4×1. In Fig. 7, box 710 corresponds to a CU. If N adjacent blocks have the same motion information, then the above template size is enlarged to 4N×1 since the MC operation can be processed at one time, which is in the same manner in ECM-OBMC. For each left block with a size of 4×4 at the left CU boundary, the left template size equals to 1×4 or 1×4N.
[0021] For each 4×4 top block (or N 4×4 blocks group) , the prediction value of boundary samples is derived according to the following steps: – Take block A as the current block and its above neighbouring block AboveNeighbour_Afor example. The operation for left blocks is conducted in the same manner. – First, three template matching costs (Cost1, Cost2, Cost3) are measured by SAD between the reconstructed samples of a template and its corresponding reference samples derived by MC process according to the following three types of motion information: Cost1 is calculated according to A’s motion information. Cost2 is calculated according to AboveNeighbour_A’s motion information. Cost3 is calculated according to weighted prediction of A’s and AboveNeighbour_A’s motion information with weighting factors as 3 / 4 and 1 / 4 respectively. – Second, choose one out of three approaches to calculate the final prediction results of boundary samples by comparing Cost1, Cost2 and Cost 3.
[0022] The original MC result using current block’s motion information is denoted as Pixel1, and the MC result using neighbouring block’s motion information is denoted as Pixel2. The final prediction result is denoted as NewPixel. - If Cost1 is minimum, then NewPixel (i, j) = Pixel1 (i, j) . - If (Cost2 + (Cost2 >> 2) + (Cost2 >> 3) ) <= Cost1, then blending mode 1 is used. For luma blocks, the number of blending pixel rows is 4. - NewPixel (i, 0) = (26×Pixel1 (i, 0) +6×Pixel2 (i, 0) +16) >>5 - NewPixel (i, 1) = (7×Pixel1 (i, 1) +Pixel2 (i, 1) +4) >>3 - NewPixel (i, 2) = (15×Pixel1 (i, 2) +Pixel2 (i, 2) +8) >>4 - NewPixel (i, 3) = (31×Pixel1 (i, 3) +Pixel2 (i, 3) +16) >>5 For chroma blocks, the number of blending pixel rows is 1. - NewPixel (i, 0) = (26×Pixel1 (i, 0) +6×Pixel2 (i, 0) +16) >>5 - If Cost1 <= Cost2, then blending mode 2 is used. For luma blocks, the number of blending pixel rows is 2. - NewPixel (i, 0) = (15×Pixel1 (i, 0) +Pixel2 (i, 0) +8) >>4 - NewPixel (i, 1) = (31×Pixel1 (i, 1) +Pixel2 (i, 1) +16) >>5 For chroma blocks, the number of blending pixel rows / columns is 1. - NewPixel (i, 0) = (15×Pixel1 (i, 0) +Pixel2 (i, 0) +8) >>4 - Otherwise, blending mode 3 is used. For luma blocks, the number of blending pixel rows is 4. - NewPixel (i, 1) = (7×Pixel1 (i, 1) +Pixel2 (i, 1) +4) >>3 - NewPixel (i, 2) = (15×Pixel1 (i, 2) +Pixel2 (i, 2) +8) >>4 - NewPixel (i, 3) = (31×Pixel1 (i, 3) +Pixel2 (i, 3) +16) >>5 For chroma blocks, the number of blending pixel rows is 1. - NewPixel (i, 0) = (7×Pixel1 (i, 0) +Pixel2 (i, 0) +4) >>3.
[0023] VVC Inter Prediction
[0024] For each inter-predicted CU, motion parameters consisting of motion vectors, reference picture indices and reference picture list usage index, and additional information needed for the new coding feature of VVC to be used for inter-predicted sample generation. The motion parameter can be signalled in an explicit or implicit manner. When a CU is coded with skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. A merge mode is specified whereby the motion parameters for the current CU are obtained from neighbouring CUs, including spatial and temporal candidates, and additional schedules introduced in VVC. The merge mode can be applied to any inter-predicted CU, not only for skip mode. The alternative to merge mode is the explicit transmission of motion parameters, where motion vector, corresponding reference picture index for each reference picture list and reference picture list usage flag and other needed information are signalled explicitly per each CU.
[0025] Affine Motion Compensated Prediction
[0026] In HEVC, only translation motion model is applied for motion compensation prediction (MCP) . While in the real world, there are many kinds of motion, e.g. zoom in / out, rotation, perspective motions and the other irregular motions. In VVC, a block-based affine transform motion compensation prediction is applied. As shown in fig. 8 A-B, the affine motion field of the block 810 is described by motion information of two control point (4-parameter) in Fig. 8A or three control point motion vectors (6-parameter) in Fig. 8B.
[0027] For 4-parameter affine motion model, motion vector at sample location (x, y) in a block is derived as:
[0028] For 6-parameter affine motion model, motion vector at sample location (x, y) in a block is derived as:
[0029] Where (mv0x, mv0y) is motion vector of the top-left corner control point, (mv1x, mv1y) is motion vector of the top-right corner control point, and (mv2x, mv2y) is motion vector of the bottom-left corner control point.
[0030] In order to simplify the motion compensation prediction, block based affine transform prediction is applied. To derive motion vector of each 4×4 luma subblock, the motion vector of the centre sample of each subblock, as shown in Fig. 9, is calculated according to above equations, and rounded to 1 / 16 fraction accuracy. Then, the motion compensation interpolation filters are applied to generate the prediction of each subblock with the derived motion vector. The subblock size of chroma-components is also set to be 4×4. The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the top-left and bottom-right luma subblocks in the collocated 8x8 luma region.
[0031] As is for translational-motion inter prediction, there are also two affine motion inter prediction modes: affine merge mode and affine AMVP mode.
[0032] Affine Merge Prediction
[0033] AF_MERGE mode can be applied for CUs with both width and height larger than or equal to 8. In this mode, the CPMVs (Control Point MVs) of the current CU is generated based on the motion information of the spatial neighbouring CUs. There can be up to five CPMVP (CPMV Prediction) candidates and an index is signalled to indicate the one to be used for the current CU. The following three types of CPVM candidate are used to form the affine merge candidate list: – Inherited affine merge candidates that are extrapolated from the CPMVs of the neighbour CUs – Constructed affine merge candidates CPMVPs that are derived using the translational MVs of the neighbour CUs – Zero MVs
[0034] In VVC, there are two inherited affine candidates at most, which are derived from the affine motion model of the neighbouring blocks, one from left neighbouring CUs and one from above neighbouring CUs. The candidate blocks of the current CU 1010 are the same as those shown in Fig. 10 . For the left predictor, the scan order is A0->A1, and for the above predictor, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No pruning check is performed between two inherited candidates. When a neighbouring affine CU is identified, its control point motion vectors are used to derived the CPMVP candidate in the affine merge list of the current CU. As shown in Fig. 11, if the neighbouring left bottom block A of the current block 1110 is coded in affine mode, the motion vectors v2 , v3 and v4 of the top left corner, above right corner and left bottom corner of the CU 1120 containing block A are attained. When block A is coded with 4-parameter affine model, the two CPMVs of the current CU (i.e., v0 and v1) are calculated according to v2, and v3. In case that block A is coded with 6-parameter affine model, the three CPMVs of the current CU are calculated according to v2 , v3 and v4.
[0035] Constructed affine candidate means the candidate is constructed by combining the neighbouring translational motion information of each control point. The motion information for the control points is derived from the specified spatial neighbours and temporal neighbour for a current block 1210 as shown in Fig. 12. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2->B3->A2 blocks are checked and the MV of the first available block is used. For CPMV2, the B1->B0 blocks are checked and for CPMV3, the A1->A0 blocks are checked. For TMVP is used as CPMV4 if it’s available.
[0036] After MVs of four control points are attained, affine merge candidates are constructed based on the motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3} , {CPMV1, CPMV2, CPMV4} , {CPMV1, CPMV3, CPMV4} , {CPMV2, CPMV3, CPMV4} , {CPMV1, CPMV2} , {CPMV1, CPMV3}
[0037] The combination of 3 CPMVs constructs a 6-parameter affine merge candidate and the combination of 2 CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling process, if the reference indices of control points are different, the related combination of control point MVs is discarded.
[0038] After inherited affine merge candidates and constructed affine merge candidate are checked, if the list is still not full, zero MVs are inserted to the end of the list.
[0039] Affine AMVP Prediction
[0040] Affine AMVP mode can be applied for CUs with both width and height larger than or equal to 16. An affine flag in the CU level is signalled in the bitstream to indicate whether affine AMVP mode is used and then another flag is signalled to indicate whether 4-parameter affine or 6-parameter affine is used. In this mode, the difference of the CPMVs of current CU and their predictors CPMVPs is signalled in the bitstream. The affine AVMP candidate list size is 2 and it is generated by using the following four types of CPVM candidate in order: – Inherited affine AMVP candidates that extrapolated from the CPMVs of the neighbour CUs – Constructed affine AMVP candidates CPMVPs that are derived using the translational MVs of the neighbour CUs – Translational MVs from neighbouring CUs – Zero MVs
[0041] The checking order of inherited affine AMVP candidates is the same as the checking order of inherited affine merge candidates. The only difference is that, for AVMP candidate, only the affine CU that has the same reference picture as current block is considered. No pruning process is applied when inserting an inherited affine motion predictor into the candidate list.
[0042] Constructed AMVP candidate is derived from the specified spatial neighbours shown in Fig. 10. The same checking order is used as that in the affine merge candidate construction. In addition, the reference picture index of the neighbouring block is also checked. In the checking order, the first block that is inter coded and has the same reference picture as in current CUs is used. When the current CU is coded with the 4-parameter affine mode, and mv0 and mv1 are both availlalbe, they are added as one candidate in the affine AMVP list. When the current CU is coded with 6-parameter affine mode, and all three CPMVs are available, they are added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set as unavailable.
[0043] If the number of affine AMVP list candidates is still less than 2 after valid inherited affine AMVP candidates and constructed AMVP candidate are inserted, mv0, mv1 and mv2 will be added as the translational MVs in order to predict all control point MVs of the current CU, when available. Finally, zero MVs are used to fill the affine AMVP list if it is still not full.
[0044] JVET-AC0164 Non-EE2: Improvements on Local Illumination Compensation in ECM7.0
[0045] In ECM-7.0, local illumination compensation (LIC) is an inter coding technique that aims at addressing the illumination variations between one block and its prediction block. The LIC is based on a linear model where a scale α and an offset β are derived from the template samples neighbouring to the current block and their corresponding prediction samples. The derived LIC parameters are then applied to adjust the prediction samples of the block as P′ [x, y] =α·P [x, y] +β
[0046] Currently, the LIC is only applicable to uni-predictive inter CUs which contains no less than 32 luma samples.
[0047] Additionally, overlapped block motion compensation (OBMC) is another inter tool in ECM7.0, which alleviates the discontinuities among the prediction samples of inter blocks by adjusting the boundary prediction samples of one inter block / sub-block using its neighbouring block’s MV. According to the existing ECM design, when the LIC is applied to one inter block, the OBMC is always disabled. Additionally, when a neighbouring block of the current CU applies the LIC, only its MVs are used to produce the corresponding prediction samples used for the OBMC process of the current CU.
[0048] The following modifications are proposed to further improve the coding efficiency of the LIC tool.
[0049] Bi-Predictive LIC
[0050] It is proposed to extend the existing LIC design to bi-predicted CUs. Specifically, when applying the proposed method to one bi-prediction block, two different linear models are derived to compensate the illumination changes that exist between the current block and its two prediction blocks. Then, the final bi-prediction of the current block is calculated as the combination of two uni-prediction blocks after the LIC adjustment, i.e., P′ [x, y] = (1-ω) ·p′0 [x, y] +ω·p′1 [x, y] , and p′0 [x, y] =α0·P0 [x, y] +β0, p′1 [x, y] =α1·P1 [x, y] +β1, where α0 and β0, and α1 and β1 indicate the scales and the offsets in L0 and L1, respectively; ω indicates the weight (as indicated by the CU-level BCW index) that is applied when combining the two uni-prediction blocks.
[0051] Same to the current LIC design, one control flag is signalled for AMVP bi-predicted CUs to indicate the enabling / disabling of the LIC while the flag is inherited from one neighbouring block for merge inter CUs (including AMVP-Merge mode) . Additionally, the LIC is disabled when decoder-side motion vector refinement (DMVR) (including multi-pass DMVR, adaptive DMVR and affine DMVR) and bi-directional optical flow (BDOF) is applied.
[0052] To reuse the linear model derivation of the existing LIC, one iterative approach is applied to alternately derive the L0 and L1 linear models. Specifically, given the two MVs of the current block, it assumes T0 and T1 are the two predictions of the current block’s template T. The method firstly derives the L0 linear model (α0 and β0) that result in the minimum difference between T0 and T; then, the L1 linear model (α1 and β1) can be calculated that minimizes the difference between T1 and the updated template. Finally, the L0 linear model is refined again in the same way.
[0053] OBMC with LIC
[0054] The following two changes are applied to better handle the interaction between the LIC and the OBMC: 1) It is proposed to enable the OBMC to the inter blocks where the LIC is applied. Additionally, to achieve a better complexity / performance trade-off, the OBMC is only applied for refining the prediction samples on the top and left boundaries of one LIC CU while the OBMC on the internal sub-block boundaries are always disabled. 2) Besides the MVs, it is proposed to also take the LIC parameters of one neighbouring block (when it is coded by the LIC) into consideration when generating its corresponding prediction samples for the OBMC of the current CU.
[0055] Multi-Hypothesis Prediction (MHP)
[0056] In the multi-hypothesis inter prediction mode (JVET-M0425) , one or more additional motion-compensated prediction signals are signalled, in addition to the conventional bi-prediction signal. The resulting overall prediction signal is obtained by sample-wise weighted superposition. With the bi / uni prediction signal pbi / uni and the first additional inter prediction signal / hypothesis h3, the resulting prediction signal p3 is obtained as follows: p3 = (1-α) pbi / uni +αh3
[0057] The weighting factor α is specified by the new syntax element add_hyp_weight_idx, according to the following mapping: add_hyp_weight_idx = 0, α => 1 / 4; and add_hyp_weight_idx = 1, α = -1 / 8.
[0058] Analogously to above, more than one additional prediction signal can be used. The resulting overall prediction signal is accumulated iteratively with each additional prediction signal. pn+1= (1-αn+1) pn+αn+1hn+1.
[0059] The resulting overall prediction signal is obtained as the last pn (i.e., the pn having the largest index n) . Up to two additional prediction signals can be used (i.e., n is limited to 2) .
[0060] The motion parameters of each additional prediction hypothesis can be signalled either explicitly by specifying the reference index, the motion vector predictor index, and the motion vector difference, or implicitly by specifying a merge index. A separate multi-hypothesis merge flag distinguishes between these two signalling modes.
[0061] For inter AMVP mode, MHP is only applied if non-equal weight in BCW is selected in bi-prediction mode.
[0062] Combination of MHP and BDOF is possible, however the BDOF is only applied to the bi-prediction signal part of the prediction signal (i.e., the ordinary first two hypotheses) .
[0063] Bi-Prediction with CU-level Weight (BCW)
[0064] In HEVC, the bi-prediction signal, Pbi-pred is generated by averaging two prediction signals, P0 and P1 obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. Pbi-pred= ( (8-w) *P0+w*P1+4) >>3.
[0065] Five weights are allowed in the weighted averaging bi-prediction, w∈ {-2, 3, 4, 5, 10} . For each bi-predicted CU, the weight w is determined in one of two ways: 1) for a non-merge CU, the weight index is signalled after the motion vector difference; 2) for a merge CU, the weight index is inferred from neighbouring blocks based on the merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256) . For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w ∈ {3, 4, 5} ) are used.
[0066] At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. The details are disclosed in the VTM software and document JVET-L0646 (Yu-Chi Su, et. al., “CE4-related: Generalized bi-prediction improvements combined from JVET-L0197 and JVET-L0296” , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 12th Meeting: Macao, CN, 3–12 Oct. 2018, Document: JVET-L0646) . - When combined with AMVR, unequal weights are only conditionally checked for 1-pel and 4-pel motion vector precisions if the current picture is a low-delay picture. - When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. - When the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked. - Unequal weights are not searched when certain conditions are met, depending on the POC distance between current picture and its reference pictures, the coding QP, and the temporal level.
[0067] The BCW weight index is coded using one context coded bin followed by bypass coded bins. The first context coded bin indicates if equal weight is used; and if unequal weight is used, additional bins are signalled using bypass coding to indicate which unequal weight is used.
[0068] Weighted prediction (WP) is a coding tool supported by the H. 264 / AVC and HEVC standards to efficiently code video content with fading. Support for WP is also added into the VVC standard. WP allows weighting parameters (weight and offset) to be signalled for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weight (s) and offset (s) of the corresponding reference picture (s) are applied. WP and BCW are designed for different types of video content. In order to avoid interactions between WP and BCW, which will complicate VVC decoder design, if a CU uses WP, then the BCW weight index is not signalled, and weight w is inferred to be 4 (i.e. equal weight is applied) . For a merge CU, the weight index is inferred from neighbouring blocks based on the merge candidate index. This can be applied to both the normal merge mode and inherited affine merge mode. For the constructed affine merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index for a CU using the constructed affine merge mode is simply set equal to the BCW index of the first control point MV.
[0069] In VVC, CIIP and BCW cannot be jointly applied for a CU. When a CU is coded with CIIP mode, the BCW index of the current CU is set to 2, (i.e., w=4 for equal weight) . Equal weight implies the default value for the BCW index.
[0070] Regression-Based GPM Blending
[0071] Regression-based GMP blending mode is designed as an additional GPM implicit mode, where the two integer blending matrices (W0 and W1) are derived from the template (1 line above, 1 column left) . The blending matrices are modelled as an affine linear function of the sample positions (x, y) in the current CU: W0 (x, y) = a. x + b. y + c and W1 (x, y) = 1 -W0 (x, y)
[0072] The parameters (a, b, c) are derived from the reference template using the same solver (MSE minimization) as the one used for CCCM. A list of pair of candidates is built from the regular GPM candidates and re-ordered with the template cost.
[0073] The GPM implicit mode is signalled by a CU-level flag (gpm_implicit_flag) . If gpm_implicit_flag is true, a merge-idx is coded to signal the pair of GPM candidates to be used. If gpm_implicit_flag is false, the regular GPM syntax elements are signalled.
[0074] In the current template-matching-based OBMC design, both the neighbouring motion and the current motion are used to generate template predictor. The cost between template predictor and neighbouring reconstruction samples is compared. When neighbouring blocks are coded with skip mode, template matching may favour the template predictor of neighbouring motion since there is no residual in the neighbouring blocks. In such case, template-matching may fail to determine the best blending lines and blending weightings in OBMC. Therefore, several new methods regarding neighbouring blocks coded in the skip mode are disclosed in the present invention. BRIEF SUMMARY OF THE INVENTION
[0075] A method and apparatus for video coding using OBMC are disclosed. According to the method, input data comprising a current block is received, a current subblock, a neighbouring block, or a neighbouring subblock. One or more OBMC (Overlapped Block Motion Compensation) settings associated with OBMC process are determined. One or more target OBMC settings are selected from said one or more OBMC settings according to one or more corresponding neighbouring blocks of the current block being the skip mode or the non-skip mode. The OBMC process is applied to at least one boundary of the current block, the current subblock, the neighbouring block and the neighbouring subblock using said one or more target OBMC settings.
[0076] In one embodiment, said one or more OBMC settings comprise one or more OBMC blending rules, one or more OBMC blending weightings and blending lines, one or more cost-based metrics, adaptive weighting decision, or a combination thereof. In one embodiment, said one or more cost-based metrics correspond to boundary matching or bilateral matching when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode. In one embodiment, said one or more OBMC settings further comprise one or more predictor generation settings when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode. In one embodiment, said one or more predictor generation settings correspond to template predictor generation settings, and the template predictor generation settings are set to the same settings as in neighbouring predictor generation when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode. In one embodiment, the template predictor generation settings comprise BCW index setting, DMVR flag setting, or BDOF setting.
[0077] In one embodiment, when template-matching is enabled for the current block, said one or more OBMC settings further comprise one or more template-matching-based OBMC decision rules. In one embodiment, only some of said one or more template-matching-based OBMC decision rules, a subset of said one or more template-matching-based OBMC decision rules, or other different kinds of template-matching-based OBMC decision rules are allowed for one of the skip mode and the non-skip mode. In one embodiment, different OBMC blending rules are used or partially different OBMC blending rules are used for the skip mode or the non-skip mode. In one embodiment, 2-line OBMC is used for said one or more corresponding neighbouring blocks of the current block being coded in the skip mode. In one embodiment, said one or more template-matching-based OBMC decision rules are the same for the skip mode and the non-skip mode except for some different thresholds being used in cost calculation. In one embodiment, said one or more template-matching-based OBMC decision rules in TM-based OBMC comprise selecting a number of OBMC lines and said selecting the number of OBMC lines is determined by neighbouring motion cost multiplied by one or more respective thresholds, and wherein said one or more respective thresholds correspond to one or more first thresholds when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode, and said one or more respective thresholds correspond to one or more second thresholds when said one or more corresponding neighbouring blocks of the current block are coded in the non-skip mode, and said one or more first thresholds are different from said one or more second thresholds.
[0078] In one embodiment, said one or more OBMC settings for the skip mode are different from the non-skip mode.BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.
[0080] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0081] Fig. 2 illustrates an example of overlapped motion compensation for geometry partitions.
[0082] Figs. 3A-B illustrate an example of OBMC for 2NxN (Fig. 3A) and Nx2N blocks (Fig. 3B) .
[0083] Fig. 4A illustrate an example of the sub-blocks that OBMC is applied, where the example includes subblocks at a CU / PU boundary.
[0084] Fig. 4B illustrate an example of the sub-blocks that OBMC is applied, where the example includes subblocks coded in the AMVP mode.
[0085] Fig. 5 illustrate an example of the OBMC processing using neighbouring blocks from above and left for the current block.
[0086] Fig. 6A illustrate an example of the OBMC processing for the right and bottom part of the current block using neighbouring blocks from right and bottom.
[0087] Fig. 6B illustrate an example of the OBMC processing for the right and bottom part of the current block using neighbouring blocks from right, bottom and bottom-right.
[0088] Fig. 7 illustrates an example of Template Matching based OBMC where, for each top block with a size of 4×4 at the top CU boundary, the above template size equals to 4×1.
[0089] Fig. 8A illustrates an example of the affine motion field of a block described by motion information of two control point (4-parameter) .
[0090] Fig. 8B illustrates an example of the affine motion field of a block described by motion information of three control point motion vectors (6-parameter) .
[0091] Fig. 9 illustrates an example of block based affine transform prediction, where the motion vector of each 4×4 luma subblock is derived from the control-point MVs.
[0092] Fig. 10 illustrates the neighbouring blocks used for deriving spatial merge candidates for VVC.
[0093] Fig. 11 illustrates an example of derivation for inherited affine candidates based on control-point MVs of a neighbouring block.
[0094] Fig. 12 illustrates an example of affine candidate construction by combining the translational motion information of each control point from spatial neighbours and temporal.
[0095] Fig. 13 illustrates an example of reasonable model check methods, where model differences at pre-defined locations are checked.
[0096] Fig. 14 illustrates an example of reasonable model check methods, where the regression weighting is checked at location (x, y) .
[0097] Fig. 15 illustrates an example of three intervals for TM-based OBMC decision, where the neighbouring motion cost are multiplied by A and B for skip mode (Fig. 15A) and non-skip mode (Fig. 15B) .
[0098] Fig. 16 illustrates an example of five intervals for TM-based OBMC decision, where the neighbouring motion cost are multiplied by A and B for skip mode (Fig. 16A) and non-skip mode (Fig. 16B) .
[0099] Fig. 17 illustrates a flowchart of an exemplary video coding system, where one or more OBMC settings for skip mode are different from non-skip mode according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0100] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0101] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other examples, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0102] Neighbouring Skip Mode in OBMC
[0103] In current template-matching-based OBMC design, both the neighbouring motion and the current motion are used to generate the template predictors and compare the costs between the template predictor and neighbouring reconstruction samples. When neighbouring blocks are coded with skip mode, template matching may favour the template predictor of neighbouring motion since there is no residual in the neighbouring blocks. In such case, template-matching may fail to determine the best blending lines and blending weightings in OBMC. Therefore, several new methods regarding neighbouring skip mode coded blocks are disclosed.
[0104] When neighbouring blocks are coded with skip mode, the template-matching-based OBMC can follow different decision rules, OBMC blending rules can be different, or predictor generation setting in motion compensation can be different.
[0105] Example 1. Template-matching-based OBMC decision rules for the case of neighbouring blocks being coded in skip mode and template-matching being enabled
[0106] In one embodiment, if template-matching is enabled and neighbouring blocks are coded in skip mode, the template-matching-based OBMC decision rules of these subblocks are different from other subblocks, where the corresponding neighbouring blocks are not coded in skip mode. For example, only certain decision rules are allowed, a subset of template-matching decision rules is allowed, or other different kinds of template-matching decision rules are allowed. For another example, only 2-line OBMC or 3-line OBMC is allowed in template-matching decision or N-line OBMC is always in template-matching decision or always use fixed OBMC weightings in template-matching decision.
[0107] Example 2. OBMC blending rules for the case of neighbouring blocks being coded in skip mode and template-matching being enabled
[0108] In one embodiment, if template-matching is enabled and neighbouring blocks are coded in skip mode, the OBMC blending rules of these subblocks are different from other subblocks, where the corresponding neighbouring blocks are not coded in skip mode. For example, different OBMC blending rules are used or partially different OBMC blending rules are used. For example, 2-line OBMC is always used when neighbouring blocks are coded in skip mode. For another example, TM-based rule is almost the same except that some different thresholds are used in cost calculation. For example, decision rules, such as more blending lines or stronger weightings compared to original OBMC blending rules, are used. For yet another example, decision rules, such as fewer blending lines or weaker weightings compared to original OBMC blending rules, are used.
[0109] Example 3. OBMC blending weightings and OBMC blending lines for the case of neighbouring blocks being coded in skip mode and template-matching being enabled
[0110] In one embodiment, if template-matching is enabled and neighbouring blocks are coded in skip mode, the OBMC blending weightings or OBMC blending lines of these subblocks are different from other subblocks, where the corresponding neighbouring blocks are not coded in skip mode. For example, stronger weightings or more blending lines are used. For another example, weaker weightings or fewer blending lines are used.
[0111] Example 4. Cost-based metrics for the case of neighbouring blocks being coded in skip mode and template-matching being enabled
[0112] In one embodiment, if template-matching is enabled and neighbouring blocks are coded in skip mode, cost-based metrics are used for these subblocks. For example, boundary matching is used for these subblocks coded with skip mode. For another example, bilateral matching is used for these subblocks coded with skip mode if they are bi-prediction coded.
[0113] Example 5. An adaptive weighting decision for the case of neighbouring blocks being coded in skip mode and template-matching being enabled
[0114] In another embodiment, if template-matching is enabled and neighbouring blocks are coded in skip mode, cost-based metrics are used for these subblocks. For example, an adaptive weighting decision rule similar to template-matching-based OBMC is used.
[0115] Example 6. OBMC blending rules for the case of neighbouring blocks being coded in skip mode and template-matching being disabled
[0116] In one embodiment, if template-matching is disabled and neighbouring blocks are coded in skip mode, the OBMC blending rules of these subblocks are different from other subblocks, where the corresponding neighbouring blocks are not coded in skip mode. For example, different OBMC blending rules are used or partially different OBMC blending rules are used. For another example, 2-line OBMC is always used when neighbouring blocks are coded in skip mode. For yet another example, TM-based rule is almost the same except that some different thresholds are used in cost calculation. For yet another example, decision rules, such as more blending lines or stronger weightings compared to original OBMC blending rules, are used.
[0117] Example 7. OBMC blending weightings and OBMC blending lines for the case of neighbouring blocks being coded in skip mode and template-matching being disabled
[0118] In one embodiment, if template-matching is disabled and neighbouring blocks are coded in skip mode, the OBMC blending weightings or OBMC blending lines of these subblocks are different from other subblocks, where the corresponding neighbouring blocks are not coded in skip mode. For example, stronger weightings or more blending lines are used. For another example, weaker weightings or fewer blending lines are used.
[0119] Example 8. Cost-based metrics for the case of neighbouring blocks being coded in skip mode and template-matching being disabled
[0120] In one embodiment, if template-matching is disabled and neighbouring blocks are coded in skip mode, cost-based metrics are used for these subblocks. For example, boundary matching is used for these subblocks coded with skip mode. For another example, bilateral matching is used for these subblocks coded with skip mode if they are bi-prediction coded.
[0121] Example 9. An adaptive weighting decision for the case of neighbouring blocks being coded in skip mode and template-matching being disabled
[0122] In another embodiment, if template-matching is disabled and neighbouring blocks are coded in skip mode, cost-based metrics are used for these subblocks. For example, an adaptive weighting decision rule similar to template-matching-based OBMC is used.
[0123] Example 10. Predictor generation setting for the case of neighbouring blocks being coded in skip mode
[0124] In one embodiment, when neighbouring blocks are coded in skip mode, template predictor generation settings are set to the same settings as in neighbouring predictor generation. For example, the template predictor generation settings may correspond to BCW index setting, DMVR flag setting, or BDOF setting.
[0125] In another embodiment, when neighbouring blocks are coded in skip mode, OBMC overlapped predictor generation settings are set to the same settings as in neighbouring predictor generation. For example, the OBMC overlapped predictor generation settings may correspond to BCW index setting, DMVR flag setting, or BDOF setting.
[0126] Example 11. Storage of template predictor for the case when the predictor is coded in skip mode for OBMC
[0127] In one embodiment, when block is coded with skip mode, the predictor is stored for OBMC usage. Instead of re-generating predictors for OBMC process, the predictor is stored in advance. For following blocks that will perform OBMC, the stored predictor is utilized to perform OBMC blending in the following blocks.
[0128] Example 12. blending rules, blending weightings or TM-based decision rules only applied to the scenarios of the current block being coded in non-skip mode and neighbouring block being coded in skip mode
[0129] In one embodiment, different blending rules, different blending weightings, or different TM-based decision rules are utilized only when neighbouring block is coded in skip mode and the current block is coded in non-skip mode.
[0130] Regression-Based Derived Weightings in TM-based OBMC
[0131] In current template-matching-based OBMC design, three kinds of motion are evaluated to determine the best match between template predictor and neighbouring reconstruction samples. However, because of various video contents and objects, limited motion options may fail to meet the optimal match between template predictor and neighbouring reconstruction samples. A regression-based derived weightings in template-matching-based OBMC is disclosed.
[0132] Derivation of Regression-Based Weightings
[0133] In template-matching-based OBMC, 3 kinds of costs are calculated, shown as follows: 1*CostCur+0*CostNei, 0*CostCur+1*CostNei,
[0134] The three terms correspond to the cost of current motion template predictor and neighbouring reconstruction samples, the cost of neighbouring motion template predictor and neighbouring reconstruction samples, and the cost of blended motion template predictor and neighbouring reconstruction samples, respectively.
[0135] It is suggested to consider a general form, where the objective function can be written as: SSD= (CostRecon- (A*CostCur+B*CostNei) ) 2.
[0136] In the above equation, A and B denote the weightings for the current motion and for the neighbouring motion, respectively.
[0137] By taking derivatives with respect to A and B, and setting the derivatives to zero, we can get:
[0138] By solving the matrix, we can get the corresponding weightings A and B for the current motion and the neighbouring motion.
[0139] In one embodiment, for one or more subblocks in template-matching-based OBMC, regression-based derived weightings are utilized.
[0140] In another embodiment, for one or more current subblocks coded in specific prediction modes in template-matching-based OBMC, regression-based derived weightings are utilized. For example, specific prediction modes can be skip mode, merge mode, affine mode, DMVR mode, AMVP mode.
[0141] In another embodiment, for some neighbouring blocks coded in specific prediction modes in template-matching-based OBMC, regression-based derived weightings are utilized. For example, specific prediction modes can be skip mode, merge mode, affine mode, DMVR mode, AMVP mode.
[0142] In another embodiment, for some blocks satisfying block size constraint, regression-based derived weightings are utilized. For example, the block size constraint can be block area, block aspect ratio, block width or block height.
[0143] Regression-Based OBMC
[0144] Several new methods regarding regression-based OBMC are proposed. During training sample collection, some samples may not be helpful to training process and can be omitted. For example, samples that are coded with skip mode are not collected since the predictor of skip mode coded samples can be identical to reconstruction samples due to no residual. Furthermore, to guarantee enough training samples, training sample amount constraint, neighbouring sample amount constraint, or current block size constraint can be imposed. Moreover, during the template predictor generation in the regression model process, certain parameters, flags, or indices can be inherited from the current block, current subblock, neighbouring block, or neighbouring subblock.
[0145] It is possible that the training process may generate unreasonable models, such as models always outputting constant weightings. Therefore, it is also proposed to exclude unreasonable models and to preserve only valid models after training process. Another proposed method is to clip regression weightings from models.
[0146] Training Sample Collection in Regression-Based OBMC
[0147] In one embodiment, one or more types of samples are omitted during training sample collection, such as skip mode coded samples or intra mode coded samples.
[0148] In another embodiment, all samples are collected during training sample collection; however, certain types, such as skip-mode coded samples or intra-mode coded samples, are not used or considered in the training process.
[0149] In another embodiment, training sample amount constraint is considered during training sample collection. For example, if the training sample amount is smaller than N samples, the training process will be terminated, where N is an integer larger than or equal to 0.
[0150] In another embodiment, the neighbouring sample amount constraint is considered during training sample collection. For example, if the training sample amount is smaller than N samples, the training process will be terminated, where N is an integer larger than or equal to 0.
[0151] In another embodiment, the current block size constraint is considered during training sample collection. For example, if the training sample amount is smaller than N samples, the training process will be terminated, where N is an integer larger than or equal to 0.
[0152] In another embodiment, the training process is executed when all neighbouring samples are inter-prediction coded.
[0153] In another embodiment, the training process is executed when all neighbouring samples are inter-prediction coded or IBC-coded.
[0154] Template Predictor Generation in Regression Model Training Process
[0155] In one embodiment, during template predictor generation for the regression model training process, certain parameters, flags, or indices are inherited or inferred from neighbouring subblocks or blocks.
[0156] In another embodiment, during template predictor generation for the regression model training process, some parameters or flags or indices are inherited or inferred from current subblocks or current blocks.
[0157] In another embodiment, during template predictor generation for the regression model training process, some parameters or flags or indices are adaptively determined from current subblocks or blocks, or neighbouring subblocks or blocks.
[0158] In another embodiment, during template predictor generation for the regression model training process, template predictor from neighbouring motion is generated using neighbouring parameters, flags, or indices.
[0159] In another embodiment, during template predictor generation for the regression model training process, template predictor from the current motion is generated using current parameters, flags, or indices.
[0160] Reasonable Regression Model Verification
[0161] In one embodiment, regression models are checked to determine whether models are reasonable or not. The models are checked using some pre-defined methods. For example, as shown in Fig. 13, some differences of derived regression weightings are calculated as follows: Diff [0] = W0 (w-1, 0) -W0 (0, 0) Diff [1] = W0 (w / 2, 0) -W0 (0, 0) Diff [2] = W0 (w-1, 0) -W0 (w / 2, 0) Diff [3] = W0 (3w / 4, 0) -W0 (w / 4, 0) . In the above equations, w corresponds to the width of the block.
[0162] If all four difference values are greater than or equal to 0, the model is determined to be unreasonable and this regression model is not used.
[0163] In another embodiment, regression models are checked to determine whether models are reasonable or not. The regression weighting is derived from W0 (x, y) = ax + by + c, as shown in Fig. 14, where two different (x, y) locations (1410 and 1420) are shown. To derive decreasing weightings, the model parameters should be smaller than or equal to 0. To derive increasing weightings, the model parameters should be larger than or equal to 0. The regression models that do not satisfy the conditions are determined to be unreasonable.
[0164] In another embodiment, regression weightings are clipped to be within a value range.
[0165] Neighbouring Skip Mode Coded Blocks and Neighbouring Non-Skip Mode Coded Blocks in Template-Matching-Based OBMC
[0166] In one embodiment, as shown in Fig. 15, there are three intervals for TM-based OBMC decision. When a neighbouring block is coded with skip mode (shown in Fig. 15A) , the decision rule in TM-based OBMC is determined by the neighbouring motion cost (i.e., Cost1) multiplied by thresholds A_skip and B_skip. When the neighbouring block is coded with non-skip mode (shown in Fig. 15B) , the decision rule in TM-based OBMC is determined by the neighbouring motion cost (i.e., Cost1) multiplied by other thresholds A_nonskip and B_nonskip. In one example, A_skip is larger than or equal to A_nonskip, and B_skip is larger than or equal to B_nonskip.
[0167] In another embodiment, as shown in Fig. 16A and Fig. 16B, there are five intervals for TM-based OBMC decision. When a neighbouring block is coded with skip mode (shown in Fig. 16A) , the decision rule in TM-based OBMC is determined by the neighbouring motion cost (i.e., Cost1) multiplied by some threshold values skip_threshold. When the neighbouring block is coded with non-skip mode (shown in Fig. 16B) , the decision rule in TM-based OBMC is determined by the neighbouring motion cost (i.e., Cost1) multiplied by some threshold values nonskip_threshold. In one example, threshold values, skip_threshold are larger than or equal to threshold values, nonskip_threshold.
[0168] In another embodiment, when checking neighbouring consecutive motion, the skip-mode flag is also considered. Specifically, if two neighbouring motions are identical but one is coded with skip mode and the other with non-skip mode, they should be treated as non-consecutive motions.
[0169] In another embodiment, the threshold values, skip_threshold or nonskip_threshold, can be signalled in the bitstream or indicated by at least one syntax element, such as a CTU-level flag, slice-level flag, picture-level flag, or sequence-level flag.
[0170] Any of the foregoing proposed methods of different OMBC setting for skip mode and non-skip mode can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in predictor derivation module of an encoder, and / or a predictor derivation module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the predictor derivation module of the encoder and / or the predictor derivation module of the decoder, so as to provide the information needed by the predictor derivation module.
[0171] With reference to the exemplary encoder in Fig. 1A and exemplary decoder in Fig. 1B, any of the proposed methods can be implemented in a predictor derivation module of an encoder, and / or a predictor derivation module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the predictor derivation module of the encoder and / or the predictor derivation module of the decoder, so as to provide the information needed by the predictor derivation module. For example, the OBMC process can be implemented in an encoder side or a decoder side, such as the Intra / Inter coding module (e.g. Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module is an encoder (e.g. Intra Pred. 110 / Inter Pred. 112 in Fig. 1A) .
[0172] Fig. 17 illustrates a flowchart of an exemplary video coding system, where one or more OBMC settings for skip mode are different from non-skip mode according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data comprising a current block is received in step 1710, a current subblock, a neighbouring block, or a neighbouring subblock. One or more OBMC (Overlapped Block Motion Compensation) settings associated with OBMC process are determined in step 1720. One or more target OBMC settings are selected from said one or more OBMC settings according to one or more corresponding neighbouring blocks of the current block being the skip mode or the non-skip mode in step 1730. The OBMC) process is applied to at least one boundary of the current block, the current subblock, the neighbouring block and the neighbouring subblock using said one or more target OBMC settings in step 1740.
[0173] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0174] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0175] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0176] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of video coding, the method comprising:receiving input data comprising a current block, a current subblock, a neighbouring block, or a neighbouring subblock;determining one or more OBMC (Overlapped Block Motion Compensation) settings associated with OBMC processselecting one or more target OBMC settings from said one or more OBMC settings according to one or more corresponding neighbouring blocks of the current block being skip mode or non-skip mode; andapplying the OBMC process to at least one boundary of the current block, the current subblock, the neighbouring block and the neighbouring subblock using said one or more target OBMC settings.2.The method of Claim 1, wherein said one or more OBMC settings comprise one or more OBMC blending rules, one or more OBMC blending weightings and blending lines, one or more cost-based metrics, adaptive weighting decision, or a combination thereof.3.The method of Claim 2, wherein said one or more cost-based metrics correspond to boundary matching or bilateral matching when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode.4.The method of Claim 2, wherein said one or more OBMC settings further comprise one or more predictor generation settings when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode.5.The method of Claim 4, wherein said one or more predictor generation settings correspond to template predictor generation settings, and the template predictor generation settings are set to the same settings as in neighbouring predictor generation when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode.6.The method of Claim 5, wherein the template predictor generation settings comprise BCW index setting, DMVR flag setting, or BDOF setting.7.The method of Claim 2, wherein when template-matching is enabled is enabled for the current block, said one or more OBMC settings further comprise one or more template-matching-based OBMC decision rules.8.The method of Claim 7, wherein only some of said one or more template-matching-based OBMC decision rules, a subset of said one or more template-matching-based OBMC decision rules, or other different kinds of template-matching-based OBMC decision rules are allowed for the skip mode or the non-skip mode.9.The method of Claim 7, wherein different OBMC blending rules are used or partially different OBMC blending rules are used for the skip mode or the non-skip mode.10.The method of Claim 7, wherein 2-line OBMC is used for said one or more corresponding neighbouring blocks of the current block being coded in the skip mode.11.The method of Claim 7, wherein said one or more template-matching-based OBMC decision rules are the same for the skip mode and the non-skip mode except for some different thresholds being used in cost calculation.12.The method of Claim 7, wherein said one or more template-matching-based OBMC decision rules in TM-based OBMC comprise selecting a number of OBMC lines and said selecting the number of OBMC lines is determined by neighbouring motion cost multiplied by one or more respective thresholds, and wherein said one or more respective thresholds correspond to one or more first thresholds when said one or more corresponding neighbouring blocks of the current block are coded in the skip mode, and said one or more respective thresholds correspond to one or more second thresholds when said one or more corresponding neighbouring blocks of the current block are coded in the non-skip mode, and said one or more first thresholds are different from said one or more second thresholds.13.The method of Claim 1, wherein said one or more OBMC settings for the skip mode are different from the non-skip mode.14.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data comprising a current block, a current subblock, a neighbouring block, or a neighbouring subblock;determine one or more OBMC (Overlapped Block Motion Compensation) settings associated with OBMC process;select one or more target OBMC settings from said one or more OBMC settings according to one or more corresponding neighbouring blocks of the current block being skip mode or non-skip mode; andapply the OBMC process to at least one boundary of the current block, the current subblock, the neighbouring block and the neighbouring subblock using said one or more target OBMC settings.
Citation Information
Patent Citations
Overlay block motion compensation using time domain neighbors
CN110858901A
Method and apparatus for video processing with overlap block motion compensation in video coding system
CN114554197A
Overlapped block motion compensation
US20220201282A1
Method and Apparatus of Overlapped Block Motion Compensation in Video Coding System
US20230328278A1