Method and apparatus of TMVP and motion-trajectory-based motion vectors for affine model derivation in video coding systems

By employing motion-trajectory-based motion vectors for SbTMVP and affine modes, the method addresses inefficiencies in video coding systems, enhancing encoding and decoding performance through optimized motion vector prediction and affine model derivation.

WO2026032350A1PCT designated stage Publication Date: 2026-02-12MEDIATEK INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/113145
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-27
Filing Date
2025-08-07
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing video coding systems, such as VVC, face challenges in efficiently encoding and decoding video data due to impairments in reconstructed video data, particularly in handling affine and subblock-based motion vector predictions, leading to suboptimal coding efficiency.

Method used

The method and apparatus utilize motion-trajectory-based motion vectors for SbTMVP and affine mode to improve coding efficiency by deriving motion information using previously coded frames and constructing motion-trajectory-based MV lists for subblocks, enabling enhanced affine model derivation and subblock motion prediction.

Benefits of technology

This approach enhances coding efficiency by optimizing motion vector prediction, particularly in affine and subblock-based modes, leading to improved video quality and reduced computational resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025113145_12022026_PF_FP_ABST
    Figure CN2025113145_12022026_PF_FP_ABST
Patent Text Reader

Abstract

Method and apparatus for video coding for SbTMVP and / or Affine mode are disclosed. According to one method, one or more MV lists for one or more subblocks of the current block are determined, wherein each MV list comprises one or more motion-trajectory-based MVs (Motion Vectors). One or more target motion-trajectory-based MVs are selected from said one or more MV lists. Motion information for the current block coded in SbTMVP mode or affine mode is derived by using said one or more target motion-trajectory-based MVs. According to another method, one or more subblock motion candidates are derived by applying chained motion derivation to one or more candidates in a subblock merge candidate list. One or more target subblock motion candidates are determined from the subblock merge candidate list.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS OF TMVP AND MOTION-TRAJECTORY-BASED MOTION VECTORS FOR AFFINE MODEL DERIVATION IN VIDEO CODING SYSTEMSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 680, 087 filed on August 7, 2024, U.S. Provisional Patent Application No. 63 / 710, 659 filed on October 23, 2024, and U.S. Provisional Patent Application No. 63 / 739, 160 filed on December 27, 2024. The U.S. Provisional Patent Applications are hereby incorporated by reference in their entireties.FIELD OF THE INVENTION

[0002] The present invention relates to video coding system. In particular, the present invention relates to using motion-trajectory-based MV to SbTMVP and / or Affine mode to improve coding efficiency.BACKGROUND

[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.

[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.

[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.

[0006] The decoder, as shown in Fig. 1B, can use some of the functional blocks as the encoder. For example, the decoder can reuse Inverse Quantization 124 and Inverse Transform 126; however, Transform 118 and Quantization 120 are not needed at the decoder. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.

[0007] In VVC, the Sequence Parameter Set (SPS) and the Picture Parameter Set (PPS) contain high-level syntax elements that apply to entire coded video sequences and pictures, respectively. The Picture Header (PH) and Slice Header (SH) contain high-level syntax elements that apply to a current coded picture and a current coded slice, respectively.

[0008] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only.

[0009] Affine Merge Prediction

[0010] AF_MERGE mode can be applied for CUs with both width and height larger than or equal to 8. In this mode, the CPMVs (Control Point MVs) of the current CU is generated based on the motion information of the spatial neighbouring CUs. There can be up to five CPMVP (CPMV Prediction) candidates and an index is signalled to indicate the one to be used for the current CU. The following three types of CPMV candidate are used to form the affine merge candidate list: – Inherited affine merge candidates that are extrapolated from the CPMVs of the neighbour CUs  – Constructed affine merge candidates CPMVPs that are derived using the translational MVs  of the neighbour CUs – Zero MVs

[0011] In VVC, there are two inherited affine candidates at most, which are derived from the affine motion model of the neighbouring blocks, one from left neighbouring CUs and one from above neighbouring CUs. The candidate blocks are the same as those shown in Fig. 2. For the left predictor, the scan order is A0→A1, and for the above predictor, the scan order is B0→B1→B2. Only the first inherited candidate from each side is selected. No pruning check is performed between two inherited candidates. When a neighbouring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidate in the affine merge list of the current CU. As shown in Fig. 3, if the neighbouring left bottom block A of the current block 310 is coded in affine mode, the motion vectors v2 , v3 and v4 of the top left corner, above right corner and left bottom corner of the CU 320 containing block A are attained. When block A is coded with 4-parameter affine model, the two CPMVs of the current CU (i.e., v0 and v1) are calculated according to v2, and v3. In case that block A is coded with 6-parameter affine model, the three CPMVs of the current CU are calculated according to v2 , v3 and v4.

[0012] Constructed affine candidate means the candidate is constructed by combining the neighbouring translational motion information of each control point. The motion information for the control points is derived from the specified spatial neighbours and temporal neighbour for a current block 410 as shown in Fig. 4. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2→B3→A2 blocks are checked and the MV of the first available block is used. For CPMV2, the B1→B0 blocks are checked and for CPMV3, the A1→A0 blocks are checked. For TMVP is used as CPMV4 if it’s available.

[0013] After MVs of four control points are attained, affine merge candidates are constructed based on the motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3} , {CPMV1, CPMV2, CPMV4} , {CPMV1, CPMV3, CPMV4} ,  {CPMV2, CPMV3, CPMV4} , {CPMV1, CPMV2} , {CPMV1, CPMV3}

[0014] The combination of 3 CPMVs constructs a 6-parameter affine merge candidate and the combination of 2 CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling process, if the reference indices of control points are different, the related combination of control point MVs is discarded.

[0015] After inherited affine merge candidates and constructed affine merge candidate are checked, if the list is still not full, zero MVs are inserted to the end of the list.

[0016] Subblock-based Temporal Motion Vector Prediction (SbTMVP)

[0017] VVC supports the subblock-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the collocated picture to improve motion vector prediction and merge mode for CUs in the current picture. The same collocated picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: - TMVP predicts motion at CU level, but SbTMVP predicts motion at sub-CU level; whereas  TMVP fetches the temporal motion vectors from the collocated block in the collocated picture (the collocated block is the bottom-right or centre block relative to the current CU) , SbTMVP applies a motion shift before fetching the temporal motion information from the collocated picture, where the motion shift is obtained from the motion vector from one of the spatial neighbouring blocks of the current CU. - The SbTMVP process is illustrated in Fig. 5 and Fig. 6. SbTMVP predicts the motion vectors  of the sub-CUs within the current CU 612 in two steps in the current picture 610. In the first step, the spatial neighbour A1 in Fig. 6 is examined. If A1 has a motion vector 630 that uses the collocated picture 620 as its reference picture, this motion vector is selected to be the motion shift to be applied. If no such motion is identified, then the motion shift is set to (0, 0) . In the second step, the motion shift identified in Step 1 is applied (i.e. added to the current block’s coordinates) to obtain sub-CU level motion information (motion vectors and reference indices) from the collocated picture 620 as shown in Fig. 6. The example in Fig. 6 assumes the motion shift is set to block A1’s motion. Then, for each sub-CU, the motion information of its corresponding block 622 (the smallest motion grid that covers the centre sample) in the collocated picture is used to derive the motion information for the sub-CU. After the motion information of the collocated sub-CU is identified, it is converted to the motion vectors and reference indices of the current sub-CU in a similar way as the TMVP process of HEVC, where temporal motion scaling is applied to align the reference pictures of the temporal motion vectors to those of the current CU. In VVC, a combined subblock based merge list which contains both SbTVMP candidate and  affine merge candidates is used for the signalling of subblock based merge mode. The SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry of the list of subblock based merge candidates, and followed by the affine merge candidates. The size of subblock based merge list is signalled in SPS and the maximum allowed size of the subblock based merge list is 5 in VVC. The sub-CU size used in SbTMVP is fixed to be 8x8, and as done for affine merge mode,  SbTMVP mode is only applicable to the CU with both width and height are larger than or equal to 8. The encoding logic of the additional SbTMVP merge candidate is the same as for the other  merge candidates, that is, for each CU in P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.

[0018] History-Parameter-Based Affine Model Inheritance and Non-Adjacent Affine Mode

[0019] History-parameter-based affine model inheritance (HAMI) allows the affine model to be inherited from a previously affine-coded block which may not be neighbouring to the current block. Similar to the enhanced regular merge mode, non-adjacent affine mode (NA-AFF) is introduced.

[0020] A first history-parameter table (HPT) is established. An entry of the first HPT stores a set of affine parameters: a, b, c and d, and each of which is represented by a 16-bit signed integer. Entries in HPT are categorized by reference list and reference index. Five reference indices are supported for each reference list in HPT. The category of HPT (denoted as HPTCat) can be calculated as: HPTCat (RefList, RefIdx) = 5×RefList + min (RefIdx, 4) , where RefList and RefIdx represent a reference picture list (0 or 1) and a reference index, respectively.  For each category, at most seven entries can be stored, resulting in a total of 70 entries in HPT. At the beginning of each CTU row, the number of entries for each category is initialized as zero. After decoding an affine-coded CU with reference list RefListcur and RefIdxcur, the affine parameters are utilized to update entries in the category HPTCat (RefListcur, RefIdxcur) in a way similar to HMVP table updating.

[0021] A history-affine-parameter-based candidate (HAPC) is derived from one of the seven neighbouring 4×4 blocks denoted as A0, A1, A2, B0, B1, B2 or B3 in Fig. 4 and a set of affine parameters stored in a corresponding entry in the first HPT. The MV of a neighbouring 4×4 block serves as the base MV. The MV of the current block at position (x, y) can be calculated as: where (mvhbase, mvvbase) represents the MV of the neighbouring 4×4 block, (xbase, ybase) represents the  centre position of the neighbouring 4×4 block. (x, y) can be the top-left, top-right and bottom-left corner of the current block to obtain the corner-position MVs (CPMVs) for the current block, or it can be the centre of the current block to obtain a regular MV for the current block.

[0022] A second history-parameter table (HPT) with base MV information is also appended. There are nine entries in the second HPT, wherein an entry comprises a base MV, a reference index and four affine parameters for each reference list, and a base position. An additional merge HAPC can be generated from the second HPT with the base MV information and the corresponding affine models stored in an entry. The difference between the first HPT and the second HPT is illustrated in Figs. 7A-B.

[0023] Moreover, pair-wised affine merge candidates are generated by two affine merge candidates, which are history-derived or not history-derived. A pair-wised affine merge candidate is generated by averaging the CPMVs of existing affine merge candidates in the list.

[0024] As a response to new HAPCs being introduced, the size of sub-block-based merge candidate list is increased from five to fifteen, which are all involved in the ARMC process.

[0025] In NA-AFF, the pattern of obtaining non-adjacent spatial neighbours is shown in Fig. 8A.Same as the existing non-adjacent regular merge candidates, the distances between non-adjacent spatial neighbours and current coding block in the NA-AFF are also defined based on the width and height of the current CU.

[0026] The motion information of the non-adjacent spatial neighbours in Fig. 8A is utilized to generate additional inherited and constructed affine merge / AMVP candidates. Specifically, for inherited candidates, the derivation process of the inherited affine merge / AMVP candidates in the VVC is kept unchanged except that the CPMVs are inherited from non-adjacent spatial neighbours. The non-adjacent spatial neighbours are checked based on their distances to the current block (i.e., from near to far) . At a specific distance, only the first available neighbour (that is coded with the affine mode) from each side (e.g. the left and above) of the current block (block 810 in Fig. 8A and block 820 in Fig. 8B) is included for inherited candidate derivation. As indicated by the dashed arrows in Fig. 8A, the checking orders of the neighbours on the left and above sides are bottom-to-up and right-to-left, respectively. Figs. 8A-B illustrate examples of non-adjacent spatial neighbours for deriving affine merge mode (NSAM) , where the pattern of obtaining non-adjacent spatial neighbours is shown in Fig. 8A for deriving inherited affine merge candidates and in Fig. 8B for deriving constructed affine merge candidates.

[0027] For the first type of constructed candidates, as shown in the Fig. 8B, the positions of left and above non-adjacent spatial neighbours are firstly determined independently; after that, the location of the top-left neighbour can be determined accordingly, which can enclose a rectangular virtual block together with the left and above non-adjacent neighbours. Then, as shown in the Fig. 9, the motion information of the three non-adjacent neighbours is used to form the CPMVs at the top-left (A) , top-right (B) and bottom-left (C) of the virtual block, which is finally projected to the current CU to generate the corresponding constructed candidates.

[0028] The NA-AFF candidates are inserted into the existing affine merge candidate list and affine AMVP candidate list according to the following orders: Affine merge mode: 1. SbTMVP candidate, if available 2. Inherited from adjacent neighbours 3. Inherited from non-adjacent neighbours 4. Constructed from adjacent neighbours 5. The first type of constructed affine candidates from non-adjacent neighbours 6. Zero MVs Affine AMVP mode: 1. Inherited from adjacent neighbours 2. Constructed from adjacent neighbours 3. Translational MVs from adjacent neighbours 4. Translational MVs from temporal neighbours 5. Inherited from non-adjacent neighbours 6. The first type of constructed affine candidates from non-adjacent neighbours 7. Zero MVs

[0029] Due to the inclusion of the additional candidates generated by NA-AFF, the size of the affine merge candidate list is increased from 5 to 15. The subgroup size of ARMC for the affine merge mode is increased from 3 to 15.

[0030] In NA-AFF: 1. The area from where the non-adjacent neighbours come is restricted to be within the current  CTU (i.e., no additional storage requirements for line buffer) . 2. The storage granularity for affine motion information, including CPMVs and reference  indexes, is reduced from 8x8 to 16x16 (i.e., only the affine motion from the top-left 8x8 block is saved) . Additionally, the saved CPMVs are projected to each 16x16 block before being stored, such that the position and size information are not needed. 3. Only the top-left and top-right CPMVs are stored (i.e., always using 4-parameter affine model  for NA-AFF) .

[0031] Regression Based Affine Candidate Derivation

[0032] The Regression based Motion Vector Field (RMVF) derivation method provides a new variety of subblock-based merge candidate. The motion vectors and centre positions from the neighbouring subblocks of the current CU, as illustrated in Fig. 10, are used as the input to the linear regression process to derive a set of linear model parameters.

[0033] The subblock motion field from a previous coded affine CU and the motion vectors from the adjacent subblocks of the current CU are used as the input for the regression process. The predicted CPMVs for current block are derived as output.

[0034] The regression based affine merge candidates are derived and added to the affine merge list. Subblock motion field from a previously coded affine CU and motion information from adjacent subblocks of the current CU are used as the input to the regression process to derive proposed affine candidates.

[0035] The previously coded affine CU can be identified from scanning through non-adjacent positions and the affine HMVP table.

[0036] Adjacent subblock information of the current CU is fetched from 4x4 sub-blocks represented by the grey zone as depicted in Fig. 10. For each sub-block, given a reference list, the corresponding motion vector and centre coordinate of the sub-block may be used.

[0037] For each affine CU, up to 2 affine candidates can be derived. One with adjacent subblock information and one without. All the linear-regression-generated candidates are pruned and collected into one candidate sub-group, and TM cost based ARMC process is applied when ARMC is enabled. Afterwards, up to N linear-regression-generated candidates are added to the affine merge list when N affine CUs are found. The number of affine candidates for ARMC is 30, the output list size is 15.

[0038] MVP Extension (JVET-AI0183)

[0039] In JVET-AI0183, additional candidates using a selected reference picture with a scaled MV is proposed. For merge candidates, it adds a candidate before the default zero MV candidates with the reference index 1 when the existing candidates in the merge list has reference index 0, otherwise it adds a candidate with the reference index 0.

[0040] For TMVP and SbTMVP, it adds a candidate with a reference picture corresponding to the collocated block’s reference picture and the collocated MV is scaled accordingly, if the picture is not in the reference picture list of the current block, it selects a reference picture between the collocated picture and the collocated reference picture with the largest POC distance to the current picture.

[0041] Finally, it adds a bi-TMVP candidate (MV0, MV1 as shown in Fig. 11) after HMVP if a motion trajectory between a block in a reference picture and its reference block crosses the current block. For a given MV shown in a dashed line, which is a reference block MV, a pair of MV0 and MV1 is constructed and if it crosses the current block then those MVs are used as bi-TMVP.

[0042] JVET-AG0091: EE2-1.8: Auto-Relocated Block Vector Prediction

[0043] In EE2-1.8, auto-relocated block vector prediction (AR-BVP) is introduced into IBC merge / AMVP candidate list construction.

[0044] As shown in Fig. 12, a guiding block vector BV0, 1 associated with the current block B0 points to a reference block B1. If B1 has a BV denoted as BV1, 2 pointing to a reference block B2, then BV0, 2, given by BV0, 2 = BV0, 1 +BV1, 2, is defined as the AR-BVP, guided by BV0, 1. Similarly, BV0, n+1 can be derived by BV0, n+1 =BV0, n+BVn, n+1 = BV0, 1+BV1, 2 +…+BVn-1, n +BVn, n+1.

[0045] Three tests are conducted in this EE. In EE2-1.8a, the length of the AR-BVP trace path is 1 (i. e, n=1) . In EE2-1.8b, the length of the AR-BVP trace path is 2 (i. e, n=2) . In EE2-1.8c, there is no constraint for the length of the AR-BVP trace path.

[0046] When deriving BVn, n+1 guided by BV0, n, all five positions including top-left (e.g., LT in Fig. 13) , top-right (e.g., RT in Fig. 13) , centre (e.g., Ctr in Fig. 13) , bottom-left (e.g., LB in Fig. 13) , and bottom-right (e.g., RB in Fig. 13) positions of Bn are checked to find BVn, n+1.

[0047] In the proposed implementation, the initial guiding block vector BV0, 1 is set to be an existing BVP already in the IBC merge / AMVP candidate list.

[0048] The AR-BVP candidates are inserted after the HBVP candidates. The IBC merge / AMVP candidate list size is kept unchanged.

[0049] JVET-AG0073: Non-EE2: Chained Motion Vector Prediction

[0050] This contribution introduces a chained MV prediction (CMVP) into inter merge candidate list construction.

[0051] As shown in Fig. 14, CMVP candidates can be derived as the sum of the recursively traced MVs and BVs based on the pre-derived MVs for the inter merge candidate list. For instance, for a CMVP candidate, a set of motion vector MVk / m and reference picture RefPick / m can be derived by: MVk / m = MVk (0) + BVk (0) + MVk (1) +MVk (2) + …+ MVk (m) , RefPick / m = RefPick (m) , where k and m indicate the number of merge index and trace depths of the CMVP.

[0052] When deriving MVk / m, MVk (m) is found by checking the existence of MVs or BVs in MV / BV storage corresponding to all five position of the current block as shown in Fig. 15 (i.e., the centre, top-left, top-right, bottom-left, and bottom-right of the current block) .

[0053] When pre-derived merge candidates targeting CMVP candidates has two MVs, a MVk / m is derived for each list (i.e., L0 and L1) and each trace depth. Up to two MVs can be derived for each list and each trace depth, and the MV set is sequentially inserted into inter merge candidate list.

[0054] The traceable reference pictures are only within the reference picture list.

[0055] CMVP candidates are inserted after HMVP candidates for the regular merge and TM merge.

[0056] When deriving CMVP candidates, hpelIfIdx, bcwIdx, licFlag, and mhpFlag are not inherited. CMVP candidates are not derived when the TMVP is disabled.

[0057] In the present invention, methods and apparatus of applying motion-trajectory-based MV to SbTMVP and / or Affine mode to improve coding efficiency are disclosed. BRIEF SUMMARY OF THE INVENTION

[0058] A method and apparatus for video coding for SbTMVP and / or Affine mode are disclosed. According to one method, input data associated with a current block is received, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. One or more MV (Motion Vector) lists for one or more subblocks of the current block are determined, wherein each MV list comprises one or more motion-trajectory-based MVs (Motion Vectors) . One or more target motion-trajectory-based MVs are selected from said one or more MV lists. Motion information for the current block coded in SbTMVP (Subblock Temporal Motion Vector Prediction) mode or affine mode is derived by using said one or more target motion-trajectory-based MVs. The current block coded in the SbTMVP mode or the affine mode is encoded or decoded by using the motion information.

[0059] In one embodiment, for a target subblock of the current block, said one or more motion-trajectory-based MVs are determined by searching previously coded frame MVs that pass through the target subblock.

[0060] In one embodiment, when the current block is coded in the SbTMVP mode, subblock MVs are derived by using said one or more target motion-trajectory-based MVs as one or more motion shift candidates, and said one or more subblocks of the current block are encoded or decoded by using the subblock MVs. In one embodiment, all of said one or more target motion-trajectory-based MVs are used as said one or more motion shift candidates to derive the subblock MVs for said one or more subblocks of the current block.

[0061] In one embodiment, when the current block uses the affine mode, an affine model is derived for the current block coded in the affine mode, and the current block is encoded or decoded by using the affine model. In one embodiment, all of said one or more target motion-trajectory-based MVs are used to derive affine candidates by deriving the affine model for the current block. In one embodiment, the affine candidates comprise constructed affine candidates, regression-based affine candidates, temporal affine candidates, or a combination thereof.

[0062] In one embodiment, multiple subblocks at specific positions in the current block are selected, and a set of motion-trajectory-based MVs is selected from corresponding MV lists of the multiple subblocks to derive the affine model as affine candidates. In one embodiment, the affine candidates comprise constructed affine candidates, regression-based affine candidates, temporal affine candidates, or a combination thereof. In one embodiment, derivation of the affine model comprise a combination of selected motion-trajectory-based MVs and one or more non-motion-trajectory-based MVs. In one embodiment, only when said one or more target motion-trajectory-based MVs referenced from a same reference picture of L0, a same reference picture of L1, or both, the target motion-trajectory-based MVs are selected s to derive the affine model as affine candidates.

[0063] In one embodiment, two rounds are performed to derive the affine model for the current block, and wherein N first subblocks are selected from the current block in a first round and M second subblocks are selected from the N first subblocks in a second round to derive the affine model.

[0064] In one embodiment, one on / off control flag is signalled at a CU level, slice level, picture level, sequence level, or a combination thereof to indicate whether said one or more target motion-trajectory-based MVs from said one or more MV lists are used to derive the motion information for the current block coded in the SbTMVP mode or the affine mode. In one embodiment, whether said one or more target motion-trajectory-based MVs from said one or more MV lists are used to derive the motion information for the current block coded in the SbTMVP mode or the affine mode is according to selected reference picture indices, temporal distance between a reference picture and a current picture, quantization parameter, coded information of the current block, prediction mode, motion vectors, motion vector resolution, residual of the current block, reference samples, or a combination thereof.

[0065] According to another method, input data associated with a current block is received, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side, and wherein the current block is partitioned into one or more subblocks. One or more subblock motion candidates are derived by applying chained motion derivation to one or more candidates in a subblock merge candidate list. One or more target subblock motion candidates are determined from the subblock merge candidate list. The current block is encoded or decoded by using said one or more target subblock motion candidates.

[0066] In one embodiment, if no valid subblock motion is found for a target subblock of said one or more subblocks after said applying the chained motion derivation, a pre-defined motion vector or an original subblock motion vector is used a subblock target motion vector for the target subblock, a pre-defined motion vector or an original subblock motion vector is added to the subblock merge candidate list.

[0067] In one embodiment, said applying the chained motion derivation to said one or more candidates in the subblock merge candidate list is disabled when the current block uses LIC (Local Illumination Compensation) refinement.

[0068] In one embodiment, said applying the chained motion derivation to said one or more candidates in the subblock merge candidate list is disabled when the current block uses BCW (Bi-prediction with CU-level Weighting) blending.

[0069] In one embodiment, a one on / off control flag is signalled or parsed at a CU level, slice level, picture level, sequence level or a combination thereof to indicate whether said applying the chained motion derivation to said one or more candidates in the subblock merge candidate list is applied.BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.

[0071] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.

[0072] Fig. 2 illustrates the neighbouring blocks used for deriving spatial merge candidates for VVC.

[0073] Fig. 3 illustrates an example of derivation for inherited affine candidates based on control-point MVs of a neighbouring block.

[0074] Fig. 4 illustrates an example of affine candidate construction by combining the translational motion information of each control point from spatial neighbours and temporal.

[0075] Fig. 5 illustrates an example of spatial neighbouring blocks used by SbTMVP.

[0076] Fig. 6 illustrates an example of deriving sub-CU motion field by applying a motion shift from spatial neighbour and scaling the motion information from the corresponding collocated sub-CUs.

[0077] Fig. 7A illustrates an example of first history-parameter table (HPT) , where each entry of the first HPT stores a set of affine parameters.

[0078] Fig. 7B illustrates an example of second history-parameter table (HPT) with base MV information also appended.

[0079] Figs. 8A-B illustrate examples of non-adjacent spatial neighbours for deriving affine merge mode (NSAM) , where the pattern of obtaining non-adjacent spatial neighbours is shown in Fig. 8A for deriving inherited affine merge candidates and in Fig. 8B for deriving constructed affine merge candidates.

[0080] Fig. 9 illustrates an example of constructed affine candidates according to non-adjacent neighbours, where the motion information of the three non-adjacent neighbours at locations A, B and C is used to form the CPMVs.

[0081] Fig. 10 illustrates an example of neighbouring 4 x 4 subblocks used for RMVF parameter derivation, where W and H are the width and height of the current CU.

[0082] Fig. 11 illustrates an example of adding a bi-TMVP candidate (MV0, MV1) after HMVP if a motion trajectory between a block in a reference picture and its reference block crosses the current block.

[0083] Fig. 12 illustrates an example of how to derive AR-BVP (Auto-Relocated Block Vector Prediction) .

[0084] Fig. 13 illustrates an example of five spatial locations checked for block Bn in order to derive block vector BVn, n+1.

[0085] Fig. 14 illustrates an example of CMVP (Chained MVP) candidates derived as the sum of the recursively traced MVs and BVs based on the pre-derived MVs for the inter merge candidate list.

[0086] Fig. 15 illustrates an example of deriving MVk (m) by checking the existence of MVs or BVs in MV / BV storage corresponding to all five positions of the current block.

[0087] Fig. 16 illustrates an example of an MV list of motion-trajectory-based MVs being constructed for each subblock, and if the current coding region is in SbTMVP mode, each subblock MV is selected from the corresponding MV list.

[0088] Fig. 17 illustrates an example of an MV list of motion-trajectory-based MVs being constructed for each subblock, and if the current coding region is in SbTMVP mode, one MV (or two MVs for bi-prediction) is selected from all subblock MV lists as the motion shift.

[0089] Fig. 18A illustrates an example of a 6-parameter constructed affine candidate derived by the 3 MVs from 3 subblocks in the current coding region, where the 3 subblocks are the left-top, right-top, and left-bottom subblocks.

[0090] Fig. 18B illustrates an example of an 8-parameter constructed affine candidate derived by 4 MVs from 4 subblocks in the current coding region, where the 4 subblocks are the left-top, right-top, left-bottom, and right-bottom subblocks.

[0091] Fig. 18C illustrates an example where some subblocks at corner may not have available subblock MVs, and other non-corner subblocks may be used to derive a constructed affine candidate.

[0092] Fig. 18D illustrates an example where the left-top and right-top MVs are from neighbouring spatial MVs, and the left-bottom MV is from motion-trajectory-based MVs. These 3 MVs are used together to derive a 6-parameter constructed affine candidate.

[0093] Fig. 19 illustrates a flowchart of an exemplary video coding system that determines one or more MV (Motion Vector) lists for one or more subblocks of the current block according to an embodiment of the present invention, wherein each MV list comprises one or more motion-trajectory-based MVs (Motion Vectors) .

[0094] Fig. 20 illustrates a flowchart of an exemplary video coding system that determines one or more subblock motion candidates by applying chained motion derivation to one or more candidates in a subblock merge candidate list according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0095] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0096] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0097] PROPOSED METHODS

[0098] Motion-Trajectory-Based Motion Vectors for SbTMVP and Affine Mode

[0099] In JVET-AI0183, a motion-trajectory-based MV for a current subblock is derived by the following process: (1) Iterate all subblock MVs in each reference picture. (2) Create an MV list for each subblock in current picture. If an MV from a reference picture  crosses a subblock in the current picture, put the MV into the corresponding MV list. (3) The MVs in the MV list of the current subblock can be the MVP candidates of a current coding  region containing the current subblock.

[0100] In the prior art, the motion-trajectory-based MVs are used to generate bi-directional TMVPs. For a current coding region, some of the motion-trajectory-based MVs from all subblocks in the current coding region are selected as MVPs based on some ranking metrics.

[0101] In this invention, we propose to extend the usage of motion-trajectory-based MV to SbTMVP and affine modes.

[0102] Embodiment 1 In one embodiment, for a current coding region, an MV list of motion-trajectory-based MVs is constructed for each subblock in the current coding region by searching reference frame MVs that pass through the subblock. If the current coding region is in SbTMVP mode, all or part of the motion-trajectory-based MVs are used to derive the subblock MVs.

[0103] Example 1:

[0104] Based on embodiment 1, an MV list of motion-trajectory-based MVs is constructed for each subblock, and if the current coding region is in SbTMVP mode, each subblock MV is selected from the corresponding MV list. An illustrative example is shown in Fig. 16.

[0105] In Fig. 16, since MVs from subblock a and subblock b in the reference picture 0 both pass through subblock c in the current picture (dashed lines) , two pairs of MVs (MVca0, MVca1, MVcb0, MVcb1) construct a motion-trajectory-based MV list for subblock c. On the other hand, since only MV from subblock b passes through subblock d in the current picture, one pair of MVs (MVdb0, MVdb1) constructs a motion-trajectory-based MV list for subblock d. If the current CU is SbTMVP, subblock c selects its subblock MV from its list (i.e., (MVca0, MVca1) , or (MVcb0, MVcb1) ) , while subblock d selects its subblock MV from its list (i.e., (MVdb0, MVdb1) ) .

[0106] If there is a subblock with an empty motion-trajectory-based MV list, a default MV is assigned to the subblock. The default MV can be the MV of a previous / neighbouring subblock, the MV of the first subblock in the current region, or the average MV of the neighbouring subblocks.

[0107] Example 2:

[0108] Based on embodiment 1, an MV list of motion-trajectory-based MVs is constructed for each subblock, and if the current coding region is in SbTMVP mode, one MV (or two MVs for bi-prediction) is selected from all subblock MV lists as the motion shift. The motion shift indicates how to fetch the motion buffer of the collocated picture (s) to obtain the subblock MVs. An illustrative example is shown in Fig. 17.

[0109] In this figure, since MV from subblock a in the reference picture 0 passes through subblock c in the current picture (dashed line) , a pair of motion-trajectory-based MVs (MVca0, MVca1) is derived. If the current CU is coded in SbTMVP, MVca0 or / and MVca1 are used as the motion shift. That is, instead of the collocated block B0 (subblock c’ in the reference picture 0 is at the same position as the subblock c in the current CU) , block B1 is used to derive subblock MVs for the current CU.

[0110] Embodiment 2 In one embodiment, for a current coding region, an MV list of motion-trajectory-based MVs is constructed for each subblock in the current coding region by searching reference frame MVs that pass through the subblock. If the current coding region is in affine mode, all or part of the motion-trajectory-based MVs are used to derive constructed affine candidates, regression-based affine candidates, temporal affine candidates, and / or the combination thereof.

[0111] Example 3:

[0112] Based on embodiment 2, an MV list of motion-trajectory-based MVs is constructed for each subblock. If the current coding region is in affine mode, some subblocks at specific positions in the current region are picked, and a set of motion-trajectory-based MVs are selected from the corresponding MV lists of the picked subblocks. The set of motion-trajectory-based MVs are used to derive a four / six / eight-parameter affine model as a constructed affine candidate. Note that the MVs used for a constructed affine candidate can be solely motion-trajectory-based MVs, or motion-trajectory-based MVs mixed with other existing MVs, such as spatial or temporal MVs. An illustrative example is shown in Fig. 18A to Fig. 18D.

[0113] In Fig. 18A, a 6-parameter constructed affine candidate is derived by the 3 MVs from 3 subblocks in the current coding region, where the 3 subblocks are the left-top, right-top, and left-bottom subblocks. In Fig. 18B, an 8-parameter constructed affine candidate is derived by 4 MVs from 4 subblocks in the current coding region, where the 4 subblocks are the left-top, right-top, left-bottom, and right-bottom subblocks. In Fig. 18C, since some subblocks at corner may not have available subblock MVs, other non-corner subblocks could be used to derive a constructed affine candidate. In these 3 examples, the subblock MVs are from motion-trajectory-based MVs.

[0114] In Fig. 18D, the left-top and right-top MVs are from neighbouring spatial MVs, and the left-bottom MV is from motion-trajectory-based MVs. These 3 MVs are used together to derive a 6-parameter constructed affine candidate.

[0115] In example 3.1, it is constrained that only if the MVs used to derive the 4 / 6 / 8-parameter constructed affine candidate are referenced from the same reference picture of L0 and reference picture of L1, they can be taken.

[0116] In example 3.2, it is constrained that only if the MVs used to derive the 4 / 6 / 8-parameter constructed affine candidate are referenced from the same reference picture of L0, they can be taken.

[0117] In example 3.3, it is constrained that only if the MVs used to derive the 4 / 6 / 8-parameter constructed affine candidate are referenced from the same reference picture of L1, they can be taken.

[0118] In example 3.4, the MVs used to derive the 4 / 6 / 8-parameter constructed affine candidate will be scaled to the same referenced picture before derivation. For example, the referenced picture can be a pre-defined picture (i.e., reference picture index 0) .

[0119] Example 4:

[0120] Based on embodiment 2, an MV list of motion-trajectory-based MVs is constructed for each subblock. If the current coding region is in affine mode, all subblocks or some subblocks at specific positions in the current region are picked, and a set of motion-trajectory-based MVs are selected from the corresponding MV lists of the picked subblocks. The set of motion-trajectory-based MVs are used to derive a four / six / eight-parameter affine model as a regression-based affine candidate. That is, the affine model parameters are calculated by minimizing the MV differences between the model predicted MVs and the subblock MVs through a regression process. Note that the MVs used for a regression-based affine candidates can be solely motion-trajectory-based MVs, or motion-trajectory-based MVs mixed with other existing MVs, such as spatial or temporal MVs.

[0121] In one embodiment, the subblocks used to derive motion-trajectory-based MVs are selected depending on the CU block size. For example, for a 64x64 block, the subblocks sub-sample step size is set to 16. For example, subblocks with top-left position at (0, 0) , (0, 16) , (0, 32) , (0, 48) , (16, 0) , (16, 16) , (16, 32) , (16, 48) , (32, 0) , (32, 16) , (32, 32) , (32, 48) , (48, 0) , (48, 16) , (48, 32) , and (48, 48) will be picked. The corresponding motion-trajectory-based MVs will be used to derive a four / six / eight-parameter affine model as a regression-based affine candidate. For another example, for a 16x16 block, the subblocks sub-sample step size is set to 8. For example, subblocks with top-left position at (0, 0) , (0, 8) , (8, 0) , and (8, 8) will be picked. The corresponding motion-trajectory-based MVs will be used to derive a four / six / eight-parameter affine model as a regression-based affine candidate. Larger step size is applied for larger CUs.

[0122] In example 4.1, regarding the target reference frame from L0 and reference frame from L1 of the derived four / six / eight-parameter affine model, a voting system is used. For example, one or more subblocks will be picked and the corresponding motion-trajectory-based MVs will be used to derive the target reference frame from L0 and reference frame from L1.

[0123] In example 4.2, two rounds will be performed to derive a four / six / eight-parameter affine model. In the first round, N subblocks will be picked and the corresponding motion-trajectory-based MVs will be used for voting. The reference frame from L0 and reference frame from L1 from highest number of voted MV will be used as a target reference frame of L0 (refIdx0) and reference frame of L1 (refIdx1) . In the second round, M subblocks will be picked and only if the corresponding motion-trajectory-based MVs are from refIdx0 or refIdx1, they can be used to derive the four / six / eight-parameter affine model. For example, for one subblock, a maximum motion-trajectory-based MV will be picked for four / six / eight-parameter affine model derivation.

[0124] In the above description, N and M are the integer values larger than 0.

[0125] The subblocks selected for voting and the subblocks selected to derive the four / six / eight-parameter affine model are the same subblocks, but not limit to.

[0126] In example 4.3, the target reference frame from L0 and reference frame from L1 of the derived four / six / eight-parameter affine model are predefined. For example, reference frame with index equal to 0 is used. M subblocks are picked, and only if the corresponding motion-trajectory-based MVs are from reference frame index 0 will be used to derive the four / six / eight-parameter affine model. For example, for one subblock, a maximum motion-trajectory-based MV will be picked for four / six / eight-parameter affine model derivation.

[0127] For another example, regarding the above example 4.2, if a picked subblock in the second round have at least one motion-trajectory-based MV, but no motion-trajectory-based MV referenced from refIdx0 and refIdx1, one motion-trajectory-based MV will be scaled. The scaled motion-trajectory-based MV will be used for four / six / eight-parameter affine model derivation.

[0128] For another example, regarding to the above example 4.3, if a picked subblock in the second round have at least one motion-trajectory-based MV, but no motion-trajectory-based MV referenced from the pre-defined target reference frame from L0 and reference frame from L1, one motion-trajectory-based MV will be scaled. The scaled motion-trajectory-based MV will be used for four / six / eight-parameter affine model derivation.

[0129] The first MV in motion-trajectory-based MV buffer will be scaled.

[0130] The MV in motion-trajectory-based MV buffer with reference picture closer to the target reference picture will be scaled.

[0131] Example 5:

[0132] Based on embodiment 2, an MV list of motion-trajectory-based MVs is constructed for each subblock, and if the current coding region is in affine mode, one MV (or two MVs for bi-prediction) is selected from all subblock MV lists as the motion shift. The motion shift indicates how to fetch the motion buffer of the collocated picture (s) to obtain a temporal affine candidate.

[0133] In the above embodiments, when selecting one or some MVs from a motion-trajectory-based MV list for a subblock, if the reference picture of the MV is not valid for the current region, the MV is discarded, or the MV is scaled to a valid reference picture.

[0134] In the above embodiments, when selecting one or some MVs from a motion-trajectory-based MV list for a subblock, if the reference picture of the MV is not the same as that of a previous MV for a previous subblock, the MV is discarded, or the MV is scaled to the same reference picture as that of the previous MV.

[0135] In the above embodiments, when selecting one or some MVs from a motion-trajectory-based MV list for a subblock as a motion shift, if the reference picture of the MV is not a valid collocated picture for the current region, the MV is discarded, or the MV is scaled to a valid collocated picture.

[0136] In another embodiment, one on / off control flag is signalled at CU level, slice level, picture level, and / or sequence level to indicate the proposed method in the above is enabled or not.

[0137] In another embodiment, the proposed method in the above is enabled or disabled, according to one or the combination of the selected reference pictures indices, temporal distance between reference picture and current picture, quantization parameter, the coded information of the current CU, prediction mode, motion vectors, motion vector resolution, residual of the current CU, and reference samples.

[0138] In some embodiments, the temporal subblock motions or subblock MVPs of the current block can be derived in some ways. The derived subblock motions or MVPs will be used to derive the affine model of the current coded block.

[0139] For example, the affine model of the current coded block is derived through regression.

[0140] For example, the subblock motions or subblock MVPs of the current block of the current block are derived by motion trajectory-based motion vectors.

[0141] For example, the subblock motions or subblock MVPs of the current block are derived by temporal motion vectors. The footprint used to search the temporal motion vector is the same as inter merge mode.

[0142] For example, the subblock motions or subblock MVPs of the current block are derived by motion trajectory-based motion vectors. Only the subblock motions within bottom-right region of the current block will be derived.

[0143] For example, the subblock motions or subblock MVPs of the current block are derived by motion trajectory-based motion vectors. Only the subblock motions within bottom-right region of the current block will be derived. Furthermore, each motion trajectory-based motion vector in the bottom-right region can be assigned a higher weight during the derivation of the affine regression model. For example, the motions with higher weights can dominate the derivation of the regression model.

[0144] For example, the subblock motions or subblock MVPs of the current block are derived by motion trajectory-based motion vectors. Only the subblock motions within bottom-right region of the current block will be derived. Furthermore, one motion trajectory-based motion vector in the bottom-right region can be treated as two exactly same motion trajectory-based motion vectors during the derivation of the affine model.

[0145] For example, the subblock motions of the current block or subblock MVPs of the current block are derived by motion trajectory-based motion vectors. And sub-sample technology is used to take motion trajectory-based motion vectors. For example, step size is 2, that is, every 2 subblock motions will be taken to derive affine model. The step size can be designed based on current CU size, QP value, or neighbouring CUs’ prediction mode.

[0146] For example, the subblock motions of the current block or subblock MVPs of the current block can be derived by temporal affine coded blocks. In that, the temporal affine coded blocks’s ubblock motions will be referenced after scaling.

[0147] The final reference picture can be the picture which let the POC distance between current picture and reference picture closest to the POC distance between collocated picture and the referenced picture of the collocated picture. Or the final reference picture is a pre-defined picture, i.e., referenced picture with reference index equal to 0.

[0148] The subblock positions of the referenced temporal affine coded blocks can be directly used for affine regression motion derivation.

[0149] The subblock positions of the referenced temporal affine coded blocks can be relocated based on the POC distance between collocated picture and the referenced picture of the collocated picture and the POC distance between current picture and reference picture.

[0150] For example, N target reference pictures can be derived by temporal affine coded blocks. N is an integer greater than 0. And then the subblock motions of the current block or subblock MVPs of the current block are derived by motion trajectory-based motion vectors with scaling. They will be scaled to the target reference picture.

[0151] Subblock Candidates for Auto-Relocated Block Vector Prediction or Chained Motion Vector Prediction.

[0152] Embodiment 1 In one embodiment, the chain motion technology is used to derive subblock mode candidates.

[0153] For example, the chain motion technology is used to derive the candidates in the subblock merge candidate list to derive more subblock candidates. For example, for each subblock, the chain motion technology is applied. If there is no valid motion found after tracing, a pre-defined motion will be used or the original subblock motion will be used.

[0154] If there is at least one tracing subblock motion available for the current coding region, the new candidate will be inserted into the subblock merge candidate list. For one example, no scaling technology is used. Not all subblocks need to be referenced by the same reference picture.

[0155] In another example, no scaling technology is used. Only if all subblocks are referenced by the same reference pictures, this new tracing subblock merge candidate will be inserted into subblock merge candidate list.

[0156] In another example, scaling technology is used. All subblock motions will be scaled to the same reference pictures, i.e., reference picture index 0.

[0157] In another example, only SbTMVP candidates in subblock merge candidate list can be used to derive new tracing subblock merge candidates, but not limit to.

[0158] For another example, only affine candidates in subblock merge candidate list, can be used to derive new tracing subblock merge candidates, but not limit to.

[0159] The new tracing subblock merge candidates cannot be enabled with LIC refinement.

[0160] The new tracing subblock merge candidates cannot be enabled with BCW blending method.

[0161] In another embodiment, an on / off control flag is signalled at CU level, slice level, picture level, and / or sequence level to indicate the above proposed method is enabled or not.

[0162] In another embodiment, the above proposed method is enabled or disabled according to one or a combination of the selected reference picture indices, temporal distance between the reference picture and the current picture, quantization parameter, the coded information of the current CU, prediction mode, motion vectors, motion vector resolution, residual of the current CU, and reference samples.

[0163] Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in inter coding of an encoder, and / or a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter coding of the encoder and / or the decoder, so as to provide the information needed by the inter coding.

[0164] With reference to the exemplary encoder and decoder in Fig. 1A and Fig 1B, the proposed methods can be implemented in the inter prediction modules. For example, in the encoder side, the required processing can be implemented as part of the Inter-Pred. unit 112 as shown in Fig. 1A. However, the encoder may also use additional processing unit to implement the required processing. For the decoder side, the required processing can be implemented as part of the MC unit 152 as shown in Fig. 1B. However, the decoder may also use additional processing unit to implement the required processing. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module. While the Inter-Pred. 112 in the encoder side and MC 152 in the decoder side are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .

[0165] Fig. 19 illustrates a flowchart of an exemplary video coding system that determines one or more MV (Motion Vector) lists for one or more subblocks of the current block according to an embodiment of the present invention, wherein each MV list comprises one or more motion-trajectory-based MVs (Motion Vectors) . The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. A method and apparatus for video coding for SbTMVP and / or Affine mode are disclosed. According to one method, input data associated with a current block is received in step 1910, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. One or more MV (Motion Vector) lists for one or more subblocks of the current block are determined in step 1920, wherein each MV list comprises one or more motion-trajectory-based MVs (Motion Vectors) . One or more target motion-trajectory-based MVs are selected from said one or more MV lists in step 1930. Motion information for the current block coded in SbTMVP (Subblock Temporal Motion Vector Prediction) mode or affine mode is derived by using said one or more target motion-trajectory-based MVs in step 1940. The current block coded in the SbTMVP mode or the affine mode is encoded or decoded by using the motion information in step 1950.

[0166] Fig. 20 illustrates a flowchart of an exemplary video coding system that determines one or more subblock motion candidates by applying chained motion derivation to one or more candidates in a subblock merge candidate list according to an embodiment of the present invention. According to this method, input data associated with a current block is received in step 2010, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. One or more subblock motion candidates are derived by applying chained motion derivation to one or more candidates in a subblock merge candidate list in step 2020. One or more target subblock motion candidates are determined from the subblock merge candidate list in step 2030. The current block is encoded or decoded by using said one or more target subblock motion candidates in step 2040.

[0167] The flowcharts shown are intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0168] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0169] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0170] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of video coding, the method comprising:receiving input data associated with a current block, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;determining one or more MV (Motion Vector) lists for one or more subblocks of the current block, wherein each MV list comprises one or more motion-trajectory-based MVs (Motion Vectors) ;selecting one or more target motion-trajectory-based MVs from said one or more MV lists;deriving motion information for the current block coded in SbTMVP (Subblock Temporal Motion Vector Prediction) mode or affine mode by using said one or more target motion-trajectory-based MVs; andencoding or decoding the current block coded in the SbTMVP mode or the affine mode by using the motion information.2.The method of Claim 1, wherein for a target subblock of the current block, said one or more motion-trajectory-based MVs are determined by searching previously coded frame MVs that pass through the target subblock.3.The method of Claim 1, wherein when the current block is coded in the SbTMVP mode, subblock MVs are derived by using said one or more target motion-trajectory-based MVs as one or more motion shift candidates, and said one or more subblocks of the current block are encoded or decoded by using the subblock MVs.4.The method of Claim 3, wherein all of said one or more target motion-trajectory-based MVs are used as said one or more motion shift candidates to derive the subblock MVs for said one or more subblocks of the current block.5.The method of Claim 1, wherein when the current block uses the affine mode, an affine model is derived for the current block coded in the affine mode, and the current block is encoded or decoded by using the affine model.6.The method of Claim 5, wherein all of said one or more target motion-trajectory-based MVs are used to derive affine candidates by deriving the affine model for the current block.7.The method of Claim 6, wherein the affine candidates comprise constructed affine candidates, regression-based affine candidates, temporal affine candidates, or a combination thereof.8.The method of Claim 5, wherein multiple subblocks at specific positions in the current block are selected, and a set of motion-trajectory-based MVs is selected from corresponding MV lists of the multiple subblocks to derive the affine model as affine candidates.9.The method of Claim 8, wherein the affine candidates comprise constructed affine candidates, regression-based affine candidates, temporal affine candidates, or a combination thereof.10.The method of Claim 8, wherein derivation of the affine model comprise a combination of selected motion-trajectory-based MVs and one or more non-motion-trajectory-based MVs.11.The method of Claim 8, wherein only when said one or more target motion-trajectory-based MVs referenced from a same reference picture of L0, a same reference picture of L1, or both, the target motion-trajectory-based MVs are selected to derive the affine model as affine candidates.12.The method of Claim 5, wherein two rounds are performed to derive the affine model for the current block, and wherein N first subblocks are selected from the current block in a first round and M second subblocks are selected from the N first subblocks in a second round to derive the affine model.13.The method of Claim 1, wherein one on / off control flag is signaled at a CU level, slice level, picture level, sequence level, or a combination thereof to indicate whether said one or more target motion-trajectory-based MVs from said one or more MV lists are used to derive the motion information for the current block coded in the SbTMVP mode or the affine mode.14.The method of Claim 13, wherein whether said one or more target motion-trajectory-based MVs from said one or more MV lists are used to derive the motion information for the current block coded in the SbTMVP mode or the affine mode is according to selected reference picture indices, temporal distance between a reference picture and a current picture, quantization parameter, coded information of the current block, prediction mode, motion vectors, motion vector resolution, residual of the current block, reference samples, or a combination thereof.15.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;determine one or more MV (Motion Vector) lists for one or more subblocks of the current block, wherein each MV list comprises one or more motion-trajectory-based MVs (Motion Vectors) ;select one or more target motion-trajectory-based MVs from said one or more MV lists;derive motion information for the current block coded in SbTMVP (Subblock Temporal Motion Vector Prediction) mode or affine mode by using said one or more target motion-trajectory-based MVs; andencode or decode the current block coded in the SbTMVP mode or the affine mode by using the motion information.16.A method of video coding, the method comprising:receiving input data associated with a current block, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side, and wherein the current block is partitioned into one or more subblocks;deriving one or more subblock motion candidates by applying chained motion derivation to one or more candidates in a subblock merge candidate list;determining one or more target subblock motion candidates from the subblock merge candidate list; andencoding or decoding the current block by using said one or more target subblock motion candidates.17.The method of Claim 16, wherein if no valid subblock motion is found for a target subblock of said one or more subblocks after said applying the chained motion derivation, a pre-defined motion vector or an original subblock motion vector is used a subblock target motion vector for the target subblock.18.The method of Claim 16, wherein said applying the chained motion derivation to said one or more candidates in the subblock merge candidate list is disabled when the current block uses LIC (Local Illumination Compensation) refinement.19.The method of Claim 16, wherein said applying the chained motion derivation to said one or more candidates in the subblock merge candidate list is disabled when the current block uses BCW (Bi-prediction with CU-level Weighting) blending.20.The method of Claim 16, wherein a one on / off control flag is signalled or parsed at a CU level, slice level, picture level, sequence level or a combination thereof to indicate whether said applying the chained motion derivation to said one or more candidates in the subblock merge candidate list is applied.21.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, wherein the input data comprise pixel data for the current block to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side, and wherein the current block is partitioned into one or more subblocks;derive one or more subblock motion candidates by applying chained motion derivation to one or more candidates in a subblock merge candidate list;determine one or more target subblock motion candidates from the subblock merge candidate list; andencode or decode the current block by using said one or more target subblock motion candidates.

Citation Information

Patent Citations

  • Motion vector derivation for subblock-based template matching of subblock-based motion vector predictor

    CN118266211A

  • Motion Field Estimation Based on Motion Trajectory Derivation

    US20210144364A1

  • Searching based motion candidate derivation for sub-block motion vector prediction

    US20210243434A1

  • Motion vector prediction in video encoding and decoding

    US20220030268A1

  • Method and apparatus for regression-based affine merge mode motion vector derivation in video coding systems

    WO2023202713A1