Method and apparatus for subblock-based motion vector prediction with reordering and refining in video coding
By using a combination of multiple motion displacement candidates and co-bit reference blocks in the video encoding and decoding system, multiple SbTMVP candidates are generated and the best candidates are selected for encoding or decoding, the problem of low encoding and decoding of the existing SbTMVP is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202380072840.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-09-26
- Publication Date
- 2025-05-23
AI Technical Summary
The existing sub-block-based time motion vector prediction (SbTMVP) is not highly coded in a video encoding and decoding system, and it is difficult to effectively improve video quality.
By determining multiple motion displacement candidates at the encoder end and positioning multiple co-bit reference blocks in the co-bit image, sub-block motion information is derived for the sub-block of the current block based on the motion information of the target co-bit reference block, multiple SbTMVP candidates are generated, and the best candidate is selected for encoding or decoding through Rate-Distortion (RD) cost.
The encoding and decoding efficiency of SbTMVP is improved. Through the combination of multiple motion displacement candidates and co-bit reference blocks, the motion vector can be predicted more accurately, thereby improving video quality and encoding and decoding performance.
Smart Images

Figure CN120035998A_ABST
Abstract
Description
[0001] [Cross reference to related applications]
[0002] This application is a non-provisional application and claims priority from U.S. Provisional Patent Application No. 63 / 379,459, filed on October 14, 2022. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety. [Technical field]
[0003] The present invention relates to a video coding and decoding system using subblock-based temporal motion vector prediction (SbTMVP for short), and in particular to a technology for improving the coding and decoding efficiency of SbTMVP. [Background technology]
[0004] Background and Related Technology
[0005] Versatile Video Coding (VVC) is the latest international video codec standard developed by the Joint Video Experts Team (JVET) of the Video Coding Experts Group (VCEG) of the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology - Coded representation of immersive media - Part 3: Versatile video codec, published in February 2021. VVC is developed on the basis of its predecessor High Efficiency Video Coding (HEVC), by adding more codec tools to improve codec efficiency and handle various types of video sources including three-dimensional (3D) video signals.
[0006] Figure 1AAn exemplary adaptive interlaced / intra video codec system is shown, which includes loop processing. For intra prediction, prediction data is derived based on previously encoded video data in the current picture. For inter prediction 112, motion estimation (ME) is performed at the encoder end, and motion compensation (MC) is performed based on the results of ME to provide prediction data derived from other pictures and motion data. Switch 114 selects intra prediction 110 or inter prediction 112, and the selected prediction data is provided to adder 116 to form a prediction error, also known as a residual. The prediction error is then processed by transform (T) 118, followed by quantization (Q) 120. The transformed and quantized residual is then encoded by entropy encoder 122 for inclusion in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packaged with side information, such as motion and codec modes associated with intra and inter prediction, and other information such as parameters associated with the loop filter applied to the underlying image area. Side information related to intra prediction 110, inter prediction 112 and loop filter 130, such as Figure 1A As shown, the video data is provided to the entropy encoder 122. When the inter-frame prediction mode is used, the reference picture or pictures must also be reconstructed at the encoder end. Therefore, the transformed and quantized residual is processed by inverse quantization (IQ) 124 and inverse transform (IT) 126 to recover the residual. The residual is then added back to the prediction data 136 and the video data is reconstructed at reconstruction (REC) 128. The reconstructed video data may be stored in a reference picture buffer 134 and used for prediction of other frames.
[0007] like Figure 1A As shown, the incoming video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subjected to various impairments due to the series of processes. Therefore, a loop filter 130 is typically applied to the reconstructed video data before it is stored in a reference picture buffer 134 to improve the video quality. For example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF) may be used. The loop filter information may need to be included in the bitstream so that the decoder can correctly recover the required information. Therefore, the loop filter information is also provided to the entropy encoder 122 for inclusion in the bitstream. Figure 1A In , a loop filter 130 is applied to the reconstructed video and the reconstructed samples are then stored in a reference picture buffer 134 . Figure 1AThe system in is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Codec (HEVC) system, VP8, VP9, H.264, or VVC.
[0008] like Figure 1B The decoder shown may use the same or partially the same functional blocks as the encoder, except for transform 118 and quantization 120, because the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses an entropy decoder 140 instead of the entropy encoder 122 to decode the video bitstream into quantized transform coefficients and required codec information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). The intra-frame prediction 150 at the decoder end does not need to perform a pattern search. Instead, the decoder only needs to generate an intra-frame prediction based on the intra-frame prediction information received from the entropy decoder 140. In addition, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from the entropy decoder 140, without performing motion estimation.
[0009] According to VVC, the input picture is divided into non-overlapping square block areas, called coding tree units (CTUs), similar to HEVC. Each CTU can be divided into one or more smaller-sized coding units (SCoding Units, CUs for short). The resulting CU partition can be square or rectangular. In addition, VVC divides CTU into prediction units (PUs for short) as units for applying prediction processes (e.g., inter-frame prediction, intra-frame prediction, etc.).
[0010] The VVC standard incorporates various new coding tools to further improve the coding efficiency relative to the HEVC standard. Among the various new coding tools, some coding tools related to the present invention are described as follows.
[0011] Subblock-based Temporal Motion Vector Prediction (SbTMVP) in VVC
[0012] VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode of CUs in the current picture. The co-located pictures used by TMVP are also used for SbTMVP. The two main differences between SbTMVP and TMVP are as follows:
[0013] TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level;
[0014] TMVP obtains the temporal motion vector from the co-located block in the co-located picture (i.e., the bottom right or center block relative to the current CU), while SbTMVP applies a motion displacement before obtaining the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0015] SbTMVP process Figure 2A -B. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, check Figure 2A The spatial neighbor A1 of the current block 222 of the current picture 220 in FIG. 1 is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0).
[0016] In the second step, if Figure 2B As shown, the motion displacement 240 determined in the first step is applied (ie, added to the coordinates of the current block) to obtain sub-CU level motion information (motion vector and reference index) from the co-located image 230 . Figure 2B The example in assumes that the motion displacement is set to the motion of block A1, and the co-located block 232 in the co-located image 230 can be located based on the co-located reference sub-block A1′. Then, for each sub-CU of the current CU 222, the motion information of its corresponding block (the minimum motion grid covering the center sample) in the co-located image is used to derive the motion information of the sub-CU of the co-located CU 232. For example, the motion information of the upper left sub-block of the co-located CU 232 is used to derive the motion information of the upper left sub-block of the current CU 222. After the motion information of the co-located sub-CU is determined, it is converted into a motion vector and reference index of the current sub-CU, which is similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU. Figure 2B In the example, the arrows in each sub-block in the co-located image 230 correspond to the motion vector of the co-located sub-block (the thick arrows represent L0 MV and the thin arrows represent L1 MV). For the current image 220, the arrows in each sub-block correspond to the scaled motion vector of the current sub-block (the thick arrows represent L0 MV and the thin arrows represent L1 MV). If there is no motion information available for the co-located sub-CU (e.g., an intra-coded sub-block), the default motion is used.
[0017] In VVC, a combined sub-block based merge list containing SbTMVP candidates and affine merge candidates is used for signaling of sub-block based merge mode. SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the sub-block based merge candidate list and is followed by the affine merge candidates. The size of the sub-block based merge list is signaled in the SPS and the maximum allowed size of the sub-block based merge list in VVC is 5.
[0018] The sub-CU size used in SbTMVP is fixed to 8x8, and like the affine merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8.
[0019] The encoding process flow of the additional SbTMVP merge candidate is the same as that of other merge candidates, that is, for each CU in a P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.
[0020] Non-adjacent spatial candidates
[0021] During the development of the VVC standard, a coding tool called non-adjacent motion vector prediction (NAMVP) was proposed in JVET-L0399 (Yu Han et al., "CE4.4.6: Improvements in Merge / Skip Mode", ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 of the Joint Video Exploration Team (JVET), 12th Meeting: Macau, China, October 3-12, 2018, Document: JVET-L0399). According to the NAMVP technique, non-adjacent spatial merge candidates are inserted after the TMVP (i.e., temporal MVP) in the regular merge candidate list. The mode of spatial merge candidates is as follows: Figure 3 The distance between the non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. Figure 3 In , each small numbered box corresponds to a NAMVP candidate, and the candidates are sorted by distance (as shown by the numbers in the boxes).
[0022] Multi-channel decoder-side motion vector refinement (MP-DMVR)
[0023] Multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, MVs in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored and used for spatial and temporal motion vector prediction.
[0024] First Pass - Block-based Bilateral Matching MV Refinement
[0025] In the first pass, a refined MV is derived by applying BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), a refined MV is searched around two initial MVs (i.e., MV0 and MV1) in the reference picture lists L0 and L1 in a bi-prediction operation. Refined MVs (i.e., MV0_pass1 and MV1_pass1) are derived based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.
[0026] BM performs a local search to derive integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern looping in the horizontal search range [-sHor, sHor] and the vertical search range [-sVer, sVer], where the values of sHor and sVer are determined by the block size and the maximum values of sHor and sVer are 8.
[0027] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to eliminate the DC effect of the distortion between the reference blocks. When the bilCost of the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search is terminated. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until the end of the search range is reached.
[0028] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MVs derived after the first pass are:
[0029] MV0_pass1=MV0+deltaMV
[0030] MV1_pass1=MV1-deltaMV
[0031] Second channel - sub-block based bilateral matching MV refinement
[0032] In the second pass, a refined MV is derived by applying BM to the 16×16 grid sub-blocks. For each sub-block, a refined MV is searched around two MVs in the reference image lists L0 and L1 (e.g., MV0_pass1 and MV1_pass1), which are obtained during the first pass. Refined MVs (i.e., MV0_pass2(sbIdx2) and MV1pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.
[0033] For each sub-block, BM performs a full search to derive integer sample precision intDeltaMV. The full search has a horizontal search range of [-sHor, sHor] and a vertical search range of [-sVer, sVer], where the values of sHor and sVer are determined by the block size and the maximum values of sHor and sVer are 8.
[0034] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference sub-blocks as follows: bilCost = satdCost * costFactor. The search area (2*sHor+1)*(2*sVer+1) is divided into Figure 4 There are five diamond-shaped search areas shown, five of which are displayed in five different shades. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. Within each area, the search points are processed in a raster scan order starting from the upper left corner to the lower right corner of the area. When the minimum bilCost in the current search area is less than or equal to the threshold of sbW*sbH, the integer-pixel full search is terminated; otherwise, the integer-pixel full search continues to the next search area until all search points are checked. In addition, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than or equal to the threshold of the block area, the search process is terminated.
[0035] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). The second pass of the refined MV is then derived as:
[0036] MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2)
[0037] MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2)
[0038] The third pass - sub-block-based bidirectional optical flow MV refinement
[0039] In the third pass, refined MVs are derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive the unclipped scaled Vx and Vy starting from the refined MV of the parent sub-block in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0040] The refinement MVs of the third pass (e.g., MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as:
[0041] MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv
[0042] MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)-bioMv
[0043] Adaptive decoder-side motion vector refinement
[0044] The adaptive decoder-side motion vector refinement method is an extension of multi-pass DMVR, which includes two new merge modes to refine the MV only in the L0 or L1 direction of bi-directional prediction for merge candidates that meet the DMVR conditions. The multi-pass DMVR process is applied to the selected merge candidates to refine the motion vector, but in 1 st In the pass (i.e. PU level) DMVR, MVD0 or MVD1 is set to zero.
[0045] The merge candidates for the new merge mode are derived from spatially neighboring coded blocks, TMVPs, non-neighboring blocks, HMVPs, pairwise candidates, similar to the regular merge mode. The difference is that only candidates that satisfy the DMVR condition are added to the candidate list. The two new merge modes use the same merge candidate list. If the BM candidate list contains inherited BCW weights and the DMVR process remains unchanged, except that if the weights are unequal and bidirectional prediction is weighted with BCW weights, the distortion is calculated using MRSAD or MRSATD. The merge index is encoded as in the regular merge mode.
[0046] Template matching for MV refinement
[0047] Template matching (TM) is a decoder-side MV derivation method for refining the motion information of the current CU by finding the closest match between the template of the current CU 512 of the current image 510 (i.e., the top 514 and / or left 516 neighboring blocks) and the blocks in the reference image 520 (i.e., the same size as the template, blocks 524 and 526), as shown in FIG. Figure 5 As shown. Figure 5In the [-8, +8]-pixel search range 522 around the position 528 in the reference image 520, a better MV is searched around the initial motion 530 of the current CU 512. The template matching method in JVET-J0021 (Yi-Wen Chen et al., "Description of the SDR, HDR and 360° Video Codec Proposal by Qualcomm and Technicolor - Low and High Complexity Versions", Joint Video Coding Team (JCT-VC) ITU-T SG16 WP3 and ISO / IEC JTC 1 / SC 29 / WG11, 10th Meeting: San Diego, USA, April 10-20, 2018, Document: JVET-J0021) uses the following modifications: the search step size is determined based on the AMVR mode, and the TM can be cascaded with the bilateral matching process in the merge mode.
[0048] In AMVP mode, MVP candidates are determined based on template matching errors to select the one that achieves the minimum difference between the current block template and the reference block template. Then only this specific MVP candidate is subjected to TM for MV refinement. TM refines this MVP candidate within a [-8, +8]-pixel search range by an iterative diamond search starting from integer-pixel MVD accuracy (or 4-pixel for 4-pixel AMVR mode). As specified in the AMVR mode, the AMVP candidate may be further refined by using cross search with integer-pixel MVD accuracy (or 4-pixel for 4-pixel AMVR mode), followed by half-pixel and quarter-pixel searches. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. During the search process, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than or equal to the threshold of the block area, the search process terminates.
[0049] Table 1 - Search patterns for AMVR and merged modes with AMVR
[0050]
[0051] In merge mode, a similar search method is applied to the merge candidates indicated by the merge index. As shown in Table 1, TM may be performed up to 1 / 8-pixel MVD accuracy, or skip those beyond half-pixel MVD accuracy depending on whether the merge motion information uses an alternative interpolation filter (for AMVR in half-pixel mode). In addition, when TM mode is enabled, template matching can be performed as an independent process or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enabling condition check.
[0052] Adaptive Re-ranking Merging Candidate and Template Matching (ARMC-TM)
[0053] The merge candidates are adaptively reordered based on their cost evaluated using Template Matching (TM). The reordering method can be applied to regular merge mode, Template Matching (TM) merge mode, and Affine merge mode (excluding SbTMVP candidates). For TM merge mode, the merge candidates are reordered before the refinement process.
[0054] After constructing the merge candidate list, the merge candidates are divided into multiple subgroups. The subgroup size for the regular merge mode and the TM merge mode is set to 5. The subgroup size for the affine merge mode is set to 3. The merge candidates in each subgroup are reordered in ascending order according to the cost value based on template matching. For ARMC-TM, candidates in a subgroup are skipped if the subgroup meets the following 2 conditions: (1) the subgroup is the last subgroup; (2) the subgroup is not the first subgroup. For simplicity, the merge candidates of the last but not the first subgroup are not reordered.
[0055] The template matching cost of the merge candidate is measured as the sum of absolute differences (SAD) between the template samples of the current block and their corresponding reference samples. The template consists of a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the merge candidate.
[0056] When the merge candidate uses bidirectional prediction, the reference sample of the template of the merge candidate is also generated by bidirectional prediction, such as Figure 6 As shown. Figure 6 , block 612 corresponds to a current block in current picture 610, and blocks 622 and 632 correspond to reference blocks in reference pictures 620 and 630 in list 0 and list 1, respectively. Templates 614 and 616 are used for current block 612, templates 624 and 626 are used for reference block 622, and templates 634 and 636 are used for reference block 632. Motion vectors 640, 642, and 644 are merge candidates in list 0, and motion vectors 650, 652, and 654 are merge candidates in list 1.
[0057] For the sub-block based merge candidate, the sub-block size is equal to Wsub×Hsub, the above template includes several sub-templates of size Wsub×1, and the left template includes several sub-templates of size 1×Hsub. Figure 7 As shown in , the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference samples of each sub-template. Figure 7, block 712 corresponds to the current block in the current picture 710, and block 722 corresponds to the co-located block in the reference picture 720. Each small square in the current block and the co-located block corresponds to a sub-block. The dot-filled areas on the left and top of the current block correspond to the template of the current block. The boundary sub-blocks are labeled from A to G. The arrows associated with each sub-block correspond to the motion vector of the sub-block. The reference sub-blocks (labeled Aref to Gref) are located according to the motion vectors associated with the boundary sub-blocks.
[0058] Merge Mode with MVD (MMVD)
[0059] In addition to the merge mode, where the implicitly derived motion information is directly used for prediction sample generation for the current CU, a merge mode with motion vector difference (MMVD) is introduced. The MMVD flag is signaled immediately after the regular merge flag is sent to specify whether the MMVD mode is used for the CU.
[0060] In MMVD, once a merge candidate is selected, it is further refined by signaled MVD information. Further information includes a merge candidate flag, an index specifying the magnitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected to be used as the MV basis. The MMVD candidate flag is signaled to specify whether the first or second merge candidate is used.
[0061] The distance index specifies the motion magnitude information and indicates the predefined offsets from the starting points (812 and 822) to the L0 reference block 810 and the L1 reference block 820, as shown in FIG. Figure 8 As shown. Figure 8 In , the offset is added to the horizontal or vertical component of the starting MV, where small circles of different styles correspond to different offsets from the center. The relationship between the distance index and the predefined offsets is specified in Table 2.
[0062] Table 2 - Relationship between distance index and predefined offset
[0063] Distance Index 0 1 2 3 4 5 6 7 Pixel distance 1 / 4-pixel 1 / 2-pixel 1-Pixel 2-Pixel 4-Pixel 8-Pixel 16-pixel 32-pixel
[0064] The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions as shown below. The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions as shown in Table 3. It should be noted that the meaning of the MVD symbol may change depending on the information of the starting MV. When the starting MV is an unpredicted MV or a dual-predicted MV, and both lists point to the same side of the current picture (that is, the POCs of both references are greater than the POC of the current picture, or are less than the POC of the current picture), the symbol in Table 3 specifies the sign of the MV offset added to the starting MV. When the starting MV is a dual-predicted MV, and the two MVs point to different sides of the current picture (that is, the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), and the POC difference in list 0 is greater than the POC difference in list 1, the symbol in Table 3 specifies the sign of the MV offset added to the list 0 MV component, and the sign of the list 1 MV has the opposite value. Otherwise, if the POC difference in List 1 is greater than List 0, then the sign in Table 3 specifies the sign of the MV offset added to the List 1 MV component, while the sign of the List 0 MV has the opposite value.
[0065] The MVD is scaled according to the POC difference in each direction. If the POC difference in the two lists is the same, no scaling is required. Otherwise, if the POC difference between the L0 (i.e. L0) reference picture and the current picture is greater than the POC difference between the L1 (i.e. L1) picture and the current picture, the MVD of list 1 is scaled according to the ratio of td and tb, where td corresponds to the POC difference between the L0 reference picture and the current picture, and tb corresponds to the POC difference between the L1 reference picture and the current picture. If the POC difference between the L1 reference picture and the current picture is greater than the POC difference between the L0 reference picture and the current picture, the MVD of list 0 is scaled in a similar manner. If the starting MV is uni-predicted, the MVD is added to the available MV.
[0066] Table 3 - MV offset symbols specified by direction index
[0067]
[0068]
[0069] In the present invention, a solution to further improve the performance of SbTMVP is disclosed. [Summary of the invention]
[0070] The present invention discloses a video encoding and decoding method and device using subblock-based temporal motion vector prediction (SbTMVP for short). According to the method, input data related to a current block is received, wherein the input data includes pixel data of the current block for encoding at the encoder end or encoding data related to the current block decoded at the decoder end. One or more motion displacement candidates are determined based on one or more spatial neighboring blocks of the current block. Two or more co-located reference blocks located in a co-located image are determined based on the one or more motion displacement candidates and the one or more spatial neighboring blocks of the current block. A target co-located reference block is determined from the two or more co-located reference blocks. Sub-block motion information is derived for a sub-block of the current block based on target motion information of a corresponding sub-block of the target co-located reference block. A SbTMVP candidate is generated for the current block based on the sub-block motion information of the sub-block of the current block. The current block is encoded or decoded using a motion prediction set including the SbTMVP candidate.
[0071] In one embodiment, the two or more co-located reference blocks located in the co-located image are determined by locating a base co-located reference block based on the position of the target spatial neighboring block of the current block and the target motion displacement candidate associated with the target spatial neighboring block of the current block, and locating one or more additional co-located reference blocks by adding one or more additional motion displacements to the base co-located reference block. In one embodiment, the target spatial neighboring block of the current block corresponds to the lower left neighboring block, and the three additional co-located reference blocks are located at the upper side, the left side and the bottom side of the base co-located reference block, respectively. In one embodiment, the target co-located reference block is determined from the two or more co-located reference blocks based on the rate-distortion cost associated with the two or more co-located reference blocks.
[0072] In one embodiment, a first syntax is signaled or parsed to indicate whether the two or more co-located reference blocks are used. In one embodiment, when the first syntax indicates that the two or more co-located reference blocks are used, a second syntax is signaled or parsed to indicate a selected target co-located reference block.
[0073] In one embodiment, when two or more motion displacement candidates are determined based on two or more spatial neighboring blocks of a current block, the two or more co-located reference blocks in the co-located image include one base co-located reference block from each of the two or more motion displacement candidates, and the one base co-located reference block is located according to the position of each spatial neighboring block of the current block and the target motion displacement candidate associated with the one base co-located reference block. In one embodiment, a first syntax is signaled or parsed to indicate whether the two or more motion displacement candidates are used. In one embodiment, when the first syntax indicates that the two or more motion displacement candidates are used, a second syntax is signaled or parsed to indicate which base co-located reference block is selected. In another embodiment, the first syntax is signaled or parsed only when affine multiple mode vector difference (MMVD) of the current block is disabled.
[0074] In one embodiment, the corresponding SbTMVP candidates associated with the two or more co-located reference blocks are reordered according to Adaptive Reordering of Merge Candidates with Template Matching (ARMC-TM). In one embodiment, among the reordered corresponding SbTMVP candidates, the N best candidates are used for further rate-distortion cost evaluation, where N is less than or equal to the total number of corresponding SbTMVP candidates. In one embodiment, a set of indexes is used to indicate the N best candidates, and a shortened codeword is used to signal a smaller index value. In one embodiment, ARMC-TM is performed using one or more templates of the current block and one or more corresponding templates of the target co-located reference block. In another embodiment, ARMC-TM is performed using one or more templates of the current block and one or more corresponding templates of the target co-located reference block, wherein the sub-blocks of the target co-located reference block are positioned based on sub-block motion.
Brief Description of the Drawings
[0075] Figure 1A An exemplary adaptive interlaced / intra video encoding and decoding system including loop processing is described.
[0076] Figure 1B Explained Figure 1A The corresponding decoder of the encoder in .
[0077] Figure 2A An example of sub-block based temporal motion vector prediction (SbTMVP) in VVC is illustrated, where spatial neighboring blocks are checked to determine the availability of motion information.
[0078] Figure 2B An example of SbTMVP is illustrated for deriving sub-CU motion fields by applying motion displacements from spatial neighbors and scaling motion information from corresponding co-located sub-blocks.
[0079] Figure 3 Exemplary patterns of non-adjacent spatial merging candidates are illustrated.
[0080] Figure 4 Illustrated are five diamond-shaped search areas for multi-pass decoder-side motion vector refinement.
[0081] Figure 5 An example of template matching for refining an initial MV by searching in a region around the initial MV is illustrated.
[0082] Figure 6 An example of templates for a current block and corresponding reference blocks for measuring the matching cost associated with a merge candidate is illustrated.
[0083] Figure 7 The offset distances of the L0 reference block and the L1 reference block in the horizontal and vertical directions according to the MMVD are illustrated.
[0084] Figure 8 An example of a merge mode with MVD (MMVD) is illustrated, where the distance index specifies the motion magnitude information and indicates a predefined offset from the starting point of the L0 reference block and the L1 reference block.
[0085] Fig. 9A An example of SbTMVP with multiple motion vector shifts (sbTMVP with Mmvs) according to an embodiment of the present invention is described, where the A1-based motion shift is used to locate A1′ in the co-located image, and additional candidates (T1, L1, and B1) are located by adding additional shifts.
[0086] Fig. 9B An example of a co-located reference block associated with candidate L1 is illustrated.
[0087] Fig.10 An example is illustrated where multiple co-located reference blocks are derived based on two different motion shifts relative to the A1 and TR1 neighboring blocks.
[0088] Fig.11 An example of a CU-based template for calculating the template matching cost is illustrated.
[0089] Fig.12 A flow chart of an exemplary video encoding and decoding system utilizing SbTMVP with multiple motion displacements according to one embodiment of the present invention is illustrated. [Specific implementation method]
[0090] It will be readily understood that the components of the present invention, as generally described and depicted in the figures herein, can be arranged and designed in a variety of different configurations. Accordingly, the following more detailed description of embodiments of the systems and methods of the present invention, as shown in the figures, is not intended to limit the scope of the present invention, as claimed, but merely represents selected embodiments of the present invention. References in this specification to "one embodiment", "an embodiment", or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places in this specification are not necessarily all referring to the same embodiment.
[0091] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. However, those skilled in the relevant art will recognize that the present invention can be practiced without one or more of the specific details, or using other methods, components, etc. In other instances, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the present invention. The embodiments of the present invention will be best understood by reference to the drawings, in which like parts are designated by like numerals. The following description is only by way of example and simply illustrates certain selected embodiments of the devices and methods consistent with the invention claimed herein.
[0092] To improve the encoding and decoding efficiency of SbTMVP for Skip, Merge, Direct, Intra mode, Inter mode, and / or IBC mode, an SbTMVP method related to multiple motion displacements is disclosed.
[0093] In one embodiment, the motion vectors of non-adjacent spatially adjacent blocks are used as motion displacements.
[0094] In one embodiment, the derivation of the motion displacement is related to TM and / or ARMC-TM and has the following variations:
[0095] Refine the motion displacement through TM and the CU template.
[0096] Reorder two or more motion vectors of adjacent and / or non-adjacent spatially adjacent blocks through ARMC-TM, and then select the first one or more as the motion displacement.
[0097] Refine two or more motion vectors of adjacent and / or non-adjacent spatially adjacent blocks through TM, then reorder them through ARMC-TM, and then select the first one or more motion vectors as the motion displacement.
[0098] Two or more motion vectors of adjacent and / or non-adjacent spatially neighboring blocks are reordered by ARMC-TM, and then the first one or more motion vectors are selected as motion displacements, which are then refined by TM.
[0099] In the process of TM or ARMC-TM, if considering the motion vector as motion displacement would make the SbTMVP candidate unavailable (ie, default motion is unavailable), the motion vector cost of TM or ARMC-TM is set to a large value or the motion vector is skipped.
[0100] In one embodiment, the default motion is further refined by TM and / or BM, with the following variants:
[0101] Refine the default motion through TM and CU templates.
[0102] Default motion is refined by BM.
[0103] The default motion is refined via TM and then BM.
[0104] The default motion is refined by BM and then TM.
[0105] In order to further improve the coding efficiency, in one embodiment, a SbTMVP with multiple motion vector displacements (sbTMVP with Mmvs) is proposed as a new mode. In other words, multiple sbTMVPs with different motion displacement candidates are tried, and the best one is selected according to the RD cost. For example, Fig. 9A As shown, a current picture 910 with a current CU 912 and a co-located picture 920, A1 is the lower left neighboring sub-block of the current block 912. The motion vector 930 points to A1′ in the co-located picture 920, and the co-located CU 922 can be located based on A1. In one embodiment, the first sbTMVP candidate with Mmvs derived based on the temporal motion of A1 is candidate 0 (i.e., A1′). The second sbTMVP candidate with Mmvs based on the temporal motion of A1 and the motion displacement of 4 pixels to the left is candidate 1 (i.e., L1). In other words, sub-block L1 is located by moving A1′ to the left by 4 pixels. As shown in FIG. Fig. 9B As shown, the corresponding co-located CU 942 can be located according to L1. The third sbTMVP candidate with Mmvs based on the temporal motion of A1 and the motion displacement of 4 pixels upward is candidate 2 (i.e., T1). The fourth sbTMVP candidate with Mmvs based on the temporal motion of A1 and the motion displacement of 4 pixels downward is candidate 3 (i.e., B1). In other words, four sbTMVPs are generated by adding additional motion displacement to the base motion displacement 930, so that in addition to the base co-located reference block (i.e., A1′), three additional co-located reference blocks (i.e., Fig. 9AThe best sbTMVP candidate with Mmvs is determined based on the Rate-Distortion (RD) cost. A flag (i.e., sbTMVP_mmvd_flag) is signaled to indicate the switch of sbTMVP with Mmvs. If sbTMVP_mmvd_flag is equal to 1, sbTMVP_merge_idx is signaled to indicate the best candidate in the sbTMVP with Mmvs candidate list.
[0106] In another embodiment, if sbTMVP_mmvd_flag is equal to 1, stmvp_base_idx may be further signaled to indicate a different initial candidate (base candidate). Fig.10 As shown, with a current picture 1010 and a co-located picture 1020, multiple base candidates are applied, and sbTMVP_base_idx is signaled to indicate the best candidate derived from base candidate 0 (i.e., A1) or base candidate 1 (i.e., TR1). If the best candidate comes from base 1 (i.e., TR1), sbTMVP_merge_idx is used to indicate that the best candidate is TT1, TB1, or TL1. If the best candidate comes from base 0 (i.e., A1), sbTMVP_merge_idx is used to indicate that the best candidate is T1, B1, or L1.
[0107] In another embodiment, if sbTMVP_mmvd_flag is equal to 1, stmvp_base_idx may be further signaled to indicate different initial candidates (ie, base candidates). For example, base candidate 0 is the motion of position A1 in collocated picture 0, and base candidate 1 is the motion of position A1 in collocated picture 1.
[0108] In another embodiment, if sbtmvp_mmvd_flag is equal to 1, stmvp_base_idx may be further signaled to indicate a different initial candidate (i.e., base candidate). For example, base candidate 0 is the motion from the A1 position of co-located image 0, and base candidate 1 is the motion from the A1 position of co-located image 1. If neither the L0 nor the L1 motion of the A1 position is from co-located image 0 or co-located image 1, motion scaling techniques may be used for the motion from L0 or from L1.
[0109] In another embodiment, a sbTMVP with multiple sub-block motion displacements (sbTMVP with sMmvs) is proposed as a new mode. This means that more than one motion displacement is added to the sub-block motion of each sbTMVP candidate to generate multiple sbTMVP candidates with sMmvs. For example, an sbTMVP candidate is generated by using the motion from A1. After that, a set of sub-block motions (i.e., an initial motion group) is derived, and multiple sub-block motion displacements are added to each sub-block motion of the initial motion group to generate more sbTMVP candidates with sMmvs. The best sbTMVP candidate with sMmvs will be selected based on the RD cost of each candidate. A flag (i.e., sbtmvp_mmvd_flag) is signaled to indicate the switch of sbTMVP with sMmvs. If sbtmvp_mmvd_flag is equal to 1, sbtmvp_merge_idx is signaled to indicate the best candidate in the sbTMVP candidate list with sMmvs.
[0110] To make the relevant syntax signaling more efficient, ARMC-TM is performed on the candidate list of sbTMVP with sMmvs or sbTMVP with Mmvs.
[0111] In another embodiment, the candidate list of sbTMVP with sMmvs or sbTMVP with Mmvs is first reordered according to ARMC-TM, and then only the best N candidates after reordering are used to calculate the RD cost for comparison. N is an integer less than or equal to the number of candidates in the candidate list. By doing so, more promising candidates will be assigned a smaller index, and the smaller index can be signaled by a shorter codeword.
[0112] Regarding TM cost calculation, the following two methods are proposed:
[0113] The TM cost is calculated by co-locating the template of the current block on the image. Fig.11 Templates for candidates A1' and L3 are shown. The TM cost is calculated based on the CU templates (1122 for A1' and 1124 for L3), where the current picture 1110 and the co-located image 1120 are shown.
[0114] TM cost is calculated by using something like Figure 7 The sub-block motion calculation of the shown technique.
[0115] In another embodiment, the availability of sbTMVP with sMmvs or sbTMVP with Mmvs candidates is checked before RD calculation. Invalid candidates (if the center position of the candidate in the current block on the co-located image is an internal mode or an IBC mode, the candidate is considered invalid) are removed, and a final candidate list is generated, which does not include invalid candidates. By doing so, the index of a more promising candidate can be signaled using a shortened codeword.
[0116] Related grammar design:
[0117] In order to improve the encoding and decoding efficiency of sbTMVP with sMmvs and sbTMVP with Mmvs mode, several syntax designs are proposed. The following methods take sbTMVP with Mmvs mode as an example, and they can all be applied to sbTMVP with sMmvs mode.
[0118] In one embodiment, a flag (e.g., sbtmvp_mmvd_flag) is signaled to indicate the switch of sbTMVP with Mmvs. If sbtmvp_mmvd_flag is equal to 1, sbtmvp_merge_idx is signaled to indicate the best candidate in the sbTMVP candidate list with Mmvs. If multiple base candidates are used, sbtmvp_base_idx is signaled, and they are signaled after the affine mmvd-related syntax. The following is an exemplary syntax design.
[0119] Grammar design 1:
[0120]
[0121] In another embodiment, SbTMVP with Mmvs can be enabled only when sbTMVP appears in the affine merge candidate list. Therefore, affine with mmvd can be enabled only when sbTMVP does not appear in the affine merge candidate list. This means that SbTMVP with Mmvs can share the same syntax data with affine with mmvd. By checking the affine merge candidate list, the decoder can know whether SbTMVP with Mmvs or affine with mmvd is enabled. If sbTMVP appears in the affine merge candidate list and subblok_mmvd_flag is equal to 1, SbTMVP with Mmvs is enabled. The following is an exemplary syntax design.
[0122] Grammar design 2:
[0123]
[0124]
[0125] In syntax design 1, if all sbTMVP candidates with Mmvs are invalid (i.e., all co-located images are internal images), sbtmvp_mmvd_flag needs to be signaled, and it must be 0. In order to make syntax signaling more efficient, in another embodiment, if all sbTMVP candidates with Mmvs are invalid, affine and mmvd can apply 2 base candidates without signaling affine_base_idx. The following is the syntax design. In the encoder, when at least one sbTMVP candidate with Mmvs is valid, the number of base candidates for affine and mmvd is equal to 1. When no sbTMVP candidate with Mmvs is valid, the number of base candidates for affine and mmvd is equal to 2. The following are exemplary semantic definitions of affine_mmvd_flag, affine_merge_idx, sbtmvp_mmvd_flag, and sbtmvp_merge_idx.
[0126] Grammar Design 3:
[0127]
[0128] affine_mmvd_flag[x0][y0] equal to 1 specifies that affine and mmvd are enabled for the current codec unit. The array index x0, y0 specifies the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the picture.
[0129] affine_merge_idx[x0][y0] specifies the merge candidate index in the affine and mmvd candidate lists, where x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the image.
[0130] sbtmvp_mmvd_flag[x0][y0] equal to 1 specifies that sbtmvp and Mmvs are enabled or base candidate 1 is used in affine and mmvd for the current codec unit. Array index x0, y0 specifies the position (x0, y0) of the top-left luma sample of the codec block under consideration relative to the top-left luma sample of the image. If none of the sbTMVP and Mmvs candidates are valid, affineBaseIdx is equal to sbtmvp_mmvd_flag; otherwise, affineBaseIdx is equal to 0. If none of the sbTMVP and Mmvs candidates are valid and sbtmvp_mmvd_flag is equal to 1, affineEnableFlag is equal to 1; otherwise, affineEnableFlag is equal to affine_mmvd_flag.
[0131] sbtmvp_merge_idx[x0][y0] specifies the merge candidate index in the SbTMVP and Mmvs candidate list or the affine and mmvd candidate list, where x0, y0 specify the position (x0, y0) of the upper left corner luma sample of the considered codec block relative to the upper left corner luma sample of the image. If no sbTMVP and Mmvs candidate is valid, affineMmvdMergeIdx is equal to sbtmvp_merge_idx and sbtmvpMmvdMergeIdx is equal to 0; otherwise, affineMmvdMergeIdx is equal to affine_merge_idx and sbtmvpMmvdMergeIdx is equal to stmvp_merge_idx.
[0132] Grammar Design 4:
[0133] In one embodiment, the motion vector displacement of the sbTMVP may be signaled in a manner similar to the MMVD scheme. In particular, after a merge candidate is selected via a merge index, it is further refined via the signaled MVDs information. Further information includes, but is not limited to, an index to specify the motion magnitude and an index to indicate the motion direction. It is noteworthy that the motion magnitude set and the motion direction set may contain any predefined values and are not limited to the sets used in the MMVD design currently used in VVC, as shown in Tables 4 and 5, respectively.
[0134] Table 4 - MmvdDistance[x0][y0] Specification Based on mmvd_distance_idx[x0][y0]
[0135]
[0136] Table 5 - MmvdSign[x0][y0] based on mmvd_direction_idx[x0][y0]
[0137] mmvddirection_idx[x0][y0] MmvdSign[x0][y0][0] MmvdSign[x0][y0][1] 0 +1 0 1 -1 0 2 0 +1 3 0 -1
[0138] Based on the proposed scheme, in another embodiment, the merge candidate list can be derived by the same merge candidate list construction process as the conventional merge candidate list construction process.Since the merge candidate can be a bidirectional merge candidate, an additional syntax element is signaled to indicate which direction is used as the motion vector displacement.
[0139] Based on the disclosed scheme, in another embodiment, the merge candidate list can be derived only by inserting the unidirectional motion vector into the candidate list. Since the merge candidate can only be a unidirectional merge candidate, the signaled MVDs information is directly added to the selected unidirectional merge candidate to derive the motion vector displacement of the sbTMVP.
[0140] Any of the SbTMVP methods described above can be implemented in an encoder and / or decoder. For example, any of the proposed SbTMVP multiple motion displacement methods can be implemented in an encoder's interactive encoding and decoding module and / or a merge / AMVP candidate derivation module (e.g., Figure 1A 112), or the motion compensation module of the decoder (e.g., Figure 1B 152) and / or the merge / AMVP candidate derivation module. Alternatively, any of the proposed methods may be implemented as a circuit connected to the inter-coding module and / or the merge / AMVP candidate derivation module of the encoder and / or the motion compensation module and / or the merge / AMVP candidate derivation module of the decoder. Although Inter-Pred.112 and MC 152 are shown as separate processing units supporting the SbTMVP method, they may correspond to executable software or firmware code stored on a medium, such as a hard disk or flash memory, for a CPU (central processing unit) or a programmable device (e.g., a DSP (digital signal processor) or an FPGA (field programmable gate array)).
[0141] Fig.12A flowchart of an exemplary video encoding and decoding system is shown, which utilizes SbTMVP with multiple motion displacements according to an embodiment of the present invention. The steps shown in the flowchart can be implemented by program code executed on one or more processors (e.g., one or more CPUs) at the encoder end. The steps shown in the flowchart can also be implemented based on hardware, for example, one or more electronic devices or processors are arranged to perform the steps in the flowchart. According to this method, in step 1210, input data related to a current block is received, wherein the input data includes pixel data of the current block for encoding at the encoder end or encoded data related to the current block for decoding at the decoder end. In step 1220, one or more motion displacement candidates are determined based on one or more spatial neighboring blocks of the current block. In step 1230, two or more co-located reference blocks located in a co-located image are respectively determined based on the one or more motion displacement candidates and the one or more spatial neighboring blocks of the current block. In step 1240, a target co-located reference block is determined from the two or more co-located reference blocks. In step 1250, sub-block motion information is derived for a sub-block of the current block based on the target motion information of the corresponding sub-block of the target co-located reference block. In step 1260, a SbTMVP (sub-block based temporal motion vector prediction) candidate is generated for the current block based on the sub-block motion information of the sub-blocks of the current block. In step 1270, the current block is encoded or decoded using a motion prediction set including the SbTMVP candidate.
[0142] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A skilled person can modify each step, rearrange the steps, split the steps, or combine the steps to practice the present invention without departing from the spirit of the present invention. In this article, specific syntax and semantics have been used to illustrate examples of implementing the present invention. A skilled person can practice the present invention by replacing the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0143] The above description is intended to enable persons of ordinary skill to practice the invention in the context of a particular application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not limited to the specific embodiments shown and described, but is intended to be given the widest scope consistent with the principles and novel features disclosed herein. In the above detailed description, various specific details are shown in order to provide a thorough understanding of the present invention. However, it will be understood by those skilled in the art that the present invention can be practiced.
[0144] Embodiments of the present invention as described above may be implemented in various hardware, software code, or a combination of both. For example, an embodiment of the present invention may be one or more circuits integrated into a video compression chip, or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code executed on a digital signal processor (DSP) to perform the processing described herein. The present invention may also relate to multiple functions executed by a computer processor, a digital signal processor, a microprocessor, or a field programmable gate array (FPGA). These processors may be configured according to the invention to perform specific tasks by executing machine-readable software code or firmware code that defines the specific methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, the different code formats, styles, and languages of the software code, as well as other configuration codes in a manner that conforms to the invention task, do not deviate from the spirit and scope of the invention.
[0145] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples should be considered illustrative rather than restrictive in all respects. Therefore, the scope of the present invention should be indicated by the appended claims rather than the above description. All changes that come within the meaning and scope of the equivalence of the claims should be embraced within their scope.
Claims
1. A video encoding and decoding method, the method include: Receiving input data related to a current block, wherein the input data includes pixel data of the current block for encoding at an encoder end or encoded data related to the current block for decoding at a decoder end; Determine one or more motion displacement candidates based on one or more spatial neighboring blocks of the current block; Determine two or more co-located reference blocks located in the co-located image based on the one or more motion displacement candidates and the one or more spatial neighboring blocks of the current block; Determine a target co-located reference block from the two or more co-located reference blocks; deriving sub-block motion information for a sub-block of the current block based on target motion information of a corresponding sub-block of the target co-located reference block; Generating a sub-block based temporal motion vector prediction (SbTMVP) candidate for the current block based on sub-block motion information of the sub-block of the current block; The current block is encoded or decoded using a motion prediction set including the SbTMVP candidate.
2. The method of claim 1 , wherein the two or more co-located reference blocks in the co-located image are determined by locating a base co-located reference block based on a position of a target spatial neighboring block of the current block and a target motion displacement candidate associated with the target spatial neighboring block of the current block, and locating one or more additional co-located reference blocks by adding one or more additional motion displacements to the base co-located reference block.
3. The method of claim 2, wherein the target spatial neighboring block of the current block corresponds to a lower left neighboring block, and the three additional co-located reference blocks are respectively located at an upper side, a left side, and a bottom side of the base co-located reference block.
4. The method of claim 2, wherein the target co-located reference block is determined from the two or more co-located reference blocks based on rate-distortion costs associated with the two or more co-located reference blocks.
5. The method of claim 1, wherein a first syntax is signaled or parsed to indicate whether the two or more co-located reference blocks are used.
6. The method of claim 5, wherein when the first syntax indicates that the two or more co-located reference blocks are used, a second syntax is signaled or parsed to indicate that the target co-located reference block is selected.
7. The method of claim 1, wherein when two or more motion displacement candidates are determined based on two or more spatial neighboring blocks of the current block, the two or more co-located reference blocks in the co-located image include a base co-located reference block from each of the two or more motion displacement candidates, and the one base co-located reference block is positioned according to the position of each spatial neighboring block of the current block and the target motion displacement candidate associated with the one base co-located reference block.
8. The method of claim 7, wherein a first syntax is signaled or parsed to indicate whether to use the two or more motion displacement candidates.
9. The method of claim 8, wherein when the first syntax indicates that the two or more motion displacement candidates are used, a second syntax is signaled or parsed to indicate which base co-located reference block is selected.
10. The method of claim 8, wherein the first syntax is signaled or parsed only when affine MMVD of the current block is disabled.
11. The method of claim 1, wherein the corresponding SbTMVP candidates associated with the two or more co-located reference blocks are reordered according to Adaptive Reorder Merging Candidate and Template Matching (ARMC-TM).
12. The method of claim 11, wherein after the corresponding SbTMVP candidates are re-ranked, N best candidates are used for further rate-distortion cost evaluation, wherein N is less than or equal to a total number of the corresponding SbTMVP candidates.
13. The method of claim 12, wherein a set of indices is used to indicate the N best candidates, and shortened codewords are used to signal smaller index values.
14. The method of claim 11, wherein the ARMC-TM is performed using one or more templates of the current block and one or more corresponding templates of the target co-located reference block.
15. The method of claim 11, wherein the ARMC-TM is performed using one or more templates of the current block and one or more corresponding templates of sub-blocks associated with the target co-located reference block, wherein the sub-blocks of the target co-located reference block are positioned based on sub-block motion.
16. An apparatus for video encoding and decoding, the apparatus comprising one or more electronic circuits or processors configured to: Receiving input data related to a current block, wherein the input data includes pixel data of the current block for encoding at an encoder end or encoded data related to the current block for decoding at a decoder end; Determine one or more motion displacement candidates based on one or more spatial neighboring blocks of the current block; Determine two or more corresponding reference blocks located in the same-position image based on the one or more motion displacement candidates and one or more spatially adjacent blocks of the current block; Determine a target corresponding reference block from the two or more corresponding reference blocks; deriving sub-block motion information of the sub-block of the current block based on target motion information of the corresponding sub-block of the target corresponding reference block; generating a sub-block based temporal motion vector prediction (SbTMVP) candidate for the sub-block of the current block based on the sub-block motion information; The current block is encoded or decoded using a motion prediction set including the SbTMVP candidate.