Methods and apparatus of BI-prediction candidates for auto-relocated block vector prediction or chained motion vector prediction
By generating combined prediction candidates through chained motion and block vectors, the method addresses inefficiencies in video coding systems, enhancing prediction accuracy and reducing complexity in bi-prediction scenarios.
Patent Information
- Application Number
- PCT/CN2025/072738
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-24
AI Technical Summary
Existing video coding systems face inefficiencies in generating prediction candidates for motion vectors and block vectors, particularly in handling bi-prediction scenarios, which affect coding performance and complexity.
The method involves deriving combined prediction candidates through chained motion vectors or block vectors, recursively combining initial candidates with other vectors to generate bi-prediction and uni-prediction candidates, and modifying or invalidating candidates that exceed available reference regions.
This approach enhances coding efficiency by improving prediction accuracy and reducing complexity in video coding systems, particularly for Intra Block Copy and Inter Prediction modes.
Smart Images

Figure CN2025072738_24072025_PF_FP_ABST
Abstract
Description
METHODS AND APPARATUS OF BI-PREDICTION CANDIDATES FOR AUTO-RELOCATED BLOCK VECTOR PREDICTION OR CHAINED MOTION VECTOR PREDICTIONCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 622,095, filed on January 18, 2024. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION
[0002] The present invention relates to video coding system. In particular, the present invention relates to generating combined prediction candidates based on chained motion vectors or block vectors in a video coding system.BACKGROUND
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0006] The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
[0007] Intra Block Copy
[0008] Intra block copy (IBC) is a tool adopted in HEVC extensions on SCC (Screen Content Coding) . It is well known that it significantly improves the coding efficiency of screen content materials. Since IBC mode is implemented as a block level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, a block vector is used to indicate the displacement from the current block to a reference block, which is already reconstructed inside the current picture. The luma block vector of an IBC-coded CU is in integer precision. The chroma block vector is rounded to integer precision as well. When combined with AMVR (Adaptive Motion Vector Resolution) , the IBC mode can switch between 1-pel and 4-pel motion vector precisions. An IBC-coded CU is treated as the third prediction mode other than intra or inter prediction modes. The IBC mode is applicable to the CUs with both width and height smaller than or equal to 64 luma samples.
[0009] At the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check for blocks with either width or height no larger than 16 luma samples. For non-merge mode, the block vector search is performed using hash-based search first. If hash search does not return a valid candidate, block matching based local search will be performed.
[0010] In the hash-based search, hash key matching (32-bit CRC) between the current block and a reference block is extended to all allowed block sizes. The hash key calculation for every position in the current picture is based on 4x4 subblocks. For the current block of a larger size, a hash key is determined to match that of the reference block when all the hash keys of all 4×4 subblocks match the hash keys in the corresponding reference locations. If hash keys of multiple reference blocks are found to match that of the current block, the block vector costs of each matched reference are calculated and the one with the minimum cost is selected.
[0011] In block matching search, the search range is set to cover both the previous and current CTUs.
[0012] At CU level, IBC mode is signalled with a flag and it can be signalled as IBC AMVP (Advanced Motion Vector Prediction) mode or IBC skip / merge mode as follows: – IBC skip / merge mode: a merge candidate index is used to indicate which of the block vectors in the list from neighbouring candidate IBC coded blocks is used to predict the current block. The merge list consists of spatial, HMVP (History based Motion Vector Prediction) , and pairwise candidates. – IBC AMVP mode: block vector difference is coded in the same way as a motion vector difference. The block vector prediction method uses two candidates as predictors, one from left neighbour and one from above neighbour (if IBC coded) . When either neighbour is not available, a default block vector will be used as a predictor. A flag is signalled to indicate the block vector predictor index.
[0013] IBC Reference Region
[0014] To reduce memory consumption and decoder complexity, the IBC in VVC allows only the reconstructed portion of the predefined area including the region of current CTU and some region of the left CTU. Fig. 2 illustrates the reference region of IBC Mode, where each block represents 64x64 luma sample unit. Depending on the location of the current coded CU within the current CTU, the following applies: – If the current block falls into the top-left 64x64 block of the current CTU (case 210 in Fig. 2) , then in addition to the already reconstructed samples in the current CTU, it can also refer to the reference samples in the top-right, bottom-left, and bottom-right 64x64 blocks of the left CTU, using current picture referencing (CPR) mode. (More details of CPR can be found in JVET-T2002 (Jianle Chen, et. al., “Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11) ” , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 20th Meeting, by teleconference, 7 –16 October 2020, Document: JVET-T2002) ) . The current block can also refer to the reference samples in the top-right, bottom-left, and bottom-left 64x64 block of the left CTU and the reference samples in the top-right 64x64 block of the current CTU, using CPR mode. – If the current block falls into the top-right 64x64 block of the current CTU (case 220 in Fig. 2) , then in addition to the already reconstructed samples in the current CTU, if luma location (0, 64) relative to the current CTU has not yet been reconstructed, the current block can also refer to the reference samples in the bottom-left 64x64 block and bottom-right 64x64 block of the left CTU, using CPR mode; otherwise, the current block can also refer to reference samples in bottom-right 64x64 block of the left CTU. – If the current block falls into the bottom-left 64x64 block of the current CTU (case 230 in Fig. 2) , then in addition to the already reconstructed samples in the current CTU, if luma location (64, 0) relative to the current CTU has not yet been reconstructed, the current block can also refer to the reference samples in the top-right 64x64 block and bottom-right 64x64 block of the left CTU, using CPR mode. Otherwise, the current block can also refer to the reference samples in the bottom-right 64x64 block of the left CTU, using CPR mode. – If current block falls into the bottom-right 64x64 block of the current CTU (case 240 in Fig. 2) , it can only refer to the already reconstructed samples in the current CTU, using CPR mode.
[0015] This restriction allows the IBC mode to be implemented using local on-chip memory for hardware implementations.
[0016] In the present invention, methods and apparatus of generating combined prediction candidates from chained motion vectors or block vectors are disclosed to improve the performance. BRIEF SUMMARY OF THE INVENTION
[0017] A method and apparatus for video coding are disclosed. According to the method, input data associated with a current block are received, wherein the input data comprises pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. An initial candidate associated with one or more target MVs (Motion Vectors) BVs (Block Vectors) is determined. One or more chained MVs or BVs for said one or more target MVs or BVs are derived, wherein each of said one or more chained MVs or BVs is derived recursively as a sum of traced vectors starting from one of said one or more target MVs or BVs. One or more combined prediction candidates are generated from said one or more chained MVs or BVs, and wherein each of said one or more combined prediction candidates is generated by combining one first candidate based on one of said one or more chained MVs or BVs and one second candidate based on another MV or BV. The current block is encoded or decoded by using coding information comprising said one or more combined prediction candidates.
[0018] In one embodiment, said one or more combined prediction candidates are used for IBC (Intra Block Copy) or IntraTMP (Intra Template Matching Prediction) bi-prediction, and wherein each of combined prediction candidates comprises a first predictor using an L0 reference picture and a second predictor using an L1 reference picture. In another embodiment, said one or more combined prediction candidates are used for IBC (Intra Block Copy) or IntraTMP (Intra Template Matching Prediction) uni-prediction, and wherein each of combined prediction candidates comprises a predictor using an L0 reference picture or an L1 reference picture.
[0019] In one embodiment, the initial candidate corresponds to one candidate already in a candidate list, a spatial candidate, a non-adjacent candidate, an HMVP (History-Based Motion Vector Prediction) candidate, a pair-wise average candidate, or zero candidate.
[0020] In one embodiment, one of said one or more combined prediction candidates is combined with one initial candidate to generate a next combined prediction candidate.
[0021] In one embodiment, if a target chained MV or BV points to a target reference block outside an available reference region, the target chained MV or BV is modified to point to inside the available reference region or the target reference block is treated as invalid.
[0022] In one embodiment, said one or more combined prediction candidates are used for inter bi-prediction, and wherein each of combined prediction candidates comprises a first predictor using an L0 reference picture and a second predictor using an L1 reference picture. In another embodiment, said one or more combined prediction candidates are used for inter based uni-prediction, and wherein each of combined prediction candidates comprises a predictor using an L0 reference picture or an L1 reference picture.
[0023] In one embodiment, some of said one or more combined prediction candidates associated with said one or more chained MVs have higher priority than others of said one or more combined prediction candidates associated with said one or more chained MVs for adding to a merge candidate list. In one embodiment, a first combined prediction candidate has higher priority than a second combined prediction candidate if the first combined prediction candidate is true-bi-prediction and the second combined prediction candidate is not.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.
[0025] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0026] Fig. 2 illustrates an example of current CTU processing order and its available reference samples in current and left CTU.
[0027] Fig. 3 illustrates an example of how to derive AR-BVP (Auto-Relocated Block Vector Prediction) .
[0028] Fig. 4 illustrates an example of five spatial locations checked for block Bn in order to derive block vector BVn, n+1.
[0029] Fig. 5 illustrates an example of CMVP (Chained MVP) candidates derived as the sum of the recursively traced MVs and BVs based on the pre-derived MVs for the inter merge candidate list.
[0030] Fig. 6 illustrates an example of deriving MVk (m) by checking the existence of MVs or BVs in MV / BV storage corresponding to all five positions of the current block.
[0031] Fig. 7 illustrates a flowchart of an exemplary video coding system that uses one or more combined prediction candidates generated from one or more chained motion vectors or block vectors according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0032] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0033] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0034] JVET-AG0091: EE2-1.8: Auto-Relocated Block Vector Prediction
[0035] In EE2-1.8, Auto-Relocated Block Vector Prediction (AR-BVP) is introduced into IBC merge / AMVP candidate list construction.
[0036] As shown in Fig. 3, a guiding block vector BV0, 1 associated with the current block B0 points to a reference block B1. If B1 has a BV denoted as BV1, 2 pointing to a reference block B2, then BV0, 2, given by BV0, 2 = BV0, 1 +BV1, 2, is defined as the AR-BVP, guided by BV0, 1. Similarly, BV0, n+1 can be derived by: BV0, n+1 =BV0, n+BVn, n+1 = BV0, 1+BV1, 2 +…+BVn-1, n +BVn, n+1.
[0037] Three tests are conducted in this EE. In EE2-1.8a, the length of the AR-BVP trace path is 1 (i.e, n=1) . In EE2-1.8b, the length of the AR-BVP trace path is 2 (i.e, n=2) . In EE2-1.8c, there is no constraint for the length of the AR-BVP trace path.
[0038] When deriving BVn, n+1 guided by BV0, n, all five positions including top-left (e.g. LT in Fig. 4) , top-right (e.g. RT in Fig. 4) , centre (e.g. Ctr in Fig. 4) , bottom-left (e.g. LB in Fig. 4) , and bottom-right (e.g. RB in Fig. 4) positions of Bn are checked to find BVn, n+1.
[0039] In our implementation, the initial guiding block vector BV0, 1 is set to be an existing BVP already in the IBC merge / AMVP candidate list.
[0040] The AR-BVP candidates are inserted after the HBVP candidates. The IBC merge / AMVP candidate list size is kept unchanged.
[0041] JVET-AG0073: Non-EE2: Chained Motion Vector Prediction
[0042] This contribution introduces a chained MV prediction (CMVP) into inter merge candidate list construction.
[0043] As shown in Fig. 5, CMVP candidates can be derived as the sum of the recursively traced MVs and BVs based on the pre-derived MVs for the inter merge candidate list. For instance, a CMVP candidate, a set of motion vectors MVk / m and reference picture RefPick / m can be derived by: MVk / m = MVk (0) + BVk (0) + MVk (1) +MVk (2) + …+ MVk (m) , RefPick / m = RefPick (m) , where k and m indicate the number of merge index and trace depths of the CMVP.
[0044] When deriving MVk / m, MVk (m) is found by checking the existence of MVs or BVs in MV / BV storage corresponding to all five positions of the current block as shown in Fig. 6 (i.e., the centre, top-left, top-right, bottom-left, and bottom-right of the current block) .
[0045] When pre-derived merge candidates targeting CMVP candidates have two MVs, a MVk / m is derived for each list (i.e., L0 and L1) and each trace depth. Up to two MVs can be derived for each list and each trace depth, and the MV set is sequentially inserted into inter merge candidate list.
[0046] The traceable reference pictures are only within the reference picture list. CMVP candidates are inserted after HMVP candidates for the regular merge and TM merge. When deriving CMVP candidates, hpelIfIdx, bcwIdx, licFlag, and mhpFlag are not inherited. CMVP candidates are not derived when the TMVP is disabled.
[0047] In prior art, auto-relocated block vector prediction (AR-BVP) or chained MV prediction (CMVP) is only used for generating uni-prediction candidates for IBC and Inter modes. In this proposal, several bi-prediction candidate generation ideas and methods are illustrated.
[0048] Bi-Prediction Candidates for IBC mode
[0049] In one embodiment, given an initial candidate with BV or BVs, one or more bi-prediction candidates can be generated by a combination of chained BV predictions wherein the reference pictures of BVs are current pictures. The “combination of chained BV predictions” means combining two or more BV predictions including at least one chained BV prediction. – Given an initial candidate being bi-prediction with BVL0 and BVL1, the reference pictures of BVL0 and BVL1 are both current pictures. BVL0 points to block A, which has motions BVAL0 and BVAL1, and the reference pictures of BVAL0 and BVAL1 are both current pictures. BVL1 points to block B, which has motions BVBL0 and BVBL1, and the reference pictures of BVBL0 and BVBL1 are both current pictures. – For example, bi-prediction candidates can be generated by setting L0 BV to (BVAL0+BVL0) , (BVAL1+BVL0) , (BVBL0+BVL1) , or (BVBL1+BVL1) . The reference pictures of all generated candidates are current pictures. For L1 MV, it can be (BVAL0+BVL0) , (BVAL1+BVL0) , (BVBL0+BVL1) , or (BVBL1+BVL1) . The reference pictures of all generated candidates are current pictures. – In the above examples, BVL0, BVL1, BVAL0, BVAL1, BVBL0, BVBL1 may not be available. For example, if BVL1 is not available, and BVBL0 and BVBL1 are not available either. the combinations with BVL1, BVBL0, or BVBL1 will be discarded in this case.
[0050] In one embodiment, given an initial candidate with BV or BVs, one or more uni-prediction candidates can be generated by a combination of chained BV predictions, where the reference pictures of BVs are current picture. – For example, L0 prediction candidates can be generated by setting L0 BV to (BVAL0+BVL0) , (BVAL1+BVL0) , (BVBL0+BVL1) , or (BVBL1+BVL1) . The reference pictures of all generated candidates are current pictures. – For example, L1 prediction candidates can be generated by setting L1 BV to (BVAL0+MVL0) , (BVAL1+BVL0) , (BVBL0+BVL1) , or (BVBL1+BVL1) . The reference pictures of all generated candidates are current pictures.
[0051] In the above embodiments, an initial candidate can be candidate already in the candidate list, a newly added candidate by the above method, or motion from spatial, non-adjacent, HMVP, pair-wise average, or zero candidates.
[0052] In the above embodiments, the initial candidate can also be one of the combinations of chained BV prediction. – Given an initial candidate being bi-prediction with BVL0 and BVL1. Both of BVs point to the reference blocks in the current picture. BVL0 points to block A which has motions BVAL0 and BVAL1. BVL1 points to block B, which has motions BVBL0 and BVBL1. – For example, uni-prediction or bi-prediction candidates can be generated by setting LX (where X = 0 or 1) BV to BVL0, BVL1, (BVAL0+BVL0) , (BVAL1+BVL0) , (BVBL0+BVL1) , or (BVBL1+BVL1) . All reference pictures of the generated BVs are current pictures.
[0053] In the above embodiments, the averaging process can be applied to combinations of chained BV prediction. – For example, uni-prediction or bi-prediction candidates can be generated by setting LX (where X = 0 or 1) BV to the averaged BV of two or more of these BVs, BVL0, BVL1, (BVAL0+BVL0) , (BVAL1+BVL0) , (BVBL0+BVL1) , or (BVBL1+BVL1) . All reference pictures of the generated BVs are the current pictures.
[0054] In the above embodiments, if the reference blocks of the derived chain BVs are out of the available reference region, they can be modified to point to valid reference blocks or directly treated as invalid. – If the reference blocks of the derived uni-prediction chain BVs are out of the available reference region, the derived uni-prediction chain BVs will be invalid. – If the reference blocks of the derived uni-prediction chain BVs are out of the available reference region, the derived uni-prediction chain BVs will be modified to point to reference blocks inside the available reference region. For example, the nearest reference blocks inside the available reference region will be referenced. For another example, the pre-defined reference blocks inside the available region will be referenced. – If one of reference blocks of the derived bi-prediction chain BV (i.e., reference block pointed by L0 BV or reference block pointed by L1 BV) is out of the available reference region, the derived bi-prediction chain BV will be invalid. – If one of reference blocks of the derived bi-prediction chain BV (i.e., reference block pointed by L0 BV or reference block pointed by L1 BV) is out of the available reference region, the derived bi-prediction chain BV will be changed to uni-prediction BV. That is, the side pointed to out of boundary BV will be discarded. – If one of reference blocks of the derived bi-prediction chain BV (i.e., reference block pointed by L0 BV or reference block pointed by L1 BV) is out of the available reference region, the BV points to the out of boundary reference block will be modified to point the nearest reference block inside the available reference region. – The available reference region can include N CTU rows above, M CTUs on the left, or N CTU rows above and M CTUs on the left. N and M can be any integer larger than 0. – The available reference region can be the decoded region within the same slice / tile as the current block. – If a chain BV is treated as invalid, it cannot be used to further derive other chain BVs and cannot be inserted into the merge candidate list. – If a chain BV is modified since the reference block isn’ t in the available region, it cannot be used to further derive other chain BVs. But the modified chain BV can be inserted into the merge candidate list.
[0055] In one embodiment, if the current block is not a RR-IBC (Reconstruction-Reordered IBC (RR-IBC) ) coded block, the reference block pointed by a chain BV cannot be a RRIBC coded block. A Reconstruction-Reordered IBC (RR-IBC) mode is allowed for IBC coded blocks. When RR-IBC is applied, the samples in a reconstruction block are flipped according to a flip type of the current block. At the encoder side, the original block is flipped before motion search and residual calculation, while the prediction block is derived without flipping. At the decoder side, the reconstruction block is flipped back to restore the original block.
[0056] In one embodiment, if the current block is a RRIBC coded block, the reference block pointed by a chain BV needs to be a RRIBC coded block. Also, the flipping type of current block and the reference blocks pointed by the chain BVs shall be the same. For example, both of them are horizontal flip type, or both of them are vertical flip type.
[0057] In one embodiment, LIC or filtering LIC process will be disabled for all derived chain BV candidates.
[0058] In one embodiment, when deriving chain BV, some positions of the current block or some adjacent positions of the current block will be checked to find the existence of BVs. The checked positions use for IBC chain BVPs and inter chain MVPs can be aligned.
[0059] Bi-Prediction Candidates for Inter Mode
[0060] In one embodiment, given an initial candidate with MV or MVs, one or more bi-prediction candidates can be generated by a combination of chained MV prediction. The “combination of chained MV predictions” means combining two or more MV predictions including at least one chained MV prediction. – Given an initial candidate being bi-prediction with MVL0 and MVL1, MVL0 points to block A which have motions MVAL0 and MVAL1, and the reference picture of MVAL0 is POCAL0 and the reference picture of MVAL1 is POCAL1. MVL1 points to block B which has motions MVBL0 and MVBL1, and the reference picture of MVBL0 is POCBL0 and the reference picture of MVBL1 is POCBL1. – For example, bi-prediction candidates can be generated by setting L0 MV to (MVAL0+MVL0) , (MVAL1+MVL0) , (MVBL0+MVL1) , or (MVBL1+MVL1) , and set L0 POC to POCAL0, POCAL1, POCBL0, or POCBL1, where the POC should be in the L0 reference pictures. For L1 MV, it can be (MVAL0+MVL0) , (MVAL1+MVL0) , (MVBL0+MVL1) , or (MVBL1+MVL1) , and set L1 POC to POCAL0, POCAL1, POCBL0, or POCBL1, where the POC should be in the L1 reference pictures. Note that, the MV and POC can be set separately, which means MV can be (MVAL0+MVL0) and POC can be one of POCAL0, POCAL1, POCBL0, or POCBL1; or POC can be POCAL0 and MV can be one of (MVAL0+MVL0) , (MVAL1+MVL0) , (MVBL0+MVL1) , or (MVBL1+MVL1) . – For example, a bi-prediction candidate can be (MVAL0+MVL0) for L0 MV with POCAL0 and (MVBL1+MVL1) for L1 MV with POCBL1. – For example, a bi-prediction candidate can be (MVAL0+MVL0) for L0 MV with POCAL0, and (MVAL1+MVL0) for L1 MV with POCAL1. For this case, only the motions from block A are needed. – In above examples, MVL0, MVL1, MVAL0, MVAL1, MVBL0, MVBL1 may not be available. For example, if MVL1 is not available, MVBL0 and MVBL1 are not available either. The combinations with MVL1, MVBL0, or MVBL1 are discarded. – In the above examples, POCAL0, POCAL1, POCBL0, or POCBL1 may not be available in L0 or L1 reference pictures. The combinations with unavailable POC are discarded.
[0061] In above embodiment, given an initial candidate with MV or MVs, one or more uni-prediction candidates can be generated by a combination of chained MV prediction. – For example, L0 prediction candidates can be generated by setting L0 MV to (MVAL0+MVL0) , (MVAL1+MVL0) , (MVBL0+MVL1) , or (MVBL1+MVL1) , and set L0 POC to POCAL0, POCAL1, POCBL0, or POCBL1, where the POC should be in the L0 reference pictures. – For example, L1 prediction candidates can be generated by setting L1 MV to (MVAL0+MVL0) , (MVAL1+MVL0) , (MVBL0+MVL1) , or (MVBL1+MVL1) , and set L1 POC to POCAL0, POCAL1, POCBL0, or POCBL1, where the POC should be in the L1 reference pictures.
[0062] In above embodiment, an initial candidate can be candidate already in candidate list, newly added candidates by above methods, or motions from spatial, non-adjacent, temporal, HMVP, pair-wise average, or zero candidates.
[0063] In above embodiment, the initial candidate can also be one of the combinations of chained MV prediction. – Given an initial candidate being bi-prediction with MVL0 and MVL1, and the reference picture of MVL0 is POCL0 and the reference picture of MVL1 is POCL1. MVL0 points to block A which the motions are MVAL0 and MVAL1, and the reference picture of MVAL0 is POCAL0 and the reference picture of MVAL1 is POCAL1. MVL1 points to block B which the motions are MVBL0 and MVBL1, and the reference picture of MVBL0 is POCBL0 and the reference picture of MVBL1 is POCBL1. – For example, uni-prediction or bi-prediction candidates can be generated by setting LX (where X = 0 or 1) MV to MVL0, MVL1, (MVAL0+MVL0) , (MVAL1+MVL0) , (MVBL0+MVL1) , or (MVBL1+MVL1) , and set LX POC to POCL0, POCL1, POCAL0, POCAL1, POCBL0, or POCBL1, where the POC should be in the LX reference pictures.
[0064] In above embodiment, the averaging process can be applied to the combinations of chained MV prediction. – For example, uni-prediction or bi-prediction candidates can be generated by setting LX (where X = or 1) MV to the averaged MV of two or more of these MVs including MVL0, MVL1, (MVAL0+MVL0) , (MVAL1+MVL0) , (MVBL0+MVL1) , (MVBL1+MVL1) , or any combination, and set LX POC to POCL0, POCL1, POCAL0, POCAL1, POCBL0, or POCBL1, where the POC should be in the LX reference pictures.
[0065] In above embodiment, the averaging process only apply to MVs with the same POC.
[0066] In the above embodiment, if the reference picture of chained MV prediction candidate is not in current L0 or L1 reference picture, MV scaling can be used to scale chained MV prediction candidate to the current L0 or L1 reference picture. – For example, a chained MV prediction candidate can be (MVAL0+MVL0) MV with POCAL0, where POCAL0 is not in the L0 or L1 reference picture. In this case, this candidate to is scaled one of L0 or L1 reference pictures.
[0067] In above embodiment, some chained MV prediction candidates have higher priority than other chained MV prediction candidates, and these candidates with higher priority will be added to merge candidate list first; or only chained MV prediction candidates with specific conditions will be added to merge candidate list. – For example, a bi-prediction candidate with true-bi-prediction (one POC small than the current POC and another POC larger than the current POC) has higher priority. – For example, a bi-prediction candidate with true-bi-prediction and with the POC differences being the same (i.e., difference of one POC and the current POC equal to difference of another POC and the current POC) has higher priority. – For example, (MVAL0+MVL0) has higher priority than (MVAL1+MVL0) due to MVL0 is L0 prediction. – For example, (MVAL0+MVL0) with POCAL0 has higher priority than (MVAL0+MVL0) with POCAL1 due to the POC of MVAL0 is POCAL1. – For example, a candidate without scaling has higher priority than scaled one.
[0068] In one embodiment, use both MV and BV to derive chained MV prediction. Given an initial candidate with MV / BV or MVs / BVs, one or more uni-prediction or bi-prediction candidates can be generated by a combination of chained MV prediction. – Given an initial candidate being bi-prediction with MVL0 / BVL0 and MVL1 / BVL1. MVL0 / BVL0 points to block A which has motions MVAL0 / BVAL0 and MVAL1 / BVAL1, and the reference pictures of MVAL0 / BVAL0 is POCAL0 and the reference picture of MVAL1 / BVAL1 is POCAL1. MVL1 / BVL1 points to block B which the motions are MVBL0 / BVBL0 and MVBL1 / BVBL1, and the reference picture of MVBL0 / BVBL0 is POCBL0 and the reference picture of MVBL1 / BVBL1 is POCBL1. – For example, uni-prediction or bi-prediction candidates can be generated by setting LX (where X = 0 or 1) MV to (MVAL0 / BVAL0+MVL0 / BVL0) , (MVAL1 / BVAL1+MVL0 / BVL0) , (MVBL0 / BVBL0+MVL1 / BVL1) , or (MVBL1 / BVBL1+MVL1 / BVL1) . And set LX POC to POCAL0, POCAL1, POCBL0, or POCBL1, where the POC should be in the LX reference pictures.
[0069] In above embodiment, the candidate can continue chaining when a BV / BVs are encountered until a MV / MVs are encountered. – Given an initial candidate being uni-prediction with MVL0, MVL0 points to block A which has motion BVAL0, and the reference picture of BVAL0 is POCAL0. – Further chain BVAL0 to another block A1 which has motion BVA1L0 and the reference picture of BVA1L0 is still POCAL0. – Further chain BVA (n-1) L0 to another block An which has motion MVAnL0 and the reference picture of MVAnL0 is POCAnL0. – For example, the candidates can be generated by setting LX (where X = 0 or 1) MV to (MVAnL0+BVA (n-1) L0+…+BVA1L0+BVAL0+MVL0) , and set LX POC to POCAnL0.
[0070] In above embodiment, a pre-defined or signalled trace depth limit for continued chaining can be used to constrain the chaining count. The continued chaining will continue when a BV / BVs are encountered until a MV / MVs are encountered. – For example, given trace depth limit equal to 3, the chaining process should stop even an MV is not encountered, candidates can be generated by setting LX (where X = 0 or 1) MV to (BVA2L0+BVA1L0+BVAL0+MVL0) , And set LX POC to POCAL0
[0071] In one embodiment, the chain MVPs can be pre-calculated and stored in buffers, and chain MVP can be fetched from the buffers when constructing the merge candidate list.
[0072] In above embodiment, the chain MVPs with different trace depths are store in different buffers.
[0073] The foregoing proposed methods can be implemented in encoders and / or decoders. For example, the proposed method can be implemented in intra / inter prediction module of an encoder or decoder.
[0074] Any of the foregoing proposed prediction candidate generation from a chained MV / BV can be implemented in encoders and / or decoders. For example, any of the proposed filtered predictor methods can be implemented in an inter / intra / predictor derivation module of an encoder, and / or an inter / intra / predictor derivation module of a decoder. Alternatively, any of the proposed methods can be implemented as circuits coupled to the inter / intra / predictor derivation module of the encoder and / or the inter / intra / predictor derivation module of the decoder, so as to provide the information needed by the inter / intra / predictor derivation module. For example, any of the proposed filtered predictor methods can be implemented in an Intra / Inter coding module (e.g. Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module is an encoder (e.g. Intra Pred. 110 / Inter Pred. 112 in Fig. 1A) . Any of the proposed filtered predictor methods can also be implemented as a circuit coupled to the intra / inter coding module at the decoder or the encoder. However, the decoder or encoder may also use additional processing unit to implement the required cross-component prediction processing. While the Intra Pred. units (e.g. unit 110 / 112 in Fig. 1A and unit 150 / 152 in Fig. 1B) are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .
[0075] Fig. 7 illustrates a flowchart of an exemplary video coding system that uses one or more combined prediction candidates generated from one or more chained motion vectors or block vectors according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data associated with a current block are received in step 710, wherein the input data comprises pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. An initial candidate associated with one or more target MVs (Motion Vectors) BVs (Block Vectors) is determined in step 720. One or more chained MVs or BVs for said one or more target MVs or BVs are derived in step 730, wherein each of said one or more chained MVs or BVs is derived recursively as a sum of traced vectors starting from one of said one or more target MVs or BVs. One or more combined prediction candidates are generated from said one or more chained MVs or BVs in step 740, and wherein each of said one or more combined prediction candidates is generated by combining one first candidate based on one of said one or more chained MVs or BVs and one second candidate based on another MV or BV. The current block is encoded or decoded by using coding information comprising said one or more combined prediction candidates in step 750.
[0076] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0077] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0078] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0079] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of video coding, the method comprising:receiving input data associated with a current block, wherein the input data comprises pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side, and wherein the current block is coded in a non-intra mode;determining an initial candidate associated with one or more target MVs (Motion Vectors) BVs (Block Vectors) ;deriving one or more chained MVs or BVs for said one or more target MVs or BVs, wherein each of said one or more chained MVs or BVs is derived recursively as a sum of traced vectors starting from one of said one or more target MVs or BVs;generating one or more combined prediction candidates from said one or more chained MVs or BVs, and wherein each of said one or more combined prediction candidates is generated by combining one first candidate based on one of said one or more chained MVs or BVs and one second candidate based on another MV or BV; andencoding or decoding the current block by using coding information comprising said one or more combined prediction candidates.2.The method of Claim 1, wherein said one or more combined prediction candidates are used for IBC (Intra Block Copy) or IntraTMP (Intra Template Matching Prediction) bi-prediction, and wherein each of combined prediction candidates comprises a first predictor using an L0 reference picture and a second predictor using an L1 reference picture.3.The method of Claim 1, wherein said one or more combined prediction candidates are used for IBC (Intra Block Copy) or IntraTMP (Intra Template Matching Prediction) uni-prediction, and wherein each of combined prediction candidates comprises a predictor using an L0 reference picture or an L1 reference picture.4.The method of Claim 1, wherein the initial candidate corresponds to one candidate already in a candidate list, a spatial candidate, a non-adjacent candidate, an HMVP (History-Based Motion Vector Prediction) candidate, a pair-wise average candidate, or zero candidate.5.The method of Claim 1, wherein one of said one or more combined prediction candidates is combined with one initial candidate to generate a next combined prediction candidate.6.The method of Claim 1, wherein if a target chained MV or BV points to a target reference block outside an available reference region, the target chained MV or BV is modified to point to inside the available reference region or the target reference block is treated as invalid.7.The method of Claim 1, wherein said one or more combined prediction candidates are used for inter bi-prediction, and wherein each of combined prediction candidates comprises a first predictor using an L0 reference picture and a second predictor using an L1 reference picture.8.The method of Claim 1, wherein said one or more combined prediction candidates are used for inter based uni-prediction, and wherein each of combined prediction candidates comprises a predictor using an L0 reference picture or an L1 reference picture.9.The method of Claim 1, wherein some of said one or more combined prediction candidates associated with said one or more chained MVs have higher priority than others of said one or more combined prediction candidates associated with said one or more chained MVs for adding to a merge candidate list.10.The method of Claim 9, wherein a first combined prediction candidate has higher priority than a second combined prediction candidate if the first combined prediction candidate is true-bi-prediction and the second combined prediction candidate is not.11.An apparatus of video coding, the apparatus comprising one or more electronic circuits or processors arranged to:receive input data associated with a current block, wherein the input data comprises pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side, and wherein the current block is coded in a non-intra mode;determine an initial candidate associated with one or more target MVs (Motion Vectors) BVs (Block Vectors) ;derive one or more chained MVs or BVs for said one or more target MVs or BVs, wherein each of said one or more chained MVs or BVs is derived recursively as a sum of traced vectors starting from one of said one or more target MVs or BVs, and wherein each of said one or more combined prediction candidates is generated by combining one first candidate based on one of said one or more chained MVs or BVs and one second candidate based on another MV or BV;generate one or more combined prediction candidates from said one or more chained MVs or BVs; andencode or decode the current block by using coding information comprising said one or more combined prediction candidates.