Subblock candidates for auto-relocated block vector or chained motion vector prediction

Cascaded vectors for subblock coding enhance motion vector prediction in video coding, addressing inefficiencies in existing standards by improving accuracy and reducing complexity, thereby enhancing compression efficiency.

WO2025152853A1PCT designated stage expired Publication Date: 2025-07-24MEDIATEK INC

Patent Information

Application Number
PCT/CN2025/071657
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-04
Filing Date
2025-01-10
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in efficiently predicting motion vectors for subblocks, particularly in handling various types of motion within video frames, leading to suboptimal compression efficiency and increased computational complexity.

Method used

The use of cascaded vectors for subblock coding, which involves deriving predictors for current subblocks based on motion information of collocated subblocks, applying position scaling when necessary, and utilizing cascaded vectors as sums of recursively traced motion vectors or block vectors to enhance prediction accuracy.

Benefits of technology

Improves motion vector prediction accuracy and reduces computational complexity by leveraging cascaded vectors, resulting in enhanced compression efficiency and improved video coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071657_24072025_PF_FP_ABST
    Figure CN2025071657_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A method of using cascaded vectors for subblock coding is provided. A video coder receives data to be encoded or decoded as a current block comprising a plurality of current subblocks. The video coder identifies a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift provided by a cascaded vector that is a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector. The video coder derives predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks. A cascaded vector is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock. The video coder encodes or decodes the current block by using the derived predictors of the plurality current subblocks.
Need to check novelty before this filing date? Find Prior Art

Description

SUBBLOCK CANDIDATES FOR AUTO-RELOCATED BLOCK VECTOR OR CHAINED MOTION VECTOR PREDICTIONCROSS REFERENCE TO RELATED PATENT APPLICATION (S)

[0001] The present disclosure is part of a non-provisional application that claims the priority benefit of U.S. Provisional Patent Application Nos. 63 / 622,097 and 63 / 690,332, filed on 18 January 2024 and 4 September 2024, respectively. Contents of above-listed applications are herein incorporated by reference.TECHNICAL FIELD

[0002] The present disclosure relates generally to video coding. In particular, the present disclosure relates to methods of coding pixel blocks by chained motion vector prediction.BACKGROUND

[0003] Unless otherwise indicated herein, approaches described in this section are not prior art to the claims listed below and are not admitted as prior art by inclusion in this section.

[0004] High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) . HEVC is based on the hybrid block-based motion-compensated DCT-like transform coding architecture. The basic unit for compression, termed coding unit (CU) , is a 2Nx2N square block of pixels, and each CU can be recursively split into four smaller CUs until the predefined minimum size is reached. Each CU contains one or multiple prediction units (PUs) .

[0005] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Expert Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11. The input video signal is predicted from the reconstructed signal, which is derived from the coded picture regions. The prediction residual signal is processed by a block transform. The transform coefficients are quantized and entropy coded together with other side information in the bitstream. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transform on the de-quantized transform coefficients. The reconstructed signal is further processed by in-loop filtering for removing coding artifacts. The decoded pictures are stored in the frame buffer for predicting the future pictures in the input video signal.

[0006] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . The leaf nodes of a coding tree correspond to the coding units (CUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors (MVs) and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only.

[0007] A CTU can be partitioned into one or multiple non-overlapped coding units (CUs) using the quadtree (QT) with nested multi-type-tree (MTT) structure to adapt to various local motion and texture characteristics. A CU can be further split into smaller CUs using one of the five split types: quad-tree partitioning, vertical binary tree partitioning, horizontal binary tree partitioning, vertical center-side triple-tree partitioning, horizontal center-side triple-tree partitioning.

[0008] Each CU contains one or more prediction units (PUs) . The prediction unit, together with the associated CU syntax, works as a basic unit for signaling the predictor information. The specified prediction process is employed to predict the values of the associated pixel samples inside the PU. Each CU may contain one or more transform units (TUs) for representing the prediction residual blocks. A transform unit (TU) is comprised of a transform block (TB) of luma samples and two corresponding transform blocks of chroma samples and each TB correspond to one residual block of samples from one color component. An integer transform is applied to a transform block. The level values of quantized coefficients together with other side information are entropy coded in the bitstream. The terms coding tree block (CTB) , coding block (CB) , prediction block (PB) , and transform block (TB) are defined to specify the 2-D sample array of one-color component associated with CTU, CU, PU, and TU, respectively. Thus, a CTU consists of one luma CTB, two chroma CTBs, and associated syntax elements. A similar relationship is valid for CU, PU, and TU.

[0009] For each inter-predicted CU, motion parameters consisting of motion vectors, reference picture indices and reference picture list usage index, and additional information are used for inter-predicted sample generation. The motion parameter can be signalled in an explicit or implicit manner. When a CU is coded with skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. A merge mode is specified whereby the motion parameters for the current CU are obtained from neighbouring CUs, including spatial and temporal candidates, and additional schedules introduced in VVC. The merge mode can be applied to any inter-predicted CU. The alternative to merge mode is the explicit transmission of motion parameters, where motion vector, corresponding reference picture index for each reference picture list and reference picture list usage flag and other needed information are signalled explicitly per each CU.

[0010] Intra block copy (IBC) or current picture referencing (CPR) refer to coding pixel blocks by referencing pixel positions within same current picture as the current block by using block vectors.

[0011] In advanced motion vector prediction (AMVP) mode, a motion vector predictor (MVP) candidate is determined based on template matching (TM) error to select the one that reaches the minimum difference between the current block template and the reference block template, and then TM is performed only for this particular MVP candidate for MV refinement. The TM process may refine this MVP candidate using iterative search according to an adaptive motion vector resolution (AMVR) mode search pattern.SUMMARY

[0012] The following summary is illustrative only and is not intended to be limiting in any way. That is, the following summary is provided to introduce concepts, highlights, benefits and advantages of the novel and non-obvious techniques described herein. Select and not all implementations are further described below in the detailed description. Thus, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended for use in determining the scope of the claimed subject matter.

[0013] Some embodiments of the disclosure provide a method of using cascaded vectors for subblock coding. A video coder receives data to be encoded or decoded as a current block of pixels of a current picture of a video. The current block comprising a plurality of current subblocks. The video coder identifies a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks. The video coder derives predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks. The video coder encodes or decodes the current block by using the derived predictors of the plurality current subblocks.

[0014] In some embodiments, the video coder applies position scaling to the collocated picture when the collocated picture has a different width or height than that of the current picture. In some embodiments, the motion shift is provided by a cascaded vector (or chain MV) that is a sum of recursively traced motion vectors and block vectors. The cascaded vector has a base vector that is a motion vector or block vector a neighboring block of the current block. In some embodiments, a list of motion shift candidates includes (i) the block vector or the motion vector of the neighboring block that is the base vector of the cascaded vector and (ii) the cascaded vector.

[0015] In some embodiments, the predictors for the plurality of current subblocks are derived by scaling the motion information of the plurality of collocated subblocks according to temporal distances among the current picture, the collocated picture, and reference pictures of the motion information of the plurality of collocated subblocks.

[0016] In some embodiments, a cascaded vector (or chained MV) is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock, the cascaded vector being a sum of recursively traced motion vectors or block vectors. In some embodiments, the motion information of the collocated subblock is temporally scaled to the current picture to be used as a base vector for deriving the cascaded vector to identify a reference block from the current subblock, with the scaling being applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the motion information of the collocated subblock (ii) a second temporal difference between the current picture and a reference picture identified by the cascaded vector.

[0017] In some embodiments, the motion information of the collocated subblock is used as a base vector for deriving the cascaded vector. The predictor of the current subblock may be derived based on a reference block that is identified by the cascaded vector from the collocated subblock. In some embodiments, the predictor of the current subblock may be derived based on a reference block that is identified by a scaled motion vector from the current subblock, the scaled motion vector derived by temporally scaling the cascaded vector to the current picture, with the scaling being applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the cascaded vector (ii) a second temporal difference between the current picture and a reference picture identified by the scaled motion vector.BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of the present disclosure. The drawings illustrate implementations of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It is appreciable that the drawings are not necessarily in scale as some components may be shown to be out of proportion than the size in actual implementation in order to clearly illustrate the concept of the present disclosure.

[0019] FIG. 1 shows the control point motion vectors (CPMVs) of a current block that is coded by affine motion field.

[0020] FIG. 2 illustrates affine motion field per subblock.

[0021] FIG. 3 illustrates the candidate neighboring blocks for inheriting affine candidates.

[0022] FIG. 4 conceptually illustrates control point motion vector inheritance.

[0023] FIGS. 5A-B illustrate spatial neighbors for deriving affine merge  / AMVP candidates.

[0024] FIG. 6 illustrates the transformation of motion information from non-adjacent neighbors to the first type of constructed affine merge / AMVP candidates.

[0025] FIG. 7 illustrates neighboring 4 x 4 subblocks that are used for RMVF parameter derivation.

[0026] FIG. 8 illustrates deriving sub-CU motion field by applying a motion shift from a spatial neighbor and scaling the motion information from the corresponding collocated sub-CUs.

[0027] FIG. 9 conceptually illustrates derivation of auto-relocated block vector.

[0028] FIG. 10 conceptually illustrates the positions to be checked for deriving auto-relocated block vector.

[0029] FIG. 11 shows an example derivation of a candidate for chained MV prediction (CMVP) .

[0030] FIG. 12 illustrates the possible sources of a base vector for CMVP.

[0031] FIG. 13 conceptually illustrates subblock motions being derived by scaling the motions from the collocated picture.

[0032] FIG. 14A conceptually illustrates using the blocks referenced by the collocated subblocks in the collocated picture for predicting the current subblocks.

[0033] FIG. 14B conceptually illustrates using collocated subblock’s reference block’s reference block as prediction for the current subblock.

[0034] FIG. 15 conceptually illustrates scaling a chain MV of a collocated picture to be the subblock motion for the current block.

[0035] FIG. 16 illustrates position scaling for a collocated picture.

[0036] FIG. 17 illustrates an example video encoder that may implement subblock motion.

[0037] FIG. 18 illustrates portions of the video encoder that implement subblock processing and cascading vectors.

[0038] FIG. 19 conceptually illustrates a process that uses cascaded vectors for encoding subblocks.

[0039] FIG. 20 illustrates an example video decoder that may implement subblock motion.

[0040] FIG. 21 illustrates portions of the video decoder that implement subblock processing and cascading vectors.

[0041] FIG. 22 conceptually illustrates a process that uses cascaded vectors for decoding subblocks.

[0042] FIG. 23 conceptually illustrates an electronic system with which some embodiments of the present disclosure are implemented.DETAILED DESCRIPTION

[0043] In the following detailed description, numerous specific details are set forth by way of examples in order to provide a thorough understanding of the relevant teachings. Any variations, derivatives and / or extensions based on teachings described herein are within the protective scope of the present disclosure. In some instances, well-known methods, procedures, components, and / or circuitry pertaining to one or more example implementations disclosed herein may be described at a relatively high level without detail, in order to avoid unnecessarily obscuring aspects of teachings of the present disclosure. I. Affine Prediction

[0044] A. Affine Motion Field

[0045] An object in a video may have different types of motion, including translation motions, zoom in / out motions, rotation motions, perspective motions and the other irregular motions. In some embodiments, a block-based affine transform motion compensation prediction is used to account for these various types of motion. A block-based affine transform motion compensation prediction may be used. Specifically, the affine motion field mvx, mvy of the current block at position (x, y) is in the form of a linear model: mvx = a*x + b*y + c mvy = d*x + e*y + f (0)

[0046] The coefficients {a, b, c, d, e, f } are parameters of the linear model. In some embodiments, the affine motion field at position (x, y) can be described by motion information of two control points (CPs) (at e.g., top-right and top-left corners of the block) (4-parameter model) or motion information of three control points (at e.g., top-right, top-left, and bottom-left corners of the block) (6-parameter model) .

[0047] For 4-parameter affine motion model, eq. 0 (motion vector at sample location (x, y) in a block) can be written as:

[0048] For 6-parameter affine motion model, eq. 0 can be written as:

[0049] Where (mv0x, mv0y) is motion vector of the top-left corner control point (top-left corner CPMV, or mv0) , (mv1x, mv1y) is the motion vector of the top-right corner control point (top-right corner CPMV, or mv1) , and (mv2x, mv2y) is the motion vector of the bottom-left corner control point (bottom-left corner CPMV, or mv2) .

[0050] In some embodiments, a 2-parameter model can be used to refine translation inter-prediction candidate (e.g., regular merge candidates) :

[0051] FIG. 1 shows the control point motion vectors (CPMVs) of a current block that is coded by affine motion field. The current block has CPMVs at top-left corner (mv0) , top-right-corner (mv1) , and bottom-left corner (mv2) . The affine motion field mv' at positions (x, y) in the current block 100 can be derived using an affine motion model such as eq. 1 (4-parameter affine model) or eq. 2 (6-parameter affine model) .

[0052] In order to simplify the motion compensation prediction, block based affine transform prediction is applied. FIG. 2 illustrates affine motion field per subblock. The figure illustrates a motion field of motion vectors for a block having 16 4x4 subblocks. To derive the motion vector of each 4×4 luma subblock, the motion vector of the center sample of each subblock. is calculated according to eq. 1 or eq. 2, and rounded to 1 / 16 fraction accuracy. A motion compensation interpolation filters can be applied to generate the prediction of each subblock with derived motion vector. The subblock size of chroma-components is also set to be 4×4. The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the top-left and bottom-right luma subblocks in the collocated 8x8 luma region.

[0053] B. Affine Merge Mode

[0054] In affine merge mode, the motion vectors at the control points (CPMVs) of the current CU are generated based on the motion information of the spatial neighboring CUs. There can be up to five CPMVP candidates, and an index is signalled to indicate the one to be used for the current CU. The following three types of CPMV candidates are used to form the affine merge candidate list: (1) inherited affine merge candidates that are extrapolated from the CPMVs of the neighbour CUs; (2) constructed affine merge candidates CPMVPs that are derived using the translational MVs of the neighbour CUs; (3) zero MVs.

[0055] An inherited affine candidate inherits an affine model from a neighboring block by directly obtaining the CPMVs from the neighboring blocks (one from left neighboring CUs and one from above neighboring CUs) . FIG. 3 illustrates the candidate neighboring blocks for inheriting affine candidates. For the left predictor, the scan order of the candidate blocks is A1→A0, and for the above predictor, the scan order for the candidate blocks is B1→B0→B2. When a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidate in the affine merge list of the current CU.

[0056] FIG. 4 conceptually illustrates control point motion vector inheritance. As illustrated, for a current block 410, if a left-bottom neighboring block A is coded in affine mode, the motion vectors mv2, mv3, and mv4 of the top left corner, above right corner and left bottom corner of a CU 420 that contains the block A can be inherited by the current block 410. When block A is coded with 4-parameter affine model, the two CPMVs of the current CU 410 can be calculated according to mv2 and mv3. In case that block A is coded with 6-parameter affine model, the three CPMVs of the current CU 410 may be calculated according to mv2, mv3, and mv4.

[0057] Constructed affine candidate means the candidate is constructed by combining the neighbor translational motion information of each control point. The motion information for the control points is derived from the specified spatial neighbors and temporal neighbor shown in FIG. 3. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2→B3→A2 blocks are checked and the MV of the first available block is used. For CPMV2, the B1→B0 blocks are checked and for CPMV3, the A1→A0 blocks are checked. For TMVP is used as CPMV4 if it’s available. After MVs of four control points are attained, affine merge candidates are constructed based on those motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3} , {CPMV1, CPMV2, CPMV4} , {CPMV1, CPMV3, CPMV4} , {CPMV2, CPMV3, CPMV4} , {CPMV1, CPMV2} , {CPMV1, CPMV3} . The combination of 3 CPMVs constructs a 6-parameter affine merge candidate and the combination of 2 CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling process, if the reference indices of control points are different, the related combination of control point MVs is discarded. After inherited affine merge candidates and constructed affine merge candidate are checked, if the list is still not full, zero MVs are inserted to the end of the list.

[0058] C. Affine AMVP Prediction

[0059] Affine AMVP mode can be applied to CUs with both width and height larger than or equal to 16. An affine flag in CU level is signalled in the bitstream to indicate whether affine AMVP mode is used and then another flag is signalled to indicate whether 4-parameter affine or 6-parameter affine. In this mode, the difference of the CPMVs of current CU and their predictors CPMVPs is signalled in the bitstream. The affine AVMP candidate list size is 2 and it is generated by using the following four types of CPMV candidate in order: (1) Inherited affine AMVP candidates that extrapolated from the CPMVs of the neighbour CUs, (2) Constructed affine AMVP candidates CPMVPs that are derived using the translational MVs of the neighbour CUs, (3) Translational MVs from neighboring CUs, and (4) Zero MVs.

[0060] The checking order of inherited affine AMVP candidates is same to the checking order of inherited affine merge candidates. The only difference is that, for AVMP candidate, only the affine CU that has the same reference picture as in current block is considered. No pruning process is applied when inserting an inherited affine motion predictor into the candidate list.

[0061] Constructed AMVP candidate is derived from the specified spatial neighbors shown in FIG. 3 above. The same checking order is used as in affine merge candidate construction. In addition, reference picture index of the neighboring block is also checked. The first block in the checking order is inter coded and uses the same reference picture as in current CUs. When the current CU is coded with 4-parameter affine mode, and mv0 and mv1 are both available, they are added as one candidate in the affine AMVP list. When the current CU is coded with 6-parameter affine mode, and all three CPMVs are available, they are added as one candidate in the affine AMVP list. Otherwise, constructed AMVP candidate is set as unavailable.

[0062] If the affine AMVP candidates list still has less than 2 candidates after valid inherited affine AMVP candidates and constructed AMVP candidate are inserted, mv0, mv1, and mv2 will be added, in order, as the translational MVs to predict all control point MVs of the current CU, when available. Finally, zero MVs are used to fill the affine AMVP candidates list if the list is still not full.

[0063] D. History-Parameter based Affine Model Inheritance and Non-Adjacent Affine Mode

[0064] History-parameter-based affine model inheritance (HAMI) allows the affine model to be inherited from a previously affine-coded block which may not be neighboring to the current block. Similar to the enhanced regular merge mode, non-adjacent affine mode (NA-AFF) is introduced. A first history-parameter table (HPT) is established. An entry of the first HPT stores a set of affine parameters: a, b, c and d, each of which is represented by a 16-bit signed integer. Entries in HPT is categorized by reference list and reference index. Five reference indices are supported for each reference list in HPT. In a formular way, the category of HPT (denoted as HPTCat) is calculated as HPTCat (RefList, RefIdx) = 5×RefList + min (RefIdx, 4) ,

[0065] wherein RefList and RefIdx represents a reference picture list (0 or 1) and a reference index, respectively. For each category, at most seven entries can be stored, resulting in 70 entries totally in HPT. At the beginning of each CTU row, the number of entries for each category is initialized as zero. After decoding an affine-coded CU with reference list RefListcur and RefIdxcur, the affine parameters are utilized to update entries in the category HPTCat (RefListcur, RefIdxcur) in a way similar to HMVP table updating.

[0066] A history-affine-parameter-based candidate (HAPC) is derived from one of the seven neighbouring 4×4 blocks denoted as A0, A1, A2, B0, B1, B2 or B3 as shown in FIG. 3 and a set of affine parameters stored in a corresponding entry in the first HPT. The MV of a neighbouring 4×4 block served as the base MV. In a formulating way, the MV of the current block at position (x, y) is calculated as:

[0067] where (mvhbase, mvvbase) represents the MV of the neighboring 4×4 block, (xbase, ybase) represents the center position of the neighboring 4×4 block. (x, y) can be the top-left, top-right and bottom-left corner of the current block to obtain the corner-position MVs (CPMVs) for the current block, or it can be the center of the current block to obtain a regular MV for the current block.

[0068] A second history-parameter table (HPT) with base MV information is also appended. There are nine entries in the second HPT, wherein an entry comprises a base MV, a reference index and four affine parameters for each reference list, and a base position. An additional merge HAPC can be generated from the second HPT with the base MV information the corresponding affine models stored in an entry.

[0069] Moreover, pair-wised affine merge candidates are generated by two affine merge candidates which are history-derived or not history-derived. A pair-wised affine merge candidates is generated by averaging the CPMVs of existing affine merge candidates in the list. As a response to new HAPCs being introduced, the size of sub-block-based merge candidate list is increased from five to fifteen, which are all involved in the ARMC process.

[0070] The patterns of obtaining non-adjacent spatial neighbors for non-adjacent affine mode (NA-AFF) is the same as the existing non-adjacent regular merge candidates. The distances between non-adjacent spatial neighbors and current coding block in the NA-AFF are also defined based on the width and height of current CU.

[0071] The motion information of the non-adjacent spatial neighbors is utilized to generate additional inherited and constructed affine merge / AMVP candidates. Specifically, for inherited candidates, the same derivation process of the inherited affine merge / AMVP candidates in the VVC is kept unchanged except that the check point MVs are inherited from non-adjacent spatial neighbors. The non-adjacent spatial neighbors are checked based on their distances to the current block, i.e., from near to far. At a specific distance, only the first available neighbor (that is coded with the affine mode) from each side (e.g., the left and above) of the current block is included for inherited candidate derivation. FIGS. 5A-B illustrate spatial neighbors for deriving affine merge  / AMVP candidates. FIG. 5A shows spatial neighbors for deriving inherited candidates. FIG. 5B shows spatial neighbors for deriving the first type of constructed candidates.

[0072] As indicated by the dash arrows in FIG. 5A, the checking orders of the neighbors on the left and above sides are bottom-to-up and right-to-left, respectively. For the first type of constructed candidates, as shown in the FIG. 5B, the positions of one left and above non-adjacent spatial neighbors are firstly determined independently. After that, the location of the top-left neighbor can be determined accordingly which can enclose a rectangular virtual block together with the left and above non-adjacent neighbors.

[0073] FIG. 6 illustrates the transformation of motion information from non-adjacent neighbors to the first type of constructed affine merge / AMVP candidates. As illustrated, the motion information of the three non-adjacent neighbors is used to form the check point MVs at the top-left (A) , top-right (B) and bottom-left (C) of the virtual block, which is finally projected to the current CU to generate the corresponding constructed candidates. The NA-AFF candidates are inserted into the existing affine merge candidate list and affine AMVP candidate list according to the following orders:Affine merge mode: (1) SbTMVP candidate, if available (2) Inherited from adjacent neighbors (3) Inherited from non-adjacent neighbors (4) Constructed from adjacent neighbors (5) The first type of constructed affine candidates from non-adjacent neighbors (6) Zero MVsAffine AMVP mode: (1) Inherited from adjacent neighbors (2) Constructed from adjacent neighbors (3) Translational MVs from adjacent neighbors (4) Translational MVs from temporal neighbors (5) Inherited from non-adjacent neighbors (6) The first type of constructed affine candidates from non-adjacent neighbors (7) Zero MVs

[0074] Due to the inclusion of the additional candidates generated by NA-AFF, the size of the affine merge candidate list is increased from 5 to 15. The subgroup size of ARMC for the affine merge mode is increased from 3 to 15. In NA-AFF: (1) The area from where the non-adjacent neighbors come is restricted to be within the current CTU (i.e., no additional storage requirements for line buffer) . (2) The storage granularity for affine motion information, including CPMVs and reference indexes, is reduced from 8x8 to 16x16 (i.e., only the affine motion from the top-left 8x8 block is saved) . Additionally, the saved CPMVs are projected to each 16x16 block before storage, such that the position and size information are not needed. (3) Only the top-left and top-right CPMVs are stored (i.e., always using 4-parameter affine model for NA-AFF) .

[0075] E. Regression Based Affine Candidate Derivation

[0076] The Regression based Motion Vector Field (RMVF) derivation method provides a new variety of subblock-based merge candidate. FIG. 7 illustrates neighboring 4 x 4 subblocks that are used for RMVF parameter derivation. W and H are the width and height of the current CU. As illustrated, the motion vectors and center positions from the neighboring subblocks of the current CU are used as the input to the linear regression process to derive a set of linear model parameters.

[0077] The subblock motion field from a previous coded affine CU and the motion vectors from the adjacent subblocks of current CU are used as the input for the regression process. The predicted CPMVs for current block are derived as output. The regression based affine merge candidates are derived and added to the affine merge list. Subblock motion field from a previously coded affine CU and motion information from adjacent subblocks of a current CU are used as the input to the regression process to derive affine candidates.

[0078] The previously coded affine CU can be identified from scanning through non-adjacent positions and the affine HMVP table. Adjacent subblock information of current CU is fetched from the neighboring 4x4 sub-blocks (represented by the grey zone in FIG. 7) . For each sub-block, given a reference list, the corresponding motion vector and center coordinate of the sub-block may be used.

[0079] For each affine CU, up to 2 affine candidates can be derived. One with adjacent subblock information and one without. All the linear-regression-generated candidates are pruned and collected into one candidate sub-group, TM cost based ARMC process is applied when ARMC is enabled. Afterwards, up to N linear-regression-generated candidates are added to the affine merge list when N affine CUs are found. The number of affine candidates for ARMC is 30, the output list size is 15. II. Subblock-based temporal motion vector prediction (SbTMVP)

[0080] VVC supports the subblock-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the collocated picture to improve motion vector prediction and merge mode for CUs in the current picture.

[0081] The same collocated picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: (1) TMVP predicts motion at CU level but SbTMVP predicts motion at sub-CU level. (2) Whereas TMVP fetches the temporal motion vectors from the collocated block in the collocated picture (the collocated block is the bottom-right or center block relative to the current CU) , SbTMVP applies a motion shift before fetching the temporal motion information from the collocated picture, where the motion shift is obtained from the motion vector from one of the spatial neighboring blocks of the current CU (e.g., spatial neighboring block A1 as shown in FIG. 3) .

[0082] SbTMVP predicts the motion vectors of the sub-CUs within the current CU in two steps. In the first step, the spatial neighbor (e.g., A1) is examined. If A1 has a motion vector that uses the collocated picture as its reference picture, this motion vector is selected to be the motion shift to be applied. If no such motion is identified, then the motion shift is set to (0, 0) . In the second step, the motion shift identified in Step 1 is applied (i.e. added to the current block’s coordinates) to obtain sub-CU level motion information (motion vectors and reference indices) from the collocated picture.

[0083] FIG. 8 illustrates deriving sub-CU motion field by applying a motion shift from a spatial neighbor and scaling the motion information from the corresponding collocated sub-CUs. In the example of FIG. 8, the motion shift is set to spatial neighboring block A1’s motion. Then, for each sub-CU of the current block (i.e., current sub-CUs) , the motion information of the sub-CU’s corresponding block in the collocated picture (i.e., collocated sub-CU; it’s the smallest motion grid that covers the center sample) is used to derive the motion information for the current sub-CU. After the motion information of the collocated sub-CUs are identified, they are converted to the motion vectors and reference indices of the current sub-CUs in a similar way as the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of each temporal motion vector to that of the current CU.

[0084] In VVC, a combined subblock based merge list which contains both SbTVMP candidate and affine merge candidates is used for the signalling of subblock based merge mode. A SbTVMP candidate (such as A1) can be referred to as a motion shift candidate. The SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry of the list of subblock based merge candidates, and followed by the affine merge candidates. The size of subblock based merge list is signalled in SPS (e.g., the maximum allowed size of the subblock based merge list is 5 in VVC. ) The sub-CU size used in SbTMVP is fixed to be 8x8. SbTMVP mode is only applicable to CUs with both width and height ≥ 8.

[0085] The encoding logic of the additional SbTMVP merge candidate is the same as for the other merge candidates, that is, for each CU in P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. III. Cascaded Vector Prediction

[0086] A. Auto Relocated Block Vector Prediction (AR-BVP)

[0087] Auto-relocated block vector prediction (AR-BVP) is part of IBC merge / AMVP candidate list construction. FIG. 9 conceptually illustrates derivation of auto-relocated block vector. As illustrated, for a current picture 900, a guiding block vector BV0, 1 associated with the current block B0 points to a reference block B1. If B1 has a BV denoted as BV1, 2 pointing to a reference block B2, then BV0, 2, given by BV0, 2 = BV0, 1 +BV1, 2, is defined as the AR-BVP, guided by BV0, 1. Similarly, BV0, n+1 can be derived by

[0088] BV0, n+1 =BV0, n+BVn, n+1 = BV0, 1+BV1, 2 +…+BVn-1, n +BVn, n+1.

[0089] In some embodiments, the length of the AR-BVP trace path is 1 (i.e., n=1) . In some embodiments, the length of the AR-BVP trace path is 2 (i.e., n=2) . In some embodiments, there is no constraint for the length of the AR-BVP trace path.

[0090] FIG. 10 conceptually illustrates the positions to be checked for deriving auto-relocated block vector. As illustrated, when deriving BVn, n+1 guided by BV0, n, all five positions including top-left ( “LT” ) , top-right ( “RT” ) , center ( “Ctr” ) , bottom-left ( “LB” ) , and bottom-right ( “RB” ) positions of Bn may be checked to find BVn, n+1. In some embodiments, the initial guiding block vector BV0, 1 is set to be an existing BVP already in the IBC merge / AMVP candidate list. In some embodiments, the AR-BVP candidates may be inserted after historical BVP candidates. The IBC merge / AMVP candidate list size may be kept unchanged.

[0091] B. Chained Motion Vector Prediction (CMVP)

[0092] In some embodiments, a chained MV prediction (CMVP) is included in the inter merge candidate list construction. FIG. 11 shows an example derivation of a candidate for chained MV prediction (CMVP) . As illustrated, CMVP candidates can be derived as the sum of the recursively traced MVs and BVs based on the pre-derived MVs for the inter merge candidate list. The illustrates a CMVP candidate that is a motion vectors MVL0k / m locating a block in a reference picture a RefPicL0k / m. The motion vector and the reference picture are determined according to: MVL0k / m = MVL0k (0) + BVk (0) + MVL0k (1) +MVL0k (2) + …+ MVL0k (m) , RefPicL0k / m = RefPicL0k (m)

[0093] More generally, for a CMVP candidate, a set of motion vectors MVk / m in a reference picture RefPick / m can be derived by MVk / m = MVk (0) + BVk (0) + MVk (1) +MVk (2) + …+ MVk (m) , RefPick / m = RefPick (m) ,

[0094] where k and m indicate the number of merge index and trace depths of the CMVP.

[0095] FIG. 12 illustrates the possible sources of a base vector for CMVP. The figure shows the operations of referencing source and destination of tracing MVs in CMVP. As illustrated, when deriving MVk / m, MVk (0) (also referred to as the base vector of the CMVP) is found by checking the existence of MVs or BVs in MV / BV storage corresponding to all five positions of the current block (i.e., the Ctr, TL, TR, BL, and BR of the current block) . In the example, a MV that is found in the MV / BV storage corresponding to the center position of the current block, and that MV is used as MVL0k (0) or base vector of the CMVP.

[0096] When pre-derived merge candidates targeting CMVP candidates has two MVs, a MVk / m is derived for each list (i.e., L0 and L1) and each trace depth. Up to two MVs can be derived for each list and each trace depth, and the MV set is sequentially inserted into inter merge candidate list. The traceable reference pictures are only within the reference picture list. CMVP candidates may be inserted after HMVP candidates for the regular merge and TM merge. When deriving CMVP candidates, hpelIfIdx, bcwIdx, licFlag, and mhpFlag may not be inherited. CMVP candidates may not be derived when the TMVP is disabled. IV. Affine Prediction with Chained Motion Vectors

[0097] A. Affine inherited candidates

[0098] For some embodiments, inherited affine candidate can be derived by recursively tracing MVs and BVs based on an initial MV or BV (as base vector) . Specifically, if a spatial candidate or no-adjacent candidate is not coded by affine, the MV or BV of the spatial candidate is set as the initial MV or BV for the chained motion vector prediction (CMVP) to derive chained motion vectors. During the recursive tracing process of CMVP, when a candidate pointed by the chained motion vector is coded by affine, two or three subblock MVs in the corners of such candidate are used to derive affine models for inheritance. The derived affine models are added into the affine merge or AMVP candidate list for affine merge or AMVP mode prediction. In some embodiments, if a spatial candidate or non-adjacent candidate or history-based candidate (e.g., a base MV stored in 2nd type affine history-parameter table) is coded by affine, the center MV of such candidate is set as the initial MV or BV (or base vector) for the chained motion vector prediction.

[0099] B. Affine constructed candidates

[0100] In some embodiments, a constructed affine candidate can be derived by recursively tracing MVs and BVs based on initial MVs or BVs. Specifically, two or three of the spatial MVs or BVs (denoted as and ) in the corners of a CU are set as the initial MVs or BVs. During the recursive tracing process of CMVP, when two or three of the chained MVs (i.e.,  and ) points to the same reference picture, an constructed affine model can be derived by the chained MVs and added into the affine merge or AMVP candidate list for affine merge or AMVP mode prediction.

[0101] C. Affine regression candidates

[0102] In some embodiments, a new type of regression-based affine candidate can be derived by recursively tracing MVs and BVs based on an initial MV or BV. Specifically, if a spatial candidate or non-adjacent candidate is not coded by affine, the MV or BV of the candidate is set as the initial MV or BV for the chained motion vector prediction. During the recursive tracing process of CMVP, when a CU pointed by the chained MV is coded by affine, the subblock MVs in the reference CU and the neighboring subblock MVs of current CU are used to derive affine models by regression based affine candidate derivation. The derived affine models are added into the affine merge or AMVP candidate list for affine merge or AMVP mode prediction.

[0103] In some embodiments, if a spatial candidate or non-adjacent candidate or history-based candidate (e.g., a base MV stored in 2nd type affine history-parameter table) is coded by affine, the center MV of such candidate is set as the initial MV or BV for the CMVP process.

[0104] D. Affine history-based candidates

[0105] In some embodiments, a new type of history-based affine candidate can be derived by recursively tracing MVs and BVs based on an initial MV or BV as the base vector. Specifically, if a spatial candidate or non-adjacent candidate or history-based candidate (e.g., a base MV stored in 2nd type affine history-parameter table) is not coded by affine, the MV or BV of spatial candidate or non-adjacent candidate or history-based candidate is set as the initial MV or BV. During the recursive tracing process of CMVP, the chained MV is used as the base MV and the position of the initial MV or BV candidate is used as base position to derive affine model by fetching parameters from the affine history-parameter table. The derived affine models are added into the affine merge or AMVP candidate list for affine merge or AMVP mode prediction.

[0106] In some embodiments, if a spatial candidate or non-adjacent candidate or history-based candidate is coded by affine, the center MV of such candidate is set as the initial MV or BV for the chained motion vector prediction.

[0107] E. Affine non-adjacent candidates

[0108] In some embodiments, a non-adjacent affine candidate can be derived by recursively tracing MVs and BVs based on initial MVs or BVs as base vector. Specifically, two or three of the non-adjacent MVs or BVs in the corners of a CU (i.e., same as the search pattern of non-adjacent affine mode) are set as the initial MVs or BVs (denoted as and ) . During the recursive tracing process of CMVP, when two or three of the chained MVs (i.e.,  and ) points to the same reference picture, an non-adjacent affine model can be derived by the chained MVs and added into the affine merge or AMVP candidate list for affine merge or AMVP mode prediction.

[0109] For the above methods, the CMVP process can be performed after the construction of affine inherited, constructed, regression, history-based and non-adjacent candidates.

[0110] In some embodiments, the CU center MV derived by affine models or check point MVs of the affine inherited, constructed, regression, history-based and non-adjacent candidates is set as the initial MV for chained MV prediction. Affine inherited, regression and history-based candidates can be derived from the CMVP process. The base position of the affine model is set as the center of the CU when using CU center MV for chained MV prediction.

[0111] In some embodiments, the subblock MVs or check point MVs in the two or three corners of the affine inherited, constructed, regression, history-based and non-adjacent candidates are set as the initial MVs (denoted as and ) for chained MV prediction. New types of affine constructed and non-adjacent candidates can be derived from the derived CMVPs (i.e.,  and ) .

[0112] For the above described methods, if a spatial candidate or non-adjacent candidate or history-based candidate is coded by affine, the CU center MV derived by affine models or check point MVs is set as the initial MV for chained MV prediction. A new type of affine candidate can be derived by inheriting the affine model (i.e., a, b, c and d) of the spatial candidate or non-adjacent candidate or history-based candidate and using a chained MV as base MV.

[0113] In some embodiments, if a candidate pointed by a chained MV is coded by affine, a new type of affine candidate can be derived by blending the affine model (i.e., a, b, c, d, e and f) or check point MVs of the pointed candidate with the affine model or check point MVs of the initial spatial candidate or non-adjacent candidate or history-based candidate. In some embodiments, if a candidate pointed by chained MV is coded by affine, a new type of affine candidate can be derived by regression (same as regression-based affine derivation, which takes the blending of subblock MVs in the pointed candidate and the subblock MVs in the initial spatial candidate or non-adjacent candidate or history-based candidate as input and minimizing the differences between the blended subblock MVs and the regression-derived subblock MVs. )

[0114] In some embodiments, a bi-predictive MV candidates can be separated into L0 and L1 MV candidates during CMVP process and the affine model of current CU can be derived from L0 and L1 chained MVs independently. In some embodiments, if the subblock MVs or check point MVs in the two or three corners of the affine candidates are set as the initial MVs for chained MV prediction, the tracing MVs of the two or three subblock MVs or check point MVs point to the same reference picture, and the new affine candidates can be derived from the tracing MVs.

[0115] In some embodiments, if the center MV derived by affine models or check point MVs of the affine candidate is set as the initial MV for the CMVP process, the check point MVs of the affine candidate are scaled to point to the reference picture which is pointed by the tracing MV of the center MV. And the new affine candidates can be derived from the tracing MV and the scaled check point MVs. V. SbTMVP with Chained Motion Vectors

[0116] In some embodiments, additional motion shift candidates of SbTMVP can be generated based on chain motion vector prediction (CMVP) . For example, in some embodiments, a regular motion shift candidate (e.g., neighboring block A1) can be used as the base vector to derive a chain MV using the CMVP process described in Section III, and the chain MV so derived becomes an additional motion shift candidate of SbTMVP.

[0117] The derived chain MVs for motion shifts of SbTMVP can be uni-prediction or bi-prediction. In some embodiments, such a derived chain MV can be used as a SbTMVP motion shift candidate only if the corresponding reference picture of the uni-prediction chain motion vector is a collocated picture, or only if one of the corresponding reference pictures of the bi-prediction chain MV is a collocated picture.

[0118] In some embodiments, the chain motion vector prediction (CMVP) can be applied to further shift the subblock motions of SbTMVP. For example, in some embodiments, the subblock motions of SbTMVP are derived by scaling the motions from the collocated picture. In that, a scaling factor is derived based on the two POC differences. One is the difference between the POC of the collocated picture and the POC of the reference picture of the motion of collocated picture. The other is the difference between the POC of the current picture and the POC of the reference picture of the current block.

[0119] FIG. 13 conceptually illustrates subblock motions being derived by scaling the motions from the collocated picture. As illustrated, a motion shift based on neighboring block A1 is applied so that a current sub-CU 1315 finds its corresponding collocated sub-CU 1355 in a collocated picture 1350. Subblock motion “ScaledrefMv” of the current sub-CU 1315 of the current block in the current picture 1310 is derived by scaling the motion “refMv” of the collocated sub-CU 1355. The motion vector “refMV” is pointing at a block in a reference picture 1360. The motion vector “ScaledrefMv” is pointing at a block in a reference picture 1320. The scaling factor for scaling “refMv” into “ScaledrefMv” is determined by (i) the POC difference “ta” between collocated picture 1350 and its reference picture 1360 and (ii) the POC difference “tb” between the current picture 1310 and its reference picture 1320.

[0120] In the figure, a chained MV “ChainMV” is derived based on the scaled subblock motion “ScaledrefMv” of the current sub-CU 1315 as its base vector, and this chained MV locates a reference block in a reference picture 1330 that may be used to generate a predictor for encoding or decoding the current sub-CU 1315.

[0121] In some embodiments, the subblock motions of SbTMVP are the collocated referencing motions in the collocated picture. That is, the referenced block of each subblock in the collocated picture are used to predict the subblocks of current block. FIG. 14A conceptually illustrates using the blocks referenced by the collocated subblocks in the collocated picture for predicting the current subblocks. As illustrated, a motion shift based on neighboring block A1 is applied so that a current sub-CU 1415 finds its corresponding collocated sub-CU 1455 in a collocated picture 1350. For the current sub-CU 1415 in the current picture 1410, its corresponding collocated sub-CU 1455 has a motion “refMv” that references a block 1465 in a reference picture 1460. The reference block 1465 can be used to generate a predictor for the sub-CU 1415 of the current block.

[0122] In some embodiments, the subblock motions of SbTMVP are the motions of the referenced blocks of the collocated referencing motions in the collocated picture. That is, the referenced block of the referenced block of each subblock in the collocated picture will be used to predict the subblocks of current block. FIG. 14B conceptually illustrates using collocated subblock’s reference block’s reference block as prediction for the current subblock. In the example, the reference block 1465 that is referenced by the collocated sub-CU 1455 has motion that further references a reference block 1475 in a further reference picture 1470. The reference block 1475 can be used to generate a predictor for the sub-CU 1415 of the current block. In this example, the motion information of the collocated sub-CU 1455 can be said to be the base vector of a chain MV 1452 that locates the reference block 1475 in the further reference picture 1470.

[0123] In some embodiments, the subblock motions of SbTMVP are the motions of the referenced blocks of the collocated referencing motions in the collocated picture (e.g., a chain MV based on a collocated subblock) . A scaling factor is applied on the chain MV of the collocated subblock to derive the current subblock motion. The scaling factor is derived based on two POC differences which is used to derive the scaled subblock motion vectors for current block prediction. One is the difference between the collocated picture POC and the POC of the referenced picture containing the reference block of the referenced block of the corresponding block in the collocated picture. The other is the difference between current picture POC and the POC of the reference picture of current block. It shall be constrained that the referenced pictures of the scaled subblock motion vectors need to be in the reference picture lists of current block.

[0124] FIG. 15 conceptually illustrates scaling a chain MV of a collocated picture to be the subblock motion for the current block. As illustrated, a motion shift based on neighboring block A1 is applied so that a current sub-CU 1515 finds its corresponding collocated sub-CU 1555 in a collocated picture 1550. The current sub-CU 1515 has a corresponding collocated sub-CU 1555 in the collocated picture 1550 based on the motion shift. The collocated sub-CU 1555 has motion that references a block 1565 in a reference picture 1560, and the reference block 1565 references a further reference block 1575 in a further reference picture 1575, thereby deriving a chain MV 1552 (with the motion of the collocated sub-CU 1555 as its base vector) . The chain MV 1552 that references the reference block 1575 from the collocated sub-CU 1555 is then scaled to become the subblock motion 1512 of the current sub-CU 1515. This scaled MV 1512 references a reference block 1545 in a reference picture 1540, which is used to generate a prediction for the sub-CU 1515. The scaling factor for scaling the chain MV 1552 to the scaled MV 1512 is calculated based on (i) POC difference between the collocated picture 1550 and the reference picture 1570 and (ii) POC difference between the current picture 1510 and the reference picture 1540.

[0125] In some embodiments, after the chain motion vector prediction is applied on subblock motions of SbTMVP, the reference blocks of each derived subblock motions are used to predict current block.

[0126] In some embodiments, the motion shift candidates are not applied on block-based shift of a SbTMVP candidates, instead, a motion shift candidate may be used to shift the subblock motions found in the collocated picture. For example, MV (0, 0) is used as the motion shift of a SbTMVP candidate. That is, the corresponding location of current block in the collocated picture will be searched for subblock motions of the SbTMVP candidate. After that, N subblock motions may be found and they will be shifted together according to the motion vector referenced from the spatial neighbor A1 of current block. N is related to the number of subblocks of the SbTMVP coded block. And then, the subblock motions will be scaled and used to predict current block. In the above example, the motion used to shift the subblock motions (the MV referenced from spatial neighbor A1 in the above example) can be replaced by other motions referenced from other spatial positions, non-adjacent positions, or from HMVP table, or the motions derived by chain MV technology. In some of these embodiments, the motion used to shift the subblock motions shall be a motion from the collocated picture. (In the above example, MV (0, 0) can be replaced by any other neighboring MVs, non-adjacnet MVs, HMVPs, or chain MVs. )

[0127] In some embodiments, a flag is contained in merge information to indicate whether SbTMVP with subblock motion shift mode is applied or not. If the flag is not enabled, traditional SbTMVP will be applied. In some embodiments, the reference picture of each subblock on a SbTMVP block is constrained to be the same, but not limit to. In some embodiments, the motions used to shift the subblock motions shall be the motions from the collocated picture. but not limit to.

[0128] In some embodiments, an additional SbTMVP list is derived including more than one SbTMVPs derived by chain motion vector prediction. In some embodiments, a TM reordering technology can be applied to the additional SbTMVP list. Only the first N candidates in the list with lower TM costs will be added to affine candidates list and compete with traditional SbTMVP and affine candidates.

[0129] In some embodiments, an additional SbTMVP list is derived including SbTMVPs derived by chain motion vector prediction and SbTMVPs derived by traditional way. In some embodiments, if one of subblock motion of a SbTMVP candidate is not available (e.g., the corresponding subblock is coded by intra or IBC, or the referenced picture of a subblock of a SbTMVP derived by chain motion vector prediction is not included in current reference picture list) , the candidate is treated as invalid.

[0130] In some embodiments, if one of subblock motion of a SbTMVP candidate is not available (e.g., the corresponding subblock is coded by intra or IBC, or the referenced picture of a subblock of a SbTMVP derived by chain motion vector prediction is not included in current reference picture list) , the center derived subblock motion will be used. In some embodiments, if one of subblock motion of a SbTMVP candidate is not available (e.g., the referenced picture (refPic1) of a subblock of a SbTMVP derived by chain motion vector prediction is not included in current reference picture list) , the corresponding referenced block in a picture with the POC closest to refPic1 and in current reference picture list will be used to predict current subblock.

[0131] In some embodiments, if one of subblock motion of a SbTMVP candidate is not available (i.e., the referenced picture (refPic1) of a subblock of a SbTMVP derived by chain motion vector prediction is not included in current reference picture list) , the motion will be scaled to the picture with the POC closest to refPic1 and in current reference picture list to predict current subblock.

[0132] For some embodiments, the above-mentioned methods are not limited for SbTMVP. For example, in some embodiments, the chain motion vector prediction can be used for inter merge mode. If the reference picture (refPic2) of the derived chain motion vectors is not in the current reference picture list, the picture (targetPic) with POC closet to refPic2 and in current reference picture list will be used to predict current block. And the corresponding motions will be scaled to targetPic according to the POC difference between current picture and targetPic. For another example, in some embodiments, the corresponding motions are not scaled for current block prediction, on the contrary, it will be directly used to get a block in targetPic for current block prediction.

[0133] In some embodiments, an on / off control flag is signaled at CU level, slice level, picture level, and / or sequence level to indicate the proposed method in the above is enabled or not. In some embodiments, the proposed method in the above is enabled or disabled, according to one or the combination of the selected reference pictures indices, temporal distance between reference picture and current picture, quantization parameter, the coded information of current CU, prediction mode, motion vectors, motion vector resolution, residual of current CU, and reference samples. VI. Reference Picture Resampling (RPR)

[0134] In some embodiment, a reference picture with different width or height compared to current picture can be a collocated picture and TMVP is enabled on the reference picture. To obtain TMVP on the collocated picture with different width or height, the position scaling on the collocated picture can be applied.

[0135] FIG. 16 illustrates position scaling for a collocated picture. As illustrated, to obtain the TMVP on the position (x, y) of current picture, the position (x’, y’) located in the collocated picture should be used. The position of (x’, y’) is derived according to the width and height ratio of current picture and collocated picture. For example, to get the TMVP on position (64, 64) of current picture with width 1920 and height 1080, the corresponding position is (128, 128) in collocated picture with width 3840 and height 2160.

[0136] In some embodiments, for SbTMVP case on collocated picture with different width or height, the position scaling method can be applied after the motion shift. For example, to obtain the subblock MVP on position (64, 64) of current picture with width 1920 and height 1080, and the derived motion shift is (4, 4) . In one case, the corresponding position is (68, 68) in collocated picture with same width and height. In another case, the corresponding position is (136, 136) in collocated picture with width 3840 and height 2160.

[0137] In some embodiments, the position scaling can be applied during the CMVP derivation process to handle the different width or height of reference pictures. Whenever the width or height of two neighboring reference pictures in the chain are different, the position scaling should be applied. For some embodiments, the position scaling may be applied to the process of obtaining any information from other pictures with different width or height.

[0138] Referring to the example of FIG. 11, if the width or height of reference picture RefPicL0k (1) is different from than that of the reference picture RefPicL0k (0) , the position in RefPicL0k (1) used to derive next motion vector component (i.e., MVL0k (2) ) should be scaled according to the width and height ratio of RefPicL0k (1) and RefPicL0k (0) .

[0139] In some embodiments, the MV candidate is considered for the chained MVP construction (or prioritized in the existing list) , if the difference between the current POC and the refence POC (e.g., RefPicL0k (0) ) has the same sign as the difference between the current POC and the POC of the reference picture to which the MVP candidate is pointing (e.g., RefPicL0k (1) ) . This means that those CMVPs of the cumulative MVP which are not changing the direction of the prediction from forward to backward, or from backward to forward, should be either used for the construction, or prioritized in the existing list.

[0140] In some embodiments, in case when the MVP candidate (MVPk (1) ) is bi-directional, the MVP which is pointing in the same direction, is chosen between the two, for the cumulative MVP construction (e.g., MVL0k (1) ) . In case when both MVPs are pointing in the same prediction direction (as the direction between the current picture and the reference picture) , the MVP pointing to the closest POC to the POC of the reference picture (e.g., RefPicL0k (0) ) shall be chosen.

[0141] The foregoing proposed methods can be implemented in encoders and / or decoders. For example, the proposed method can be implemented in an in-loop filtering module of an encoder, and / or an in-loop filtering module of a decoder. VII. Example Video Encoder

[0142] FIG. 17 illustrates an example video encoder 1700 that may implement subblock motion. As illustrated, the video encoder 1700 receives input video signal from a video source 1705 and encodes the signal into bitstream 1795. The video encoder 1700 has several components or modules for encoding the signal from the video source 1705, at least including some components selected from a transform module 1710, a quantization module 1711, an inverse quantization module 1714, an inverse transform module 1715, an intra-picture estimation module 1724, an intra-prediction module 1725, a motion compensation module 1730, a motion estimation module 1735, an in-loop filter 1745, a reconstructed picture buffer 1750, a MV buffer 1765, and a MV prediction module 1775, and an entropy encoder 1790. The motion compensation module 1730 and the motion estimation module 1735 are part of an inter-prediction module 1740. The intra-prediction module 1725 and the intra-prediction estimation module 1724 are part of a current picture prediction module 1720, which uses current picture reconstructed samples as reference samples for prediction of the current block.

[0143] In some embodiments, the modules 1710 –1790 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device or electronic apparatus. In some embodiments, the modules 1710 –1790 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Though the modules 1710 –1790 are illustrated as being separate modules, some of the modules can be combined into a single module.

[0144] The video source 1705 provides a raw video signal that presents pixel data of each video frame without compression. A subtractor 1708 computes the difference between the raw video pixel data of the video source 1705 and the predicted pixel data 1713 from the motion compensation module 1730 or intra-prediction module 1725 as prediction residual 1709. The transform module 1710 converts the difference (or the residual pixel data or residual signal 1708) into transform coefficients (e.g., by performing Discrete Cosine Transform, or DCT) . The quantization module 1711 quantizes the transform coefficients into quantized data (or quantized coefficients) 1712, which is encoded into the bitstream 1795 by the entropy encoder 1790.

[0145] The inverse quantization module 1714 de-quantizes the quantized data (or quantized coefficients) 1712 to obtain transform coefficients 1718, and the inverse transform module 1715 performs inverse transform on the transform coefficients 1718 to produce reconstructed residual 1719. The reconstructed residual 1719 is added with the predicted pixel data 1713 to produce reconstructed pixel data 1717. In some embodiments, the reconstructed pixel data 1717 is temporarily stored in a line buffer 1727 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction. The reconstructed pixels are filtered by the in-loop filter 1745 and stored in the reconstructed picture buffer 1750. In some embodiments, the reconstructed picture buffer 1750 is a storage external to the video encoder 1700. In some embodiments, the reconstructed picture buffer 1750 is a storage internal to the video encoder 1700.

[0146] The intra-picture estimation module 1724 performs intra-prediction based on the reconstructed pixel data 1717 to produce intra prediction data. The intra-prediction data is provided to the entropy encoder 1790 to be encoded into bitstream 1795. The intra-prediction data is also used by the intra-prediction module 1725 to produce the predicted pixel data 1713.

[0147] The motion estimation module 1735 performs inter-prediction by producing MVs to reference pixel data of previously decoded frames stored in the reconstructed picture buffer 1750. These MVs are provided to the motion compensation module 1730 to produce predicted pixel data.

[0148] Instead of encoding the complete actual MVs in the bitstream, the video encoder 1700 uses MV prediction to generate predicted MVs, and the difference between the MVs used for motion compensation and the predicted MVs is encoded as residual motion data and stored in the bitstream 1795.

[0149] The MV prediction module 1775 generates the predicted MVs based on reference MVs that were generated for encoding previously video frames, i.e., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 1775 retrieves reference MVs from previous video frames from the MV buffer 1765. The video encoder 1700 stores the MVs generated for the current video frame in the MV buffer 1765 as reference MVs for generating predicted MVs.

[0150] The MV prediction module 1775 uses the reference MVs to create the predicted MVs. The predicted MVs can be computed by spatial MV prediction or temporal MV prediction. The difference between the predicted MVs and the motion compensation MVs (MC MVs) of the current frame (residual motion data) are encoded into the bitstream 1795 by the entropy encoder 1790.

[0151] The entropy encoder 1790 encodes various parameters and data into the bitstream 1795 by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding. The entropy encoder 1790 encodes various header elements, flags, along with the quantized transform coefficients 1712, and the residual motion data as syntax elements into the bitstream 1795. The bitstream 1795 is in turn stored in a storage device or transmitted to a decoder over a communications medium such as a network.

[0152] The in-loop filter 1745 performs filtering or smoothing operations on the reconstructed pixel data 1717 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 1745 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.

[0153] FIG. 18 illustrates portions of the video encoder 1700 that implement subblock processing and cascading vectors. As illustrated, a subblock processing module 1810 receives a motion shift indication, which is selected from a list of motion shift candidates 1820 by the entropy encoder 1790. The motion shift candidates 1820 include neighboring blocks of the current block (e.g., A1), as well as cascaded vectors that are derived based on the motion information of those neighboring blocks as base vectors. The motion information of the neighboring blocks are provided by the MV buffer 1765, which also provides MVs and BVs to a vector cascader 1830 to generate the cascaded vectors.

[0154] The subblock processing module 1810 performs several functions: collocated subblocks identification 1811, vector cascading 1812, temporal scaling 1813, and positional scaling 1814. The collocated subblock identification function 1811 uses the selected motion shift to identify the collocated picture and the collocated subblocks that correspond to the current subblocks. The vector cascading function 1812 creates cascaded vectors (or chain MVs) based on certain base vectors (the base vector may be the motion information of the collocated subblocks, or the motion information of the collocated subblocks temporally scaled to the current picture, as described by reference to FIGS. 13-15 above. ) The temporal scaling function 1813 may temporally scale motion information based on scale factors that are computed based on (i) temporal differences between the current picture and its subblocks’ reference picture (s) and (ii) temporal difference between the collocated picture and its subblocks’ reference picture (s) . The positional scaling function 1814 may positionally scale the collocated picture if the collocated picture is of a different height  / width than the current picture.

[0155] The subblock processing module 1810 outputs a motion field for the current subblocks (i.e., subblocks of the current block) to the inter prediction module 1740. The motion field is determined by the subblocks identification 1811, the vector cascading 1812, the temporal scaling 1813 functions. The inter prediction module 1740 fetches samples from the reconstructed picture buffer 1750 and the line buffer 1727 according to the motion field to generate the predictors for the subblocks of the current block to become the prediction pixel data 1713.

[0156] FIG. 19 conceptually illustrates a process 1900 that uses cascaded vectors for encoding subblocks. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the encoder 1700 performs the process 1900 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the encoder 1700 performs the process 1900.

[0157] The encoder receives (at block 1910) data to be encoded as a current block of pixels of a current picture of a video. The current block includes a plurality of current subblocks.

[0158] The encoder identifies (at block 1920) a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift. In some embodiments, the encoder applies position scaling to the collocated picture when the collocated picture has a different width or height than that of the current picture. In some embodiments, the motion shift is provided by a cascaded vector (or chain MV) that is a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector. The cascaded vector has a base vector that is a motion vector or block vector a neighboring block of the current block. In some embodiments, a list of motion shift candidates includes (i) the block vector or the motion vector of the neighboring block that is the base vector of the cascaded vector and (ii) the cascaded vector.

[0159] The encoder derives (at block 1930) predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks. In some embodiments, the predictors for the plurality of current subblocks are derived by scaling the motion information of the plurality of collocated subblocks according to temporal distances among the current picture, the collocated picture, and reference pictures of the motion information of the plurality of collocated subblocks.

[0160] In some embodiments, a cascaded vector (or chained MV) is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock, the cascaded vector being a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector.

[0161] In some embodiments, the motion information of the collocated subblock is temporally scaled to the current picture to be used as a base vector for deriving the cascaded vector to identify a reference block from the current subblock, as described by reference to FIG. 13 above, with the scaling being applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the motion information of the collocated subblock (ii) a second temporal difference between the current picture and a reference picture identified by the cascaded vector.

[0162] In some embodiments, the motion information of the collocated subblock is used as a base vector for deriving the cascaded vector. The predictor of the current subblock may be derived based on a reference block that is identified by the cascaded vector from the collocated subblock, as described by reference to FIG. 14B. In some embodiments, the predictor of the current subblock may be derived based on a reference block that is identified by a scaled motion vector from the current subblock, the scaled motion vector derived by temporally scaling the cascaded vector to the current picture, as described by reference to FIG. 15 above, with the scaling being applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the cascaded vector (ii) a second temporal difference between the current picture and a reference picture identified by the scaled motion vector.

[0163] The encoder encodes (at block 1940) the current block by using the derived predictors of the plurality current subblocks to produce prediction residuals. VIII. Example Video Decoder

[0164] In some embodiments, an encoder may signal (or generate) one or more syntax element in a bitstream, such that a decoder may parse said one or more syntax element from the bitstream.

[0165] FIG. 20 illustrates an example video decoder 2000 that may implement subblock motion. As illustrated, the video decoder 2000 is an image-decoding or video-decoding circuit that receives a bitstream 2095 and decodes the content of the bitstream into pixel data of video frames for display. The video decoder 2000 has several components or modules for decoding the bitstream 2095, including some components selected from an inverse quantization module 2014, an inverse transform module 2015, an intra-prediction module 2025, a motion compensation module 2030, an in-loop filter 2045, a decoded picture buffer 2050, a MV buffer 2065, a MV prediction module 2075, and a parser 2090. The motion compensation module 2030 is part of an inter-prediction module 2040. The intra-prediction module 2025 is part of a current picture prediction module 2020, which uses current picture reconstructed samples as reference samples for prediction of the current block.

[0166] In some embodiments, the modules 2014 –2090 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device. In some embodiments, the modules 2014 –2090 are modules of hardware circuits implemented by one or more ICs of an electronic apparatus. Though the modules 2014 –2090 are illustrated as being separate modules, some of the modules can be combined into a single module.

[0167] The parser 2090 (or entropy decoder) receives the bitstream 2095 and performs initial parsing according to the syntax defined by a video-coding or image-coding standard. The parsed syntax element includes various header elements, flags, as well as quantized data (or quantized coefficients) 2012. The parser 2090 parses out the various syntax elements by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding.

[0168] The inverse quantization module 2014 de-quantizes the quantized data (or quantized coefficients) 2012 to obtain transform coefficients, and the inverse transform module 2015 performs inverse transform on the transform coefficients 2018 to produce reconstructed residual signal 2019. The reconstructed residual signal 2019 is added with predicted pixel data 2013 from the intra-prediction module 2025 or the motion compensation module 2030 to produce decoded pixel data 2017. The decoded pixels data are filtered by the in-loop filter 2045 and stored in the decoded picture buffer 2050. In some embodiments, the decoded picture buffer 2050 is a storage external to the video decoder 2000. In some embodiments, the decoded picture buffer 2050 is a storage internal to the video decoder 2000.

[0169] The intra-prediction module 2025 receives intra-prediction data from bitstream 2095 and according to which, produces the predicted pixel data 2013 from the decoded pixel data 2017 stored in the decoded picture buffer 2050. In some embodiments, the decoded pixel data 2017 is also stored in a line buffer 2027 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction.

[0170] In some embodiments, the content of the decoded picture buffer 2050 is used for display. A display device 2005 either retrieves the content of the decoded picture buffer 2050 for display directly, or retrieves the content of the decoded picture buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 2050 through a pixel transport.

[0171] The motion compensation module 2030 produces predicted pixel data 2013 from the decoded pixel data 2017 stored in the decoded picture buffer 2050 according to motion compensation MVs (MC MVs) . These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 2095 with predicted MVs received from the MV prediction module 2075.

[0172] The MV prediction module 2075 generates the predicted MVs based on reference MVs that were generated for decoding previous video frames, e.g., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 2075 retrieves the reference MVs of previous video frames from the MV buffer 2065. The video decoder 2000 stores the motion compensation MVs generated for decoding the current video frame in the MV buffer 2065 as reference MVs for producing predicted MVs.

[0173] The in-loop filter 2045 performs filtering or smoothing operations on the decoded pixel data 2017 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 2045 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.

[0174] FIG. 21 illustrates portions of the video decoder 2000 that implement subblock processing and cascading vectors. As illustrated, a subblock processing module 2110 receives a motion shift indication, which is selected from a list of motion shift candidates 2120 by the entropy decoder 2090. The motion shift candidates 2120 include neighboring blocks of the current block (e.g., A1), as well as cascaded vectors that are derived based on the motion information of those neighboring blocks as base vectors. The motion information of the neighboring blocks are provided by the MV buffer 2065, which also provides MVs and BVs to a vector cascader 2130 to generate the cascaded vectors.

[0175] The subblock processing module 2110 performs several functions: collocated subblocks identification 2111, vector cascading 2112, temporal scaling 2113, and positional scaling 2114. The collocated subblock identification function 2111 uses the selected motion shift to identify the collocated picture and the collocated subblocks that correspond to the subblocks of the current block. The vector cascading function 2112 creates cascaded vectors (or chain MVs) based on certain base vectors (the base vector may be the motion information of the collocated subblocks, or the motion information of the collocated subblocks temporally scaled to the current picture, as described by reference to FIGS. 13-15 above. ) The temporal scaling function 2113 may temporally scale motion information based on scale factors that are computed based on (i) temporal differences between the current picture and its subblocks’ reference picture (s) and (ii) temporal difference between the collocated picture and its subblocks’ reference picture (s) . The positional scaling function 2114 may positionally scale the collocated picture if the collocated picture is of a different height  / width than the current picture.

[0176] The subblock processing module 2110 outputs a motion field for the current subblocks (i.e., subblocks of the current block) to the inter prediction module 2040. The motion field is determined by the subblocks identification 2111, the vector cascading 2112, the temporal scaling 2113 functions. The inter prediction module 2040 fetches samples from the decoded picture buffer 2050 and the line buffer 2027 according to the motion field to generate the predictors for the subblocks of the current block to become the prediction pixel data 2013.

[0177] FIG. 22 conceptually illustrates a process 2200 that uses cascaded vectors for decoding subblocks. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the decoder 2000 performs the process 2200 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the decoder 2000 performs the process 2200.

[0178] The decoder receives (at block 2210) data to be decoded as a current block of pixels of a current picture of a video. The current block includes a plurality of current subblocks.

[0179] The decoder identifies (at block 2220) a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift. In some embodiments, the decoder applies position scaling to the collocated picture when the collocated picture has a different width or height than that of the current picture. In some embodiments, the motion shift is provided by a cascaded vector (or chain MV) that is a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector. The cascaded vector has a base vector that is a motion vector or block vector a neighboring block of the current block. In some embodiments, a list of motion shift candidates includes (i) the block vector or the motion vector of the neighboring block that is the base vector of the cascaded vector and (ii) the cascaded vector.

[0180] The decoder derives (at block 2230) predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks. In some embodiments, the predictors for the plurality of current subblocks are derived by scaling the motion information of the plurality of collocated subblocks according to temporal distances among the current picture, the collocated picture, and reference pictures of the motion information of the plurality of collocated subblocks.

[0181] In some embodiments, a cascaded vector (or chained MV) is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock, the cascaded vector being a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector.

[0182] In some embodiments, the motion information of the collocated subblock is temporally scaled to the current picture to be used as a base vector for deriving the cascaded vector to identify a reference block from the current subblock, as described by reference to FIG. 13 above, with the scaling being applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the motion information of the collocated subblock (ii) a second temporal difference between the current picture and a reference picture identified by the cascaded vector.

[0183] In some embodiments, the motion information of the collocated subblock is used as a base vector for deriving the cascaded vector. The predictor of the current subblock may be derived based on a reference block that is identified by the cascaded vector from the collocated subblock, as described by reference to FIG. 14B. In some embodiments, the predictor of the current subblock may be derived based on a reference block that is identified by a scaled motion vector from the current subblock, the scaled motion vector derived by temporally scaling the cascaded vector to the current picture, as described by reference to FIG. 15 above, with the scaling being applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the cascaded vector (ii) a second temporal difference between the current picture and a reference picture identified by the scaled motion vector.

[0184] The decoder reconstructs (at block 2240) the current block by using the derived predictors of the plurality current subblocks and corresponding prediction residuals. The decoder may then provide the reconstructed current block for display as part of the reconstructed current picture. IX. Example Electronic System

[0185] Many of the above-described features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer readable storage medium (also referred to as computer readable medium) . When these instructions are executed by one or more computational or processing unit (s) (e.g., one or more processors, cores of processors, or other processing units) , they cause the processing unit (s) to perform the actions indicated in the instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, random-access memory (RAM) chips, hard drives, erasable programmable read only memories (EPROMs) , electrically erasable programmable read-only memories (EEPROMs) , etc. The computer readable media does not include carrier waves and electronic signals passing wirelessly or over wired connections.

[0186] In this specification, the term “software” is meant to include firmware residing in read-only memory or applications stored in magnetic storage which can be read into memory for processing by a processor. Also, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while remaining distinct software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement a software invention described here is within the scope of the present disclosure. In some embodiments, the software programs, when installed to operate on one or more electronic systems, define one or more specific machine implementations that execute and perform the operations of the software programs.

[0187] FIG. 23 conceptually illustrates an electronic system 2300 with which some embodiments of the present disclosure are implemented. The electronic system 2300 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc. ) , phone, PDA, or any other sort of electronic device. Such an electronic system includes various types of computer readable media and interfaces for various other types of computer readable media. Electronic system 2300 includes a bus 2305, processing unit (s) 2310, a graphics-processing unit (GPU) 2315, a system memory 2320, a network 2325, a read-only memory 2330, a permanent storage device 2335, input devices 2340, and output devices 2345.

[0188] The bus 2305 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of the electronic system 2300. For instance, the bus 2305 communicatively connects the processing unit (s) 2310 with the GPU 2315, the read-only memory 2330, the system memory 2320, and the permanent storage device 2335.

[0189] From these various memory units, the processing unit (s) 2310 retrieves instructions to execute and data to process in order to execute the processes of the present disclosure. The processing unit (s) may be a single processor or a multi-core processor in different embodiments. Some instructions are passed to and executed by the GPU 2315. The GPU 2315 can offload various computations or complement the image processing provided by the processing unit (s) 2310.

[0190] The read-only-memory (ROM) 2330 stores static data and instructions that are used by the processing unit (s) 2310 and other modules of the electronic system. The permanent storage device 2335, on the other hand, is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 2300 is off. Some embodiments of the present disclosure use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as the permanent storage device 2335.

[0191] Other embodiments use a removable storage device (such as a floppy disk, flash memory device, etc., and its corresponding disk drive) as the permanent storage device. Like the permanent storage device 2335, the system memory 2320 is a read-and-write memory device. However, unlike storage device 2335, the system memory 2320 is a volatile read-and-write memory, such a random access memory. The system memory 2320 stores some of the instructions and data that the processor uses at runtime. In some embodiments, processes in accordance with the present disclosure are stored in the system memory 2320, the permanent storage device 2335, and / or the read-only memory 2330. For example, the various memory units include instructions for processing multimedia clips in accordance with some embodiments. From these various memory units, the processing unit (s) 2310 retrieves instructions to execute and data to process in order to execute the processes of some embodiments.

[0192] The bus 2305 also connects to the input and output devices 2340 and 2345. The input devices 2340 enable the user to communicate information and select commands to the electronic system. The input devices 2340 include alphanumeric keyboards and pointing devices (also called “cursor control devices” ) , cameras (e.g., webcams) , microphones or similar devices for receiving voice commands, etc. The output devices 2345 display images generated by the electronic system or otherwise output data. The output devices 2345 include printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD) , as well as speakers or similar audio output devices. Some embodiments include devices such as a touchscreen that function as both input and output devices.

[0193] Finally, as shown in FIG. 23, bus 2305 also couples electronic system 2300 to a network 2325 through a network adapter (not shown) . In this manner, the computer can be a part of a network of computers (such as a local area network ( “LAN” ) , a wide area network ( “WAN” ) , or an Intranet, or a network of networks, such as the Internet. Any or all components of electronic system 2300 may be used in conjunction with the present disclosure.

[0194] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as computer-readable storage media, machine-readable media, or machine-readable storage media) . Some examples of such computer-readable media include RAM, ROM, read-only compact discs (CD-ROM) , recordable compact discs (CD-R) , rewritable compact discs (CD-RW) , read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM) , a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc. ) , flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc. ) , magnetic and / or solid state hard drives, read-only and recordable  discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media may store a computer program that is executable by at least one processing unit and includes sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter.

[0195] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, many of the above-described features and applications are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) . In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In addition, some embodiments execute software stored in programmable logic devices (PLDs) , ROM, or RAM devices.

[0196] As used in this specification and any claims of this application, the terms “computer” , “server” , “processor” , and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. For the purposes of the specification, the terms display or displaying means displaying on an electronic device. As used in this specification and any claims of this application, the terms “computer readable medium, ” “computer readable media, ” and “machine readable medium” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.

[0197] While the present disclosure has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the present disclosure can be embodied in other specific forms without departing from the spirit of the present disclosure. In addition, a number of the figures (including FIG. 19 and FIG. 22) conceptually illustrate processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in one continuous series of operations, and different specific operations may be performed in different embodiments. Furthermore, the process could be implemented using several sub-processes, or as part of a larger macro process. Thus, one of ordinary skill in the art would understand that the present disclosure is not to be limited by the foregoing illustrative details, but rather is to be defined by the appended claims. Additional Notes

[0198] The herein-described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are merely examples, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively "associated" such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as "associated with" each other such that the desired functionality is achieved, irrespective of architectures or intermediate components. Likewise, any two components so associated can also be viewed as being "operably connected" , or "operably coupled" , to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being "operably couplable" , to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable and / or physically interacting components and / or wirelessly interactable and / or wirelessly interacting components and / or logically interacting and / or logically interactable components.

[0199] Further, with respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.

[0200] Moreover, it will be understood by those skilled in the art that, in general, terms used herein, and especially in the appended claims, e.g., bodies of the appended claims, are generally intended as “open” terms, e.g., the term “including” should be interpreted as “including but not limited to, ” the term “having” should be interpreted as “having at least, ” the term “includes” should be interpreted as “includes but is not limited to, ” etc. It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim recitation to implementations containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an, " e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more; ” the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number, e.g., the bare recitation of "two recitations, " without other modifiers, means at least two recitations, or two or more recitations. Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. In those instances where a convention analogous to “at least one of A, B, or C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B. ”

[0201] From the foregoing, it will be appreciated that various implementations of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various implementations disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1.A video coding method comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video, the current block comprising a plurality of current subblocks;identifying a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift provided by a cascaded vector that is a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector;deriving predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks; andencoding or decoding the current block by using the derived predictors of the plurality current subblocks.2.The video coding method of claim 1, wherein the cascaded vector has a base vector that is a motion vector or block vector of a neighboring block of the current block.3.The video coding method of claim 2, wherein a list of motion shift candidates comprises (i) the block vector or the motion vector of the neighboring block that is base vector of the cascaded vector and (ii) the cascaded vector.4.The video coding method of claim 1, wherein the predictors for the plurality of current subblocks are derived by scaling the motion information of the plurality of collocated subblocks according to temporal distances among the current picture, the collocated picture, and reference pictures of the motion information of the plurality of collocated subblocks.5.The video coding method of claim 1, further comprising applying position scaling to the collocated picture when the collocated picture has a different width or height than that of the current picture.6.A video coding method comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video, the current block comprising a plurality of current subblocks;identifying a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift;deriving predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks, wherein a cascaded vector is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock, the cascaded vector being a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector; andencoding or decoding the current block by using the derived predictors of the plurality current subblocks.7.The video coding method of claim 6, wherein the motion information of the collocated subblock is used as a base vector for deriving the cascaded vector.8.The video coding method of claim 7, wherein the predictor of the current subblock is derived based on a reference block that is identified by the cascaded vector from the collocated subblock.9.The video coding method of claim 7, wherein the predictor of the current subblock is derived based on a reference block that is identified by a temporally scaled motion vector from the current subblock, the scaled motion vector derived by temporally scaling the cascaded vector to the current picture.10.The video coding method of claim 9, wherein said scaling is applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the cascaded vector (ii) a second temporal difference between the current picture and a reference picture identified by the scaled motion vector.11.The video coding method of claim 6, wherein the motion information of the collocated subblock is temporally scaled to the current picture to be used as a base vector for deriving the cascaded vector to identify a reference block from the current subblock.12.The video coding method of claim 11, wherein said scaling is applied according to a scaling factor that is calculated based on (i) a first temporal difference between the collocated picture and a reference picture identified by the motion information of the collocated subblock (ii) a second temporal difference between the current picture and a reference picture identified by the cascaded vector.13.An electronic apparatus comprising:a video coder circuit configured to perform operations comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video, the current block comprising a plurality of current subblocks;identifying a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift;deriving predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks, wherein a cascaded vector is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock, the cascaded vector being a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector; andencoding or decoding the current block by using the derived predictors of the plurality current subblocks.14.A video decoding method comprising:receiving data to be decoded as a current block of pixels of a current picture of a video, the current block comprising a plurality of current subblocks;identifying a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift;deriving predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks, wherein a cascaded vector is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock, the cascaded vector being a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector; andreconstructing the current block by using the derived predictors of the plurality current subblocks.15.A video encoding method comprising:receiving data to be encoded as a current block of pixels of a current picture of a video, the current block comprising a plurality of current subblocks;identifying a plurality of collocated subblocks in a collocated picture that correspond to the plurality of current subblocks based on a motion shift;deriving predictors for the plurality of current subblocks based on motion information of the plurality of collocated subblocks, wherein a cascaded vector is derived based on the motion information of a collocated subblock to derive a predictor for a current subblock, the cascaded vector being a sum of at least two recursively traced vectors, each traced vector being a motion vector or a block vector; andencoding the current block by using the derived predictors of the plurality current subblocks.

Citation Information

Patent Citations

  • Subblock-based temporal motion vector predictor with motion vector offset

    CN117256144A

  • Parking site management system using raider sensor and operating method thereof

    KR1020220138448A

  • Subblock level temporal motion vector prediction with multiple displacement vector predictors and an offset

    WO2023229664A1

Cited By

  • Scaling and reordering of chained motion vector prediction for video coding

    US20250317570A1