Motion vector prediction for video coding
By deriving unscaled motion vector predictors from spatial neighboring blocks and optimizing the MVP candidate list, the method addresses inefficiencies in video coding, enhancing prediction accuracy and reducing complexity.
Patent Information
- Application Number
- JP2025090014
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-20
AI Technical Summary
Existing video coding techniques face challenges in efficiently compressing video data while maintaining video quality, particularly in reducing the complexity and improving the accuracy of motion vector prediction.
The method involves obtaining spatial neighboring blocks to derive at most one unscaled motion vector predictor (MVP) from left and above blocks, constructing an MVP candidate list by reducing the likelihood of selecting scaled MVPs, and selecting the best MVP based on a cost value to improve motion vector prediction.
This approach enhances the efficiency and accuracy of motion vector prediction, reducing the complexity and improving coding efficiency in video coding processes.
Smart Images

Figure 2025122190000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to U.S. Provisional Patent Application No. 62 / 861,315, filed June 13, 2019, the entire disclosure of which is incorporated herein by reference in its entirety.
[0002] FIELD OF THE DISCLOSURE This disclosure relates to video encoding and compression, and more particularly, to methods and apparatus for motion vector prediction in video encoding. [Background technology]
[0003] Various video coding techniques may be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), Joint Exploration Model (JEM), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, etc. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit redundancy present in a video image or sequence. An important goal of video coding techniques is to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality. Summary of the Invention [Problem to be solved by the invention]
[0004] Examples of this disclosure provide methods and apparatus for motion vector prediction in video coding. [Means for solving the problem]
[0005] According to a first aspect of the present disclosure, a method of decoding a video signal is provided. The method may include, at a decoder, obtaining a video block from the video signal. The method may further include, at the decoder, obtaining spatial neighboring blocks based on the video block. The spatial neighboring blocks may include a plurality of left spatial neighboring blocks and a plurality of above spatial neighboring blocks. The method may include, at the decoder, obtaining at most one unscaled left motion vector predictor (MVP) based on the plurality of left spatial neighboring blocks. The method may also include, at the decoder, obtaining at most one unscaled above MVP based on the plurality of above spatial neighboring blocks. The method may include, at the decoder, deriving an MVP candidate list based on the video block, the plurality of left spatial neighboring blocks, and the plurality of above spatial neighboring blocks by reducing a likelihood of selecting a scaled MVP derived from the spatial neighboring blocks. The MVP candidate list may include at most one unscaled left MVP and at most one unscaled above MVP. The method may further include, at the decoder, receiving a best MVP based on the MVP candidate list. The encoder selects the best MVP from the MVP candidate list. The method may include, at a decoder, obtaining a prediction signal for the video block based on the best MVP.
[0006] According to a second aspect of the present disclosure, a computing device for decoding a video signal is provided. The computing device may include one or more processors and a non-transitory computer-readable memory storing instructions executable by the one or more processors. The one or more processors may be configured to obtain a video block from the video signal. The one or more processors may be further configured to obtain spatial neighboring blocks based on the video block. The spatial neighboring blocks may include a plurality of left spatial neighboring blocks and a plurality of above spatial neighboring blocks. The one or more processors may be configured to obtain up to one unscaled left motion vector predictor (MVP) based on the plurality of left spatial neighboring blocks. The one or more processors may also be configured to obtain up to one unscaled above MVP based on the plurality of above spatial neighboring blocks. The one or more processors may be configured to derive an MVP candidate list based on the video block, the plurality of left spatial neighboring blocks, and the plurality of above spatial neighboring blocks by reducing the likelihood of selecting a scaled MVP derived from the spatial neighboring blocks. The MVP candidate list may include up to one unscaled left MVP and up to one unscaled above MVP. The one or more processors may be further configured to receive a best MVP based on the MVP candidate list. The encoder selects the best MVP from the MVP candidate list. The one or more processors may be configured to derive a prediction signal for the video block based on the best MVP.
[0007] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium having instructions stored thereon is provided. When executed by one or more processors of the device, the instructions cause the device to, at a decoder, obtain a video block from a video signal. The instructions cause the device to, at the decoder, obtain spatial neighboring blocks based on the video block. The spatial neighboring blocks may include a plurality of left spatial neighboring blocks and a plurality of above spatial neighboring blocks. The instructions further cause the device to, at the decoder, obtain up to one unscaled left motion vector predictor (MVP) based on the plurality of left spatial neighboring blocks. The instructions further cause the device to, at the decoder, obtain up to one unscaled above MVP based on the plurality of above spatial neighboring blocks. The instructions further cause the device to, at the decoder, derive an MVP candidate list based on the video block, the plurality of left spatial neighboring blocks, and the plurality of above spatial neighboring blocks by reducing a likelihood of selecting a scaled MVP derived from the spatial neighboring blocks. The MVP candidate list may include at most one unscaled left MVP and at most one unscaled top MVP. The instructions may also cause the apparatus to receive, at a decoder, a best MVP based on the MVP candidate list. The encoder selects the best MVP from the MVP candidate list. The instructions may also cause the apparatus to obtain, at the decoder, a scaled MVP derived from spatial neighboring blocks after obtaining the history-based MVP.
[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 2 is a block diagram of an encoder according to an example of the present disclosure. [Figure 2]FIG. 2 is a block diagram of a decoder according to an example of the present disclosure. [Figure 3] FIG. 10 illustrates candidate positions for spatial MVP candidates and temporal MVP candidates according to an example of the present disclosure. [Figure 4] FIG. 10 illustrates motion vector scaling for spatial motion vector candidates according to an example of the present disclosure. [Figure 5] FIG. 10 illustrates motion vector scaling for temporal merge candidates according to an example of the present disclosure. [Figure 6] FIG. 2 is a diagram of a method for decoding a video signal according to an example of the present disclosure. [Figure 7] FIG. 1 is a diagram of a method for deriving an MVP candidate list according to an example of the present disclosure. [Figure 8] FIG. 1 illustrates a computing environment coupled to a user interface according to an example of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0010] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of the embodiments do not represent all implementations consistent with the present disclosure. Instead, these implementations are merely examples of apparatus and methods consistent with aspects related to the disclosure recited in the appended claims.
[0011] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. When used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will also be understood that the term "and / or," as used herein, means and is intended to include any and all possible combinations of one or more of the associated listed items.
[0012] Terms such as "first," "second," and "third" may be used herein to describe various pieces of information, but it will be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another category of information. For example, first information may be referred to as second information, and similarly, second information may be referred to as first information, without departing from the scope of this disclosure. As used herein, the term "if" will be understood to mean "when" or "upon" or "in response to a determination," depending on the context.
[0013] Video Encoding System
[0014] Conceptually, the video coding standards mentioned above are similar, e.g., they all use block-based processing and share similar video coding block diagrams to achieve video compression.
[0015] Figure 1 shows a schematic diagram of a block-based video encoder for VVC. Specifically, Figure 1 shows a typical encoder 100. The encoder 100 includes a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction-related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy coding 138, and bitstream 144.
[0016] At encoder 100, a video frame is partitioned into video blocks for processing. For each given video block, a prediction is formed based on either an inter-prediction or an intra-prediction technique.
[0017] A prediction residual, which represents the difference between a current video block, which is part of video input 110, and its predictor, which is part of block predictor 140, is sent from summer 128 to transform 130. The transform coefficients are then sent from transform 130 to quantization 132 for entropy reduction. The quantized coefficients are then sent to entropy coding 138 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 142 from intra / inter mode decision 116, such as video block partition information, motion vectors (MVs), reference picture indexes, and intra-prediction modes, is also sent through entropy coding 138 and stored in compressed bitstream 144. Compressed bitstream 144 comprises the video bitstream.
[0018] Decoder-related circuitry is also required in encoder 100 to reconstruct pixels for prediction purposes. First, a prediction residual is reconstructed by inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with block predictor 140 to generate unfiltered reconstructed pixels for the current video block.
[0019] Spatial prediction (or "intra prediction") predicts a current video block using pixels from samples of already coded neighboring blocks (called reference samples) within the same video frame as the current video block.
[0020] Temporal prediction (also known as "inter-prediction") uses reconstructed pixels from an already coded video picture to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in video signals. Typically, the temporal prediction signal for a given coding unit (CU) or coding block is transmitted by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal references. Additionally, if multiple reference pictures are supported, a reference picture index is also transmitted, which is used to identify which reference picture in the reference picture storage the temporal prediction signal comes from.
[0021] Motion estimation 114 takes in video input 110 and signals from picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 takes in video input 110, signals from picture buffer 120, and a motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.
[0022] After spatial and / or temporal prediction is performed, intra / inter mode decision 116 within encoder 100 chooses the best prediction mode, for example, based on a rate-distortion optimization method. Block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is de-correlated using transform 130 and quantization 132. The resulting quantized residual coefficients are inversely quantized by inverse quantization 134 and inversely transformed by inverse transform 136 to form a reconstructed residual, which is then added back to the predictive block to form a reconstructed signal for the CU. Furthermore, in-loop filtering 122, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), may be applied to the reconstructed CU, which is then placed in reference picture storage in picture buffer 120 and used to encode future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 138, where they are further compressed and packed to form the bitstream.
[0023] In an encoder, a video frame is partitioned into blocks for processing. For each given video block, a prediction is formed based on inter-prediction or intra-prediction. In inter-prediction, a predictor may be formed by motion estimation and motion compensation based on pixels of a previously reconstructed frame. In intra-prediction, a predictor may be formed based on reconstructed pixels in a current frame. A mode decision may select the best predictor to predict the current block.
[0024] The prediction residual (i.e., the difference between the current block and its predictor) is sent to a transform module. The transform coefficients are then sent to a quantization module for entropy reduction. The quantized coefficients are sent to an entropy coding module to generate a compressed video bitstream. As shown in Figure 1, prediction-related information from the inter and / or intra prediction modules, such as block partition information, motion vectors, reference picture indexes, and intra prediction modes, also passes through the entropy coding module and is saved in the bitstream.
[0025] In the encoder, a decoder-related module is also required to reconstruct pixels for prediction purposes. First, a prediction residual is reconstructed by inverse quantization and inverse transformation. The reconstructed prediction residual is then combined with a block predictor to generate unfiltered reconstructed pixels for the current block.
[0026] In-loop filters are commonly used to improve coding efficiency and visual quality. For example, deblocking filters are available in AVC, HEVC, and now VVC. In HEVC, an additional in-loop filter called SAO (Sample Adaptive Offset) is defined to further improve coding efficiency. In the latest VVC, yet another in-loop filter called ALF (Adaptive Loop Filter) is actively investigated and is likely to be included in the final standard.
[0027] Figure 2 shows a schematic block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of an exemplary decoder 200. The decoder 200 includes a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction-related information 234, and video output 232.
[0028] The decoder 200 is similar to the reconstruction-related portion present in the encoder 100 of FIG. 1. In the decoder 200, an incoming video bitstream 210 is first decoded by entropy decoding 212 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by inverse quantization 214 and inverse transform 216 to obtain reconstructed prediction residuals. A block predictor mechanism implemented in an intra / inter mode selector 220 is configured to perform intra prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by adding the reconstructed prediction residual from the inverse transform 216 and the prediction output generated by the block predictor mechanism using an adder 218.
[0029] The reconstructed blocks may further pass through an in-loop filter 228 and then be stored in a picture buffer 226, which serves as a reference picture store. The reconstructed video in the picture buffer 226 may be sent to drive a display device as well as used to predict future video blocks. In situations where the in-loop filter 228 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 232.
[0030] Figure 2 shows the block diagram of a typical decoder for these standards, which can be seen to be almost identical to the reconstruction-related parts present in the encoder.
[0031] In the decoder, the bitstream is first decoded by an entropy decoding module to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by an inverse quantization and inverse transform module to obtain reconstructed prediction residuals. Block predictors are formed by intra-prediction or motion compensation processes based on the decoded prediction information. Unfiltered reconstructed pixels are obtained by adding the reconstructed prediction residuals and the block predictors. If an in-loop filter is turned on, a filtering operation is performed on these pixels to derive the final reconstructed video for output.
[0032] Versatile Video Coding (VVC)
[0033] At the 10th JVET Conference (April 10-20, 2018, San Diego, USA), JVET defined the draft of the Versatile Video Coding (VVC) and VVC Test Model 1 (VTM1) coding method. JVET decided to use the 2-part and 3K-part coding block structures as the first new coding feature of VVC, including a quadtree with nested multiple tree types. Since then, the reference software VTM for implementing the coding method and the draft VVC decoding process has been developed during the JVET conference.
[0034] The picture partitioning structure divides the input video into blocks called coding tree units (CTUs). The CTUs are then divided into coding units (CUs) using a quadtree, a nested tree structure with multiple types of trees. Leaf coding units (CUs) define regions that share the same prediction mode (e.g., intra or inter). In this specification, the term "unit" defines a region of an image that contains all components, while the term "block" defines a region that contains a specific component (e.g., luma), which may have different spatial locations when considering chroma sampling schemes such as 4:2:0.
[0035] Motion Vector Prediction in VVC
[0036] Motion vector prediction exploits the spatial and temporal correlation of motion vectors of neighboring CUs, which are used for explicit transmission of motion parameters. Motion vector prediction first builds a motion vector predictor (MVP) candidate list by checking the left, upper, and temporally neighboring block positions (shown in Figure 3 and described below) as well as the availability of history-based motion vectors. Then, redundant candidates are removed and zero vectors are added to make the candidate list a certain length. The encoder can then select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. The index of the best motion vector candidate is then coded into the bitstream. Details regarding the MVP candidate derivation process are provided in the following sections.
[0037] FIG. 3 shows candidate locations for spatial MVP candidates A0, A1, B0, B1, and B2, and temporal MVP candidates T0 and T1.
[0038] Deriving motion vector predictor candidates
[0039] In VVC, the MVP candidate list is constructed by including the following candidates, in order:
[0040] (1) Derive at most one unscaled MVP from the left spatial neighboring CU (A0 → A1).
[0041] (2) If no unscaled MVP from the left is available, derive at most one scaled MVP from the left spatial neighboring CU (A0 → A1).
[0042] (3) Derive at most one unscaled MVP from the upper spatial neighboring CUs (B0 → B1 → B2).
[0043] (4) If both neighboring blocks A0 and A1 are unavailable or coded in intra-mode, derive at most one scaled MVP from the upper spatial neighboring CUs (B0 → B1 → B2).
[0044] (5) Conditional pruning.
[0045] (6) Derive at most one MVP from the time array CU (T0 → T1).
[0046] (7) Derive up to two history-based MVPs from the FIFO table.
[0047] (8) Derive up to two zero MVs.
[0048] Item 5, "Conditional Pruning," occurs when two MVP candidates are derived from a spatial block. The two MVP candidates are compared with each other, and if they are identical, one is removed from the MVP candidate list. The size of the inter MVP list is fixed at 2. For each CU coded in inter prediction but non-merge mode, a flag is coded to indicate which MVP candidate is used.
[0049] Three types of MVP candidates are considered for motion vector prediction: spatial MVP candidates, temporal MVP candidates, and history-based MVP candidates. For spatial MVP candidate derivation, up to two MVP candidates are derived based on the motion vectors of each block located at five different positions, as shown in Figure 3. For temporal MVP candidate derivation, up to one MVP candidate is selected from two candidates derived based on two different co-located positions. If the number of MVP candidates is less than two, an additional zero motion vector candidate is added to the list to make the number of MVP candidates two.
[0050] Spatial motion vector predictor candidates
[0051] In deriving spatial motion vector predictor candidates, up to two candidates may be selected from up to five potential candidates derived from blocks arranged as shown in FIG. 3. The block positions are the same as those used to derive MVP candidates in the merge mode. For blocks arranged to the left of the current block, including A0 and A1, the MV checking order is defined as A0, A1, and scaled A0, scaled A1. Here, scaled A0 and scaled A1 refer to the scaled MVs of blocks A0 and A1, respectively. For blocks arranged above the current block, including B0, B1, and B2, the MV checking order is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. For each side, there are four cases to be checked when deriving motion vector candidates; in two cases, spatial MV scaling is not required, and in the other two cases, spatial MV scaling is required. The four different cases are summarized as follows:
[0052] Cases where spatial MV scaling is not required
[0053] (1) MVs with the same reference picture list and the same reference picture index (same POC).
[0054] (2) MVs with different reference picture lists but the same reference pictures (same POC).
[0055] Cases where spatial MV scaling is needed
[0056] (3) MVs with the same reference picture list but different reference pictures (different POC).
[0057] (4) MVs with different reference picture lists and different reference pictures (different POCs).
[0058] First, the cases without spatial MV scaling (cases 1 and 2 above) are checked, followed by the cases where spatial MV scaling is required (cases 3 and 4 above). More specifically, for MVs, the checking order is as follows:
[0059] By sequentially checking A0 Case 1 → A0 Case 2 → A1 Case 1 → A1 Case 2 → A0 Case 3 → A0 Case 4 → A1 Case 3 → A1 Case 4, a maximum of one MV candidate is derived from blocks A0 and A1.
[0060] By sequentially checking B0 Case 1 → B0 Case 2 → B1 Case 1 → B1 Case 2 → B2 Case 1 → B2 Case 2, a maximum of one MV candidate is derived from blocks B0, B1, and B2.
[0061] If no MV candidate is derived in step 1, a maximum of one MV candidate is derived from blocks B0, B1, and B2 by sequentially checking B0 Case 3 → B0 Case 4 → B1 Case 3 → B1 Case 4 → B2 Case 3 → B2 Case 4.
[0062] If the number of MVP candidates is less than two, an additional zero motion vector candidate is added to the list to bring the number of MVP candidates to two.
[0063] Regardless of the reference picture list, if the POC of the reference picture of a neighboring block is different from the POC of the current block, spatial MV scaling is required. As can be seen in the above steps, if the MV of the block to the left of the current block is not available (for example, the block is not available or all of these blocks are intra-coded), MV scaling is enabled for the motion vector of the block above the current block. Otherwise, spatial MV scaling is not enabled for the motion vector of the block above. As a result, only at most one spatial MV scaling is required in the entire MV prediction candidate derivation process.
[0064] In the spatial MV scaling process, similar to the temporal MV scaling, the motion vectors of neighboring blocks are scaled as shown in FIG. 4 (described below).
[0065] FIG. 4 shows a diagram of motion vector scaling for spatial motion vector candidates. This diagram includes 410neigh_ref, 420curr_ref, 430curr_pic, 440neighbor_PU, 450curr_PU, 460tb, and 470td. 410neigh_ref is the neighbor reference used to scale the motion vector. 420curr_ref is the current reference used to scale the motion vector. 430 is the current picture used to scale the motion vector. 440neighbor_PU is the neighbor prediction unit used to scale the motion vector. 450curr_PU is the current prediction unit used to scale the motion vector. 460tb is the POC difference between the reference picture of the current picture and the current picture. 470td is the POC difference between the neighbor reference picture and the current picture.
[0066] Temporal motion vector predictor candidates
[0067] In the derivation of temporal MVP candidates, scaled motion vectors are derived from the MVs of alignment blocks in previously coded pictures in the reference picture list. In the following description, a picture containing alignment blocks for deriving temporal MVP candidates is referred to as an "aligned picture." To derive temporal motion candidates, an explicit flag (collocated_from_L0_flag) in the slice header is first sent to the decoder to indicate whether the alignment picture is selected from list 0 or list 1. An alignment reference index (collocated_ref_idx) is also sent to indicate which picture in that list is selected as the alignment picture for deriving the temporal motion vector predictor candidate. The L0 and L1 MVs of temporal MVP candidates are derived independently according to a predefined order, as shown in Table 1.
[0068] In Table 1, CollocatedPictureList is a syntax signaled to indicate from which prediction list the collocated picture is located. As shown by the dotted line in Figure 5 (described below), a scaled motion vector for the temporal MVP candidate is obtained. This scaled motion vector is scaled from the selected motion vector of the co-located block using POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero. The practical implementation of the scaling process is described in the HEVC specification. For a B-slice, two motion vectors are obtained: a motion vector for reference picture list 0 and a motion vector for reference picture list 1, and they are combined to create a bi-predictive merge candidate. [Table 1]
[0069] Figure 5 shows a diagram of motion vector scaling for temporal merge candidates. This diagram includes 510neigh_ref, 520curr_ref, 530curr_pic, 540col_pic, 550curr_PU, 560col_PU, 570tb, and 580td. 510neigh_ref is a neighboring reference picture used to scale the motion vector. 520curr_ref is a current reference picture used to scale the motion vector. 530curr_pic is a current picture used to scale the motion vector. 540col_pic is a contiguous picture used to scale the motion vector. 550curr_PU is a current prediction unit used to scale the motion vector. 560col_PU is a contiguous prediction unit used to scale the motion vector. 570tb is a POC difference between the reference picture of the current picture and the current picture. 580td is a POC difference between the neighboring reference picture and the current picture.
[0070] In a constellation picture, the constellation block for deriving temporal candidates is selected between T0 and T1, as shown in Figure 3. If the block at position T0 is not available, or is intra-coded, or is outside the current CTU, position T1 is used. Otherwise, position T0 is used in deriving temporal merge candidates.
[0071] History-Based Merge Candidate Deriving
[0072] History-based MVP (HMVP) is added to the merge list after spatial MVP and TMVP. In this method, the motion information of a previously coded block stored in a table may be used as the MVP for the current block. Such an MVP is called an HMVP. During the encoding / decoding process, a table containing multiple HMVP candidates is maintained. When a new CTU column is encountered, this table is reset (emptied). Whenever a CU is coded in an inter prediction mode other than subblock, the CU's associated motion information is added as a new HMVP candidate to the last entry in the table.
[0073] Spatial motion vector scaling for motion vector prediction.
[0074] In the current design of MVP, at most one scaled spatial MV candidate may be derived. Since the MV scaling process is non-trivial in terms of its complexity, a new scheme is designed to further reduce the possibility of using spatial MV scaling in deriving MVP.
[0075] Improved motion vector prediction
[0076] In this disclosure, several methods are proposed to improve motion vector prediction in terms of complexity or coding efficiency. It should be noted that the proposed methods may be applied independently or in combination.
[0077] 6 illustrates a method for decoding a video signal according to the present disclosure, which may be applied, for example, to a decoder.
[0078] In step 610, the decoder may obtain a video block from the video signal.
[0079] The decoder may obtain spatial neighboring blocks based on the video block at step 612. For example, the spatial neighboring blocks may include multiple left spatial neighboring blocks and multiple above spatial neighboring blocks.
[0080] In step 614, the decoder may obtain at most one unscaled left motion vector predictor (MVP) based on multiple left spatial neighboring blocks.
[0081] In step 616, the decoder may obtain at most one unscaled upper MVP based on multiple upper spatial neighboring blocks.
[0082] At step 618, the decoder may derive an MVP candidate list based on the video block, the plurality of left spatial neighboring blocks, and the plurality of above spatial neighboring blocks by reducing the likelihood of selecting a scaled MVP derived from the spatial neighboring blocks. For example, the MVP candidate list may include at most one unscaled left MVP and at most one unscaled above MVP.
[0083] In step 620, the decoder can receive a best MVP based on the MVP candidate list, and the encoder selects the best MVP from the MVP candidate list. In one example, the best MVP may be selected by the encoder based on a cost value. For example, in some encoder designs, the encoder can use each MVP candidate to derive an MV difference between the selected MVP and an MV derived by motion estimation. For each MVP candidate, a cost value is calculated using, for example, the MV difference and bits of the MVP index of the selected MVP. The MVP candidate with the best cost value (e.g., the smallest cost) is selected as the final MVP, and its MVP index is encoded into the bitstream.
[0084] In step 622, the decoder may obtain a prediction signal for the video block based on the best MVP.
[0085] According to the first aspect, in order to reduce the possibility of selecting / deriving a scaled MVP, it is proposed to put the scaled spatial MVP in a later position when constructing the MVP candidate list. The conditions for deriving a scaled MVP from spatial neighboring blocks may be the same as or different from those of the current VVC.
[0086] In one example, the conditions for deriving a scaled MVP from spatially neighboring blocks remain the same as those for the current VVC. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0087] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0088] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0089] 3) Conditional pruning.
[0090] 4) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0091] 5) Derive up to two history-based MVPs from the FIFO table.
[0092] 6) If no unscaled MVP from the left neighboring block is available, derive at most one scaled MVP from the left spatial neighboring block (A0 → A1).
[0093] 7) If neither of the neighboring blocks A0 and A1 are available or are coded in intra-mode, derive at most one scaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0094] 8) Derive up to two zero MVs.
[0095] In another example, the requirement for deriving the scaled MVP from spatially neighboring blocks is removed. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0096] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0097] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0098] 3) Conditional pruning.
[0099] 4) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0100] 5) Derive up to two history-based MVPs from the FIFO table.
[0101] 6) Derive at most one scaled MVP from the left spatial neighboring block (A0 → A1).
[0102] 7) Derive at most one scaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0103] 8) Derive up to two zero MVs.
[0104] In the third example, the conditions for deriving a scaled MVP from spatially neighboring blocks are modified from those of the current VVC. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0105] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0106] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0107] 3) Conditional pruning.
[0108] 4) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0109] 5) Derive up to two history-based MVPs from the FIFO table.
[0110] 6) Derive at most one scaled MVP from the left spatial neighboring block (A0 → A1).
[0111] 7) If neither of the neighboring blocks A0 and A1 are available or are coded in intra-mode, derive at most one scaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0112] 8) Derive up to two zero MVs.
[0113] In the fourth example, the conditions for deriving a scaled MVP from spatially neighboring blocks are modified from those of the current VVC. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0114] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0115] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0116] 3) Conditional pruning.
[0117] 4) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0118] 5) Derive up to two history-based MVPs from the FIFO table.
[0119] 6) Derive at most one scaled MVP from the left spatial neighboring block A0.
[0120] 7) Derive at most one scaled MVP from the left spatial neighboring block A1.
[0121] 8) Derive at most one scaled MVP from the upper spatial neighboring block B0.
[0122] 9) Derive at most one scaled MVP from the upper spatial neighboring block B1.
[0123] 10) Derive at most one scaled MVP from the upper spatial neighboring block B2.
[0124] 11) Derive up to two zero MVs.
[0125] According to the second aspect, in order to reduce the possibility of selecting / deriving a scaled MVP, it is proposed to put both the temporal MVP and the scaled spatial MVP in a later position when constructing the MVP candidate list. The conditions for deriving a scaled MVP from spatial neighboring blocks may be the same as or different from those of the current VVC.
[0126] In one example, the conditions for deriving a scaled MVP from spatially neighboring blocks remain the same as those for the current VVC. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0127] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0128] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0129] 3) Conditional pruning.
[0130] 4) Derive up to two history-based MVPs from the FIFO table.
[0131] 5) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0132] 6) If no unscaled MVP from the left neighboring block is available, derive at most one scaled MVP from the left spatial neighboring block (A0 → A1).
[0133] 7) If neither of the neighboring blocks A0 and A1 are available or are coded in intra-mode, derive at most one scaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0134] 8) Derive up to two zero MVs.
[0135] In another example, the requirement for deriving the scaled MVP from spatially neighboring blocks is removed. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0136] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0137] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0138] 3) Conditional pruning.
[0139] 4) Derive up to two history-based MVPs from the FIFO table.
[0140] 5) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0141] 6) Derive at most one scaled MVP from the left spatial neighboring block (A0 → A1).
[0142] 7) Derive at most one scaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0143] 8) Derive up to two zero MVs.
[0144] In the third example, the conditions for deriving a scaled MVP from spatially neighboring blocks are modified from those of the current VVC. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0145] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0146] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0147] 3) Conditional pruning.
[0148] 4) Derive up to two history-based MVPs from the FIFO table.
[0149] 5) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0150] 6) Derive at most one scaled MVP from the left spatial neighboring block (A0 → A1).
[0151] 7) If neither of the neighboring blocks A0 and A1 are available or are coded in intra-mode, derive at most one scaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0152] 8) Derive up to two zero MVs.
[0153] In the fourth example, the conditions for deriving a scaled MVP from spatially neighboring blocks are modified from those of the current VVC. The MVP candidate list is constructed by performing the following operations using one or more processors:
[0154] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0155] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0156] 3) Conditional pruning.
[0157] 4) Derive up to two history-based MVPs from the FIFO table.
[0158] 5) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0159] 6) Derive at most one scaled MVP from the left spatial neighboring block A0.
[0160] 7) Derive at most one scaled MVP from the left spatial neighboring block A1.
[0161] 8) Derive at most one scaled MVP from the upper spatial neighboring block B0.
[0162] 9) Derive at most one scaled MVP from the upper spatial neighboring block B1.
[0163] 10) Derive at most one scaled MVP from the upper spatial neighboring block B2.
[0164] 11) Derive up to two zero MVs.
[0165] According to a third aspect, in order to reduce the possibility of selecting / deriving a scaled MVP, it is proposed to exclude scaled spatial MVP candidates from the MVP candidate list.
[0166] 7 illustrates a method for deriving an MVP candidate list according to the present disclosure, which may be applied, for example, to a decoder.
[0167] In step 710, the decoder may prune one of the two MVP identity candidates from the MVP candidate list when two MVP identity candidates are derived from spatially neighboring blocks.
[0168] In step 712, the decoder can obtain at most one MVP from the time ordering block, which includes ordering block T0 and ordering block T1.
[0169] In step 714, the decoder can retrieve up to two history-based MVPs from the FIFO table, which contains previously coded blocks.
[0170] In step 716, the decoder can obtain up to two MVs equal to zero.
[0171] In one example, the MVP candidate list is constructed by performing the following operations using one or more processors.
[0172] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0173] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0174] 3) Conditional pruning.
[0175] 4) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0176] 5) Derive up to two history-based MVPs from the FIFO table.
[0177] 6) Derive up to two zero MVs.
[0178] In another example, the MVP candidate list is constructed by performing the following operations using one or more processors.
[0179] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0180] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0181] 3) Conditional pruning.
[0182] 4) Derive up to two history-based MVPs from the FIFO table.
[0183] 5) Derive at most one MVP from the time-sequenced block (T0 → T1).
[0184] 6) Derive up to two zero MVs.
[0185] According to a fourth aspect, in order to reduce the possibility of selecting / deriving a scaled MVP, it is proposed to exclude both scaled spatial MVP candidates and TMVPs from the MVP candidate list.
[0186] In one example, the MVP candidate list is constructed by performing the following operations using one or more processors.
[0187] 1) Derive at most one unscaled MVP from the left spatial neighboring block (A0 → A1).
[0188] 2) Derive at most one unscaled MVP from the upper spatial neighboring blocks (B0 → B1 → B2).
[0189] 3) Conditional pruning.
[0190] 4) Derive up to two history-based MVPs from the FIFO table.
[0191] 5) Derive up to two zero MVs.
[0192] It should be noted that in all of the embodiments and / or examples shown above, the derivation process terminates when the MVP candidate list is full.
[0193] 8 illustrates a computing environment 810 coupled to a user interface 860. The computing environment 810 may be part of a data processing server. The computing environment 810 includes a processor 820, a memory 840, and an I / O interface 850.
[0194] The processor 820 typically controls the overall operation of the computing environment 810, such as operations related to display, data acquisition, data communication, and image processing. The processor 820 may include one or more processors to execute instructions for performing all or some of the steps of the methods described above. Additionally, the processor 820 may include one or more modules that facilitate interaction between the processor 820 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.
[0195] The memory 840 is configured to store various types of data to support the operation of the computing environment 810. The memory 840 may include predefined software 842. Examples of such data include instructions for any application or method operated on the computing environment 810, video data sets, image data, etc. The memory 840 may be implemented using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disks, or a combination thereof.
[0196] The I / O interface 850 provides an interface between the processor 820 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and an end scan button. The I / O interface 850 may be coupled to an encoder and a decoder.
[0197] In some embodiments, a non-transitory computer-readable storage medium comprising a plurality of programs executing the aforementioned methods, executable by processor 820 in computing environment 810, such as that contained in memory 840, is also provided. For example, the non-transitory computer-readable storage medium can be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0198] A non-transitory computer-readable storage medium has stored thereon a plurality of programs for execution by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, causing the computing device to perform the above-described method for motion prediction.
[0199] In some embodiments, the computing environment 810 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0200] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the present disclosure. Many modifications, variations and alternative implementations will be apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the accompanying drawings.
[0201] These examples are chosen and described to explain the principles of the disclosure and to enable those skilled in the art to understand the disclosure of various implementations and to best utilize the underlying principles and various implementations with various modifications suited to the particular applications contemplated. Therefore, it should be understood that the scope of the disclosure is not limited to the specific examples of implementations disclosed, and that modifications and other implementations are also intended to be encompassed within the scope of the disclosure.
Claims
1. 1. A method for decoding a video signal, comprising: obtaining, at a decoder, video blocks from the video signal; obtaining, at the decoder, spatial neighboring blocks based on the video block, the spatial neighboring blocks including a plurality of left spatial neighboring blocks and a plurality of above spatial neighboring blocks; At the decoder, obtaining at most one unscaled left motion vector predictor (MVP) based on the plurality of left spatial neighboring blocks; At the decoder, obtaining at most one unscaled upper MVP based on the plurality of upper spatial neighboring blocks; deriving, at the decoder, an MVP candidate list based on the video block, the plurality of left spatial neighboring blocks, and the plurality of above spatial neighboring blocks by reducing a likelihood of selecting a scaled MVP derived from the spatial neighboring blocks, wherein the MVP candidate list includes at most one unscaled left MVP and at most one unscaled above MVP; receiving, at the decoder, a best MVP based on the MVP candidate list, wherein an encoder selects the best MVP from the MVP candidate list; and at the decoder, deriving a prediction signal for the video block based on the best MVP.
2. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; The method of claim 1 , further comprising, at the decoder, obtaining the scaled MVP derived from the spatial neighboring blocks after obtaining a history-based MVP.
3. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; retrieving, at the decoder, up to two history-based MVPs from a first-in-first-out (FIFO) table, the FIFO table containing previously encoded blocks; At the decoder, when no unscaled left spatial neighboring blocks are available, obtaining at most one scaled MVP from the plurality of left spatial neighboring blocks; At the decoder, when none of the plurality of left spatial neighboring blocks are available, obtaining at most one scaled MVP from the plurality of above spatial neighboring blocks; and taking at most two MVs equal to 0 at the decoder.
4. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; At the decoder, obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; and at the decoder, obtaining at most one scaled MVP based on the plurality of upper spatial neighboring blocks.
5. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; At the decoder, obtaining at most one scaled MVP based on the plurality of upper spatial neighboring blocks; and taking at most two MVs equal to 0 at the decoder.
6. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; At the decoder, obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; At the decoder, when none of the plurality of left spatial neighboring blocks are available, deriving at most one scaled MVP based on the plurality of above spatial neighboring blocks; and taking at most two MVs equal to 0 at the decoder.
7. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; obtaining, at the decoder, at most one scaled MVP from a first left spatial neighboring block A0, wherein the plurality of left spatial neighboring blocks includes the first left spatial neighboring block A0 and a second left spatial neighboring block A1; obtaining at most one scaled MVP from the second left spatial neighboring block A1 at the decoder; obtaining, at the decoder, at most one scaled MVP from a first upper spatial neighboring block B0, wherein the plurality of upper spatial neighboring blocks includes the first upper spatial neighboring block B0, a second upper spatial neighboring block B1, and a third upper spatial neighboring block B2; obtaining, at the decoder, at most one scaled MVP from the second upper spatial neighboring block B1; obtaining at most one scaled MVP from the third upper spatial neighboring block B2 at the decoder; and taking at most two MVs equal to 0 at the decoder.
8. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; The method of claim 1 , comprising, at the decoder, excluding scaled spatial MVP candidates from the MVP candidate list.
9. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; and taking at most two MVs equal to 0 at the decoder.
10. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; and taking at most two MVs equal to 0 at the decoder.
11. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; The method of claim 1 , comprising, at the decoder, excluding scaled spatial MVP candidates and scaled temporal MVP candidates from the MVP candidate list.
12. deriving the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks at the decoder; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; and taking at most two MVs equal to 0 at the decoder.
13. 1. A computing device for decoding a video signal, comprising: one or more processors; a non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors: obtaining a video block from the video signal; obtaining spatial neighboring blocks based on the video block, the spatial neighboring blocks including a plurality of left spatial neighboring blocks and a plurality of above spatial neighboring blocks; obtaining at most one unscaled left motion vector predictor (MVP) based on the plurality of left spatial neighboring blocks; obtaining at most one unscaled upper MVP based on the plurality of upper spatial neighboring blocks; deriving an MVP candidate list based on the video block, the plurality of left spatial neighboring blocks, and the plurality of above spatial neighboring blocks by reducing a likelihood of selecting a scaled MVP derived from the spatial neighboring blocks, wherein the MVP candidate list includes at most one unscaled left MVP and at most one unscaled above MVP; receiving a best MVP based on the MVP candidate list, wherein an encoder selects the best MVP from the MVP candidate list; and deriving a prediction signal for the video block based on the best MVP.
14. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; The computing device of claim 13 , further configured to, after obtaining a history-based MVP, obtain the scaled MVP derived from the spatially neighboring blocks.
15. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list; Obtaining at most one MVP from a time sequence block, the time sequence block including a sequence block T0 and a sequence block T1; Obtaining up to two history-based MVPs from a first-in-first-out (FIFO) table, the FIFO table containing previously encoded blocks; obtaining at most one scaled MVP from the plurality of left spatial neighboring blocks when no unscaled left spatial neighboring blocks are available; When none of the plurality of left spatial neighboring blocks is available, obtaining at most one scaled MVP from the plurality of above spatial neighboring blocks; 14. The computing device of claim 13, further configured to: take a maximum of two MVs equal to 0.
16. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; The computing device of claim 13 , further configured to: obtain at most one scaled MVP based on the plurality of upper spatial neighboring blocks.
17. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list; Obtaining at most one MVP from a time sequence block, the time sequence block including a sequence block T0 and a sequence block T1; Obtaining up to two history-based MVPs from a FIFO table, said FIFO table containing previously encoded blocks; obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; obtaining at most one scaled MVP based on the plurality of upper spatial neighboring blocks; 14. The computing device of claim 13, further configured to: take a maximum of two MVs equal to 0.
18. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list; Obtaining up to two history-based MVPs from a FIFO table, said FIFO table containing previously encoded blocks; Obtaining at most one MVP from a time sequence block, the time sequence block including a sequence block T0 and a sequence block T1; obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; When none of the plurality of left spatial neighboring blocks are available, obtaining at most one scaled MVP based on the plurality of above spatial neighboring blocks; 14. The computing device of claim 13, further configured to: take a maximum of two MVs equal to 0.
19. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list; Obtaining up to two history-based MVPs from a FIFO table, said FIFO table containing previously encoded blocks; Obtaining at most one MVP from a time sequence block, the time sequence block including a sequence block T0 and a sequence block T1; obtaining at most one scaled MVP from a first left spatial neighboring block A0, wherein the plurality of left spatial neighboring blocks includes the first left spatial neighboring block A0 and a second left spatial neighboring block A1; obtaining at most one scaled MVP from the second left spatial neighboring block A1; Obtaining at most one scaled MVP from a first upper spatial neighboring block B0, wherein the plurality of upper spatial neighboring blocks includes the first upper spatial neighboring block B0, a second upper spatial neighboring block B1, and a third upper spatial neighboring block B2; Obtaining at most one scaled MVP from the second upper spatial neighboring block B1; obtaining at most one scaled MVP from the third upper spatial neighboring block B2; 14. The computing device of claim 13, further configured to: take a maximum of two MVs equal to 0.
20. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; The computing device of claim 13 , further configured to remove scaled spatial MVP candidates from the MVP candidate list.
21. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list; Obtaining at most one MVP from a time sequence block, the time sequence block including a sequence block T0 and a sequence block T1; Obtaining up to two history-based MVPs from a FIFO table, said FIFO table containing previously encoded blocks; 14. The computing device of claim 13, further configured to: take a maximum of two MVs equal to 0.
22. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list; Obtaining up to two history-based MVPs from a FIFO table, said FIFO table containing previously encoded blocks; Obtaining at most one MVP from a time sequence block, the time sequence block including a sequence block T0 and a sequence block T1; 14. The computing device of claim 13, further configured to: take a maximum of two MVs equal to 0.
23. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; The computing device of claim 13 , further configured to remove scaled spatial MVP candidates and scaled temporal MVP candidates from the MVP candidate list.
24. the one or more processors configured to derive the MVP candidate list by reducing the likelihood of selecting the scaled MVP derived from the spatial neighboring blocks; When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list; Obtaining up to two history-based MVPs from a FIFO table, said FIFO table containing previously encoded blocks; 14. The computing device of claim 13, further configured to: take a maximum of two MVs equal to 0.
25. 1. A non-transitory computer-readable storage medium storing a plurality of programs for execution by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, causing the computing device to: obtaining, at a decoder, video blocks from the video signal; obtaining, at the decoder, spatial neighboring blocks based on the video block, the spatial neighboring blocks including a plurality of left spatial neighboring blocks and a plurality of above spatial neighboring blocks; At the decoder, obtaining at most one unscaled left motion vector predictor (MVP) based on the plurality of left spatial neighboring blocks; At the decoder, obtaining at most one unscaled upper MVP based on the plurality of upper spatial neighboring blocks; deriving, at the decoder, an MVP candidate list based on the video block, the plurality of left spatial neighboring blocks, and the plurality of above spatial neighboring blocks by reducing a likelihood of selecting a scaled MVP derived from the spatial neighboring blocks, wherein the MVP candidate list includes at most one unscaled left MVP and at most one unscaled above MVP; receiving, at the decoder, a best MVP based on the MVP candidate list, wherein an encoder selects the best MVP from the MVP candidate list; and after obtaining a history-based MVP, obtaining, at the decoder, the scaled MVP derived from the spatial neighboring blocks.
26. The plurality of programs are installed on the computing device.
26. The non-transitory computer-readable storage medium of claim 25, further causing the computer to perform operations including, after obtaining a history-based MVP, obtaining the scaled MVP derived from the spatially neighboring blocks.
27. The plurality of programs are installed on the computing device. When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; retrieving, at the decoder, up to two history-based MVPs from a first-in-first-out (FIFO) table, the FIFO table containing previously encoded blocks; At the decoder, when no unscaled left spatial neighboring blocks are available, obtaining at most one scaled MVP from the plurality of left spatial neighboring blocks; At the decoder, when none of the plurality of left spatial neighboring blocks are available, obtaining at most one scaled MVP from the plurality of above spatial neighboring blocks; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining a maximum of two MVs equal to zero.
28. The plurality of programs are installed on the computing device. At the decoder, obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining at most one scaled MVP based on the plurality of upper spatial neighboring blocks.
29. The plurality of programs are installed on the computing device. When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; At the decoder, obtaining at most one scaled MVP based on the plurality of upper spatial neighboring blocks; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining a maximum of two MVs equal to zero.
30. The plurality of programs are installed on the computing device. When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; At the decoder, obtaining at most one scaled MVP based on the plurality of left spatial neighboring blocks; At the decoder, when none of the plurality of left spatial neighboring blocks are available, deriving at most one scaled MVP based on the plurality of above spatial neighboring blocks; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining a maximum of two MVs equal to zero.
31. The plurality of programs are installed on the computing device. When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; obtaining, at the decoder, at most one scaled MVP from a first left spatial neighboring block A0, wherein the plurality of left spatial neighboring blocks includes the first left spatial neighboring block A0 and a second left spatial neighboring block A1; obtaining at most one scaled MVP from the second left spatial neighboring block A1 at the decoder; obtaining, at the decoder, at most one scaled MVP from a first upper spatial neighboring block B0, wherein the plurality of upper spatial neighboring blocks includes the first upper spatial neighboring block B0, a second upper spatial neighboring block B1, and a third upper spatial neighboring block B2; obtaining, at the decoder, at most one scaled MVP from the second upper spatial neighboring block B1; obtaining at most one scaled MVP from the third upper spatial neighboring block B2 at the decoder; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining a maximum of two MVs equal to zero.
32. The plurality of programs are installed on the computing device.
26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including removing scaled spatial MVP candidates from the MVP candidate list.
33. The plurality of programs are installed on the computing device. When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining a maximum of two MVs equal to zero.
34. The plurality of programs are installed on the computing device. When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; At the decoder, obtaining at most one MVP from a time alignment block, the time alignment block including an alignment block T0 and an alignment block T1; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining a maximum of two MVs equal to zero.
35. The plurality of programs are installed on the computing device.
26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including removing scaled spatial MVP candidates and scaled temporal MVP candidates from the MVP candidate list.
36. The plurality of programs are installed on the computing device. When two MVP same candidates are derived from the spatial neighboring blocks, pruning one of the two MVP same candidates from the MVP candidate list in the decoder; retrieving, at the decoder, up to two history-based MVPs from a FIFO table, the FIFO table containing previously encoded blocks; 26. The non-transitory computer-readable storage medium of claim 25, further causing the decoder to perform operations including: obtaining a maximum of two MVs equal to zero.