Spatial predictor scanning order

By adjusting the scanning order of motion vector predictors based on coding modes, the inefficiencies in existing video coding technologies are addressed, leading to improved compression efficiency and reduced complexity in video coding systems.

JP7840263B2Active Publication Date: 2026-04-03INTERDIGITALCE PATENT HLDG SAS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in leveraging spatial and temporal redundancy for motion vector prediction, particularly in the construction and scanning order of motion vector predictors, which affects compression efficiency and computational complexity.

Method used

Adjusting the scanning order of motion vector predictors based on different coding modes, such as merge modes and AMVP modes, by prioritizing spatial predictors from specific neighbor blocks, including upper and left neighbors, to optimize the construction of motion vector predictor lists.

Benefits of technology

Improves compression efficiency and reduces computational complexity by optimizing the selection and scanning order of motion vector predictors, enhancing the overall performance of video coding systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007840263000014
    Figure 0007840263000014
  • Figure 0007840263000015
    Figure 0007840263000015
  • Figure 0007840263000016
    Figure 0007840263000016
Patent Text Reader

Abstract

To encode coding units in inter modes, there are multiple inter modes that signal motion information using motion vector predictor (MVP) lists. The MVP lists for different coding modes include spatial predictors. In one embodiment, for all coding modes that use MVP lists, the top spatial predictor is added to the motion vector predictor list first, before the left spatial predictor. In another embodiment, a subset of coding modes adds the top spatial predictor before the left spatial predictor, and another subset of coding modes adds the left spatial predictor before the top spatial predictor. For example, all merge modes add the top spatial predictor first, and all AMVP modes add the left spatial predictor first. In another example, all non-subblock modes add the top spatial predictor first, and all subblock modes add the left spatial predictor first.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This embodiment relates, in general terms, to a method and apparatus for encoding or decoding video, and more specifically, to a method and apparatus for adjusting motion vector predictors in a predictor list. [Background technology]

[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to leverage spatial and temporal redundancy within video content. Generally, intra-picture or inter-picture correlation is used to utilize intra-picture or inter-picture correlation, and the difference between the original and predicted blocks, often called the prediction error or prediction residual, is then transformed, quantized, and entropicoded. To reconstruct the video, the compressed data is decoded by the reverse processes corresponding to entropicoding, quantization, transformation, and prediction. [Overview of the Initiative]

[0003] According to one embodiment, a video decoding method is provided, comprising generating a list of motion vector predictors for a block of picture in a video sequence, the block being decoded using one of a plurality of coding modes, the list of motion vector predictors comprising a plurality of spatial predictors, the first spatial predictor in the list being from an upper neighbor block and the second spatial predictor following the first spatial predictor being from a left neighbor block, and decoding a motion vector associated with the block based on the list of motion vector predictors.

[0004] According to another embodiment, a video coding method is provided, which includes: accessing a block of pictures in a video sequence; generating a list of motion vector predictors for the block, wherein the block is coded using one of a plurality of coding modes, the list of motion vector predictors comprises a plurality of spatial predictors, and for each of the plurality of coding modes, a first spatial predictor in the list is from an upper neighbor block, and a second spatial predictor following the first spatial predictor is from a left neighbor block; and coding a motion vector associated with the block based on the list of motion vector predictors.

[0005] According to another embodiment, a video decoding method is provided, comprising generating a list of motion vector predictors for a block, the block being decoded using one of a plurality of coding modes, the list of motion vector predictors comprising a plurality of spatial predictors, wherein for each of one or more merge modes in the plurality of coding modes, a first spatial predictor in the list is from an upper neighbor block, a second spatial predictor following the first spatial predictor is from a left neighbor block, and for each of one or more AMVP modes in the plurality of coding modes, a first spatial predictor in the list is from a left neighbor block, and a second spatial predictor following the first spatial predictor is from an upper neighbor block; and decoding a motion vector associated with the block based on the list of motion vector predictors.

[0006] According to another embodiment, a video coding method is provided, comprising: accessing a block of pictures in a video sequence; generating a list of motion vector predictors for the block, wherein the block is coded using one of a plurality of coding modes, and the list of motion vector predictors comprises a plurality of spatial predictors, wherein for each of one or more merge modes in the plurality of coding modes, a first spatial predictor in the list is from an upper neighbor block, a second spatial predictor following the first spatial predictor is from a left neighbor block, and for each of one or more AMVP modes in the plurality of coding modes, a first spatial predictor in the list is from a left neighbor block, and a second spatial predictor following the first spatial predictor is from an upper neighbor block; and coding a motion vector associated with the block based on the list of motion vector predictors.

[0007] According to another embodiment, a video decoding method is provided, comprising generating a list of motion vector predictors for a block, the block being decoded using one of a plurality of coding modes, the list of motion vector predictors comprising a plurality of spatial predictors, wherein for each of one or more non-subblock modes in the plurality of coding modes, a first spatial predictor in the list is from the upper neighbor block, a second spatial predictor following the first spatial predictor is from the left neighbor block, and for each of one or more subblock modes in the plurality of coding modes, a first spatial predictor in the list is from the left neighbor block, and a second spatial predictor following the first spatial predictor is from the upper neighbor block; and decoding a motion vector associated with the block based on the list of motion vector predictors.

[0008] According to another embodiment, a video coding method is provided, comprising: accessing a block of pictures in a video sequence; generating a list of motion vector predictors for the block, wherein the block is coded using one of a plurality of coding modes, and the list of motion vector predictors comprises a plurality of spatial predictors, wherein for each of one or more non-sub-modes in the plurality of coding modes, a first spatial predictor in the list is from the upper neighbor block, a second spatial predictor following the first spatial predictor is from the left neighbor block, and for each of one or more sub-block modes in the plurality of coding modes, a first spatial predictor in the list is from the left neighbor block, and a second spatial predictor following the first spatial predictor is from the upper neighbor block; and coding a motion vector associated with the block based on the list of motion vector predictors.

[0009] According to another embodiment, a device for video coding is provided, comprising one or more processors configured to: access a block of picture in a video sequence; generate a list of motion vector predictors for the block, wherein the block is coded using one of a plurality of coding modes, the list of motion vector predictors comprises a plurality of spatial predictors, and for each of the plurality of coding modes, a first spatial predictor in the list is from an upper neighbor block, and a second spatial predictor following the first spatial predictor is from a left neighbor block; and encode a motion vector associated with the block based on the list of motion vector predictors.

[0010] According to another embodiment, a device for video decoding is provided, comprising one or more processors configured to generate a list of motion vector predictors for a block of picture in a video sequence, the block being decoded using one of a plurality of coding modes, the list of motion vector predictors comprising a plurality of spatial predictors, the first spatial predictor in the list being from an upper neighbor block and the second spatial predictor following the first spatial predictor being from a left neighbor block, and decoding a motion vector associated with the block based on the list of motion vector predictors.

[0011] One embodiment provides a computer program that, when executed by one or more processors, includes instructions causing one or more processors to perform an encoding or decoding method according to any of the embodiments described above. One or more embodiments also provide a computer-readable storage medium storing instructions for encoding or decoding video data according to the methods described above. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving a bitstream generated according to the methods described above. [Brief explanation of the drawing]

[0012] [Figure 1] A block diagram illustrating a system in which an embodiment of this model may be implemented is provided as an example.

[0013] [Figure 2] A block diagram illustrating one embodiment of a video encoder is provided as an example.

[0014] [Figure 3] A block diagram illustrating one embodiment of a video decoder is provided as an example.

[0015] [Figure 4] An example of the concept of Coding Tree Unit (CTU) and Coding Tree (CT) for representing compressed HEVC pictures is illustrated.

[0016] [Figure 5] An example of the division of a Coding Tree Unit (CTU) into Coding Units (CUs), Prediction Units (PUs), and Transform Units (TUs) is illustrated.

[0017] [Figure 6] The partitioning used in VVC is illustrated.

[0018] [Figure 7] An example of partitioning a picture having a quad tree and a binary / trinary tree in VVC is illustrated.

[0019] [Figure 8] The positions of spatial and temporal predictors are illustrated.

[0020] [Figure 9] An example of a 4×4 sub-CU-based affine motion vector field used in VVC is illustrated.

[0021] [Figure 10] An example of affine model inheritance is illustrated.

[0022] [Figure 11] An example of SbTMVP prediction is illustrated.

[0023] [Figure 12] A method of adjusting the spatial predictor scanning order based on whether the coding mode is the merge mode or the AMVP mode is illustrated according to one embodiment.

[0024] [Figure 13] According to one embodiment, we illustrate a method for adjusting the spatial predictor scan order based on whether the coding mode is a subblock mode. [Modes for carrying out the invention]

[0025] Figure 1 illustrates a block diagram of an example of a system in which various embodiments and forms can be implemented. System 100 may be embodied as a device comprising various components described below and configured to perform one or more of the embodiments described herein. Examples of such a device include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 100 may be embodied individually or in combination as a single integrated circuit, a plurality of ICs, and / or individual components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of System 100 are distributed across a plurality of ICs and / or individual components. In various embodiments, System 100 is communicably coupled to other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, System 100 is configured to implement one or more of the embodiments described in this application.

[0026] System 100 includes, for example, at least one processor 110 configured to execute internally loaded instructions in order to implement various embodiments described in this application. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include, but is not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drives, and / or optical disk drives, as well as non-volatile and / or volatile memory. The storage device 140 may, in non-limiting examples, include an internal storage device, a mounted storage device, and / or a network-accessible storage device.

[0027] System 100 includes, for example, an encoder / decoder module 130 configured to process data and provide encoded or decoded video, the encoder / decoder module 130 of which may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, the device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 130 may be implemented as a separate element of System 100, or may be incorporated into the processor 110 as a combination of hardware and software, as is known to those skilled in the art.

[0028] Program code loaded onto the processor 110 or the encoder / decoder 130 to perform the various embodiments described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video, or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0029] In some embodiments, internal memory of the processor 110 and / or encoder / decoder module 130 is used to store instructions and to provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory of the processing device (for example, the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or storage device 140, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video coding and decoding operations such as MPEG-2, HEVC, or VVC.

[0030] Inputs to the elements of system 100 can be provided through various input devices, as shown in block 105. Such input devices include, but are not limited to, (i) an RF unit for receiving RF signals wirelessly transmitted by a broadcasting station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0031] In various embodiments, the input device of block 105 has associated input processing elements, as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band in order to select a signal frequency band that may be referred to as a channel in a particular embodiment, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) multiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, e.g., frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs these various functions, e.g., down-converting a received signal to a lower frequency (e.g., an intermediate frequency or adjacent baseband frequency) or to the baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering it to a desired frequency band. Various embodiments may involve rearranging the order of the elements described above (and others), removing some of these elements, and / or adding other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0032] Additionally, USB and / or HDMI terminals may include their respective interface processors for connecting System 100 to other electronic devices across the entire USB and / or HDMI connection. It should be understood that various forms of input processing, such as Reed-Solomon error correction, may be implemented, for example, within separate input processing ICs or within Processor 110, as needed. Similarly, forms of USB or HDMI interface processing may be implemented, for example, within separate interface ICs or within Processor 110, as needed. Demodulated, error-corrected, and multiplexed streams are provided to various processing elements, including, for example, Processor 110, and an encoder / decoder 130, which works in conjunction with memory and storage elements to process the data streams necessary for presentation on an output device.

[0033] Various elements of system 100 may be provided within an integrated housing, where the various elements are interconnected using internal buses known in the art, such as a suitable connection configuration 115, including an I2C bus, wiring, and a printed circuit board, and can transmit data between them.

[0034] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, transceivers configured to transmit and receive data via the communication channel 190. The communication interface 150 may also include, but is not limited to, a modem or network card, and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.

[0035] In various embodiments, the data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11. In these embodiments, the Wi-Fi signal is received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. In these embodiments, the communication channel 190 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In another embodiment, streaming data is provided to system 100 using a set-top box that distributes data via an HDMI connection on input block 105. In yet another embodiment, streaming data is provided to system 100 using an RF connection on input block 105.

[0036] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various embodiments, the other peripheral devices 185 include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of System 100. In various embodiments, control signals are communicated between System 100 and the display 165, speaker 175, or other peripheral devices 185 using signal transmission such as AV.Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. The output devices may be communicably coupled to System 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to System 100 via a communication interface 150 and a communication channel 190. The display 165 and speaker 175 may be integrated into a single unit with other components of System 100, for example, in an electronic device such as a television. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0037] The display 165 and speaker 175 can, alternatively, be separated from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments where the display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, for example, including an HDMI port, a USB port, or a COMP output.

[0038] Figure 2 illustrates an example of a video encoder 200, such as a VVC (Versatile Video Coding) encoder. In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” and “coded” may be used interchangeably, and the terms “image,” “picture,” and “frame” may be used interchangeably. Typically, though not always, the term “reconstructed” is used on the encoder side, and the term “decoded” is used on the decoder side.

[0039] Before encoding, a video sequence may undergo preprocessing, such as applying a color conversion to the input color picture (e.g., converting from RGB4:4:4 to YCbCr4:2:0), or remapping the input picture components to obtain a more compression-resistant signal distribution (e.g., using histogram equalization of one of the color components). Metadata may be associated with its preprocessing and attached to the bitstream.

[0040] In encoder 200, the picture is encoded by the encoder elements described below. The input signal is mapped (201). The mapping in 201 may correspond to the forward mapping in 291, or may further include other mappings for preprocessing. The picture to be encoded is processed by units of CU (202). Each CU is encoded using either intra-mode or inter-mode. When a CU is encoded in intra-mode, intra-prediction (260) is performed. In inter-mode, motion estimation (275) and motion compensation (270) are performed. Forward mapping (291) is applied to the predicted signal. The encoder device determines whether to use intra-mode or inter-mode to encode the CU (205), and the intra / inter-mode decision is indicated by a prediction mode flag. The prediction residual is calculated by subtracting the mapped prediction block (from step 291) from the mapped original image block (from step 201) (210).

[0041] Next, the predicted residual is transformed (225) and quantized (230). The quantized transformed coefficients, as well as the motion vector and other syntax elements, are entropi-coded (245) and output as a bitstream. The encoder can also skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can also bypass both transformation and quantization, i.e., the residual is coded directly without applying any transformation or quantization process. In direct PCM coding, no prediction is applied, and the coded unit samples are coded directly into the bitstream.

[0042] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transformation coefficients are quantized (240), inversely transformed (250), and the prediction residuals are decoded. Combining the decoded prediction residuals and prediction blocks (255), the image blocks are reconstructed. Inverse mapping (290) and in-loop filtering (265) are applied to the reconstructed signal to perform deblocking / SAO (sample adaptive offset) filtering, for example, to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280).

[0043] Figure 3 illustrates a block diagram of an example of a video decoder 300, such as a VVC decoder. In the decoder 300, the bitstream is decoded by the decoder elements described below. Generally, the video decoder 300 performs video decoding as part of encoding the video data, executing a decoding path that is contrary to the encoding path, as described in Figure 2.

[0044] In particular, the input to the decoder includes a video bitstream that may be generated by the video encoder 200. Firstly, the bitstream is entropy-decoded (330) to obtain transformation coefficients, motion vectors, picture partitioning information, and other coded information. The picture partitioning information indicates the size of the CTU and how the CTU is partitioned into CUs and, where applicable, into PUs. Thus, the decoder may partition the picture into CTUs and each CTU into CUs according to the decoded picture partitioning information (335). The transformation coefficients are dequantized (340), inversely transformed (350), and the predicted residuals are decoded.

[0045] The decoded prediction residuals and prediction blocks are combined (355) to reconstruct an image block. The prediction block can be obtained from intra-predictions (360) or from motion-compensated predictions (i.e., inter-predictions) (375) (370). Forward mapping (395) is also applied to the prediction signal. In the case of bi-predictions, two motion-compensated predictions can be combined with a weighted sum. Inverse mapping (396) and in-loop filtering (365) are applied to the reconstructed signal. The filtered image is stored in a reference picture buffer (380).

[0046] The output from the loop filter may pass through inverse mapping (390), which performs the reverse of the mapping process (201) performed in preprocessing. The decoded picture may then pass through other post-decoded processing, such as inverse color transformation (e.g., conversion from YCbCr4:2:0 to RGB4:4:4). Post-decoded processing may use metadata derived in pre-encoding and signaled in the bitstream.

[0047] In HEVC, motion-compensated time prediction is employed to leverage the redundancy present between consecutive pictures of video, where motion vectors are associated with each prediction unit (PU). Each CTU is represented by a coding tree within the compression domain. This is a quartic tree partition of CTUs, where each leaf is called a coding unit (CU), as illustrated in Figure 4. Each CU is then given several intra or inter-prediction parameters as prediction information. To do so, a CU may be spatially partitioned into one or more prediction units (PUs), each PU being assigned several prediction parameters. The intra or intercoding mode is assigned at the CU level. These concepts are further illustrated in Figure 5.

[0048] The new video compression tool in VVC includes a coding tree unit representation within the compression domain, which allows for a more flexible representation of picture data. VVC uses bipartite and ternary segmentation structures; the quartic tree with nested multi-type trees (MTTs) replaces the concept of multiple partitioning unit types, i.e., VVC eliminates the separation of CU, PU, ​​and TU concepts, except in a few special cases. In the VVC coding tree structure, the CU can have either a square or rectangular shape. A coding tree unit (CTU) is first partitioned by a quartic tree structure. Then, the quartic tree leaf nodes can be further partitioned by a multi-type tree structure.

[0049] Specifically, the tree decomposition of the CTU proceeds in different stages. First, the CTU is divided in a quaternary tree manner, and then each quaternary tree leaf can be further divided in a bipartite or tripartite manner. As shown in Figure 6, there are four types of division in the multi-type tree structure: vertical bipartite (VER), horizontal bipartite (HOR), vertical tripartite (VER_TRIPLE), and horizontal tripartite (HOR_TRIPLE). HOR_TRIPLE or VER_TRIPLE division (horizontal or vertical tripartite tree division mode) consists of dividing a coding unit (CU) into three sub-coding units (sub-CUs), the size of which is equal to 1 / 4, 1 / 2, and 1 / 4 of the parent CU size in the direction of the spatial division considered. The multi-type tree leaf nodes are called coding units (CUs), and except in some special cases, this segmentation is used for prediction and transformation processing without any further partitioning. Figure 7 shows an example of a CTU divided into multiple CUs having quartic tree and ternary / binary tree coding block structures.

[0050] Prediction list in VTM (VVC test model)

[0051] In VVC, prediction lists exist within different intercoding modes: normal merge, MMVD (merge mode with MVD), CIIP (combined intra and inter prediction), TPM / GEO, IBC (intra-block copy), normal AMVP (advanced motion vector prediction), affine AMVP, and sub-block merge. These prediction lists contain spatial predictors selected from spatial neighbor blocks at positions {A0, A1, A2, B0, B1, B2, B3}, as shown in Figure 8. Predictors from left neighbor blocks consider left spatial predictors, and predictors from upper neighbor blocks consider upper spatial predictors. Some blocks, such as upper-left block B2, are considered as both left and upper neighbor blocks.

[0052] In VTM-6.0 (see "Algorithm description for Versatile Video Coding and Test Model 6 (VTM 6)", JVET-O2002, 15th Meeting: Gothenburg, SE, 3-12 July 2019), all list spatial candidates are generally scanned from left to top. For example, in normal merge mode, the first spatial predictor is from A1 (left side), and the second spatial predictor is from B1 (top side) if it differs from the A1 predictor. In normal AMVP mode, the first spatial predictor is selected from A0 or A1 (left side), and the second spatial predictor is from B0, B1, or B2 (top side). Different lists are described in Table 1. [Table 1] [Table 2] [Table 3]

[0053] Normal merge mode

[0054] In VTM-6.0, the predictive list for the normal merge mode, also called the merge candidate list, merge list, or normal merge list, is constructed by including the following types of predictors (also called candidates): 1) Spatial MVP (motion vector predictor) from spatial neighbor CU, 2) Time MVP from sequence CU, 3) History-based MVP from FIFO table, 4) Pairwise average MVP, and 5) No MVP.

[0055] For each CU coded in merge mode, the index of the selected predictor is encoded. The generation process for each category of merge candidates is described below.

[0056] Derivation of candidate spaces

[0057] The derivation of spatial merge candidates within VVC is the same as in HEVC. Up to four merge candidates are selected from among the candidates located within the positions illustrated in Figure 8. The order of derivation is A1, B1, B0, A0, and B2. Position B2 is considered only if any of the CUs at positions A1, B1, B0, or A0 are unavailable (e.g., belonging to a different slice or tile) or are intracoded. After a candidate is added at position A1, the remaining candidates are subjected to redundancy checks, which ensure that candidates with the same motion information are excluded from the list, resulting in improved coding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only the pairs in Table 1 are considered, and a candidate is added to the list only if the corresponding candidate used in the redundancy check does not have the same motion information. For example, in Table 1, "Spatial B1 + Pruned A1" indicates that a candidate at position B1 is added to the merge list only if it is different from the candidate at position A1.

[0058] Time candidate derivation (“TMVP C0 / C1”)

[0059] Only one time candidate is added to the merge list. In particular, in the derivation of this time merge candidate, the scaled motion vector is derived based on the array CU belonging to the merged reference picture. As illustrated in Figure 8, the position relative to the time candidate is selected between candidates C0 and C1. If CU at position C0 is not available, if CU is intracoded, or if CU is outside the current row of CTU, position C1 is used. Alternatively, position C0 is used in the derivation of the time merge candidate.

[0060] History-based merge candidate derivation

[0061] History-based MVP (HMVP) merge candidates are added to the merge list after spatial MVPs and TMVPs. To use HMVP, the movement information of previously coded blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained throughout the coding / decoding process. The table is reset (empty) when a new CTU row is encountered. Whenever a non-subblock coded CU exists, the associated movement information is added to the last entry in the table as a new HMVP candidate.

[0062] HMVP candidates can be used in the merge candidate list construction process. The most recent HMVP candidates in the table are examined in order and inserted into the candidate list after the TMVP candidates. Redundancy checks are applied to several spatial merge candidates on the HMVP candidates. Each of the last two HMVP candidates in the FIFO (the more recent one) is compared with spatial candidates A1 and B1 and added to the merge list only if they do not have the same motion information. The merge candidate list construction process from HMVPs ends when the total number of available merge candidates reaches the maximum allowed merge candidate minus 1.

[0063] Derivation of pairwise mean merge candidates

[0064] Pairwise average candidates are generated by averaging the candidates of the first pair in the existing merge candidate list. If the merge list is incomplete after the pairwise average merge candidates are added, zero MVPs are inserted at the ends until the maximum number of merge candidates is reached.

[0065] Merge mode with MVD (MMVD)

[0066] In addition to merge mode, if implicitly derived motion information is used directly for predictive sample generation of the current CU, a merge mode with motion vector difference (MMVD) is introduced within the VVC.

[0067] In MMVD, after a merge candidate is selected, it is further refined by signaled MVD information. This additional information includes merge candidate flags, an index to specify the magnitude of the movement, and an index indicating the direction of the movement. In MMVD mode, one of the first two candidates in the merge list is selected to be used as the MV criterion. The merge candidate flags are signaled to specify which one is used.

[0068] Triangular partitioning (TPM) for interpretation

[0069] When the triangular partitioning mode is used, the CU is uniformly divided into two triangular partitions using either diagonal or opposite-angled matrix partitioning. Each triangular partition within the CU is mutually predicted using its own movement, and only a single prediction is allowed for each partition. The single prediction candidate list is derived directly from the merge candidate list constructed according to the merge mode described above.

[0070] When the triangular partitioning mode is used for the current CU, a flag indicating the orientation of the triangular partition (diagonal or opposite-angled matrix) and two merge indices (one for each partition) are further transmitted.

[0071] Intra Block Copy mode (IBC)

[0072] The Intrablock Copy (IBC) mode is derived from HEVC's Screen Content Extension. This mode emulates interpretation with the decoded current frame. It supports only the IBC AMVP and IBC Merge modes, which are identical to the corresponding normal mode except for the simpler predictor list.

[0073] In VVC, the IBC merge predictor list (also known as the IBC merge candidate list or IBC merge list) is constructed by including the following types of candidates: 1) Spatial MVP from spatial neighbor CU, 2) History-based MVP from IBC FIFO tables, and 3) No MVPs.

[0074] Derivation of candidate spaces

[0075] The derivation of spatial IBC merge candidates is the same as for normal merges. Only a maximum of two spatial IBC merge candidates are selected from the candidates located within positions A1 and B1. After a candidate is added at position A1, the addition of the remaining candidates is subjected to redundancy checks, which ensures that candidates with the same motion information are excluded from the list, resulting in improved code efficiency.

[0076] History-based merge candidate derivation

[0077] History-based MVP (HMVP) IBC merge candidates are added to the IBC merge list after a spatial IBC MVP, in the same way as for a regular merge list. However, a separate FIFO is used for IBC mode.

[0078] If the IBC merge list is incomplete after HMVP candidates have been added, zero IBC MVPs are inserted at the ends until the maximum number of IBC merge candidates is reached.

[0079] In IBC AMVP mode, since the only available reference frame is the current one, only one predictor list should be built. This predictor list reuses an IBC merge list, limited to two candidates, without pruning.

[0080] Normal AMVP mode

[0081] The AMVP mode of VVC (also known as regular AMVP) is similar to one of the HEVC modes. The AMVP mode predictor list, also called the AMVP candidate list, AMVP list, or regular AMVP list, is constructed for each reference frame in each reference frame list by including candidates of the following types: 1) Spatial MV from spatial neighbor CU, 2) Time MV from sequence CU, 3) History-based MV from FIFO table, and 4) Zero MV.

[0082] The accuracy of the motion memory unit (1 / 16-pel) is finer than the acceptable accuracy of the motion vector difference, which can be 1 / 4-pel, 1 / 2-pel, 1-pel, or 4-pel depending on the AMVR (Adaptive Motion Vector Resolution) mode. To comply with this constraint, each predictor is rounded in its acceptable accuracy of the motion vector difference.

[0083] For each CU coded in AMVP mode and each reference frame list, the index of the selected reference frame, the index of the selected predictor (also referred to as the AMVP candidate), and the motion vector difference (as the difference between the selected motion vector and the predictor estimated through motion estimation) are encoded. The generation process for each category of AMVP candidates is performed in the reference frame list y, ref xy The reference frame x, as indicated, is explained below.

[0084] Derivation of candidate spaces

[0085] Ref within VVC xy The derivation of spatial AMVP candidates is the same as for HEVC. Up to two AMVP candidates are selected from among the candidates located within the positions shown in Figure 8. One candidate is selected from the left position A0 or A1, and the other is selected from the top position B0, B1, or B2. It should be considered that the candidates are also referenced in the frame ref xy You must use this. Specifically, A0 is not intercoded, or the reference frame ref xy When not used, the left predictor is selected from position A1. After the left candidate is added, the top candidate is subjected to redundancy checks, which ensure that candidates with the same motion vector are excluded from the list, resulting in improved code efficiency.

[0086] Time candidate derivation (“TMVP C0 / C1”)

[0087] By using the same derivation process as for normal merge list construction, only one time candidate is added to the AMVP list, but ref xy Usage is restricted.

[0088] History-based candidate derivation

[0089] Several HMVP candidates in the FIFO table were examined in reverse order, ref xy When used, and without redundancy checks, it is inserted into the candidate list after the TMVP candidates. The AMVP candidate list construction process from HMVP ends when the total number of available AMVP candidates reaches the maximum allowed number of AMVP candidates.

[0090] If the AMVP list is incomplete after HMVP candidates have been added, zero MVs are inserted at the ends until the maximum number of AMVP candidates is reached.

[0091] Affine AVP Mode

[0092] Affine AMVP, a subblock mode, is similar to standard AMVP, but uses an affine motion model as shown in Figure 9 instead of translation.

[0093] One flag signals that the AMVP mode is affine, and another flag indicates whether a 4- or 6-parameter affine model is used. Then, for each reference frame list, the index of the selected reference frame, the index of the selected predictor (also referred to as the affine AMVP candidate), and the difference between two motion vectors (as the difference between the affine motion estimate and the best CPMV estimated through the CPMV predictor) are encoded.

[0094] The predictors are selected within a predictor list, which is constructed for each reference frame in each reference frame list, by including candidates of the following types: 1) Spatial inherited affine CPMV (control point motion vector) from spatial affine neighborhood CU, 2) Affine CPMV constructed from spatially near-space CU, 3) Translational CPMV constructed from spatial and temporal neighborhood CU, and 4) Zero Affine CPMV.

[0095] The accuracy of the motion memory unit (1 / 16-pel) is finer than the acceptable accuracy of the motion vector difference, which can be 1 / 16-pel, 1 / 4-pel, or 1-pel depending on the AMVR (Adaptive Motion Vector Resolution) mode. To comply with this constraint, each predictor is rounded in its acceptable accuracy of the motion vector difference.

[0096] Derivation of candidate spatial inheritance affines

[0097] Reference frame in VVC xyThe derivation of the spatial inheritance affine candidates is based on the affine candidates positioned within the location illustrated in FIG. 10. One candidate is selected at the left positions A0, A1 and another candidate is selected at the upper positions B0, B1, B2. It should be considered that the candidates are affine coded and the reference frame ref xy must be used.

[0098] The affine model (CPMV) for the current CU is derived from the stored models of the considered affine neighborhood.

[0099] Constructed affine candidate derivation

[0100] One affine model is constructed by extracting the CPMV from the upper-left (LT) positions (A2, B2, B3), lower-left (LB) positions (A0, A1), and upper-right (RT) positions (B0, B1) using ref xy . The current affine model is then defined as a combination of these extracted CPMV.

[0101] Derivation of constructed translation candidates

[0102] When the affine AMVP list is not complete after the constructed affine candidates are added, the translation model is inserted into the list. All CPMV of the current affine model are continuously set to the translation motion vectors from the upper, left, upper-left, and temporal positions.

[0103] When the affine AMVP list is still not complete afterwards, zero CPMV are inserted at the ends until the maximum number of affine AMVP candidates is reached.

[0104] Sub-block merge mode

[0105] The subblock merge mode includes two different sets of predictor candidates: SbTMVP and affine merge. For the affine AMVP mode, a flag signals that the merge mode is subblock. Then, for the normal merge mode, the index of the selected predictor is encoded for each CU encoded in the subblock merge mode.

[0106] The prediction list for subblock merge mode, also known as the subblock merge candidate list or affine merge list, is constructed by including the following types of candidates: 1) SbTMVP candidate, 2) Spatial inheritance affine CPMV from spatial affine neighborhood CU, 3) Affine CPMV constructed from spatial and temporal neighborhood CU, and 4) Zero CPMV.

[0107] Derivation of the Subblock Time Motion Vector Predictor (SbTMVP)

[0108] The SbTMVP candidate uses left space candidate A1 (or zero MV if A1 is available) as input. This input motion vector is used on an array reference frame to obtain motion information for the current CU, on an 8x8 basis, as shown in Figure 11.

[0109] Derivation of candidate spatial inheritance affines

[0110] The derivation of spatial inheritance affine candidates is the same as the affine AMVP mode derived from the stored models of the left- and upper affine neighborhoods, where the current affine model (CPMV) is the affine model for the CU.

[0111] Derivation of constructed affine candidates

[0112] The affine model is constructed by extracting CPMVs from the top-left positions (A2, B2, B3), left-side positions (A0, A1), top positions (B0, B1), and time C0 position. The current affine model is then defined as a combination of these extracted CPMVs.

[0113] Subsequently, if the subblock merge list is incomplete, zero CPMVs are inserted at the ends until the maximum number of subblock merge candidates is reached.

[0114] In VTM-7.0 (see "Algorithm description for Versatile Video Coding and Test Model 7 (VTM 7)", JVET-P2002, 16th Meeting: Geneva, CH, 1-11 October 2019), top-left scanning is employed for normal merge, MMVD, and TPM / GEO modes, while left-top scanning is maintained for other coding modes as shown in Table 2. Differences in merge list construction for normal merge mode are described in Table 3. [Table 4] [Table 5]

[0115] The top-left scan of VTM-7.0 affects only the merge, MMVD, and triangular modes. Therefore, it is possible to extend this top-left scan concept to all other coding modes as IBC, AMVP, affine AMVP, and subblock merge modes. In one embodiment, it is proposed to consider the top-left scan of the spatial predictor in all applicable coding modes.

[0116] Extension to IBC

[0117] As presented in Table 4, it is proposed to use an upper-left scan of the spatial predictor for IBC list construction. As shown in Table 4, spatial predictor B1 is considered before spatial predictor A1 in predictor list construction. This affects IBC merging and IBC AMVP because both modes use the same list. [Table 6]

[0118] Extension to standard AMVP

[0119] In AMVP, one set of spatial neighbors is explored. However, in VTM-7.0, the left-hand set is scanned before the upper set. Therefore, as shown in Table 5, it is proposed to explore the upper positions (B0 / B1 / B2) before the left-hand set (A0 / A1). [Table 7]

[0120] In the modified example, as shown in Table 6, it is also possible to move B2, i.e., the top-left candidate, in the set of candidates on the left. [Table 8]

[0121] Extension to Affine AMVP

[0122] In affine AMVP, candidates from the same left-hand pair are searched before the top candidate, in the case of AMVP. It is also possible to reverse the scan order, as shown in Table 7. [Table 9]

[0123] In the modified example, as shown in Table 8, it is also possible to move B2, i.e., the top-left candidate, in the set of candidates on the left. [Table 10]

[0124] Extended to subblock merging

[0125] The subblock merge list in VTM-7.0 consists of five candidates of different properties. These may be SbTMVP candidates, inherited or constructed affine models. The proposed top-left scan of the spatial candidates can be extended to the derivation of inherited affine models, as presented in Table 9. [Table 11]

[0126] In the modified example, as shown in Table 10, it is also possible to move B2, i.e., the top-left candidate, in the set of candidates on the left. [Table 12] Table 10: Top of subblock merge list with B2 as candidate on the left - Left-side scanning extension

[0127] In another variation, the modification may also affect the SbTMVP candidate by using the upper spatial candidate as input instead of the left-hand candidate, as shown in Table 11. [Table 13]

[0128] In the above, it is proposed to use left-top scanning of the spatial neighborhood for all coding modes. In another embodiment, it is proposed to apply this order only to a subset of coding modes.

[0129] As explained above, when motion vector predictor lists are used, there are generally two categories of coding modes: merge mode and AMVP mode. For coding units using merge mode, a predictor list is generated, and an index to the list (e.g., mvp_idx) is sent to the decoder, which can generate the same predictor list and select a motion vector predictor as the motion vector for the current coding unit based only on the decoded index of the selected motion vector predictor. Typically, other motion information such as the reference picture index (ref_idx) or motion vector difference (MVD) does not need to be sent.

[0130] For coding units using AMVP mode, a predictor list is generated, and its index to the list (e.g., mvp_idx) is sent to the decoder. In addition, other motion information such as the reference picture index (ref_idx) and the motion vector difference (MVD) is also sent. The decoder generates the same predictor list, selects a motion vector predictor from the predictor list as the motion vector predictor for the current coding unit, and can then decode the motion vector based on the motion vector predictor, the decoded MVD, and the reference picture index.

[0131] In VVC, MMVD is typically considered as a merged mode, but it transmits MVD information via the signal.

[0132] In one embodiment, the inventors propose using an upper-left scan (1210, e.g., upper spatial predictor is the first predictor in the predictor list, followed by left spatial predictors) for all coding units with merge modes (1205, e.g., normal merge, MMVD, TPM, IBC, subblock merge), and a left-top scan (1220, e.g., left spatial predictor is the first predictor in the predictor list, followed by upper spatial predictors) for all coding units with AMVP modes (1205, e.g., normal AMVP, affine AMVP). Other predictors may also be added to the predictor list (1230). After the predictor list is constructed, motion information may be encoded or decoded based on the predictor list (1240). It should also be noted that the upper-left state may be reversed for AMVP, and the left-top state may be reversed for merge modes.

[0133] Another method of classification is that when a motion vector predictor list is used, the coding mode can be classified as non-subblock mode and subblock mode. In subblock mode, while the coding mode is indicated at the coding unit level, motion compensation is performed at the subblock level, and subblocks of a CU can have different motion vectors. In VVC, only affine candidates (from affine AMVP or subblock merge) and SbTMVP candidates exist. However, other potential subblock tools can be included, such as FRUC (frame rate up conversion) and OBMC (overlapping block motion compensation).

[0134] In one embodiment, the inventors propose using an upper-left scan (1310, e.g., upper spatial predictor is the first predictor in the predictor list, followed by left spatial predictors) for all coding units with non-subblock modes (1305, e.g., merge, MMVD, TPM, IBC, AMVP), and a left-top scan (1320, e.g., left spatial predictor is the first predictor in the predictor list, followed by upper spatial predictors) for all coding units with subblock modes (1305, affine AMVP, subblock merge). Other predictors may also be added to the predictor list (1330). After the predictor list is constructed, motion information can be encoded or decoded based on the predictor list (1340). It should also be noted that the upper-left state can be reversed for subblock modes, and the left-top state can be reversed for non-subblock modes.

[0135] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a particular order of steps or actions is required for the proper operation of the method, any particular order of steps and / or actions and / or use may be modified or combined.

[0136] Modules of the video encoder 200 and video decoder 300, such as the motion estimation and compensation modules (270, 275, 375) shown in Figures 2 and 3, can be modified using the various methods and other embodiments described in this application. Furthermore, these embodiments are not limited to VVC or HEVC and can be applied to other standards and recommendations, and any extensions of such standards and recommendations. Unless otherwise indicated or technically excluded, the embodiments described in this application may be used individually or in combination.

[0137] One embodiment provides a computer program that, when executed by one or more processors, includes instructions causing one or more processors to perform an encoding or decoding method according to any of the embodiments described above. One or more embodiments also provide a computer-readable storage medium storing instructions for encoding or decoding video data according to the methods described above. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving a bitstream generated according to the methods described above.

[0138] Various implementations include decoding. As used in this application, “decoding” may encompass all or part of the processes performed on the received encoded sequence to produce a final output suitable for display. In various embodiments, such processing may include one or more of the processes commonly performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase “decoding processing” is intended to refer specifically to a subset of operations or to a broader decoding processing in general will become clear from the context of the particular description and will be well understood by those skilled in the art.

[0139] Various implementations involve encoding. Similar to the above considerations regarding "decoding," "encoding" as used in this application may encompass all or part of the process performed on an input video sequence to generate an encoded bitstream, for example.

[0140] It should be noted that, as used herein, the syntax elements—for example, the syntax used to indicate the RST kernel index, and the syntax used to indicate whether or not the planar intra-prediction mode is being used—are descriptive terms. Therefore, these do not preclude the use of other syntax element names.

[0141] The implementations and embodiments described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if considered only in the context of a single form of implementation (e.g., considered only as a method), the implementations of the considered features may also be implemented in other forms (e.g., apparatus or programs). Apparatus may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented in apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.

[0142] The references to “one embodiment,” “one implementation,” or “implementation,” and other variations thereof, mean that certain features, structures, characteristics, etc., described in relation to the embodiments are included in at least one embodiment. Therefore, the appearances of the phrases “in one embodiment,” “in one embodiment,” or “in one implementation,” and any other variations, appearing in various places in this specification, do not necessarily all refer to the same embodiment.

[0143] In addition, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.

[0144] Furthermore, this application may refer to “accessing” various types of information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0145] In addition, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory) one or more of these actions. Furthermore, "receiving" is typically involved in some way during an action, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0146] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of any of the following " / ", "and / or", and "at least one of" is intended to encompass the selection of only the first listed option (A), only the second listed option (B), or both options (A and B). As a further embodiment, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such expressions are intended to encompass the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to the number of listed items, as would be obvious to those with the ordinary art in the art and related art.

[0147] As will be apparent to those skilled in the art, the implementation can generate various signals formatted to carry information that can be stored or transmitted. The information may include, for example, instructions for performing the method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal can be transmitted by various different wired or wireless links, as is known. The signal can be stored in a processor-readable medium.

Claims

1. A method for video encoding, Accessing blocks of pictures within a video sequence, To generate a first list of motion vector predictors for the block, wherein the block is encoded using one of a plurality of subblock coding modes, the first list of motion vector predictors comprises a first plurality of spatial predictors associated with the plurality of subblock coding modes, the first spatial predictors in the first list are from the upper neighbor block, and the second spatial predictors in the first list, following the first spatial predictor, are from the left neighbor block. To generate a second list of motion vector predictors for the block, wherein a plurality of non-merge coding modes are available for coding the block, the second list of motion vector predictors comprises a plurality of second spatial predictors associated with the plurality of non-merge coding modes, the first spatial predictor in the second list is from the left neighbor block, and the second spatial predictor in the second list, following the first spatial predictor, is from the upper neighbor block. A method comprising encoding a motion vector associated with the block based on the first list of motion vector predictors or the second list of motion vector predictors.

2. The process further includes generating a third list of motion vector predictors for the aforementioned block, The method according to claim 1, wherein a plurality of advanced motion vector prediction (AMVP) coding modes are available for coding the block, and a third list of motion vector predictors comprises a third plurality of spatial predictors associated with the plurality of AMVP coding modes, wherein a first spatial predictor in the third list is from the left neighborhood block, and a second spatial predictor in the third list, following the first spatial predictor, is from the upper neighborhood block, and the motion vector is coded based further on the third list of motion vector predictors.

3. The process further includes generating a fourth list of motion vector predictors for the aforementioned block, The method according to claim 2, wherein a plurality of merge coding modes are available for coding the block, and the fourth list of motion vector predictors comprises a fourth plurality of spatial predictors associated with the plurality of merge coding modes, the first spatial predictor in the fourth list being from the left neighborhood block, and the second spatial predictor in the fourth list following the first spatial predictor being from the upper neighborhood block, and the motion vector is coded based further on the fourth list of motion vector predictors.

4. A method for video decoding, To generate a first list of motion vector predictors for a block of picture in a video sequence, wherein the block is decoded using one of a plurality of subblock coding modes, the first list of motion vector predictors comprises a plurality of first spatial predictors associated with the plurality of subblock coding modes, the first spatial predictors in the first list are from the upper neighbor block, and the second spatial predictors in the first list, following the first spatial predictor, are from the left neighbor block. To generate a second list of motion vector predictors for the block, wherein a plurality of non-merge coding modes are available for decoding the block, the second list of motion vector predictors comprises a plurality of second spatial predictors associated with the plurality of non-merge coding modes, the first spatial predictor in the second list is from the left neighborhood block, and the second spatial predictor in the second list, following the first spatial predictor, is from the upper neighborhood block. A method comprising decoding the motion vector associated with the block based on a first list of motion vector predictors or a second list of motion vector predictors.

5. The process further includes generating a third list of motion vector predictors for the aforementioned block, The method according to claim 4, wherein a plurality of advanced motion vector prediction (AMVP) coding modes are available for decoding the block, and a third list of motion vector predictors comprises a third plurality of spatial predictors associated with the plurality of AMVP coding modes, wherein a first spatial predictor in the third list is from the left neighborhood block, and a second spatial predictor in the third list following the first spatial predictor is from the upper neighborhood block, and the motion vector is further decoded based on the third list of motion vector predictors.

6. The process further includes generating a fourth list of motion vector predictors for the aforementioned block, The method according to claim 5, wherein a plurality of merge coding modes are available for decoding the block, and a fourth list of motion vector predictors comprises a fourth plurality of spatial predictors associated with the plurality of merge coding modes, the first spatial predictor in the fourth list being from the left neighborhood block, and the second spatial predictor in the fourth list following the first spatial predictor being from the upper neighborhood block, and the motion vector is further decoded based on the fourth list of motion vector predictors.

7. A device for video encoding comprising at least one memory and one or more processors, wherein the one or more processors Accessing blocks of pictures within a video sequence, To generate a first list of motion vector predictors for the block, wherein the block is encoded using one of a plurality of subblock coding modes, the first list of motion vector predictors comprises a first plurality of spatial predictors associated with the plurality of subblock coding modes, the first spatial predictors in the first list are from the upper neighbor block, and the second spatial predictors in the first list, following the first spatial predictor, are from the left neighbor block. To generate a second list of motion vector predictors for the block, wherein a plurality of non-merge coding modes are available for coding the block, the second list of motion vector predictors comprises a plurality of second spatial predictors associated with the plurality of non-merge coding modes, the first spatial predictor in the second list is from the left neighbor block, and the second spatial predictor in the second list, following the first spatial predictor, is from the upper neighbor block. A device configured to encode a motion vector associated with a block based on a first list of motion vector predictors or a second list of motion vector predictors.

8. The one or more processors are further configured to generate a third list of motion vector predictors for the block, Apparatus according to claim 7, wherein a plurality of advanced motion vector prediction (AMVP) coding modes are available for coding the block, and a third list of motion vector predictors comprises a third plurality of spatial predictors associated with the plurality of AMVP coding modes, wherein a first spatial predictor in the third list is from the left neighborhood block, and a second spatial predictor in the third list, following the first spatial predictor, is from the upper neighborhood block, and the motion vector is coded based further on the third list of motion vector predictors.

9. The one or more processors are further configured to generate a fourth list of motion vector predictors for the block, Apparatus according to claim 8, wherein a plurality of merge coding modes are available for coding the block, and the fourth list of motion vector predictors comprises a fourth plurality of spatial predictors associated with the plurality of merge coding modes, the first spatial predictor in the fourth list being from the left neighborhood block, and the second spatial predictor in the fourth list following the first spatial predictor being from the upper neighborhood block, and the motion vector is coded further based on the fourth list of motion vector predictors.

10. A device for video decoding comprising at least one memory and one or more processors, wherein the one or more processors To generate a first list of motion vector predictors for a block of picture in a video sequence, wherein the block is decoded using one of a plurality of subblock coding modes, the first list of motion vector predictors comprises a plurality of first spatial predictors associated with the plurality of subblock coding modes, the first spatial predictors in the first list are from the upper neighbor block, and the second spatial predictors in the first list, following the first spatial predictor, are from the left neighbor block. To generate a second list of motion vector predictors for the block, wherein a plurality of non-merge coding modes are available for decoding the block, the second list of motion vector predictors comprises a plurality of second spatial predictors associated with the plurality of non-merge coding modes, the first spatial predictor in the second list is from the left neighborhood block, and the second spatial predictor in the second list, following the first spatial predictor, is from the upper neighborhood block. A device configured to decode a motion vector associated with a block based on a first list of motion vector predictors or a second list of motion vector predictors.

11. The one or more processors are further configured to generate a third list of motion vector predictors for the block, Apparatus according to claim 10, wherein a plurality of advanced motion vector prediction (AMVP) coding modes are available for decoding the block, and a third list of motion vector predictors comprises a third plurality of spatial predictors associated with the plurality of AMVP coding modes, wherein a first spatial predictor in the third list is from the left neighborhood block, and a second spatial predictor in the third list following the first spatial predictor is from the upper neighborhood block, and the motion vector is further decoded based on the third list of motion vector predictors.

12. The one or more processors are further configured to generate a fourth list of motion vector predictors for the block, Apparatus according to claim 11, wherein a plurality of merge coding modes are available for decoding the block, and a fourth list of motion vector predictors comprises a fourth plurality of spatial predictors associated with the plurality of merge coding modes, wherein a first spatial predictor in the fourth list is from the left neighborhood block, and a second spatial predictor in the fourth list following the first spatial predictor is from the upper neighborhood block, and the motion vector is decoded based further on the fourth list of motion vector predictors.

13. The method according to claim 1, wherein the plurality of subblock coding modes include a subblock time motion vector prediction (SbTMVP) mode.

14. The method according to claim 13, wherein the plurality of subblock coding modes further include at least one of frame rate upconversion (FRUC) coding modes or overlapped block motion compensation (OBMC) coding modes.

15. The method according to claim 4, wherein the plurality of subblock coding modes include a subblock time motion vector prediction (SbTMVP) mode.

16. The method according to claim 15, wherein the plurality of subblock coding modes further include at least one of frame rate upconversion (FRUC) coding modes or overlapped block motion compensation (OBMC) coding modes.

17. The apparatus according to claim 7, wherein the plurality of subblock coding modes include at least one of a subblock time motion vector prediction (SbTMVP) mode, a frame rate up conversion (FRUC) coding mode, or an overlapped block motion compensation (OBMC) coding mode.

18. The apparatus according to claim 10, wherein the plurality of subblock coding modes include at least one of a subblock time motion vector prediction (SbTMVP) mode, a frame rate up conversion (FRUC) coding mode, or an overlapped block motion compensation (OBMC) coding mode.

Citation Information

Patent Citations

  • Image decoding apparatus, image decoding method and image decoding program

    JP2013016931A

  • Affine motion prediction for video coding

    US20170332095A1