Displacement-based temporal motion vector predictor

By using additional displacement vectors and flexible block identification methods, the method addresses the limitations of predefined positions in TMVP, enhancing accuracy and efficiency in video coding.

JP2026074084APending Publication Date: 2026-05-01TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2026-01-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video coding technologies, such as HEVC and VVC, lack flexibility and efficiency in deriving motion vectors for temporal motion vector prediction (TMVP) due to the use of predefined fixed positions, which limits the accuracy and adaptability of motion information prediction.

Method used

The proposed method introduces additional displacement vectors or offsets to identify blocks in a reference picture, allowing for more flexible and efficient TMVP by signaling displacement vectors using methods like MMVD and AMVR, and deriving motion vectors from multiple block positions based on template matching costs.

Benefits of technology

This approach enhances the flexibility and efficiency of TMVP by improving the accuracy of motion vector prediction, reducing computational complexity, and increasing coding efficiency in video coding systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074084000001_ABST
    Figure 2026074084000001_ABST
Patent Text Reader

Abstract

A method, apparatus, and non-temporary storage medium for coding and decoding video data using temporal motion vector prediction (TMVP) are provided. [Solution] The method may include the steps of: receiving a video bitstream containing one or more pictures; and determining that the one or more pictures should be predicted in normal merge mode or adaptive motion vector prediction (AMVP) mode. Displacement vectors related to the current block in the current picture are obtained, and the displacement vectors are signaled in the video bitstream to identify a reference block in the current picture. A list of TMVP candidates containing the motion information is generated, and the motion vector of the current block is derived using the TMVP candidate list. The current block is then decoded using the derived motion vector for prediction in normal merge mode or adaptive motion vector prediction (AMVP) mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Related Applications] This application claims priority to U.S. Provisional Patent Application No. 63 / 391,219, filed on July 21, 2022, and U.S. Patent Application No. 18 / 080,450, filed on December 13, 2022, the entire disclosures of which are hereby incorporated by reference in their entirety.

[0002] [Technical Field] Embodiments of the present disclosure relate to image and video coding techniques. More specifically, embodiments of the present disclosure relate to the derivation of a temporal motion vector predictor (TMVP) using displacement vectors.

Background Art

[0003] The H.265 / HEVC (High Efficiency Video Coding) standard was published by ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standards organizations jointly formed JVET (Joint Video Exploration Team) to explore the possibility of developing a next-generation video coding standard beyond HEVC. In October 2017, they issued the Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses had been submitted for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for 360 video categories. In April 2018, all received CfP responses were evaluated at the 122MPEG / 10th JVET meeting. As a result of this meeting, JVET officially launched the standardization process for next-generation video coding beyond HEVC, the new standard was named Versatile Video Coding (VVC), and JVET was renamed Joint Video Experts Team. In 2020, ITU-TVCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the VVC video coding standard (version 1). [Overview of the project]

[0004] According to the embodiment, a method for coding or decoding video data using temporal motion vector prediction (TMVP) can be provided. The method is executable by a processor, The steps include receiving a video bitstream containing one or more pictures, The steps include determining that one or more of the aforementioned pictures should be predicted in normal merge mode or adaptive motion vector prediction (AMVP) mode, A step of obtaining a displacement vector associated with the current block in the current picture, wherein the displacement vector is signaled in the video bitstream to identify a reference block in the current picture. A step of determining motion information associated with the reference block based on the displacement vector, wherein the motion information is used as a motion vector predictor (MVP) from a candidate for a temporal motion vector predictor (TMVP), The steps include generating a list of TMVP candidates that includes the aforementioned motion information, The steps include: deriving the motion vector of the current block using the TMVP candidate list; The steps include decoding the current block using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode, It may include.

[0005] According to one embodiment, a device can be provided for coding or decoding video data using temporal motion vector prediction (TMVP). The device may include at least one memory configured to store program code, and at least one processor configured to access the program code and operate as directed by the program code. The program code is: A receive code configured to cause at least one processor to receive a video bitstream containing one or more pictures, A decision code configured to cause at least one processor to determine whether one or more pictures are predicted in normal merge mode or adaptive motion vector prediction (AMVP) mode, Acquisition code configured to cause at least one processor to acquire a displacement vector associated with the current block in the current picture, wherein the displacement vector is signaled in the video bitstream to identify a reference block in the current picture; A motion information code configured to cause at least one processor to determine motion information associated with the reference block based on the displacement vector, wherein the motion information is used as a motion vector predictor (MVP) from a candidate for a temporal motion vector predictor (TMVP), Generation code configured to cause at least one processor to generate a list of TMVP candidates including the motion information, Derivation code configured to cause at least one processor to derive the motion vector of the current block using the TMVP candidate list, The at least one processor is configured to decode the current block using the derived motion vector for prediction in the normal merge mode or adaptive motion vector prediction (AMVP) mode, It may include.

[0006] According to the embodiment, a non-temporary computer-readable medium for storing instructions can be provided. When the instructions are executed by one or more processors of a device for coding video data using Temporal Motion Vector Prediction (TMVP), the one or more processors will be provided with Receive a video bitstream containing one or more pictures, Determine that one or more of the aforementioned pictures should be predicted in normal merge mode or adaptive motion vector prediction (AMVP) mode. The displacement vector associated with the current block in the current picture is obtained, and the displacement vector is signaled within the video bitstream to identify a reference block in the current picture. Based on the displacement vector, motion information associated with the reference block is determined, and the motion information is used as a motion vector predictor (MVP) from a candidate for a temporal motion vector predictor (TMVP). A list of TMVP candidates including the aforementioned motion information is generated. Using the aforementioned TMVP candidate list, derive the motion vector of the current block. The current block is decoded using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode. It may include one or more instructions that can be performed. [Brief explanation of the drawing]

[0007] [Figure 1A] An example of the location of a special merge candidate according to one embodiment of the present disclosure is shown.

[0008] [Figure 1B] An example of candidate pairs considered for checking the redundancy of spatial merge candidates according to one embodiment of this disclosure is shown.

[0009] [Figure 1C] An example of motion vector scaling of a temporal merge candidate according to one embodiment of the present disclosure is shown.

[0010] [Figure 1D] An example of a location for a temporal merge candidate according to one embodiment of this disclosure is shown.

[0011] [Figure 1E] This invention illustrates an exemplary process for merging using motion vector difference (MMVD) search according to one embodiment of this disclosure.

[0012] [Figure 1F] An example of merging using motion vector difference search points according to one embodiment of this disclosure is shown.

[0013] [Figure 1G] An example of an additional direction along a diagonal angle according to an embodiment of the present disclosure is shown.

[0014] [Figure 1H] An example of a spatial neighborhood block used by ATVMP according to an embodiment of the present disclosure is shown. <>

[0015] [Figure 1I] An exemplary process for deriving a sub-CU motion field based on a motion shift from a spatial neighborhood according to an embodiment of the present disclosure is shown.

[0016] [Figure 2] An exemplary block diagram of a plurality of displacement vectors used for coding or decoding video data using temporal motion vector prediction (TMVP) using displacement vectors according to an embodiment of the present disclosure is shown.

[0017] [Figure 3] A flowchart of an exemplary process for coding and / or decoding video data using temporal motion vector prediction (TMVP) using displacement vectors according to an embodiment of the present disclosure is shown.

[0018] [Figure 4] A simplified block diagram of a communication system according to an embodiment of the present disclosure is shown.

[0019] [Figure 5] A diagram of the arrangement of a video encoder and a video decoder in a streaming environment is shown.

[0020] <000> [Figure 6] A functional block diagram of a video decoder according to an embodiment of the present disclosure is shown.

[0021] [Figure 7] A functional block diagram of a video encoder according to an embodiment of the present disclosure is shown.

[0022] [Figure 8] This is a diagram of a computer system according to one embodiment of the present disclosure. [Modes for carrying out the invention]

[0023] The proposed methods and processes can be used individually or in combination. Embodiments of this disclosure relate to methods and systems for coding or decoding video data using temporal motion vector prediction (TMVP) with displacement vectors.

[0024] In related technologies, the block positions used to fetch motion vectors from TMVP candidates are predefined and fixed. Embodiments of this disclosure relate to additional motion offsets used to derive motion vectors for TMVPs in order to improve the flexibility and efficiency of TMVPs.

[0025] According to aspects of this disclosure, instead of using predefined fixed positions to fetch motion information to be used as MVP from TMVP candidates for TMVP candidate derivation used in normal merge mode or AMVP mode, additional or additional offsets, i.e., displacement offsets, can be signaled to identify blocks in a reference picture, and motion information associated with these identified blocks can be used as MVP from TMVP candidates. For example, for a current block, one or more displacement vectors can be added to the current block to identify multiple block positions. Motion vectors associated with these identified block positions in the reference picture can be used as temporal motion vector predictors.

[0026] In one embodiment, displacement vectors can be signaled by an index using the merged motion vector difference (MMVD) method. In one embodiment, displacement vectors can be signaled using a method similar to motion vector difference signaling with Adaptive Motion Vector Resolution (AMVR). The displacement vector resolution may be N samples, for example, N may be 1, 4, or 8. In one embodiment, displacement vector resolution can be signaled by a high-level syntax such as sequence level, picture level, slice level, or tile / tile group level. In one embodiment, displacement vector resolution can be signaled by a block-level resolution index. The resolution index can be used to retrieve displacement vectors in a resolution table. In some embodiments, the resolution table can be predefined. In some embodiments, the resolution table can be signaled at a high level, such as sequence level, picture level, etc.

[0027] In one embodiment, the displacement offset index can be sorted using ascending or descending template matching costs by applying sorting of the displacement vector index based on template matching. In one example, a first N candidates with ascending template matching costs can be used, where N is greater than or equal to 1 and less than or equal to the total number of available candidates.

[0028] In one embodiment, candidate positions indicated by different displacement vectors can be scanned in a predefined order, and a first N candidate positions associated with a coded block using motion vectors can be identified, and indices between these N candidate positions can be signaled to indicate which one of the candidates will be used as the TMVP candidate block position. In some embodiments, the predefined scan order can be determined by the relative distance between the candidate positions and the starting point positions.

[0029] In one embodiment, the initial position may refer to a candidate position having a displacement vector of zero, and the starting position may be a default position, e.g., C0 in Figure 2, or it may be implicitly derived from coded information including, but not limited to, a selected candidate block position of a neighboring block coded using TMVP, and the motion vector of the neighboring block.

[0030] In one embodiment, the motion vector used as TMVP can be derived as the average or weighted average of motion vectors fetched from multiple block positions in the reference picture. In another embodiment, the motion vector used as TMVP can be derived as the motion vector value with the highest count among all motion vectors fetched from multiple block positions in the reference picture.

[0031] It can be understood that the methods and processes disclosed herein can be extended to multiple identically located pictures by extending the overlapping subblocks in the movement field from one identically located picture to multiple identically located pictures.

[0032] Interpretation in VVC

[0033] For each interpredictive coding unit (CU), motion parameters may include motion vectors, reference picture indexes and reference picture list usage indexes, and additional information required for new coding functions in the VVC to be used for generating interpredicted samples. Motion parameters can be signaled explicitly or implicitly. If a CU is coded in skip mode, it can be associated with one PU and may not have significant residual coefficients, coded motion vector deltas, or reference picture indexes. A merge mode can be specified, where the motion parameters of the current CU are taken from neighboring CUs, including spatial and temporal candidates, and from additional schedules introduced in the VVC. Merge mode can be applied to interpredicted CUs as well as skip mode. An alternative to merge mode is the explicit transmission of motion parameters, where motion vectors, corresponding reference picture indexes for each reference picture list, reference picture list usage flags, and other necessary information can be explicitly signaled for each CU.

[0034] Extended Merge Prediction

[0035] In VTM4, the merge candidate list for a mode consists of the following five types of candidates, in order: (1) spatial MVP from spatially neighboring CUs, (2) temporal MVP from identically located CUs, (3) history-based MVP from a FIFO table, (4) average MVP per pair, and (5) zero MVP. The size of the merge list can be signaled by the slice header, and the maximum allowable size of the merge list is 6 in VTM4. For each CU coded in merge mode, an index of the best merge candidate can be coded using truncated unary binarization (TU). The first bin of the merge index can be coded in context, and bypass coding can be used for other bins.

[0036] Derivation of candidate spaces

[0037] The derivation of spatial merge candidates in VVC is similar to that in HEVC. Up to four merge candidates can be selected from the candidates. Figure 1A shows current block 1100 with exemplary locations for merge candidates B1, A1, B0, A0, and B2. In some embodiments, the derivation order can be B1, A1, B0, A0, and B2. Location B2 can only be considered if any of the CUs for locations A0, B0, B1, or A1 are unavailable (e.g., because the CU belongs to a different slice or tile), or if it is intra-coded. After the candidate for location A1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the candidate list, thus improving coding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only pairs linked by arrows (where the error of 1B is not allowed as a reference source) are considered, and candidates are added to the candidate list only if the corresponding candidates used in the redundancy check do not have the same motion information.

[0038] Derivation of temporal candidates

[0039] In some embodiments, when deriving temporal candidates, only one candidate can be added to the list. In particular, in the derivation of temporal merge candidates, the scaled motion vector can be derived based on co-position CUs belonging to co-position reference pictures. The list of reference pictures used to derive co-position CUs can be explicitly signaled in the slice header. The scaled motion vector of the temporal merge candidate can be obtained as shown in Figure 1C and scaled from the motion vector of the co-position CU. As shown in Figure 1C, the scaled motion vector of the temporal merge candidate can be obtained and scaled from the motion vector of the co-position CU based on Picture Order Count (POC) distances tb and td, where tb is the POC difference between the current picture's reference picture and the current picture, and td is the POC difference between the co-position picture's reference picture and the co-position picture. The reference picture index of the temporal merge candidate can be set to 0.

[0040] As shown in Figure 1D, the position of the temporal candidate is selected between candidates C0 and C1. In some embodiments, if the CU at position C0 is unavailable, if it is intracoded, or if it is outside the current row of the Coding Tree Unit (CTU), position C1 can be used. Otherwise, position C0 can be used to derive the temporal merge candidate.

[0041] Merge with Motion Vector Difference (MMVD)

[0042] The merge mode in MMVD is used for either a skip mode or a merge mode based on the motion vector representation method. MMVD can reuse merge candidates in VVC. Candidates can be selected from the merge candidates and further extended by the proposed motion vector representation method as shown in Figures 1E and 1F. MMVD can provide a new motion vector representation with simplified signaling. The representation method can include the starting point, the magnitude of the motion, and the direction of the motion.

[0043] The MMVD technology can utilize the merge candidate list in VVC. However, only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered in MMVD deployment. A starting point is defined by the basic candidate index. The basic candidate index indicates the best candidate from the list of candidates, as shown in Table 1. [Table 1] Basic candidate IDX [Table 1]

[0044] If the number of base candidates is equal to 1, the base candidate IDX does not need to be signaled. The distance index is information about the magnitude of the movement. The distance index indicates a predefined distance from the starting point information. The predefined distance may be as shown in Table 2. [Table 2] Distance IDX [Table 2]

[0045] The direction index can represent the direction of MMVD relative to the starting point. As shown in Table 3, the direction index can represent four directions. [Table 3] Direction IDX [Table 3]

[0046] In some embodiments, the MMVD flag may be signaled immediately after the skip and merge flags are sent. If the skip and merge flags are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is AFFINE mode, but if it is not 1, the skip / merge index is parsed for the VTM's skip / merge mode.

[0047] Candidate sorting based on template matching on MMVD and Affine MMVD.

[0048] In related technologies, the MMVD offset can be extended for MMVD and affine MMVD modes. Additional refinement positions along a diagonal angle of k × π / 8 can be added as shown in Figure 1G, thus increasing the number of directions from 4 to 16. Furthermore, all possible MMVD refinement positions (16 × 6) for each basic candidate can be sorted based on the sum of absolute different (SAD) cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. In some embodiments, the top 1 / 8 refinement positions with the smallest template SAD cost are retained as available positions and, as a result, retained for MMVD index coding. The MMVD index is binarized by a rice code with a parameter equal to 2.

[0049] In aspects of this disclosure, an affine MMVD sort can also be extended on top of the MMVD extension described herein, where additional subdivision positions along a diagonal angle of k × π / 4 can be added. After sorting, the top half of the subdivision positions with the smallest template SAD cost can be retained.

[0050] Subblock-based TMVP (SbTMVP)

[0051] To improve coding efficiency and reduce motion vector transmission overhead, subblock-level motion vector subdivision can be applied to extend CU-level temporal motion vector prediction (TMVP). Subblock-based TMVP (SbTMVP) allows for the inheritance of motion information at the subblock level from co-position reference pictures. Each subblock of a large-sized CU can have its own motion information without explicitly transmitting block partition structure or motion information. SbTMVP can obtain motion information for each subblock as follows: First, SbTMVP can include the derivation of the displacement vector (DV) of the current CU. Next, it derives the motion of the center based on the availability of SbTMVP candidates. Finally, SbTMVP can include deriving the motion information of a subblock from the corresponding subblock by its DV. Unlike the derivation of TMVP candidates, which always derives a temporal motion vector from a co-located block within a reference frame, SbTMVP can apply a DV derived from the motion vector (MV) of the left neighboring CU of the current CU to find the corresponding subblock within the co-located picture of each subblock of the current CU. If the corresponding subblock is not encoded, the motion information of the current subblock can be set as the central motion.

[0052] VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to HEVC's temporal motion vector prediction (TMVP), SbTMVP uses the motion field of a co-position picture to improve the current picture motion vector prediction and CU merge mode. The co-position picture used by TMVP is used by SbTMVP. SbTMVP differs from TMVP in the following two main aspects:

[0053] (1) TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. (2) TMVP fetches temporal motion vectors from colocation blocks within a colocation picture (colocation blocks are the blocks to the lower right or center relative to the current CU), while SbTMVP applies a motion shift before fetching temporal motion information from colocation pictures, where the motion shift (also called displacement vector or DV) is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.

[0054] Figure 1H shows an exemplary SbTMVP candidate selection using spatial neighbor blocks. SbTMVP predicts the motion vector of a subCU within the current CU in two parts. As the first part, it examines the spatial neighbor A1 in Figure 1H. If A1 has a motion vector that uses the same-position picture as the reference picture, this motion vector is selected as the motion shift (or displacement vector) to be applied. If no such motion is identified, the motion shift is set to (0,0).

[0055] In the second part, the motion shift identified in the first part is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vectors and reference indices) from the co-position picture, as shown in Figure 1I. Assume that the motion shift is set to the motion of block A1, as shown in Figure 1I. Next, for each sub-CU, the motion information of the sub-CU is derived using the motion information of the corresponding block (the smallest motion grid covering the central sample) within the co-position picture. After the motion information of the co-position sub-CU is identified, it is converted into the motion vectors and reference indices of the current sub-CU in a manner similar to HEVC's TMVP processing. Here, temporal motion scaling is applied to align the reference picture of the temporal motion vector with that of the current CU.

[0056] In VVC, a merge list based on merged subblocks, containing both SbTMVP candidates and affine merge candidates, is used for signaling the subblock-based merge mode. The SbTMVP mode is enabled / disabled by the Sequence Parameter Set (SPS) flag. When SbTMVP mode is enabled, SbTMVP predictors are added as the first entry in the list of subblock-based merge candidates, followed by affine merge candidates. The size of the subblock-based merge list is signaled by the SPS, and the maximum allowed size of the subblock-based merge list is 5 in VVC.

[0057] In VVC, the sub-CU size used in SbTMVP is fixed at 8x8, and SbTMVP mode is only applicable to CUs where both width and height are 8 or greater, as is done in affine merge mode. Sub-block sizes may be configurable to other sizes, such as 4x4, when using ECM software models for studies after VVC.

[0058] Figure 2 shows an exemplary block diagram 200 of multiple displacement vectors used to code or decode video data using temporal motion vector prediction (TMVP) with displacement vectors, according to one embodiment of the present disclosure.

[0059] Embodiments of this disclosure relate to additional motion offsets used to derive motion vectors of a TMVP in order to improve the flexibility and efficiency of the TMVP.

[0060] According to aspects of this disclosure, instead of using predefined fixed positions to fetch motion information to be used as MVP from TMVP candidates for TMVP candidate derivation used in normal merge mode or AMVP mode, additional or additional offsets, i.e., displacement offsets, can be signaled to identify blocks in a reference picture, and motion information associated with these identified blocks can be used as MVP from TMVP candidates. For example, for current block C0, one or more displacement vectors (shown as solid arrows in Figure 2) can be added to current block C0 to identify multiple block positions (shown as dashed boxes in Figure 2). Motion vectors associated with these identified block positions in the reference picture can be used as temporal motion vector predictors.

[0061] In one embodiment, displacement vectors can be signaled by an index using the merged motion vector difference (MMVD) method. In one embodiment, displacement vectors can be signaled using a method similar to motion vector difference signaling with Adaptive Motion Vector Resolution (AMVR). The displacement vector resolution may be N samples, for example, N may be 1, 4, or 8. In one embodiment, displacement vector resolution can be signaled by a high-level syntax such as sequence level, picture level, slice level, or tile / tile group level. In one embodiment, displacement vector resolution can be signaled by a block-level resolution index. The resolution index can be used to retrieve displacement vectors in a resolution table. In some embodiments, the resolution table can be predefined. In some embodiments, the resolution table can be signaled at a high level, such as sequence level, picture level, etc.

[0062] In one embodiment, the displacement offset index can be sorted using ascending or descending template matching costs by applying sorting of the displacement vector index based on template matching. In one example, a first N candidates with ascending template matching costs can be used, where N is greater than or equal to 1 and less than or equal to the total number of available candidates.

[0063] In one embodiment, candidate positions indicated by different displacement vectors can be scanned in a predefined order, and a first N candidate positions associated with a coded block using motion vectors can be identified, and indices among these N candidate positions can be signaled to indicate which one of the candidates will be used as the TMVP candidate block position. In some embodiments, the predefined scan order can be determined by the relative distance between the candidate positions and the starting position. As an example, the initial position may represent a candidate position with a displacement vector of zero, e.g., C0 in Figure 2.

[0064] In one embodiment, the initial position may refer to a candidate position having a displacement vector of zero, and the starting position may be a default position, e.g., C0 in Figure 2, or it may be implicitly derived from coded information including, but not limited to, a selected candidate block position of a neighboring block coded using TMVP, and the motion vector of the neighboring block.

[0065] In one embodiment, the motion vector used as TMVP can be derived as the average or weighted average of motion vectors fetched from multiple block positions in the reference picture. In another embodiment, the motion vector used as TMVP can be derived as the motion vector value with the highest count among all motion vectors fetched from multiple block positions in the reference picture.

[0066] Figure 3 is a flowchart illustrating an exemplary process for coding and / or decoding video data using temporal motion vector prediction (TMVP) with displacement vectors, according to one embodiment of the present disclosure.

[0067] As shown in Figure 3, in operation 305, displacement vectors related to the current block in the current picture are obtained, and these displacement vectors may be signaled in the video bitstream to identify the reference block in the current picture. As an example, multiple displacement vectors can be obtained and added to block C0, as shown in Figure 2 (displacement vectors can be indicated using arrows).

[0068] In some embodiments, operation 305 may include the step of receiving a video bitstream containing one or more pictures. Operation 305 may also include the step of determining that one or more pictures should be predicted in normal merge mode or adaptive motion vector prediction (AMVP) mode. In some embodiments, the displacement vector or offset indicates the position of at least one of the motion vector predictors in the list of candidate motion vector predictors. In some embodiments, the displacement vector or offset indicates at least one of the displacement vectors among a plurality of displacement vectors associated with each candidate in the list of candidate motion vector predictors.

[0069] In operation 310, motion information associated with a reference block can be determined based on a displacement vector, and this motion information is used as a motion vector predictor (MVP) from a candidate for a temporal motion vector predictor (TMVP).

[0070] According to one aspect of this disclosure, the list of candidate time motion vector predictors may be sorted based on template matching costs. In some embodiments, the list of candidate time motion vector predictors can be generated using a predefined scan order, which may be based on the magnitude of a displacement vector among a plurality of displacement vectors.

[0071] In operation 315, a list of TMVP candidates containing motion information can be generated. In operation 320, the motion vector can be derived for the current block using the TMVP candidate list.

[0072] In operation 325, the current block can be decoded using the derived motion vector for prediction in normal merge mode or adaptive motion vector prediction (AMVP) mode.

[0073] In some embodiments, during encoding, a displacement offset associated with at least one motion vector predictor in the list of candidate temporal motion vector predictors used when deriving a motion vector for the current block using TMVP can be signaled. For example, the displacement offset can be signaled as an index using motion vector differences by motion vector representation techniques or by motion vector differences by adaptive motion vector resolution techniques. When using motion vector differences by adaptive motion vector resolution techniques, multiple displacement vectors may have a displacement vector resolution at a given number of samples, and the displacement vector resolution can be signaled with a high level of syntax.

[0074] Figure 3 shows an exemplary block of process 300, but in some implementations, process 300 may include more blocks, fewer blocks, or blocks in a different arrangement than those shown in Figure 3. Additionally or alternatively, two or more blocks of process 300 may be executed in parallel.

[0075] Furthermore, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-temporary computer-readable medium to perform one or more of the proposed methods.

[0076] Figure 4 shows a simplified block diagram of a communication system 400 according to an embodiment of the present disclosure. The communication system 400 may include at least two terminals 410-420 interconnected via a network 450. In one-way data transmission, the first terminal 410 may encode video data at its local location for transmission to the second terminal 420 via the network 450. The second terminal 420 may receive the encoded video data from the other terminal via the network 450, decode the encoded data, and display the restored video data. One-way data transmission may be common in media serving applications, etc.

[0077] Figure 4 shows a second pair of terminals 430, 440 applied to support bidirectional transmission of coded video, which may occur, for example, during a video conference. In bidirectional data transmission, each terminal 430, 440 may code locally captured video data for transmission to other terminals via the network 450. Each terminal 430, 440 may also receive coded video data transmitted by other terminals, decode the coded data, and display the restored video data on a local display device.

[0078] In Figure 4, terminal devices 410-440 may be shown as servers, personal computers, and smartphones, but the principles of this disclosure are not limited to these. Embodiments of this disclosure include applications by laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 450 represents any number of networks that carry coded video data between terminals 410-440, including, for example, wired and / or wireless communication networks. Communication network 450 may exchange data over circuit switching and / or packet switching channels. Typical networks include electronic communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of discussing this invention, the architecture and topology of network 450 may not be important to the operation of this disclosure unless otherwise specified below.

[0079] Figure 5 shows an example of the application of the subject matter of disclosure, illustrating the arrangement of a video encoder and video decoder in a streaming environment, such as a streaming system 500. The subject matter of disclosure is equally applicable to, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc., and other video-enabled applications.

[0080] The streaming system may include a capture subsystem 513, which may include a video source 501, such as a digital camera, that generates an uncompressed video sample stream 502. The sample stream 502 is shown in thick lines to emphasize its high data capacity compared to the encoded video bitstream and can be processed by an encoder 503 coupled to the camera 501. The encoder 503 may include hardware, software, or a combination thereof, and can enable or implement aspects of the subject of disclosure as detailed below. The encoded video bitstream 504 is shown in thin lines to emphasize its low data capacity compared to the sample stream and can be stored in the streaming server 505 for future use. One or more streaming clients 506, 508 can access the streaming server 505 and read video bitstreams 507, 509, which may be copies of the encoded video bitstream 504, for example. Client 506 may include a video decoder 510. The video decoder 510 decodes an incoming copy of the encoded video bitstream 507 to generate an output video sample stream 511 that can be rendered on the display 512 or other rendering device (not shown). In some streaming systems, the video bitstreams 504, 507, and 509 can be encoded according to specific video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. Video coding standards under development are informally known as VVC (Versatile Video Coding). The subject matter of this disclosure may be used in the context of VVC.

[0081] Figure 6 may be a functional block diagram of the video decoder 510 according to an embodiment of the present disclosure.

[0082] Receiver 610 may receive one or more coded video sequences to be decoded by decoder 510, or in the same or different embodiments, one coded video sequence at a time, where the decoding of each coded video sequence is independent of other coded video sequences. Coded video sequences may be received from channel 612, which may be a hardware / software link to a storage device that stores coded video data. Receiver 610 may receive coded video data together with other data, e.g., coded audio data and / or auxiliary data streams that may be transferred to each not-illustrated usage entity. Receiver 610 may isolate coded video sequences from other data. To eliminate network jitter, a buffer 615, which may be, for example, a buffer memory, may be coupled between receiver 610 and entropy decoder / parser 620, hereafter "parser". If receiver 610 is receiving data controllably from a storage / transfer device with sufficient bandwidth or from an isosynchronous network, buffer 615 may be unnecessary or small. When used in best-effort packet networks such as the internet, a buffer of 615 may be required, which can be relatively large and, advantageously, can be adapted to an adaptive size.

[0083] The video decoder 510 may include a parser 620 to reconstruct symbols 621 from the entropy-coded video sequence. These symbol categories may include information used to manage the operation of the decoder 510, and information for controlling a rendering device, such as a display 512, which may not be part of the decoder's integration but may be coupled to the decoder, as shown in Figure 6. The control information for the rendering device may be in the form of SEI (Supplementary Enhancement Information) messages or VUI (Video Usability Information) parameter set fragments, but are not shown. The parser 620 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow video coding techniques or standards, and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser 620 may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder, based on at least one parameter corresponding to that group. Subgroups may include GOP (Groups of Picture), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The entropy decoder / parser may also extract information from the coded video sequence such as transformation coefficients, quantizer parameter (QP) values, motion vectors, etc.

[0084] The parser 620 may perform an entropy decoding / parse operation on the video sequence received from the buffer 615 to generate a symbol 621. The parser 620 may receive encoded data and selectively decode a particular symbol 621. Furthermore, the parser 620 may determine whether a particular symbol 621 should be provided to the motion compensation prediction unit 653, the scaler / inverse transform unit 651, the intra prediction unit 652, or the loop filter unit 656.

[0085] The reconstruction of symbol 621 may include multiple different units, depending on the type of coded video picture or part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how they are included can be controlled by subgroup control information parsed from the video sequence coded by parser 620. The flow of such subgroup control information between parser 620 and the following multiple units is not shown for clarity.

[0086] Beyond the functional blocks already mentioned, the decoder 510 may be conceptually subdivided into numerous functional units, as described below. In actual implementations operating under commercial constraints, many of these units may interact closely with each other and be at least partially integrated. However, for the purpose of illustrating the subject of this disclosure, the following conceptual subdivision into functional units is appropriate.

[0087] The first unit is the scaler / inverse unit 651. The scaler / inverse unit 651 receives the quantized transformation coefficients and control information, including which transformation should be used, block size, quantization coefficients, quantization scaling matrix, etc., from the parser 620 as symbol 621. This can output a block containing sample values ​​that can be input to the aggregator 655.

[0088] In some cases, the output samples of the scaler / inverse transformer 651 may relate to intracoded blocks, i.e., blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from portions previously reconstructed of the picture. The intrapicture prediction unit 652 can provide such predictive information. In some cases, the intrapicture prediction unit 652 generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the currently partially reconstructed picture 658. In some cases, the aggregator 655 adds the predictive information generated by the intrapredictive unit 652 to the output sample information provided by the scaler / inverse transformer 651, sample by sample.

[0089] In other cases, the output samples of the scaler / inverse unit 651 may be associated with an interconnected, and possibly motion-compensated, block. In such cases, the motion-compensated prediction unit 653 can access the reference picture memory 657 to fetch samples to be used for prediction. After motion-compensating the fetched samples according to the symbols, the aggregator 655 may add 621 associated with the block to the output of the scaler / inverse unit. In this case, these sample conversion units are called residual samples or residual signals and generate output sample information. The address in the reference picture memory from which the motion-compensated prediction unit fetches the predicted samples can be controlled by the motion vector available to the motion-compensated prediction unit, for example, in the form of a symbol 621 which may have X, Y and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory when the exact motion vector of a subsample is in use, a motion vector prediction mechanism, etc.

[0090] The output samples of the aggregator 655 can undergo various loop filtering techniques in the loop filter unit 656. The video compression technique is controlled by parameters contained in the coded video bitstream and made available to the loop filter unit 656 as symbols 621 from the parser 620, but may include in-loop filtering techniques that can also respond to metadata obtained during decoding of previous parts in the decoding order of the coded picture or coded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.

[0091] The output of the loop filter unit 656 may be a sample stream that can be output to a display 512, which may be a renderer, and can be stored in a reference picture memory for use in future interpicture prediction.

[0092] A specific coded picture, once fully reconstructed, can be used as a reference picture for future predictions. Once a coded picture is fully reconstructed and identified as a reference picture by, for example, parser 620, the current reference picture 658 can become part of reference picture memory 657, which may be, for example, a reference picture buffer, and can be reallocated as fresh current picture memory before starting the reconstruction of subsequent coded pictures.

[0093] The video decoder 510 may perform decoding operations in accordance with a predetermined video compression technology, which may be defined in a standard such as ITU-T Rec. H.265. The coded video sequence may follow the syntax specified by the video compression technology or standard in use, in the sense that the coded video sequence conforms to the syntax of the video compression technology or standard, specifically as specified in the profile document therein. Furthermore, compliance may require that the complexity of the coded video sequence be within the limits set by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate measured in megasamples / second, maximum reference picture size, etc. The limits set by the level may, in some cases, be further limited through the HRD (Hypothetical Reference Decoder) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0094] In the embodiment, receiver 610 may receive additional redundant data along with the encoded video. The additional data may be included as a portion of the coded video sequence. The additional data may be used by video decoder 510 to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) extension layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0095] Figure 7 may be a functional block diagram of a video encoder 503 according to one embodiment of the present disclosure.

[0096] The encoder 503 may receive video samples from a video source 501 that is not part of the encoder capable of capturing video images to be coded by the encoder 503.

[0097] The video source 501 may provide a source video sequence to be coded by the encoder 503 in the form of a digital video sample stream of any suitable bit depth, e.g., 8 bits, 10 bits, 12 bits, ..., any color space, e.g., BT.601 Y CrCb, RGB, ..., and any suitable sampling structure, e.g., Y CrCb 4:2:0, Y CrCb 4:4:4. In a media delivery system, the video source 501 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 501 may be a camera that captures local image information as a video sequence. The video data may be provided as a series of individual pictures that give motion when viewed sequentially. The pictures themselves may be organized as a spatial array of pixels. Each pixel may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will immediately understand the relationship between pixels and samples. The following description will focus on samples.

[0098] According to one embodiment, the encoder 503 may encode and compress the pictures of the source video sequence into a coded video sequence 743 in real time or under any other time constraints required by the application. Implementing an appropriate coding speed is one function of the control unit 750. The control unit controls and is functionally coupled to other functional units, as described later. The coupling is not shown for clarity. Parameters set by the control unit may include rate control-related parameters such as picture skip, quantizer, lambda value of rate distortion optimization technique, ..., picture size, GOP (group of pictures) layout, maximum motion vector search range, etc. Those skilled in the art will immediately recognize other functions of the control unit 750 when relating to a video encoder 503 optimized for a particular system design.

[0099] Some video encoders operate within what a person skilled in the art would immediately recognize as a “coding loop.” In very simplified terms, the coding loop includes, for example, an encoding portion of a source coder 730, which may be an encoder that generates symbols based on an input picture to be coded and a reference picture, and a local decoder 733 embedded in the encoder 503. The local decoder 733 reconstructs the symbols to generate sample data. This sample data is also generated by a remote decoder, provided that the compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject. The reconstructed sample stream is input to a reference picture memory 734. The contents of the reference picture buffer are also bit-accurate between the local encoder and the remote encoder, where decoding of the symbol stream yields bit-accurate results, independently of whether the decoder position is local or remote. In other words, the predictive portion of the encoder “sees” the exact same sample values ​​as the decoder “sees” when using predictions during decoding, “sees” the reference picture samples. This fundamental principle of reference picture synchronization, and the resulting drift when synchronization cannot be maintained, for example due to channel errors, are well known to those skilled in the art.

[0100] The operation of the local decoder 733 may be the same as that of the remote decoder 510, as detailed above in relation to Figure 6. However, as also briefly referring to Figure 7, the entropy decoding portion of decoder 510, including channel 612, receiver 610, buffer 615, and parser 620, does not need to be fully implemented in the local decoder 733, since symbols are available and the encoding / decoding of symbols to the coded video sequence by entropy coder 745 and parser 620 may be lossless.

[0101] The consideration here is that any decoder techniques within the decoder, excluding parse / entropy decoding, must exist in substantially the same functional form as those within the corresponding encoder. The description of encoder techniques can be omitted, as they are the inverse of the decoder techniques over which they are comprehensively described. More detailed explanations are necessary only in specific areas, and are provided below.

[0102] During operation, in some examples, the source coder 730 may perform motion-compensated predictive coding. This predictively codes the input frame by referencing one or more previously coded frames from a video sequence designated as “reference frames”. In this method, the coding engine 732 codes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame which may be selected as the prediction criterion for the input frame.

[0103] The local video decoder 733 may decode the coded video data of a frame that may be designated as a reference frame, based on the symbols generated by the source coder 730. The operation of the coding engine 732 may, advantageously, be lossy. When the coded video data can be decoded by a video decoder not shown in Figure 7, the reconstructed video sequence may, as a standard, be a copy of the source video sequence with some errors. The local video decoder 733 may duplicate the decoding process that may be performed by the video decoder on the reference frame, resulting in a reconstructed reference frame to be stored in a reference picture memory 734, which may be a reference picture cache, for example. Thus, the encoder 503 may locally store a copy of the reconstructed reference frame that has the same content as the reconstructed reference frame obtained by the far-end video decoder, provided there are no transmission errors.

[0104] The predictor 735 may perform a predictive search for the coding engine 732. That is, for a new frame to be coded, the predictor 735 may search the reference picture memory 734 for sample data such as candidate reference pixel blocks or specific metadata such as reference picture motion vectors, block shapes, etc., which can serve as appropriate predictive criteria for the new picture. The predictor 735 may operate sample block-pixel block by sample block to find appropriate predictive criteria. In some examples, the input picture may have predictive criteria drawn from multiple reference pictures stored in the reference picture memory 734, as determined by the search results obtained by the predictor 735.

[0105] The control unit 750 may manage the coding operation of the source coder 730, including, for example, setting parameters and subgroup parameters used for encoding video data.

[0106] The outputs of all the aforementioned functional units may undergo entropy coding in the entropy coder 745. The entropy coder converts the symbols generated by the various functional units into coded video sequences by lossless compression of the symbols according to techniques well known to those skilled in the art, such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0107] The transmitter 740 may buffer the coded video sequence generated by the entropy coder 745 in preparation for transmission over a communication channel 760, which may be a hardware / software link to a storage device capable of storing coded video data. The transmitter 740 may merge the coded video data from the source coder 730 with other data to be transmitted, such as coded audio data and / or auxiliary data stream sources (not shown).

[0108] The control unit 750 may manage the operation of the encoder 503. During coding, the control unit 750 may assign each coded picture a specific coded picture type that may affect the coding techniques that can be applied to each picture. For example, a picture may often be assigned as one of the following picture types:

[0109] An intra-picture (I-picture) may be a picture that can be coded and decoded without using any other frames in the sequence as a source for prediction. Some video codecs allow different types of intra-pictures, including, for example, IDR (Independent Decoder Refresh) pictures. A person skilled in the art will recognize variations of I-pictures and their individual applications and characteristics.

[0110] A predictive picture (P-picture) may, in most cases, be a picture that can be coded and decoded using intra-prediction or inter-prediction with a single motion vector and reference index to predict the sample values ​​of each block.

[0111] A bidirectionally predictive picture (B-picture) may be a picture that can be coded and decoded using intra-prediction or inter-prediction with up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predictive picture may use two or more reference pictures and associated metadata for the reconstruction of a single block.

[0112] A source picture may generally be spatially subdivided into multiple sample blocks, for example, blocks of 4x4, 8x8, 4x8, or 16x16 samples each, and each block may be coded. Blocks may be coded predictively by references to other already coded blocks, determined by the coding assignment applied to each picture in the block. For example, blocks of picture I may be coded unpredictably, or they may be coded predictively by referencing already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of picture P may be coded predictively via spatial prediction or temporal prediction by referencing one previously coded reference picture. Blocks of picture B may be coded unpredictably via spatial prediction or temporal prediction by referencing one or two previously coded reference pictures.

[0113] The encoder 503 may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec.H.265. In these operations, the encoder 503 may perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. The coded video data may, therefore, conform to the syntax specified by the video coding technique or standard being used.

[0114] In one embodiment, the transmitter 740 may transmit additional data along with the encoded video. The source coder 730 may include such data as part of the encoded video sequence. The additional data may include time / space / SNR extension layers, other forms of redundant data such as redundant pictures and slices, SEI (Supplementary Enhancement Information) messages, VUI (Visual Usability Information) parameter set fragments, and the like.

[0115] The techniques described above are implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 9 shows a computer system 900 suitable for implementing a particular embodiment of the subject matter of this disclosure.

[0116] Computer software can be coded using any suitable machine code or computer language that can be processed by mechanisms such as assembly, compilation, and linking to generate code containing instructions that can be executed directly or through interpretation, microcode execution, etc., by a computer's central processing unit (CPU), graphics processing unit (GPU), etc.

[0117] The instructions can be executed on various computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things devices, etc.

[0118] The components shown in Figure 8 of the computer system 800 are illustrative and do not imply any limitation on the scope of use or functionality of computer software implementing embodiments of this disclosure. Furthermore, the configuration of the components should not be construed as having any dependencies or requirements relating to any one or combination of the components shown in the exemplary embodiments of the computer system 800.

[0119] The computer system 800 may include certain human interface input devices. Such human interface input devices may respond to input from one or more human users through, for example, sensory input (e.g., keystrokes, swipes, data grab actions), voice input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices can also be used to capture certain media that do not necessarily need to be directly related to conscious human input, such as voice (e.g., conversation, music, ambient sounds), images (e.g., scanned images, photographic images taken from a digital camera), and video (e.g., including 2D video, 3D video, and stereoscopic video).

[0120] The input human interface device may include one or more of the following (only one is shown): a keyboard 801, a mouse 802, a trackpad 803, a screen 810 which may be a touchscreen, a data grab 1204, a joystick 805, a microphone 806, a scanner 807, and a camera 808.

[0121] The computer system 800 may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through sensory output, sound, light, and smell / taste. Such human interface output devices may include sensory output devices (e.g., sensory feedback via a screen 810, data grab 1204, or joystick 805, although sensory feedback devices that do not function as input devices may also exist), audio output devices (e.g., speaker 809, headphones (not shown)), and visual output devices (e.g., screen 810, cathode ray tube (CRT) screen, liquid crystal display (LCD) screen, plasma screen, organic light-emitting diode (OLED) screen, each having or not having touchscreen input capability, each having or not having sensory feedback capability, some of which may be capable of outputting more, for example, stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown)).

[0122] The computer system 800 may also include human-accessible storage devices, and related media such as optical media like CD / DVDROM / RW820 with media 821 such as CD / DVD, a thumb drive 822, a removable hard drive or solid state drive 823, legacy magnetic media such as tape and floppy disks (not shown), and devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown).

[0123] Those skilled in the art should also understand that the term “computer-readable medium” as used in connection with the subject matter of this disclosure does not include a transmission medium, carrier wave, or other transient signal.

[0124] The computer system 800 may also include an interface to one or more communication networks 855. Network 855 may be, for example, wireless, wired, or optical. Network 855 may further be local, wide-area, urban, vehicle and industrial, real-time, latency-tolerant, etc. Examples of network 855 include local area networks such as Ethernet, cellular networks including wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial networks including CANBus, etc. Certain networks 855 generally require a specific general-purpose data port or peripheral bus 849 (for example, an external network interface connected to a USB port on a computer system 800). Others are generally integrated into the core of the computer system 800 by being connected to a system bus 1248, as described later (for example, an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using these networks 855, the computer system 800 can communicate with other entities. Such communication may be unidirectional reception only (e.g., broadcast TV), unidirectional transmission only (e.g., CANBus to a specific CANBus device), or bidirectional to other computer systems using, for example, a local or wide-area digital network. Specific protocols and protocol stacks may be used with each of these networks 855 and network interfaces such as the external network interface adapter 854 described above.

[0125] The aforementioned human interface device, human-accessible storage device, and network interface can be mounted on the core 840 of the computer system 800.

[0126] The core 840 may include one or more central processing units (CPUs) 841, graphics processing units (GPUs) 842, dedicated programmable processing units in the form of FPGAs 843, hardware accelerators 844 for specific tasks, etc. These devices may be connected via a system bus 1248, along with read-only memory (ROM) 845, random access memory (RAM) 846, and internal mass storage devices 847 such as internal, user-inaccessible hard drives, SSDs, etc. In some computer systems, the system bus 1248 is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripherals can be attached directly to the core's system bus 1248 or via a peripheral bus 849. The architecture of the peripheral bus includes peripheral component interconnects (PCI), USB, etc.

[0127] The CPU 841, GPU 842, FPGA 843, and accelerator 844 can execute specific instructions that, when combined, can generate the aforementioned computer code. This computer code can be stored in ROM 845 or RAM 846. Temporary data can also be stored in RAM 846, while permanent data can be stored, for example, in the built-in mass storage device 847. High-speed storage and retrieval to any of the memory devices can be enabled through the use of cache memory that may be closely associated with one or more of the CPU 841, GPU 842, mass storage device 847, ROM 845, RAM 846, etc.

[0128] Computer-readable media may contain computer code for performing actions performed by various computers. The media and computer code may be specifically designed and configured for the purposes of this disclosure, or they may be of a type well known and available to those skilled in the computer software field.

[0129] As an example and not limited thereto, a computer system having the architecture of computer system 800, and specifically the core 840, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be specific storage devices of the core 840 having non-transient characteristics, such as the core-integrated mass storage device 847 or ROM 845, and media associated with the user-accessible mass storage devices described above. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 840. The computer-readable media may include one or more memory devices or chips, depending on the specific needs. The software can cause the core 840 and specifically the processors therein (including a CPU, GPU, FPGA, etc.) to execute specific processes or specific parts of specific processes described herein, including defining and modifying data structures stored in RAM 846 according to software-defined processes. As an addition or alternative, a computer system may provide functionality as a result of a logic hardwired or other circuit implementation (e.g., accelerator 844) that can operate together with or in place of the software to perform the specific processes or specific parts of the specific processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include, where appropriate, circuits (such as integrated circuits (ICs)) that house software for execution, circuits that implement logic for execution, or both. This disclosure includes any appropriate combination of hardware and software.

[0130] While this disclosure describes several exemplary embodiments, alternatives, substitutions, and various equivalents exist and are included within the scope of this disclosure. As will be apparent to those skilled in the art, numerous systems and methods can be devised to implement the principles of this disclosure and thus fall within the spirit and scope of this disclosure, although these are not expressly shown or described herein.

Claims

[Claim 1] A method for coding video data using temporal motion vector prediction (TMVP), wherein the method is performed by one or more processors, The steps include receiving a video bitstream containing one or more pictures, The steps include determining that one or more of the aforementioned pictures should be predicted in normal merge mode or adaptive motion vector prediction (AMVP) mode, A step of obtaining a displacement vector associated with the current block in the current picture, wherein the displacement vector is signaled in the video bitstream to identify a reference block in the current picture. A step of determining motion information associated with the reference block based on the displacement vector, wherein the motion information is used as a motion vector predictor (MVP) from a candidate for a temporal motion vector predictor (TMVP), The steps include generating a list of TMVP candidates that includes the aforementioned motion information, The steps include: deriving the motion vector of the current block using the TMVP candidate list; The steps include decoding the current block using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode, A method that includes this.