Displacement-based temporal motion vector predictor
By employing additional displacement vectors and innovative signaling methods, the flexibility and efficiency of TMVP are improved, addressing limitations in existing video coding standards to enhance motion vector prediction accuracy and coding efficiency.
Patent Information
- Application Number
- JP2024547231
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-12-13
- Filing Date
- 2022-12-14
- Publication Date
- 2026-02-12
- Estimated Expiration
- 2042-12-14
AI Technical Summary
Existing video coding standards, such as HEVC and VVC, lack flexibility and efficiency in deriving motion vectors for temporal motion vector prediction (TMVP) due to the use of predefined block positions, which limits the accuracy and adaptability of motion vector prediction.
Introduce additional displacement vectors to identify reference blocks in the reference picture, allowing for flexible and efficient TMVP by signaling displacement offsets to derive motion vectors from multiple block positions, using methods like MMVD and AMVR, and applying template matching-based displacement vector index sorting.
Enhances the flexibility and efficiency of TMVP by improving the accuracy of motion vector prediction, reducing computational complexity, and optimizing coding efficiency in video coding processes.
Smart Images

Figure 0007813376000004 
Figure 0007813376000005 
Figure 0007813376000006
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims priority to U.S. Provisional Patent Application No. 63 / 391,219, filed July 21, 2022, and U.S. Provisional Patent Application No. 18 / 080,450, filed December 13, 2022, with the U.S. Patent and Trademark Office, the entire disclosures of which are incorporated herein by reference.
[0002] [Technical field] FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate to image and video coding techniques, and more particularly, to deriving a temporal motion vector predictor (TMVP) using displacement vectors. [Background technology]
[0003] The H.265 / HEVC (High Efficiency Video Coding) standard was published by ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standardization organizations formed the Joint Video Exploration Team (JVET) to explore the possibility of developing a next-generation video coding standard beyond HEVC. In October 2017, they announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for 360 video categories had been submitted. In April 2018, all received CfP responses were evaluated at the 122 MPEG / 10th JVET meeting. As a result of this meeting, JVET officially launched the standardization process for next-generation video coding beyond HEVC. The new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Experts Team. In 2020, ITU-TVCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the VVC video coding standard (Version 1). Summary of the Invention
[0004] According to an embodiment, there is provided a method for coding or decoding video data using temporal motion vector prediction (TMVP), the method being executable by a processor, comprising: receiving a video bitstream including one or more pictures; determining that the one or more pictures should be predicted in a normal merge mode or an adaptive motion vector prediction (AMVP) mode; obtaining a displacement vector associated with a current block in a current picture, the displacement vector being signaled in the video bitstream to identify a reference block in the current picture; determining motion information associated with the reference block based on the displacement vector, the motion information being used as a motion vector predictor (MVP) from temporal motion vector predictor (TMVP) candidates; generating a TMVP candidate list including the motion information; deriving a motion vector for the current block using the TMVP candidate list; decoding the current block using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode; may include:
[0005] According to embodiments, an apparatus for coding or decoding video data using temporal motion vector prediction (TMVP) may be provided. The apparatus may include at least one memory configured to store program code and at least one processor configured to access the program code and operate as directed by the program code. The program code may include: receiving code configured to cause the at least one processor to receive a video bitstream including one or more pictures; decision code configured to cause the at least one processor to determine that the one or more pictures are predicted in a normal merge mode or an adaptive motion vector prediction (AMVP) mode; retrieval code configured to cause the at least one processor to retrieve a displacement vector associated with a current block in a current picture, the displacement vector being signaled in the video bitstream to identify a reference block in the current picture; and a motion information code configured to cause the at least one processor to determine motion information associated with the reference block based on the displacement vector, the motion information being used as a motion vector predictor (MVP) from temporal motion vector predictor (TMVP) candidates; and generation code configured to cause the at least one processor to generate a TMVP candidate list that includes the motion information; derivation code configured to cause the at least one processor to derive a motion vector for the current block using the TMVP candidate list; decoding code configured to cause the at least one processor to decode the current block using the derived motion vector for prediction in the normal merge mode or an adaptive motion vector prediction (AMVP) mode; may include:
[0006] According to an embodiment, there may be provided a non-transitory computer-readable medium storing instructions that, when executed by one or more processors of an apparatus for coding video data using temporal motion vector prediction (TMVP), cause the one or more processors to: receiving a video bitstream including one or more pictures; determining that the one or more pictures should be predicted in a normal merge mode or an adaptive motion vector prediction (AMVP) mode; obtaining a displacement vector associated with a current block in a current picture, the displacement vector being signaled in the video bitstream to identify a reference block in the current picture; determining motion information associated with the reference block based on the displacement vector, the motion information being used as a motion vector predictor (MVP) from temporal motion vector predictor (TMVP) candidates; generating a TMVP candidate list including the motion information; deriving a motion vector for the current block using the TMVP candidate list; decoding the current block using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode; The instruction may include one or more instructions that can [Brief explanation of the drawings]
[0007] [Figure 1A] 10 illustrates an example of special merge candidate locations, according to one embodiment of the present disclosure.
[0008] [Figure 1B] 10 illustrates an example of candidate pairs considered for redundancy check of spatial merge candidates according to one embodiment of the present disclosure.
[0009] [Figure 1C] 10 illustrates an example of motion vector scaling for temporal merge candidates, according to one embodiment of the present disclosure.
[0010] [Figure 1D] 1 illustrates an example of the locations of temporal merge candidates, according to one embodiment of the present disclosure.
[0011] [Figure 1E] 1 illustrates an exemplary process for merging by motion vector difference (MMVD) search, according to one embodiment of the present disclosure.
[0012] [Figure 1F] 10 illustrates an example of merging by motion vector difference search points, according to one embodiment of the present disclosure.
[0013] [Figure 1G] 10 illustrates examples of additional directions along diagonal angles according to one embodiment of the present disclosure.
[0014] [Figure 1H] 1 illustrates an example of spatial neighborhood blocks used by the ATVMP, according to one embodiment of the present disclosure.
[0015] [Figure 1I] 10 illustrates an example process for deriving a sub-CU motion field based on motion shifts from spatial neighbors, according to one embodiment of the present disclosure.
[0016] [Figure 2] 1 shows an example block diagram of a plurality of displacement vectors used to code or decode video data using temporal motion vector prediction (TMVP) using displacement vectors, according to one embodiment of the present disclosure.
[0017] [Figure 3] 1 is a flowchart of an exemplary process for coding and / or decoding video data using temporal motion vector prediction (TMVP) using displacement vectors, in accordance with one embodiment of the present disclosure.
[0018] [Figure 4] 1 shows a simplified block diagram of a communication system according to one embodiment of the present disclosure.
[0019] [Figure 5] FIG. 1 is a diagram of the arrangement of a video encoder and a video decoder in a streaming environment.
[0020] [Figure 6] FIG. 2 is a functional block diagram of a video decoder according to one embodiment of the present disclosure.
[0021] [Figure 7] FIG. 2 is a functional block diagram of a video encoder according to one embodiment of the present disclosure.
[0022] [Figure 8] FIG. 1 is a diagram of a computer system according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0023] The proposed methods and processes can be used individually or in combination.Embodiments of the present disclosure relate to methods and systems for coding or decoding video data using temporal motion vector prediction (TMVP) using displacement vectors.
[0024] In related art, the block positions used to fetch motion vectors from TMVP candidates are predefined and fixed. Embodiments of the present disclosure relate to additional motion offsets used to derive motion vectors for TMVP to improve the flexibility and efficiency of TMVP.
[0025] According to aspects of the present disclosure, for TMVP candidate derivation used in normal merge mode or AMVP mode, instead of using a predefined fixed location to fetch motion information to be used as MVP from a TMVP candidate, an additional or supplemental offset, i.e., a displacement offset, can be signaled to identify a block in a reference picture, and motion information associated with this identified block can be used as MVP from the TMVP candidate. As an example, for a current block, one or more displacement vectors can be added to the current block to identify multiple block positions. The motion vectors associated with these identified block positions in the reference picture can be used as temporal motion vector predictors.
[0026] In one embodiment, the displacement vector can be signaled by an index using a merged motion vector difference (MMVD) method. In one embodiment, the displacement vector can be signaled using a method similar to motion vector difference signaling with Adaptive Motion Vector Resolution (AMVR). The displacement vector resolution can be N samples, where N may be 1, 4, or 8, for example. In one embodiment, the displacement vector resolution can be signaled by a high-level syntax, such as at the sequence level, picture level, slice level, or tile / tile group level. In one embodiment, the displacement vector resolution can be signaled by a block-level resolution index. The resolution index can be used to look up the displacement vector in a resolution table. In some embodiments, the resolution table can be predefined. In some embodiments, the resolution table can be signaled at a high level, such as at the sequence level, picture level, etc.
[0027] In one embodiment, template matching-based displacement vector index sorting can be applied to sort displacement offset indexes using ascending or descending template matching costs. In one example, the first N candidates with ascending template matching costs can be used, where N is greater than or equal to 1 and less than or equal to the total number of available candidates.
[0028] In one embodiment, the candidate locations indicated by the different displacement vectors can be scanned in a predefined order, the motion vector can be used to identify the first N candidate locations associated with the coded block, and an index among these N candidate locations can be signaled to indicate which one of the candidates is to be used as the TMVP candidate block location. In some embodiments, the predefined scan order can be determined by the relative distance between the candidate locations and the starting point location.
[0029] In one embodiment, the initial position may refer to a candidate position with a displacement vector of zero, and the starting point position may be a default position, e.g., C0 in FIG. 2, or may be implicitly derived by coded information, including but not limited to, selected candidate block positions of neighboring blocks coded using TMVP, motion vectors of neighboring blocks.
[0030] In one embodiment, the motion vector used as the TMVP may be derived as the average or weighted average of MVs fetched from multiple block locations in the reference picture. In one embodiment, the motion vector used as the TMVP may be derived as the motion vector value with the highest count among all motion vectors fetched from multiple block locations in the reference picture.
[0031] It can be appreciated that the methods and processes disclosed in this specification can be extended to multiple co-located pictures by extending the overlapping sub-blocks in the motion field from one co-located picture to multiple co-located pictures.
[0032] Inter Prediction in VVC
[0033] For each inter-prediction coding unit (CU), the motion parameters can include a motion vector, a reference picture index, a reference picture list usage index, and additional information required for the new coding features of VVC to be used for generating inter-predicted samples. Motion parameters can be signaled explicitly or implicitly. If a CU is coded in skip mode, the CU can be associated with one PU and can have no significant residual coefficients, coded motion vector deltas, or reference picture indexes. A merge mode can be specified, whereby the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates, and an additional schedule introduced in VVC. The merge mode can be applied not only to skip mode but also to inter-predicted CUs. An alternative to the merge mode is explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index of each reference picture list, the reference picture list usage flag, and other necessary information can be explicitly signaled for each CU.
[0034] Enhanced Merge Prediction
[0035] In VTM4, the merge candidate list for a mode is constructed by including five types of candidates, in order: (1) spatial MVPs from spatially neighboring CUs, (2) temporal MVPs from co-located CUs, (3) history-based MVPs from a FIFO table, (4) pairwise average MVPs, and (5) zero MVPs. The size of the merge list can be signaled in the slice header, and the maximum allowed size of a merge list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate can be coded using truncated unary binarization (TU). The first bin of the merge index can be coded with context, and bypass coding can be used for the other bins.
[0036] Deriving spatial candidates
[0037] The derivation of spatial merge candidates in VVC is similar to that in HEVC. Up to four merge candidates can be selected from among the candidates. FIG. 1A shows a current block 1100 showing exemplary positions of merge candidates B1, A1, B0, A0, and B2. In some embodiments, the derivation order can be B1, A1, B0, A0, and B2. Position B2 can be considered only if any of the CUs in positions A0, B0, B1, or A1 are unavailable (e.g., because the CU belongs to another slice or tile) or are intra-coded. After the candidate in position A1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are removed from the candidate list, thus improving coding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only pairs linked by arrows (where the error in 1B is not allowed as a reference source) are considered, and a candidate is added to the candidate list only if the corresponding candidate used in the redundancy check does not have the same motion information.
[0038] Temporal candidate derivation
[0039] In some embodiments, when deriving a temporal candidate, only one candidate can be added to a list. In particular, in deriving a temporal merge candidate, a scaled motion vector can be derived based on a co-located CU belonging to a co-located reference picture. The reference picture list used to derive the co-located CU can be explicitly signaled in the slice header. The scaled motion vector of the temporal merge candidate can be obtained as shown in FIG. 1C and scaled from the motion vector of the co-located CU. As shown in FIG. 1C, the scaled motion vector of the temporal merge candidate can be obtained and scaled from the motion vector of the co-located CU based on Picture Order Count (POC) distances tb and td, where tb is the POC difference between the reference picture of the current picture and the current picture, and td is the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate can be set to 0.
[0040] 1D, a temporal candidate location is selected between candidates C0 and C1. In some embodiments, location C1 can be used if the CU at location C0 is unavailable, intra-coded, or outside the current row of the Coding Tree Unit (CTU). Otherwise, location C0 can be used to derive the temporal merge candidate.
[0041] Merge with Motion Vector Difference (MMVD)
[0042] The merge mode with MMVD is used either in skip mode or merge mode with the motion vector representation method. MMVD can reuse merge candidates in VVC. A candidate can be selected from the merge candidates and further extended by the proposed motion vector representation method as shown in Figures 1E and 1F. MMVD can provide a new motion vector representation with simplified signaling. The representation method can include the starting point, the motion magnitude, and the motion direction.
[0043] The MMVD technique can utilize the merge candidate list in VVC. However, only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered for MMVD deployment. The base candidate index defines the starting point. The base candidate index indicates the best candidate from the list, as shown in Table 1. [Table 1] Basic candidate IDX [Table 1]
[0044] If the number of base candidates is equal to 1, the base candidate IDX may not be signaled. The distance index is information of the magnitude of the motion. The distance index indicates a predefined distance from the starting point information. The predefined distance may be as shown in Table 2. [Table 2] Distance IDX [Table 2]
[0045] The direction index can represent the direction of the MMVD based on the starting point. The direction index can represent four directions as shown in Table 3. [Table 3] Direction IDX [Table 3]
[0046] In some embodiments, the MMVD flag may be signaled immediately after sending the skip and merge flags. If the skip and merge flags are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is AFFINE mode, but if it is not 1, the skip / merge index is parsed for the VTM's skip / merge mode.
[0047] Candidate reordering based on template matching on MMVD and affine MMVD
[0048] In related art, the MMVD offset can be extended for MMVD and affine MMVD modes. Additional refinement positions along the k × π / 8 diagonal angle can be added as shown in FIG. 1G, thus increasing the number of directions from 4 to 16. Furthermore, all possible MMVD refinement positions (16 × 6) for each base candidate can be sorted based on the sum of absolute difference (SAD) cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. In some embodiments, the top 1 / 8 refinement positions with the smallest template SAD cost are retained as available positions and, therefore, for MMVD index coding. The MMVD index is binarized by a Rice code with a parameter equal to 2.
[0049] In embodiments of the present disclosure, on top of the MMVD extensions described herein, affine MMVD reordering can also be extended, where additional subdivision positions along the k×π / 4 diagonal angle can be added. After reordering, the top half of the subdivision positions with the smallest template SAD cost can be retained.
[0050] Subblock-based TMVP (SbTMVP)
[0051] To improve coding efficiency and reduce motion vector transmission overhead, subblock-level motion vector refinement can be applied to extend CU-level temporal motion vector prediction (TMVP). Subblock-based TMVP (SbTMVP) enables subblock-level motion information inheritance from co-located reference pictures. Each subblock of a large-sized CU can have its own motion information without explicitly transmitting block partition structure or motion information. SbTMVP can obtain the motion information of each subblock as follows: First, SbTMVP can include deriving the displacement vector (DV) of the current CU. Second, it derives the center motion based on the availability of SbTMVP candidates. Finally, SbTMVP can include deriving the motion information of a subblock from the corresponding subblock via the DV. Unlike TMVP candidate derivation, which always derives temporal motion vectors from co-located blocks in a reference frame, SbTMVP can apply DV derived from the motion vector (MV) of the current CU's left neighboring CU to find the corresponding sub-block in the co-located picture for each sub-block of the current CU. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set as the central motion.
[0052] VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to HEVC's temporal motion vector prediction (TMVP), SbTMVP uses motion fields from co-located pictures to improve the motion vector prediction of the current picture and the CU merge mode. The co-located pictures used by TMVP are used for SbTMVP. SbTMVP differs from TMVP in two main aspects:
[0053] (1) TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. (2) TMVP fetches temporal motion vectors from co-located blocks in the co-located picture (the co-located block is the bottom-right or center block relative to the current CU), while SbTMVP applies a motion shift before fetching temporal motion information from the co-located picture, where the motion shift (also called displacement vector or DV) is obtained from a motion vector from one of the spatial neighboring blocks of the current CU.
[0054] Figure 1H shows an example SbTMVP candidate selection using spatial neighboring blocks. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two parts. As the first part, it examines the spatial neighbor A1 in Figure 1H. If A1 has a motion vector that uses the co-located picture as a reference picture, it selects this motion vector as the motion shift (or displacement vector) to be applied. If no such motion is identified, the motion shift is set to (0,0).
[0055] In the second part, the motion shift identified in the first part is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vectors and reference indexes) from the co-located picture, as shown in FIG. 1I. Assume that the motion shift is set to the motion of block A1, as shown in FIG. 1I. Next, for each sub-CU, the motion information of the sub-CU is derived using the motion information of the corresponding block (the smallest motion grid covering the center sample) in the co-located picture. After the motion information of the co-located sub-CU is identified, it is converted into the motion vectors and reference indexes of the current sub-CU in a manner similar to the TMVP processing of HEVC. Here, temporal motion scaling is applied to align the reference picture of the temporal motion vector with that of the current CU.
[0056] In VVC, a combined sub-block-based merge list containing both SbTMVP candidates and affine merge candidates is used to signal the sub-block-based merge mode. SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. When SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block-based merge candidates, followed by the affine merge candidates. The size of the sub-block-based merge list is signaled in the SPS, and the maximum allowed size of the sub-block-based merge list is 5 in VVC.
[0057] In VVC, the sub-CU size used in SbTMVP is fixed at 8x8, and as is done in affine merge mode, SbTMVP mode is only applicable to CUs whose width and height are both 8 or greater. The sub-block size may be configurable to other sizes, such as 4x4, in the use of ECM software models for subsequent VVC studies.
[0058] FIG. 2 illustrates an example block diagram 200 of multiple displacement vectors used to code or decode video data using temporal motion vector prediction (TMVP) using displacement vectors, according to one embodiment of the present disclosure.
[0059] Embodiments of the present disclosure relate to additional motion offsets used to derive motion vectors for TMVP to improve the flexibility and efficiency of TMVP.
[0060] According to aspects of the present disclosure, for TMVP candidate derivation used in normal merge mode or AMVP mode, instead of using a predefined fixed position to fetch motion information to be used as MVP from a TMVP candidate, an additional or supplemental offset, i.e., a displacement offset, can be signaled to identify a block in a reference picture, and motion information associated with this identified block can be used as MVP from the TMVP candidate. As an example, for a current block C0, one or more displacement vectors (indicated by solid arrows in FIG. 2) can be added to the current block C0 to identify multiple block positions (indicated by dashed boxes in FIG. 2). The motion vectors associated with these identified block positions in the reference picture can be used as temporal motion vector predictors.
[0061] In one embodiment, the displacement vector can be signaled by an index using a merged motion vector difference (MMVD) method. In one embodiment, the displacement vector can be signaled using a method similar to motion vector difference signaling with Adaptive Motion Vector Resolution (AMVR). The displacement vector resolution can be N samples, where N may be 1, 4, or 8, for example. In one embodiment, the displacement vector resolution can be signaled by a high-level syntax, such as at the sequence level, picture level, slice level, or tile / tile group level. In one embodiment, the displacement vector resolution can be signaled by a block-level resolution index. The resolution index can be used to look up the displacement vector in a resolution table. In some embodiments, the resolution table can be predefined. In some embodiments, the resolution table can be signaled at a high level, such as at the sequence level, picture level, etc.
[0062] In one embodiment, template matching-based displacement vector index sorting can be applied to sort displacement offset indexes using ascending or descending template matching costs. In one example, the first N candidates with ascending template matching costs can be used, where N is greater than or equal to 1 and less than or equal to the total number of available candidates.
[0063] In one embodiment, the candidate locations indicated by the different displacement vectors can be scanned in a predefined order, and the motion vector can be used to identify the first N candidate locations associated with the coded block, and an index among these N candidate locations can be signaled to indicate which one of the candidates is to be used as the TMVP candidate block location. In some embodiments, the predefined scan order can be determined by the relative distance between the candidate locations and the starting point location. As an example, the initial location may represent a candidate location with a displacement vector of zero, e.g., C0 in FIG. 2.
[0064] In one embodiment, the initial position may refer to a candidate position with a displacement vector of zero, and the starting point position may be a default position, e.g., C0 in FIG. 2, or may be implicitly derived by coded information, including but not limited to, selected candidate block positions of neighboring blocks coded using TMVP, motion vectors of neighboring blocks.
[0065] In one embodiment, the motion vector used as the TMVP may be derived as the average or weighted average of MVs fetched from multiple block locations in the reference picture. In one embodiment, the motion vector used as the TMVP may be derived as the motion vector value with the highest count among all motion vectors fetched from multiple block locations in the reference picture.
[0066] FIG. 3 is a flowchart of an exemplary process for coding and / or decoding video data using temporal motion vector prediction (TMVP) using displacement vectors, according to one embodiment of this disclosure.
[0067] 3, in operation 305, a displacement vector associated with a current block in a current picture is obtained, and the displacement vector may be signaled in a video bitstream to identify a reference block in the current picture. As an example, as shown in FIG. 2, multiple displacement vectors may be obtained and added to block C0 (the displacement vectors may be indicated using arrows).
[0068] In some embodiments, operation 305 may include receiving a video bitstream including one or more pictures. Operation 305 may also include determining that the one or more pictures should be predicted in a normal merge mode or an adaptive motion vector prediction (AMVP) mode. In some embodiments, the displacement vector or offset indicates a position of at least one respective one of the at least one motion vector predictor in a temporal motion vector predictor candidate list. In some embodiments, the displacement vector or offset indicates at least one respective displacement vector among a plurality of displacement vectors associated with each candidate in the temporal motion vector predictor candidate list.
[0069] In operation 310, motion information associated with the reference block can be determined based on the displacement vector, and this motion information is used as a motion vector predictor (MVP) from the temporal motion vector predictor (TMVP) candidates.
[0070] According to one aspect of the present disclosure, the temporal motion vector predictor candidate list may be sorted based on template matching cost. In some embodiments, the temporal motion vector predictor candidate list may be generated using a predefined scan order, and the predefined scan order may be based on the magnitude of the displacement vectors among the plurality of displacement vectors.
[0071] A TMVP candidate list including motion information may be generated in operation 315. In operation 320, the TMVP candidate list may be used to derive a motion vector for the current block.
[0072] In operation 325, the current block can be decoded using the derived motion vector for prediction in normal merge mode or adaptive motion vector prediction (AMVP) mode.
[0073] In some embodiments, during encoding, a displacement offset associated with at least one motion vector predictor in a temporal motion vector predictor candidate list used in deriving a motion vector for a current block using TMVP can be signaled. As an example, the displacement offset can be signaled as an index using a motion vector differential with a motion vector representation technique or a motion vector differential with an adaptive motion vector resolution technique. When using a motion vector differential with an adaptive motion vector resolution technique, multiple displacement vectors can have a displacement vector resolution in a specific number of samples, and the displacement vector resolution can be signaled in a high-level syntax.
[0074] Although Figure 3 illustrates example blocks of process 300, process 300 may, in some implementations, include more, fewer, or differently arranged blocks than those illustrated in Figure 3. Additionally or alternatively, two or more of the blocks of process 300 may be performed in parallel.
[0075] Additionally, the proposed methods may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.
[0076] 4 shows a simplified block diagram of a communication system 400 according to an embodiment of the present disclosure. The communication system 400 may include at least two terminals 410-420 interconnected via a network 450. In a unidirectional data transmission, a first terminal 410 may locally code video data for transmission to a second terminal 420 via the network 450. The second terminal 420 may receive the coded video data of another terminal from the network 450, decode the coded data, and display the recovered video data. Unidirectional data transmission may be common in media serving applications, etc.
[0077] 4 shows a second pair of terminals 430, 440 adapted to support two-way transmission of coded video, such as may occur during a video conference. In the two-way transmission of data, each terminal 430, 440 may code locally captured video data for transmission to the other terminal over network 450. Each terminal 430, 440 may also receive coded video data transmitted by the other terminal, decode the coded data, and display the recovered video data on a local display device.
[0078] In FIG. 4 , terminal devices 410-440 may be depicted as servers, personal computers, and smartphones, although the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure may also be applied to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 450 represents any number of networks carrying coded video data between terminals 410-440, including, for example, wired and / or wireless communication networks. Communication network 450 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include electronic communication networks, local area networks, wide area networks, and / or the Internet. For purposes of discussion of the present invention, the architecture and topology of network 450 may not be important to the operation of the present disclosure, unless otherwise noted below.
[0079] 5 illustrates, as an example of an application of the disclosed subject matter, the arrangement of a video encoder and a video decoder in a streaming environment, e.g., streaming system 500. The disclosed subject matter is equally applicable to, e.g., video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc., other video-enabled applications, etc.
[0080] The streaming system may include a video source 501, e.g., a capture subsystem 513 that may include, e.g., a digital camera, generating an uncompressed video sample stream 502. The sample stream 502 can be processed by an encoder 503, shown in bold to emphasize its high data volume when compared to an encoded video bitstream, coupled to the camera 501. The encoder 503 may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 504, shown in thin to emphasize its low data volume when compared to the sample stream, can be stored on a streaming server 505 for future use. One or more streaming clients 506, 508 can access the streaming server 505 to retrieve video bitstreams 507, 509, which may be copies of the encoded video bitstream 504, for example. The client 506 may include a video decoder 510. A video decoder 510 decodes an incoming copy of the coded video bitstream 507 and generates an output video sample stream 511 that can be rendered on a display 512 or other rendering device (not shown). In some streaming systems, the video bitstreams 504, 507, 509 may be coded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. The video coding standard under development is known informally as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.
[0081] FIG. 6 may be a functional block diagram of a video decoder 510 according to an embodiment of the present disclosure.
[0082] Receiver 610 may receive one or more coded video sequences to be decoded by decoder 510, one coded video sequence at a time in the same or another embodiment, where decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from channel 612, which may be a hardware / software link to a storage device that stores the coded video data. Receiver 610 may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to respective using entities (not shown). Receiver 610 may separate the coded video sequences from the other data. To remove network jitter, buffer 615, which may be a buffer memory, may be coupled between receiver 610 and entropy decoder / parser 620 (hereafter referred to as the “parser”). If receiver 610 is receiving data controllably from a storage / forwarding device with sufficient bandwidth or from an isochronous network, buffer 615 may not be necessary or may be small. For use in a best-effort packet network such as the Internet, buffer 615 may be necessary and may be relatively large, and advantageously may be adaptively sized.
[0083] The video decoder 510 may include a parser 620 to reconstruct symbols 621 from the entropy-coded video sequence. These symbol categories include information used to manage the operation of the decoder 510 and information for controlling a rendering device, such as a display 512, which may be coupled to the decoder but not an integral part of the decoder, as shown in FIG. 6. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment, not shown. The parser 620 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser 620 may extract a set of subgroup parameters from the coded video sequence for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the subgroup. Subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter (QP) values, motion vectors, etc. from the coded video sequence.
[0084] Parser 620 may perform entropy decoding / parsing operations on the video sequence received from buffer 615 to generate symbols 621. Parser 620 may receive the encoded data and selectively decode particular symbols 621. Furthermore, parser 620 may determine whether a particular symbol 621 should be provided to motion compensated prediction unit 653, scaler / inverse transform unit 651, intra prediction unit 652, or loop filter unit 656.
[0085] The reconstruction of symbols 621 may include several different units, depending on the type of coded video picture or portion thereof (e.g., inter and intra picture, inter and intra block) and other factors. Which units are included and how they are included can be controlled by subgroup control information parsed from the coded video sequence by parser 620. The flow of such subgroup control information between parser 620 and the following units is not shown for clarity.
[0086] Beyond the functional blocks already mentioned, the decoder 510 may be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0087] The first unit is a scalar / inverse transform unit 651. The scalar / inverse transform unit 651 receives quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as symbols 621 from a parser 620. It can output blocks containing sample values that can be input to an aggregator 655.
[0088] In some examples, the output samples of the scaler / inverse transform unit 651 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. The intra-picture prediction unit 652 may provide such prediction information. In some cases, the intra-picture prediction unit 652 generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current partially reconstructed picture 658. The aggregator 655, in some cases, adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 652 to the output sample information provided by the scaler / inverse transform unit 651.
[0089] In other cases, the output samples of the scalar / inverse transform unit 651 may relate to an inter-coded, possibly motion-compensated, block. In such cases, the motion-compensated prediction unit 653 can access the reference picture memory 657 to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols, the aggregator 655 can add the block-related sample transforms 621 to the output of the scalar / inverse transform unit. In this case, these sample transforms, called residual samples or residual signals, generate output sample information. The addresses in the reference picture memory from which the motion-compensated prediction unit fetches prediction samples can be controlled by the motion-compensated prediction unit's available motion vectors, e.g., in the form of symbols 621 that may have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory when sub-sample accurate motion vectors are in use, motion vector prediction mechanisms, etc.
[0090] The output samples of aggregator 655 may be subjected to various loop filtering techniques in loop filter unit 656. The video compression techniques are controlled by parameters included in the coded video bitstream and made available to loop filter unit 656 as symbols 621 from parser 620, but may also include in-loop filter techniques that are responsive to meta-information obtained during decoding of previous portions in decoding order of the coded picture or coded video sequence, and that may also be responsive to previously reconstructed, loop-filtered sample values.
[0091] The output of the loop filter unit 656 may be a sample stream that can be output to a display 512, which may be, for example, a render device, and stored in a reference picture memory for use in future inter-picture prediction.
[0092] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture, for example, by parser 620, current reference picture 658 can become part of reference picture memory 657, which may be, for example, a reference picture buffer, and fresh current picture memory can be reallocated before starting reconstruction of a subsequent coded picture.
[0093] The video decoder 510 may perform decoding operations in accordance with a predetermined video compression technology, which may be defined in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard in use, in the sense that the coded video sequence conforms to the syntax of the video compression technology or standard, as specified in the video compression technology or standard, specifically in a profile document therein. Compliance may also require that the complexity of the coded video sequence be within limits established by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate, e.g., measured in megasamples per second, maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled within the coded video sequence.
[0094] In an embodiment, receiver 610 may receive additional redundant data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by video decoder 510 to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0095] FIG. 7 may be a functional block diagram of a video encoder 503 according to one embodiment of the present disclosure.
[0096] The encoder 503 may receive video samples from a video source 501 that is not part of the encoder and may capture video images to be coded by the encoder 503 .
[0097] The video source 501 may provide a source video sequence to be coded by the encoder 503 in the form of a digital video sample stream of any suitable bit depth, e.g., 8-bit, 10-bit, 12-bit, ..., any color space, e.g., BT.601 Y CrCb, RGB, ..., and any suitable sampling structure, e.g., Y CrCb 4:2:0, Y CrCb 4:4:4. In a media presentation system, the video source 501 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 501 may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed sequentially, give the appearance of motion. The pictures themselves may be organized as a spatial array of pixels. Each pixel may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following discussion focuses on samples.
[0098] According to one embodiment, the encoder 503 may code and compress pictures of a source video sequence into a coded video sequence 743 in real time or under any other time constraints required by the application. Enforcing the appropriate coding rate is one function of the controller 750. The controller controls and is operatively coupled to other functional units, as described below. The coupling is not shown for clarity. Parameters set by the controller may include rate control-related parameters picture skip, quantizer, lambda value for rate-distortion optimization techniques, picture size, GOP (group of pictures) layout, maximum motion vector search range, etc. Those skilled in the art will readily identify other functions of the controller 750 as they may be associated with optimizing the video encoder 503 for a particular system design.
[0099] Some video encoders operate in what those skilled in the art would immediately recognize as a "coding loop." As a highly simplified explanation, the coding loop includes an encoding portion of a source coder 730, which may be, for example, an encoder responsible for generating symbols based on an input picture to be coded and a reference picture, and a local decoder 733 embedded in the encoder 503. The local decoder 733 reconstructs the symbols to generate sample data. This sample data is also generated by a remote decoder when the compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter. The reconstructed sample stream is input to a reference picture memory 734. When decoding the symbol stream yields bit-exact results independent of whether the decoder is local or remote, the contents of the reference picture buffer are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" exactly the same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization and the resulting drift when synchronization cannot be maintained, for example due to channel errors, are well known to those skilled in the art.
[0100] The operation of the local decoder 733 may be the same as that of the remote decoder 510, detailed above in connection with Figure 6. However, and referring briefly to Figure 7 as well, the entropy decoding portion of the decoder 510, including the channel 612, receiver 610, buffer 615, and parser 620, may not be fully implemented in the local decoder 733 because symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder 745 and parser 620 may be lossless.
[0101] A consideration to be made at this point is that any decoder technology, other than parsing / entropy decoding, present in the decoder must also be present in substantially the same functional form as in the corresponding encoder. Descriptions of encoder technologies can be omitted, as they are the inverse of the decoder technology, which is described generically. Only in certain areas is more detailed description necessary, and is provided below.
[0102] In operation, in some examples, the source coder 730 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from a video sequence designated as "reference frames." In this method, the coding engine 732 codes differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as prediction references for the input frame.
[0103] The local video decoder 733 may decode the coded video data of frames, which may be designated as reference frames, based on symbols generated by the source coder 730. The operation of the coding engine 732 may advantageously be lossy. When the coded video data is decoded in a video decoder not shown in FIG. 7, the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 733 may replicate the decoding process that may be performed by a video decoder on the reference frames, resulting in reconstructed reference frames to be stored in a reference picture memory 734, which may be, for example, a reference picture cache. In this way, the encoder 503 may locally store copies of reconstructed reference frames that have content in common with reconstructed reference frames that would be obtained by the far-end video decoder in the absence of transmission errors.
[0104] The predictor 735 may perform a predictive search for the coding engine 732. That is, for a new frame to be coded, the predictor 735 may search the reference picture memory 734 for sample data, such as candidate reference pixel blocks, or specific metadata, such as reference picture motion vectors, block shapes, etc., that can serve as suitable prediction references for the new picture. The predictor 735 may operate on a sample block-pixel block basis to find a suitable prediction reference. In some examples, an input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 734, as determined by the search results obtained by the predictor 735.
[0105] The control unit 750 may manage the coding operations of the source coder 730, including, for example, setting the parameters and subgroup parameters used for encoding the video data.
[0106] The output of all the aforementioned functional units may undergo entropy coding in entropy coder 745. The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0107] The transmitter 740 may buffer the coded video sequence generated by the entropy coder 745 for transmission over a communication channel 760, which may be a hardware / software link to a storage device that may store the coded video data. The transmitter 740 may merge the coded video data from the source coder 730 with other data to be transmitted, such as coded audio data and / or an auxiliary data stream source, not shown.
[0108] A control unit 750 may manage the operation of the encoder 503. During coding, the control unit 750 may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0109] An intra-picture (I-picture) may be a picture that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will recognize variations of I-pictures and their respective applications and characteristics.
[0110] A predictive picture (P-picture) may be a picture that can be coded and decoded using intra- or inter-prediction, in most cases using one motion vector and reference index to predict the sample values of each block.
[0111] A bidirectionally predictive picture (B picture) may be a picture that can be coded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0112] A source picture is generally spatially subdivided into multiple sample blocks, e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each, and may be coded block by block. Blocks may be predictively coded with reference to other already coded blocks determined by the coding assignment applied to each picture of the block. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0113] The encoder 503 may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the encoder 503 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard being used.
[0114] In one embodiment, the transmitter 740 may transmit additional data along with the coded video. The source coder 730 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0115] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 9 illustrates a computer system 900 suitable for implementing certain embodiments of the subject matter of this disclosure.
[0116] Computer software can be coded using any suitable machine code or computer language that can be processed by mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly or through interpretation, microcode execution, etc.
[0117] The instructions may be executed by a variety of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0118] 8 of computer system 800 are exemplary in nature and do not suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Furthermore, the arrangement of components should not be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 800.
[0119] The computer system 800 may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users, for example, through sensory input (e.g., keystrokes, swipes, data grabbing actions), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a digital camera), and video (including, for example, two-dimensional video, three-dimensional video, and stereoscopic video).
[0120] The input human interface devices may include one or more of a keyboard 801, a mouse 802, a trackpad 803, a screen 810 which may be, for example, a touchscreen, a data grab 1204, a joystick 805, a microphone 806, a scanner 807, and a camera 808 (only one of which is shown).
[0121] The computer system 800 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses through, for example, sensory output, sound, light, and smell / taste. Such human interface output devices may include sensory output devices (e.g., sensory feedback via the screen 810, data glove 1204, or joystick 805, although there may also be sensory feedback devices that do not function as input devices), audio output devices (e.g., speakers 809, headphones (not shown)), visual output devices (e.g., including the screen 810, a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, and an organic light-emitting diode (OLED) screen, each with or without touchscreen input capability, each with or without sensory feedback capability, some of which may be capable of outputting more than one output, such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke generator tanks (not shown), and printers (not shown)).
[0122] The computer system 800 may also include human-accessible storage and associated media such as optical media such as CD / DVD ROM / RW 820 with media 821 such as CD / DVD, thumb drive 822, removable hard drive or solid state drive 823, legacy magnetic media such as tape and floppy disk (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.
[0123] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0124] The computer system 800 may also include interfaces to one or more communications networks 855. The networks 855 may be, for example, wireless, wired, optical, and may be local, wide area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks 855 include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, etc. Particular networks 855 generally require external network interfaces that are attached to particular general-purpose data ports or peripheral buses 849 (e.g., USB ports on the computer system 800). Others are generally integrated into the core of the computer system 800 by attachment to the system bus 1248 as described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using these networks 855, the computer system 800 can communicate with other entities. Such communications may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., the CANBus to a particular CANBus device), or two-way to other computer systems, using, for example, local or wide-area digital networks. Particular protocols and protocol stacks may be used with each of these networks 855 and network interfaces, such as the external network interface adapter 854 described above.
[0125] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to core 840 of computer system 800 .
[0126] The core 840 may include one or more central processing units (CPUs) 841, graphics processing units (GPUs) 842, dedicated programmable processing units 843 in the form of FPGAs, hardware accelerators 844 for specific tasks, etc. These devices may be connected through a system bus 1248, along with read-only memory (ROM) 845, random-access memory (RAM) 846, and internal mass storage device 847 such as an internal non-user-accessible hard drive, SSD, etc. In some computer systems, the system bus 1248 is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripherals can be attached directly to the core's system bus 1248 or through a peripheral bus 849. Peripheral bus architectures include peripheral component interconnect (PCI), USB, etc.
[0127] The CPU 841, GPU 842, FPGA 843, and accelerator 844 may combine to execute specific instructions that may generate the aforementioned computer code. The computer code may be stored in ROM 845 or RAM 846. Temporary data may also be stored in RAM 846, while permanent data may be stored, for example, in internal mass storage device 847. Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 841, GPU 842, mass storage device 847, ROM 845, RAM 846, etc.
[0128] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0129] As an example and not by way of limitation, a computer system having the architecture of computer system 800, and specifically core 840, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be specific storage of core 840 of a non-transitory nature, such as core's internal mass storage 847 or ROM 845, as well as media associated with user-accessible mass storage devices such as those described above. Software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 840. The computer-readable media may include one or more memory devices or chips, depending on particular needs. The software can cause core 840, and specifically the processor therein (including a CPU, GPU, FPGA, etc.), to perform specific processes or portions of specific processes described herein, including defining and modifying data structures stored in RAM 846 according to software-defined operations. Additionally or alternatively, a computer system may provide functionality as a result of implementation in hardwired or other circuitry (e.g., accelerator 844) that can operate in conjunction with or in place of software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include, where appropriate, circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that implements logic for execution, or both. The present disclosure includes any appropriate combination of hardware and software.
[0130] While this disclosure has described several exemplary embodiments, alterations, permutations, and various substitute equivalents exist, and are encompassed within the scope of this disclosure. Those skilled in the art will appreciate that numerous systems and methods can be devised that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore are within the spirit and scope of the present disclosure.
Claims
1. 1. A method for coding video data using temporal motion vector prediction (TMVP), the method being executed by one or more processors, the method comprising: receiving a video bitstream including one or more pictures; determining that the one or more pictures should be predicted in a normal merge mode or an adaptive motion vector prediction (AMVP) mode; obtaining a displacement vector associated with a current block in a current picture, the displacement vector being signaled in the video bitstream to identify a reference block in the current picture; determining motion information associated with the reference block based on the displacement vector, the motion information being used as a motion vector predictor (MVP) from temporal motion vector predictor (TMVP) candidates; generating a TMVP candidate list including the motion information; deriving a motion vector for the current block using the TMVP candidate list; decoding the current block using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode; Including, The TMVP candidate list is sorted based on template matching cost.
2. The method of claim 1 , wherein the displacement vector indicates the position of at least one respective one of at least one motion vector predictor in the temporal motion vector predictor candidate list.
3. The method of claim 1 , wherein the displacement vector indicates at least one respective displacement vector among a plurality of displacement vectors associated with each candidate in the temporal motion vector predictor candidate list.
4. The method of claim 1 , wherein the step of obtaining the displacement vector is based on an index indicating a motion vector difference according to a motion vector representation technique.
5. The method of claim 1 , wherein the step of obtaining the displacement vector is based on an index indicative of a motion vector difference according to an adaptive motion vector resolution technique.
6. The method of claim 5 , wherein the displacement vector has a displacement vector resolution of a particular number of samples.
7. The method of claim 6 , wherein the displacement vector resolution is signaled in a high-level syntax.
8. The method of claim 2 , wherein the temporal motion vector predictor candidate list is generated using a predefined scan order, the predefined scan order being based on the magnitude of the displacement vector.
9. 1. An apparatus for coding video data using temporal motion vector prediction (TMVP), the apparatus comprising: at least one memory configured to store program code; at least one processor configured to read the program code and to act according to the instructions of the program code; the program code comprising: receiving code configured to cause the at least one processor to receive a video bitstream including one or more pictures; decision code configured to cause the at least one processor to determine that the one or more pictures are predicted in a normal merge mode or an adaptive motion vector prediction (AMVP) mode; retrieval code configured to cause the at least one processor to retrieve a displacement vector associated with a current block in a current picture, the displacement vector being signaled in the video bitstream to identify a reference block in the current picture; and a motion information code configured to cause the at least one processor to determine motion information associated with the reference block based on the displacement vector, the motion information being used as a motion vector predictor (MVP) from temporal motion vector predictor (TMVP) candidates; and generation code configured to cause the at least one processor to generate a TMVP candidate list that includes the motion information; derivation code configured to cause the at least one processor to derive a motion vector for the current block using the TMVP candidate list; decoding code configured to cause the at least one processor to decode the current block using the derived motion vector for prediction in the normal merge mode or an adaptive motion vector prediction (AMVP) mode; Including, The TMVP candidate list is sorted based on template matching cost.
10. The apparatus of claim 9 , wherein the displacement vector indicates a position of at least one respective one of at least one motion vector predictor in the temporal motion vector predictor candidate list.
11. The apparatus of claim 9 , wherein the displacement vector indicates at least one respective displacement vector among a plurality of displacement vectors associated with each candidate in the temporal motion vector predictor candidate list.
12. The apparatus of claim 9 , wherein the step of obtaining the displacement vector is based on an index indicating a motion vector difference according to a motion vector representation technique.
13. The apparatus of claim 9 , wherein obtaining the displacement vector is based on an index indicative of a motion vector difference according to an adaptive motion vector resolution technique.
14. The apparatus of claim 13 , wherein the displacement vector has a displacement vector resolution of a particular number of samples.
15. The apparatus of claim 10 , wherein the temporal motion vector predictor candidate list is generated using a predefined scan order, the predefined scan order being based on the magnitude of the displacement vector.
16. 1. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of an apparatus for coding video data using temporal motion vector prediction (TMVP), cause the one or more processors to: receiving a video bitstream including one or more pictures; determining that the one or more pictures should be predicted in a normal merge mode or an adaptive motion vector prediction (AMVP) mode; obtaining a displacement vector associated with a current block in a current picture, the displacement vector being signaled in the video bitstream to identify a reference block in the current picture; determining motion information associated with the reference block based on the displacement vector, the motion information being used as a motion vector predictor (MVP) from temporal motion vector predictor (TMVP) candidates; generating a TMVP candidate list including the motion information; deriving a motion vector for the current block using the TMVP candidate list; decoding the current block using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode; It contains one or more instructions, The TMVP candidate list is sorted based on template matching cost.
17. The non-transitory computer-readable medium of claim 16 , wherein the displacement vector indicates a position of at least one respective one of at least one motion vector predictor in the temporal motion vector predictor candidate list.
18. 17. The non-transitory computer-readable medium of claim 16, wherein the displacement vector indicates at least one respective displacement vector among a plurality of displacement vectors associated with each candidate in the temporal motion vector predictor candidate list.
19. 1. A method for coding video data using temporal motion vector prediction (TMVP), the method being executed by one or more processors, the method comprising: generating a video bitstream including one or more pictures, the video bitstream including a displacement vector signaled to identify a reference block in a current picture; The displacement vector is determining that the one or more pictures should be predicted in normal merge mode or adaptive motion vector prediction (AMVP) mode; obtaining the displacement vector associated with a current block in a current picture; determining motion information associated with the reference block based on the displacement vector, the motion information being used as a motion vector predictor (MVP) from temporal motion vector predictor (TMVP) candidates; generating a TMVP candidate list including the motion information; deriving a motion vector for the current block using the TMVP candidate list; decoding the current block using the derived motion vector for prediction in the normal merge mode or the adaptive motion vector prediction (AMVP) mode; Used for, The TMVP candidate list is sorted based on template matching cost.
Citation Information
Patent Citations
MMVD and Combining SMVD with Motion and Prediction Models
JP2022515088A
Method and device for image encoding and decoding
JP2022526276A
Offset vector identification of temporal motion vector predictor
US20180084260A1
Decoding method and apparatus, encoding method and apparatus, and device
WO2021190591A1
Template matching in video coding
WO2022146833A1