Method, apparatus, and computer program for video coding
By fetching and combining multiple motion vectors from overlapping sub-blocks and pictures, the method enhances the accuracy and efficiency of motion vector prediction in video coding, addressing the limitations of existing SbTMVP and TMVP techniques.
Patent Information
- Application Number
- JP2024518249
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-26
- Filing Date
- 2022-11-03
- Publication Date
- 2025-08-01
AI Technical Summary
Existing video coding techniques, such as SbTMVP and TMVP, are suboptimal as they only consider a single motion vector when multiple related motion vectors are available, leading to errors and limited accuracy due to the use of overlapping sub-blocks and collocated pictures.
The method involves fetching multiple motion vectors from overlapping collocated sub-blocks and pictures to derive a final motion vector predictor, using weighted averages or the motion vector associated with the most overlapping sub-block, enhancing accuracy and efficiency.
This approach improves motion vector prediction accuracy and efficiency by utilizing multiple motion vectors from overlapping sub-blocks and pictures, reducing errors and enhancing coding performance.
Smart Images

Figure 2025524748000001_ABST
Abstract
Description
Technical Field
[0001] This application claims priority from U.S. Provisional Patent Application No. 63 / 391,193, filed on July 21, 2022, and U.S. Patent Application No. 17 / 973,663, filed on October 26, 2022, the disclosures of which are hereby incorporated by reference in their entireties.
[0002] Embodiments of the present disclosure relate to image and video coding techniques. More specifically, embodiments of the present disclosure relate to improvements in motion vector predictor fusion for SbTMVP and TMVP in coding and decoding video data.
Background Art
[0003] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) issued the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standardization bodies jointly formed the JVET (Joint Video Exploration Team) to explore the possibility of developing the next video coding standard beyond HEVC. In October 2017, they jointly called for proposals on video compression with capabilities beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for the 360 video category were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET meeting. As a result of this meeting, JVET officially started the standardization process for the next-generation video coding beyond HEVC, and this new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the VVC video coding standard (version 1).
Summary of the Invention
[0004] According to an embodiment, a method for fusing a plurality of sub-block motion vector predictors into one sub-block motion vector predictor in video coding may be provided. The method includes, for a sub-block of a current block, deriving a first displacement vector for identifying a collocated sub-block of a collocated picture, determining that the collocated sub-block overlaps with one or more sub-blocks related to a motion field grid in the collocated picture, based on the determination that the collocated sub-block overlaps with the one or more sub-blocks, extracting one or more sub-block motion vectors respectively related to the one or more sub-blocks, and based on the extracted one or more sub-block motion vectors, deriving a final motion vector predictor for the sub-block of the current block.
[0005] According to an embodiment, an apparatus for fusing a plurality of sub-block motion vector predictors into one sub-block motion vector predictor in video coding may be provided. The apparatus may include at least one memory configured to store program code, and at least one processor configured to read the program code and operate as instructed by the program code. The program code may include a first derivation code configured to cause the at least one processor to derive a first displacement vector for identifying a collocated sub-block of a collocated picture for a sub-block of a current block, a first determination code configured to cause the at least one processor to determine that the collocated sub-block overlaps with one or more sub-blocks related to a motion field grid in the collocated picture, an extraction code configured to cause the at least one processor to extract one or more sub-block motion vectors respectively related to the one or more sub-blocks based on the determination that the collocated sub-block overlaps with the one or more sub-blocks, and a second derivation code configured to cause the at least one processor to derive a final motion vector predictor for the sub-block of the current block based on the one or more extracted sub-block motion vectors.
[0006] According to an embodiment, a non-transitory computer-readable medium storing instructions may be provided. When the instructions are executed by at least one processor of an apparatus for fusing a plurality of sub-block motion vector predictors into one sub-block motion vector predictor in video coding, the at least one processor causes one or more processors to derive a first displacement vector for identifying a collocated sub-block of a collocated picture for a sub-block of a current block, determine that the collocated sub-block overlaps with one or more sub-blocks related to a motion field grid in the collocated picture, based on the determination that the collocated sub-block overlaps with the one or more sub-blocks, cause the one or more sub-block motion vectors respectively related to the one or more sub-blocks to be retrieved, and based on the retrieved one or more sub-block motion vectors, derive a final motion vector predictor for the sub-block of the current block.
Brief Description of Drawings
[0007]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 1F
Figure 1G
Figure 1H
Figure 1I
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0008] The proposed methods and processes can be used separately or in combination. Embodiments of the present disclosure relate to the fusion of one or more sub-block predictors into a plurality of sub-block predictors in a Sub-block Temporal Motion Predictor (SbTMVP). Embodiments of the present disclosure relate to the fusion of one or more motion vector predictors into a plurality of predictors in a Temporal Motion Vector Predictor (TMVP).
[0009] A displacement vector (also referred to as a "motion shift") is used to identify a block or sub-block within a reference picture in order to derive or determine a motion vector prediction for the current block. In the context of SbTMVP, improved SbTMVP, or TMVP, the displacement vectors used to derive or determine the motion vector of the current block may overlap with a plurality of sub-blocks within the reference picture. However, in SbTMVP, improved SbTMVP, or TMVP, only the motion vector associated with one sub-block can be considered or used as a predictor. Using only one motion vector in this way when multiple related motion vectors are available for determining a predictor is not optimal and often results in errors.
[0010] Also, in order to derive or determine motion prediction, a plurality of collocated pictures and / or reference pictures may be used in SbTMVP, improved SbTMVP, or TMVP, but only one of those plurality of collocated pictures can be used at a time, which limits the improvement in accuracy and efficiency obtained by using a plurality of collocated pictures.
[0011] Accordingly, to solve the above technical problems and deficiencies in the related art, embodiments of the present disclosure relate to using one or more overlapping collocated sub-blocks and / or a plurality of collocated pictures to determine displacement vectors and / or motion predictors in SbTMVP, improved SbTMVP, or TMVP.
[0012] It can be understood that the methods and processes disclosed herein are not limited to SbTMVP, improved SbTMVP, or TMVP, and can also be extended to use in other techniques for motion vector prediction.
[0013] According to an embodiment of the present disclosure, when fetching motion vectors from another sub-block in a collocated picture using displacement vectors to derive a motion predictor for one or more sub-blocks of a current coding block, a plurality of motion vectors can be fetched, and a final motion predictor can be derived based on those plurality of motion vectors.
[0014] According to an aspect of the present disclosure, when fetching motion vectors from another sub-block in a collocated picture specified by a displacement vector to derive a motion predictor for one or more sub-blocks of a current coding block, a plurality of motion vectors can be fetched, and a weighted average of a plurality of prediction blocks related to the fetched plurality of motion vectors can be used to derive a final prediction block for the current sub-block.
[0015] It can be understood that the methods and processes disclosed herein can be extended to a plurality of collocated pictures by expanding overlapping sub-blocks in a motion field from one collocated picture to a plurality of collocated pictures.
[0016] Inter Prediction in VVC
[0017] For each inter-predicted coding unit (CU), the motion parameters can consist of a motion vector, a reference picture index and a reference picture list use index, and additional information required for the new coding mechanism of VVC to be used for inter-predicted sample generation. The motion parameters can be signaled in an explicit or implicit way. When a CU is coded in skip mode, the CU can be associated with one PU, have no significant residual coefficients, and can have no coded motion vector delta or reference picture index. A merge mode can be specified, whereby the motion parameters for the current CU are obtained from neighboring CUs including spatial and temporal candidates and additional schedules introduced in VVC. The merge mode can be applied not only to the skip mode but also to any inter-predicted CU. An alternative to the merge mode is the explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list use flag for each reference picture list, and other necessary information are explicitly signaled for each CU.
[0018] Extended merge prediction
[0019] In VTM4, the merge candidate list is constructed by including the following five types of candidates in order: (1) spatial MVP from spatial neighbor CUs, (2) temporal MVP from collocated CUs, (3) history-based MVP from the FIFO table, (4) pairwise-average MVP, and (5) zero MV. The size of the merge list can be signaled in the slice header, and the maximum allowable size of the merge list can be 6 in VTM4. For each CU code in the merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first bin of the merge index is coded using context, and bypass coding is used for the other bins.
[0020] Spatial candidate derivation
[0021] The derivation of spatial merge candidates in VVC is the same as that in HEVC. Up to four merge candidates are selected from the candidates. FIG. 1A shows exemplary positions of merge candidates B1, A1, B0, A0, and B2 and shows the current block 1100. In some embodiments, the order of derivation can be B1, A1, B0, A0, and then B2. The position B2 can be considered only when any of the CUs at positions A0, B0, B1, A1 are not available (e.g., because it belongs to another slice or tile) or are intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates can be subjected to a redundancy check to ensure that the coding efficiency is improved by excluding candidates having the same motion information from the list. To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs connected by the arrows shown in FIG. 1B are considered, and a candidate can be added to the list only if the corresponding candidates used in the redundancy check do not have the same motion information.
[0022] Temporal candidate derivation
[0023] In some embodiments, when deriving time candidates, only one candidate may be added to the list. In particular, in the derivation of time merge candidates, scaled motion vectors may be derived based on collocated CUs belonging to the collocated reference picture. The reference picture list to be used for the derivation of collocated CUs may be explicitly signaled in the slice header. The scaled motion vectors for time merge candidates can be obtained as shown in FIG. 1C and can be scaled from the motion vectors of the collocated CUs. As shown in FIG. 1C, the scaled motion vectors for time merge candidates can be obtained and scaled from the motion vectors of the collocated CUs based on the picture order count (POC) distances tb and td, where tb is the POC difference between the reference picture of the current picture and the current picture, and td is the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the time merge candidate can be set equal to 0.
[0024] As shown in FIG. 1D, the position of the time candidate is selected between candidates C0 and C1. In some embodiments, position C1 may be used when the CU at position C0 is not available, or is intra-coded, or is outside the current row of the coding tree unit (CTU). Otherwise, position C0 may be used in the derivation of time merge candidates.
[0025] Merge with Motion Vector Difference (MMVD)
[0026] MMVD can be used in either skip mode or merge mode using a motion vector representation method. MMVD can reuse merge candidates in VVC. From among the merge candidates, a candidate is selected and can be further extended by the proposed motion vector representation method as shown in FIGS. 1E and 1F. MMVD can provide a new motion vector representation using simplified signaling. The representation method may include a starting point, a magnitude of motion, and a direction of motion.
[0027] The MMVD technique can use the merge candidate list in VVC. However, only candidates that are of the default merge type (MRG_TYPE_DEFAULT_N) can be considered for the extension of MMVD. The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list as shown in Table 1 here.
Table 1
[0028] When the number of base candidates is equal to 1, the base candidate IDX need not be signaled. The distance index is the motion magnitude information. The distance index indicates a predetermined distance from the starting point information. The predetermined distance can be as shown in Table 2 here.
Table 2
[0029] The direction index can represent the direction of MMVD with respect to the starting point. The direction index can represent four directions as shown in Table 3 here.
Table 3
[0030] In some embodiments, the MMVD flag may be signaled immediately after sending the skip flag and the merge flag. When the skip and merge flags are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is in the AFFINE mode, but if it is not 1, the skip / merge index is parsed for the skip / merge mode of VTM.
[0031] Template matching based candidate reordering for MMVD and affine MMVD
[0032] In the related art, the MMVD offset can be extended for the MMVD mode and the affine MMVD mode. As shown in FIG. 1G, additional refinement positions can be added along the diagonal angle of k×π / 8, and the number of directions can be increased from 4 to 16. Further, based on the sum of absolute differences (SAD) cost between the template (one row above and one column to the left of the current block) and its reference at each refinement position, all possible MMVD refinement positions (16×6) for each base candidate can be reordered. In some embodiments, the top 1 / 8 of the refinement positions with the minimum template SAD cost are retained as available positions for MMVD index coding. The MMVD index can be binarized by a Rice code with a parameter equal to 2.
[0033] In one aspect of the present disclosure, in addition to the MMVD extension described herein, the affine MMVD reordering can also be extended, and additional refinement positions along the diagonal angle of k×π / 4 can be added. After reordering, the top 1 / 2 of the refinement positions with the minimum template SAD cost can be retained.
[0034] Sub-block based TMVP (SbTMVP)
[0035] To improve coding efficiency and reduce the transmission overhead of motion vectors, sub-block level motion vector refinement can be applied to extend the CU level temporal motion vector prediction (TMVP). Sub-block based TMVP (SbTMVP) enables inheriting motion information at the sub-block level from the collocated reference picture. Each sub-block of a large-sized CU can have its own motion information without explicitly transmitting the block partition structure or motion information. SbTMVP can obtain the motion information of each sub-block as follows. First, SbTMVP can include deriving the displacement vector (DV) of the current CU. Then, based on the availability of SbTMVP candidates, the central motion is derived. Finally, SbTMVP can include deriving the sub-block motion information from the corresponding sub-block by the DV. Different from the TMVP candidate derivation that always derives the temporal motion vector from the collocated block within the reference frame, SbTMVP can apply the DV derived from the motion vector (MV) of the left adjacent CU of the current CU to find the corresponding sub-block within the collocated picture for each sub-block of the current CU. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set to the central motion.
[0036] VVC supports the sub-block based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the collocated picture to improve the motion vector prediction and merge mode for the CUs in the current picture. The same collocated picture used by TMVP is used for SbTMVP. SbTMVP is mainly different from TMVP in the following two aspects.
[0037] (1) TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. (2) TMVP fetches the temporal motion vector from the collocated blocks within the collocated picture (the collocated blocks are the blocks at the lower right or center with respect to the current CU), while SbTMVP applies a motion shift before fetching the temporal motion information from the collocated picture, and the motion shift (also referred to as the displacement vector or DV) is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0038] Figure 1H shows an exemplary SbTMVP candidate selection using spatial neighboring blocks. SbTMVP predicts the motion vectors of the sub-CUs within the current CU in two parts. As the first part, the spatial neighbor A1 in Figure 1H is examined. If A1 has a motion vector using the collocated picture as its reference picture, this motion vector is selected to be the applied motion shift (or displacement vector). If such a motion is not identified, the motion shift is set to (0, 0).
[0039] As the second part, as shown in Figure 1I, by applying the motion shift identified in the first part (i.e., adding it to the coordinates of the current block), the sub-CU level motion information (motion vector and reference index) can be obtained from the collocated picture. As shown in Figure 1I, assume that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information for that sub-CU is derived using the motion information of its corresponding block (the smallest motion grid covering the central sample) within the collocated picture. After the motion information of the collocated sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a similar manner as the TMVP process in HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.
[0040] In VVC, a combined sub-block based merge list that includes both SbTMVP candidates and affine merge candidates is used for signaling the sub-block based merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. The size of the sub-block based merge list is signaled within the SPS, and the maximum allowable size of the sub-block based merge list is 5 in VVC.
[0041] In VVC, the sub-CU size used for SbTMVP is fixed at 8×8, and like the affine merge mode, the SbTMVP mode is applicable only to CUs where both the width and height are 8 or more. The sub-block size may be configurable to other sizes such as 4×4 in the use of the ECM software model in post-VVC exploration. To provide temporal motion information for SbTMVP and TMVP in the AMVP mode, two collocated frames are proposed and utilized.
[0042] Sub-block based TMVP using motion vector offset
[0043] To obtain the best match, previously an additional motion vector offset (MVO) was added to the displacement vector (DV). By using MVO(xo,yo), the position of the MV field within the collocated CU can be adjusted. When MVO(xo,yo) is not a zero motion offset, DV’, which is the sum of DV and MVO, is used as the displacement vector indicating the position of the collocated CU for deriving the sub-block level temporal motion vector prediction (SbTMVP). Also, a template matching (TM) method has been proposed, in which the displacement vector (DV) of SbTMVP is used as the motion vector for the template matching process.
[0044] As described above, in the context of SbTMVP, improved SbTMVP, or TMVP, the displacement vectors used to derive or determine the motion vector of the current block may overlap with multiple sub-blocks in the reference picture. However, in SbTMVP, improved SbTMVP, or TMVP, only the motion vector associated with one sub-block can be considered or used as a predictor. Using only one motion vector in this way when multiple related motion vectors are available to determine the predictor is not optimal and often results in errors.
[0045] Also, to derive or determine motion prediction, multiple collocated pictures and / or reference pictures may be used in SbTMVP, improved SbTMVP, or TMVP, but only one of those multiple collocated pictures can be used at a time, which limits the improvement in accuracy and efficiency obtained by using multiple collocated pictures.
[0046] Accordingly, to solve the above technical problems and deficiencies in the related art, embodiments of the present disclosure relate to using one or more overlapping collocated sub-blocks and / or multiple collocated pictures to determine displacement vectors and / or motion predictors in SbTMVP, improved SbTMVP, or TMVP.
[0047] According to an embodiment of the present disclosure, when fetching motion vectors from another sub-block in a collocated picture using displacement vectors to derive a motion vector predictor for one or more sub-blocks of the current coding block, multiple motion vectors can be fetched, and a final motion vector predictor can be derived based on those multiple motion vectors.
[0048] In one embodiment, for one sub-block of the current prediction block (i.e., sub-block T), the DV can be derived first, and it can be used to identify another sub-block in the collocated picture (i.e., sub-block T’). If the identified sub-block T’ overlaps with the motion field grid (e.g., 4×4 or 8×8) in the collocated picture, that is, if the identified sub-block overlaps with a plurality of sub-blocks storing different motion information in the collocated picture, a plurality of motion vectors associated with those overlapping sub-blocks can be fetched and used to derive the final motion vector predictor for the current sub-block of the current prediction block. The motion vector field can be a set of motion vectors associated with sub-blocks in the picture.
[0049] As an example, referring to FIG. 2, the identified sub-block T’ in the collocated picture 215 (which can be a collocated block in some embodiments) overlaps with four sub-blocks A, B, C, and D (overlapping sub-blocks 225) in the collocated picture 215 aligned with the motion vector field grid 220. As another example, referring to FIG. 3, the identified sub-block T’ overlaps with nine sub-blocks A, B, C, D, E, F, G, H, and I (overlapping sub-blocks 325) in the collocated picture 315 (which can be referred to as a collocated block in some embodiments) aligned with the motion vector field grid 320.
[0050] In one embodiment of the present disclosure, the motion vector associated with the sub-block that most overlaps with the area of sub-block T’ can be used as the final MVP for sub-block T in SbTMVP. As an example, referring to FIG. 2, if T’ most overlaps with sub-block D, the motion information associated with sub-block D can be used as the MVP for sub-block T in SbTMVP.
[0051] In one embodiment of the present disclosure, the weighted average of the motion vectors associated with a plurality of sub-blocks overlapping with sub-block T' can be used as the final MVP of sub-block T in SbTMVP. As an example, referring to FIG. 2, T' overlaps with A, B, C, and D in the collocated picture aligned with the motion vector field grid, and the weighted average of the motion vectors associated with A, B, C, and D can be derived as the MVP of sub-block T in SbTMVP. In some embodiments, the weighting in the weighted average may depend on how much area overlaps between the associated sub-block in the collocated picture and sub-block T'. In some embodiments, the weighting in the weighted average may depend on the distance between the center pixel position of T' and the center position of the generated motion vector predictor A, B, C, or D. In some embodiments, when the overlapping sub-blocks in the collocated picture have MVs pointing outside the picture boundary, those MVs can be excluded in the weighted average process.
[0052] In one embodiment of the present disclosure, the motion vector associated with the sub-block overlapping the center point position of sub-block T' can be used as the final MVP of sub-block T in SbTMVP. As an example, referring to FIG. 2, the center point of T' overlaps with sub-block D, and the motion information associated with sub-block D can be used as the MVP of sub-block T in SbTMVP.
[0053] According to one aspect of the present disclosure, when fetching motion vectors from another sub-block in the collocated picture specified by the displacement vector to derive a motion vector predictor for one or more sub-blocks of the current coding block, a plurality of motion vectors can be fetched, and the weighted average of the plurality of prediction blocks associated with the fetched plurality of motion vectors can be used to derive the final prediction block for the current sub-block.
[0054] In one embodiment, for one sub-block of the current prediction block (i.e., sub-block T), first, a DV can be derived, and using it, another sub-block in the collocated picture (i.e., sub-block T’) can be identified. If the identified sub-block T’ overlaps with a motion field grid (e.g., 4×4 or 8×8) in the collocated picture, that is, if the identified sub-block overlaps with a plurality of sub-blocks storing separate motion information in the collocated picture, a plurality of prediction blocks generated by the motion vectors associated with those overlapping sub-blocks are generated, and a final prediction block for the current sub-block is generated as the weighted average of those plurality of prediction blocks. As an example, referring to FIG. 2, if the identified sub-block T’ overlaps with four sub-blocks A, B, C, and D (overlapping sub-blocks 225) in the collocated picture 215 aligned with the motion vector field grid 220, four prediction blocks can be generated using the motion vectors associated with A, B, C, and D, and a final prediction block for sub-block T is generated as the weighted average of these four prediction blocks. In some embodiments, the weighting depends on how much area can overlap between the relevant sub-block in the collocated picture and sub-block T’. In some embodiments, each pixel in sub-block T’ has a weighted average corresponding to itself of the motion vector predictors generated from A, B, C, and D. In some embodiments, the corresponding weighted average can be derived from the distance between its pixel position and the center position of the generated motion vector predictor A, B, C, or D.
[0055] In one embodiment, when the overlapping sub-blocks in the collocated picture have an MV pointing outside the picture boundary, the prediction from that MV can be excluded in the weighted average process.
[0056] It can be understood that the methods and processes disclosed herein can be extended to a plurality of collocated pictures by expanding overlapping sub-blocks in the motion field from one collocated picture to a plurality of collocated pictures.
[0057] FIG. 4 illustrates a simplified block diagram of a communication system 400 according to an embodiment of the present disclosure. The communication system 400 may include at least two terminals 410-420 interconnected via a network 450. In one-way data transmission, the first terminal 410 may code video data at a local location for transmission to the second terminal 420 via the network 450. The second terminal 420 may receive the coded video data of the other terminal from the network 450, decode the coded data, and display the restored video data. One-way data transmission may be common in media service providing applications and the like.
[0058] FIG. 4 illustrates a second pair of terminals 430, 440 provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. In bidirectional data transmission, each terminal 430, 440 may code video data captured at a local location for transmission to the other terminal via the network 450. Each terminal 430, 440 may also receive the coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.
[0059] In FIG. 4, terminals 410-440 can be exemplified as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited as such. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 450 represents any number of networks that transmit coded video data between terminals 410-440, including, for example, a wired communication network and / or a wireless communication network. Communication network 450 can exchange data over a circuit-switched channel and / or a packet-switched channel. Representative networks include long-distance communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network 450 may not be considered important for the operation of the present disclosure, unless otherwise described below.
[0060] FIG. 5 illustrates the arrangement of video encoders and decoders in a streaming environment, such as a streaming system 500, as an example of an application related to the matters disclosed. The matters disclosed can be equally applicable to other uses where video can be used, including, for example, videoconferencing, digital TV, and the storage of compressed video on digital media such as CDs, DVDs, memory sticks, and the like.
[0061] The streaming system can include a capture subsystem 513, which can include a video source 501 such as, for example, a digital camera that creates an uncompressed video sample stream 502. The sample stream 502 is drawn as a thick line to emphasize that it has a high data volume compared to the encoded video bitstream, and can be processed by an encoder 503 coupled to the video source 501, which can be, for example, a camera. The encoder 503 can include hardware, software, or a combination thereof to enable or implement aspects of the disclosure described in more detail later. The encoded video bitstream 504 is drawn as a thin line to emphasize that it has a low data volume compared to the sample stream, and can be stored in a streaming server 505 for later use. One or more streaming clients 506, 508 can access the streaming server 505 to retrieve video bitstreams 507, 509, which can be, for example, copies of the encoded video bitstream 504. The client 506 can include a video decoder 510 that decodes an incoming copy 507 of the encoded video bitstream and creates an outgoing video sample stream 511, and the outgoing video sample stream 511 can be rendered on a display 512 or other rendering device not shown. In some streaming systems, the video bitstreams 504, 507, 509 can be encoded according to a particular video coding / compression standard. Examples of those standards include ITU-T Recommendation H.265. A video coding standard, informally known as VVC (Versatile Video Coding), is under development. The subject matter of the disclosure may be used in the context of VVC.
[0062] FIG. 6 can be a functional block diagram of a video decoder 510 according to an embodiment of the present disclosure.
[0063] The receiver 610 can receive one or more codec video sequences decoded by the decoder 510 and, in the same or other embodiments, can receive one coded video sequence at a time, and the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence can be received from a channel 612 that can be a hardware / software link to a storage device storing the encoded video data. The receiver 610 may receive the encoded video data together with other data such as, for example, coded audio data and / or auxiliary data streams, and those data can be transferred to their respective using entities (not shown). The receiver 610 can separate the coded video sequence from other data. To counter network jitter, a buffer 615, which can be, for example, a buffer memory, can be coupled between the receiver 610 and the entropy decoder / parser 620 hereinafter referred to as the "parser". When the receiver 610 is receiving data from a storage / transfer device with sufficient bandwidth and controllability or from a synchronous network, the buffer 615 may not be needed or can be made smaller. For use on a best-effort packet network such as the Internet, the buffer 615 may be needed, be made relatively large, and advantageously be of an adaptable size.
[0064] Video decoder 510 may include a parser 620 for reconstructing symbol 621 from an entropy-coded video sequence. The categories of those symbols include information used to manage the operation of decoder 510 and may potentially include information for controlling a rendering device such as display 512. A rendering device such as display 512 is not an integral part of the decoder but can be coupled to the decoder as shown in FIG. 6. Control information for the (one or more) rendering devices may be in the form of an SEI (Supplementary Enhancement Information) message or a VUI (Video Usability Information) parameter set fragment (not shown). Parser 620 may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence can be according to a video coding technology or standard and can follow principles well-known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context dependence, etc. Parser 620 can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to a group. Subgroups can include group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter (QP) values, motion vectors, etc. from the coded video sequence information.
[0065] Parser 620 may perform entropy decoding / analysis processing on the video sequence received from buffer 615 to produce symbol 621. The parser 620 may receive the encoded data and selectively decode a specific symbol 621. Further, the parser 620 may determine whether a specific symbol 621 should be provided to the motion compensation prediction unit 653, the scaler / inverse transform unit 651, the intra prediction unit 652, or the loop filter unit 656.
[0066] For the reconstruction of symbol 621, multiple different units may be involved depending on, for example, the type of coded video picture or its part such as inter-picture and intra-picture, inter-block and intra-block, and other factors. How each unit is involved can be controlled by subgroup control information parsed by parser 620 from the coded video sequence. Such a flow of subgroup control information between parser 620 and the following multiple units is not illustrated for clarity.
[0067] Beyond the above-described functional blocks, decoder 510 can conceptually be subdivided into a number of functional units as described later. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for explaining the matters related to the disclosure, a conceptual subdivision into the following functional units is appropriate.
[0068] The first unit is the scaler / inverse transform unit 651. The scaler / inverse transform unit 651 receives the quantized transform coefficients along with control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as (one or more) symbols 621 from the parser 620. This can output a block with sample values that can be input to aggregator 655.
[0069] In some cases, the output samples of the scaler / inverse transform 651 may relate to blocks that are intra-coded, i.e., blocks that do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit 652. In some cases, the intra-picture prediction unit 652 generates blocks of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current partially reconstructed picture 658. The aggregator 655, in some cases, adds, for each sample, the prediction information generated by the intra prediction unit 652 to the output sample information provided by the scaler / inverse transform unit 651.
[0070] In other cases, the output samples of the scaler / inverse transform unit 651 may relate to inter-coded blocks that may be motion compensated. In such cases, the motion compensation prediction unit 653 can access the reference picture memory 657 to fetch the samples used for prediction. After motion compensating the fetched samples according to the symbols 621 related to the block, these samples can be added by the aggregator 655 to the output of the scaler / inverse transform unit (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference picture memory from which the motion compensation unit fetches the prediction samples can be controlled, for example, by the motion vectors available to the motion compensation unit in the form of symbols 621 that may have X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory when exact sub-sample motion vectors are used, and motion vector prediction mechanisms, etc.
[0071] The output samples of the aggregator 655 can be subjected to various loop filtering techniques in the loop filter unit 656. The video compression technique can include in-loop filter techniques, which are controlled by parameters made available to the loop filter unit 656 as symbols 621 from the parser 620 included in the coded video bitstream, but can also respond to meta information obtained during the decoding of a previous part in decoding order of the coded picture or coded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.
[0072] The output of the loop filter unit 656 can be made into a sample stream that can be output to, for example, the display 512 which can be a rendering device, and this can also be stored in the reference picture memory for use in future inter-picture prediction.
[0073] A certain coded picture, when fully reconstructed, can be used as a reference picture for future prediction. When a coded picture is fully reconstructed and that coded picture is identified as a reference picture, for example, by the parser 620, the current picture 658 can become part of the reference picture memory 657 which can be a reference picture buffer, for example, and a new current picture memory can be reallocated before starting the reconstruction of the next coded picture.
[0074] The video decoder 510 can perform decoding processing according to a predetermined video compression technique that can be documented by a standard such as ITU-T Recommendation H.265. The coded video sequence can conform to the syntax defined by the video compression technique or standard being used, in the sense of faithfully following the syntax of the video compression technique or standard as defined in the video compression technique document or standard, particularly the profile document therein. Also, for compliance, it is also necessary that the complexity of the coded video sequence be within the limits determined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The restrictions set by the level can, in some cases, be further constrained through the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled within the coded video sequence.
[0075] In one embodiment, the receiver 610 can receive additional redundant data along with the encoded video. The additional data can be included as part of the (one or more) coded video sequences. The additional data can be used by the video decoder 510 for properly decoding the data and / or for more accurately reconstructing the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0076] FIG. 7 can be a functional block diagram of a video encoder 503 according to an embodiment of the present disclosure.
[0077] The encoder 503 can receive video samples from a video source 501 (not part of the encoder) that can capture the video images to be coded by the encoder 503.
[0078] The video source 501 may provide a source video sequence to be coded by the encoder 503 in the form of a digital video sample stream having an arbitrary suitable bit depth, such as 8 bits, 10 bits, 12 bits, …, an arbitrary color space, such as BT.601 Y CrCB, RGB, …, and an arbitrary suitable sampling structure, such as Y CrCb 4:2:0, Y CrCb 4:4:4. In a media service providing system, the video source 501 may be a storage device storing pre-prepared videos. In a video conferencing system, the video source 501 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. Those pictures themselves can be organized as a spatial array of pixels, and each pixel can have one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can immediately understand the relationship between pixels and samples. The following description focuses on samples.
[0079] According to one embodiment, the encoder 503 can code and compress the pictures of the source video sequence into a coded video sequence 743 in real time or under other time constraints required by the application. Enforcing an appropriate coding speed is one function of the controller 750. The controller controls other functional units as described later and is functionally coupled to those units. That coupling is not shown for clarity. The parameters set by the controller may include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can immediately identify other functions of the controller 750 as being related to a video encoder 503 optimized for a specific system design.
[0080] Some video encoders operate in a manner that those skilled in the art would immediately recognize as a "coding loop". As an overly simplified explanation, the coding loop can be composed of, for example, the encoding part of the source coder 730 which can be an encoder (responsible for creating symbols based on the input picture to be coded and the (one or more) reference pictures), and the local decoder 733 embedded in the encoder 503. The local decoder 733, in the video compression technology being considered in the matters related to the disclosure, since any compression between the symbol and the coded video bitstream is reversible, reconstructs the symbol and generates sample data that can also be created by the remote decoder. The reconstructed sample stream is input into the reference picture memory 734. Since the decoding of the symbol stream yields a bit-exact result independent of the decoder position (local or remote), the content of the reference picture buffer is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift in case synchronization cannot be maintained, for example due to channel errors) is well known to those skilled in the art.
[0081] The operation of the local decoder 733 can be considered the same as that of the remote decoder 510, which has already been described in detail above in relation to Figure 6. However, also briefly referring to Figure 7, since the symbol is available and the encoding / decoding of the symbol into the coded video sequence by the entropy coder 745 and the parser 620 can be considered reversible, the entropy decoding part of the decoder 510 including the channel 612, the receiver 610, the buffer 615, and the parser 620 does not have to be fully implemented in the local decoder 733.
[0082] What can be noticed at this point is that any decoder technology, except for the parsing / entropy decoding existing in the decoder, must necessarily exist in the corresponding encoder in a substantially the same functional form. Since the description of the encoder technology is the reverse of the thoroughly described decoder technology, it can be omitted. Only in specific fields is a more detailed description required and is provided below.
[0083] As part of its operation, source coder 730 can perform motion compensation predictive coding that predictively codes an input frame with respect to one or more previously coded frames designated as "reference frames" from the video sequence. Thus, coding engine 732 codes the difference between a pixel block of the input frame and a pixel block of the (one or more) reference frames that can be selected as the (one or more) prediction references for the input frame.
[0084] Local video decoder 733 can decode the coded video data of a frame that can be designated as a reference frame based on the symbols created by source coder 730. The operation of coding engine 732 can advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 7), the reconstructed video sequence can typically be a replica of the source video sequence with some error. Local video decoder 733 can replicate the decoding process that can be performed by the video decoder on the reference frame and cause the reconstructed reference frame to be stored in reference picture memory 734 that can be, for example, a reference picture cache. Thus, encoder 503 can locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame that would be obtained by the far-end video decoder in the absence of transmission errors.
[0085] Predictor 735 may perform prediction search for the coding engine 732. That is, for a new frame to be coded, predictor 735 may search the reference picture memory 734 for sample data as candidate reference pixel blocks that may serve as appropriate prediction references for the new picture, or for specific metadata such as reference picture motion vectors and block shapes. Predictor 735 may operate on a per pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have prediction references drawn from a plurality of reference pictures stored in the reference picture memory 734 as determined by the search results obtained by predictor 735.
[0086] Controller 750 may manage the coding process of source coder 730, including, for example, setting parameters and subgroup parameters used to code video data.
[0087] The outputs of all the foregoing functional units may be subjected to entropy coding in entropy coder 745. The entropy coder converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as, for example, Huffman coding, variable length coding, arithmetic coding, and the like.
[0088] Transmitter 740 may buffer the coded video sequence generated by entropy coder 745 and prepare it for transmission via communication channel 760. Communication channel 760 may be a hardware / software link to a storage device that stores the encoded video data. Transmitter 740 may merge the coded video data from source coder 730 with other data to be transmitted, such as, for example, coded audio data and / or auxiliary data streams (sources not shown).
[0089] Controller 750 can manage the operation of encoder 503. In encoding, controller 750 can assign to each encoded picture a specific encoded picture type that can affect the encoding technique applicable to that picture. For example, a picture can often be assigned as one of the following frame types.
[0090] An intra picture (I picture) can be coded and decoded without using other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example including independent decoder refresh pictures. Those skilled in the art are aware of those variants of I pictures, as well as their respective uses and characteristics.
[0091] A predicted picture (P picture) can be coded and decoded using intra prediction or inter prediction, using at most one motion vector and a reference index to predict the sample values of each block.
[0092] A bi - directionally predicted picture (B picture) can be coded and decoded using intra prediction or inter prediction, using at most two motion vectors and a reference index to predict the sample values of each block. Similarly, a multi - predicted picture can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0093] The source picture is generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively with reference to other (already coded) blocks determined by the coding assignment applied to each of those blocks in their respective pictures. For example, blocks of an I picture can be coded non-predictively, or they can be coded predictively with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded non-predictively or via spatial or temporal prediction with reference to a reference picture coded one ahead. Blocks of a B picture can be coded non-predictively or via spatial or temporal prediction with reference to one or two reference pictures coded ahead.
[0094] The encoder 503 can perform coding processing according to a predetermined video coding technology or standard such as ITU-T Recommendation H.265. In its operation, the encoder 503 can perform various compression processes including predictive coding processes that utilize temporal and spatial redundancies in the input video sequence. The coded video data can thus conform to the syntax defined by the video coding technology or standard being used.
[0095] In one embodiment, the transmitter 740 can transmit additional data along with the encoded video. The source coder 730 can include such data as part of the coded video sequence. The additional data can have temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI (Supplementary Enhancement Information) messages, VUI (Visual Usability Information) parameter set fragments, and the like.
[0096] Figure 8 shows an exemplary process 800 for fusing a plurality of sub-block motion vector predictors into one sub-block motion vector predictor in video coding and decoding according to an embodiment of the present disclosure.
[0097] In operation 805, for a sub-block of a current block, a first displacement vector for identifying a collocated sub-block of a collocated block in a collocated picture may be derived. As an example, as shown in FIGS. 2 and 3, a sub-block T' that is a collocated sub-block of a collocated picture (215 or 315) may be identified based on a first displacement vector of a current block (210 or 310).
[0098] In operation 810, it may be determined whether one or more sub-blocks related to a motion field grid in a collocated picture overlap with the collocated sub-block. As an example, it may be determined whether sub-block T' overlaps with any of the collocated sub-blocks in a collocated picture (215 or 315).
[0099] In operation 815, based on determining that the collocated sub-block overlaps with one or more sub-blocks, one or more sub-block motion vectors respectively related to the one or more overlapping sub-blocks may be retrieved. As an example, sub-block motion vectors related to sub-blocks A, B, C, and D (overlapping sub-block 225 in FIG. 2) or sub-blocks A, B, C, D, E, F, G, H, and I (overlapping sub-block 325 in FIG. 3) may be determined or retrieved.
[0100] At operation 820, a final motion vector predictor for sub - blocks of the current block may be derived based on the one or more sub - block motion vectors. As an example, the final motion vector for sub - block T may be derived based on the sub - block motion vectors associated with sub - blocks A, B, C, and D (overlapping sub - blocks 225 in FIG. 2) or sub - blocks A, B, C, D, E, F, G, H, and I (overlapping sub - blocks 325 in FIG. 3).
[0101] In some embodiments, deriving a final motion vector predictor for sub - blocks of the current block may include determining a first motion vector associated with a first sub - block that has the largest overlap with a collocated sub - block among one or more sub - blocks associated with a motion field grid in a collocated picture.
[0102] In some embodiments, the step of deriving a final motion vector predictor for sub - blocks of the current block may include determining a weighted average of the one or more retrieved sub - block motion vectors. The weighted average may be based on the area of overlap between the collocated sub - block and each of the one or more sub - blocks associated with a motion field grid in a collocated picture. The weighted average may also be based on the distance between the position of the center pixel of the collocated sub - block and the position of the center pixel of each of the one or more sub - blocks associated with a motion field grid in a collocated block. In some embodiments, the weighted average may be based on the direction associated with the one or more sub - block motion vectors. As an example, based on determining that the direction associated with at least one of the one or more sub - block motion vectors points outside the picture boundary, the at least one of the one or more sub - block motion vectors may be excluded from being used in the weighted average calculation.
[0103] In some embodiments, deriving the final motion vector predictor for a sub-block of the current block may include determining a first motion vector associated with a first sub-block that overlaps the position of the central pixel of the collocated sub-block among one or more sub-blocks related to the motion field grid in the collocated picture.
[0104] FIG. 8 shows a block example of process 800, but in some implementations, process 800 may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in FIG. 8. Additionally, or alternatively, two or more of the blocks of process 800 may be executed in parallel.
[0105] Also, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium to execute one or more of the proposed methods.
[0106] The above-described techniques can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 9 shows a computer system 900 suitable for implementing a particular embodiment of the matters disclosed.
[0107] The computer software can be coded using any suitable machine code or computer language such that, when subjected to assembly, compilation, linking, or similar mechanisms, it can produce code having instructions that can be executed directly or via interpretation, microcode execution, and the like, by a computer central processing unit (CPU), a graphics processing unit (GPU), and the like.
[0108] The command can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0109] The components shown in FIG. 9 with respect to computer system 900 are essentially exemplary and are not intended to imply any limitation as to the use or functionality of the computer software implementing the embodiments of the present disclosure. Also, the component configuration should not be construed as having any dependency or requirement regarding any one or combination of the components shown in this exemplary embodiment of computer system 900.
[0110] Computer system 900 may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, moving a data glove, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures, etc.), olfactory input (not shown). The human interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as, for example, audio (e.g., conversation, music, ambient sound, etc.), images (e.g., scanned images, photographic images obtained from a still camera, etc.), video (e.g., 2D video, 3D video including stereoscopic video, etc.).
[0111] The input human interface device may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a screen 910 that can be a touch screen, a data glove, a joystick 905, a microphone 906, a scanner 907, a camera 908 (only one of each is shown).
[0112] Computer system 900 may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by screen 910, data glove, or joystick 905, although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speaker 909, headphones (not shown), etc.), visual output devices (e.g., screen 910 including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light emitting diode (OLED) screen (each may or may not have a touch screen input function. Each may or may not have a tactile feedback function. Some of these can output two-dimensional visual output or output higher than four-dimensional output through means such as stereoscopic output, etc.), virtual reality glasses (not shown), holographic display, and smoke tank (not shown), etc.), and a printer (not shown).
[0113] Computer system 900 may also include human-accessible storage devices and their associated media, such as optical media including a CD / DVD ROM / RW 920 having a CD / DVD or similar medium 921, a thumb drive 922, a removable hard drive or solid state drive 923, legacy magnetic media such as tapes and floppy disks (registered trademark, not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0114] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the matters disclosed herein does not include transmission media, carrier waves, or other transient signals.
[0115] The computer system (900) may also include an interface to one or more communication networks (955). The network (955) can be, for example, wireless, wired, or optical. The network (955) can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of the network (955) include local area networks such as Ethernet (registered trademark), wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE and the like, TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANBus, and the like. A particular network (955) generally requires an external network interface adapter (954) attached to a particular general-purpose data port or peripheral bus (949) (e.g., a USB port of the computer system (900)), while others are generally integrated into the core of the computer system (900) by attachment to the system bus 1248 described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using any of these networks (955), the computer system (900) can communicate with other entities. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANbus to a particular CANbus device), or bidirectional, for example, to other computer systems using a local or wide area digital network. Particular protocols and protocol stacks can be used on each of the network (955) and network interfaces such as the external network interface adapter (954) as described above.
[0116] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core 940 of the computer system 900.
[0117] The core 940 may include one or more central processing units (CPUs) 941, a graphics processing unit (GPU) 942, a special programmable processing unit in the form of a field programmable gate array (FPGA) 943, a hardware accelerator 944 for specific tasks, and the like. These devices can be connected via a system bus 1248 together with a read-only memory (ROM) 945, a random access memory (RAM) 946, and an internal mass storage 947 such as an internal hard drive, a solid state drive (SSD), and the like that may not be accessible to internal users. In some computer systems, the system bus 1248 can be made accessible in the form of one or more physical plugs to allow for expansion by additional CPUs, GPUs, and the like. Peripheral devices may be attached either directly to the system bus 1248 of the core or via a peripheral bus 949. The architecture of the peripheral bus includes a peripheral component interconnect (PCI), USB, and the like.
[0118] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute specific instructions that can be combined to form the aforementioned computer code. The computer code can be stored in the ROM 945 or RAM 946. Transient data can also be stored in the RAM 946, and permanent data can be stored, for example, in the internal mass storage 947. Fast storage and retrieval to any of the memory devices can be enabled by the use of cache memory that may be associated near one or more of the CPUs 941, GPUs 942, mass storage 947, ROM 945, RAM 946, and the like.
[0119] A computer-readable medium can have computer code thereon for performing various computer-implemented processes. The medium and the computer code may be specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well-known and available to those having skill in the computer software arts.
[0120] As an example, and not by way of limitation, a computer system having the architecture of computer system 900, particularly core 940, can provide functionality as a result of software embodied on one or more tangible computer-readable media being executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, and the like). Such computer-readable media can be specific storage of core 940 that is non-transitory in nature, such as mass storage 947 or ROM 945 inside the core, and media associated with user-accessible mass storage as introduced above. The software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 940. The computer-readable media can include one or more memory devices or chips, depending on specific needs. The software can cause core 940 and particularly the processors (including CPUs, GPUs, FPGAs, and the like) therein to execute the specific processes described herein or specific portions of the specific processes, including defining data structures stored in RAM 946 and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of logic wired or otherwise embodied in a circuit (e.g., accelerator 944) that operates instead of or in conjunction with software to execute the specific processes described herein or specific portions of the specific processes. References to software include logic, and vice versa where appropriate. References to computer-readable media can include circuits (e.g., integrated circuits (ICs), etc.) storing software for execution, circuits embodying logic for execution, or both where appropriate. The present disclosure includes suitable combinations of hardware and software.
[0121] Although this disclosure describes several exemplary embodiments, there are changes, substitutions, and various equivalent alternatives that fall within the scope of the disclosure. Accordingly, it is understood that those skilled in the art, although not explicitly illustrated or described herein, can embody the principles of the disclosure and, therefore, devise numerous systems and methods that are within its spirit and scope.
Claims
Claim 1 A method for fusing a plurality of sub-block motion vector predictors into one sub-block motion vector predictor in video coding, the method being executed by one or more processors, the method comprising: deriving a first displacement vector for identifying a collocated sub-block of a collocated block for a sub-block of a current block; determining that the collocated sub-block overlaps with two or more sub-blocks related to a motion field grid in the collocated block; extracting two or more sub-block motion vectors respectively related to the two or more sub-blocks based on the determination that the collocated sub-block overlaps with the two or more sub-blocks; deriving a final motion vector predictor for the sub-block of the current block based on the two or more extracted sub-block motion vectors; A method having the above. Claim 2 The step of deriving the final motion vector predictor for the sub-block of the current block includes determining a first motion vector related to a first sub-block having the largest overlap with the collocated sub-block among the two or more sub-blocks related to the motion field grid in the collocated block. The method according to Claim 1. Claim 3 The step of deriving the final motion vector predictor for the sub-block of the current block includes determining a weighted average of the two or more extracted sub-block motion vectors. The method according to Claim 1. Claim 4 The weighted average is based on the overlapping area between the collocated sub-block and each of the two or more sub-blocks related to the motion field grid in the collocated block. The method according to Claim 3. Claim 5 The weighted average is based on the distance between the position of the central pixel of the collocated sub-block and the position of the central pixel of each of the two or more sub-blocks related to the motion field grid in the collocated block. The method according to Claim 3. Claim 6 The method according to claim 3, wherein the weighted average is based on a direction related to the two or more sub-block motion vectors. **Claim 7** The method according to claim 6, wherein at least one of the two or more sub-block motion vectors is excluded from the weighted average based on a determination that a direction related to at least one of the two or more sub-block motion vectors points outside a picture boundary. **Claim 8** The step of deriving the final motion vector predictor for the sub-blocks of the current block comprises determining a first motion vector related to a first sub-block that overlaps a position of a center pixel of the collocated sub-block among the two or more sub-blocks related to the motion field grid in the collocated block. The method according to claim 1. **Claim 9** One or more processors; One or more memories storing a computer program; comprising The computer program causes the one or more processors to execute the method according to any one of claims 1 to 8. An apparatus. **Claim 10** A computer program for causing a computer to execute the method according to any one of claims 1 to 8.