Method and apparatus for temporal interpolation prediction in a video bitstream

The method improves video encoding by refining motion vectors through decoder-side motion vector refinement based on bilateral matching, addressing the challenge of accurate motion field prediction in existing technologies like AV1.

JP2025516414APending Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024513451
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-07
Filing Date
2022-11-09
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing video encoding technologies, such as AV1, face challenges in accurately predicting motion fields, particularly through temporal interpolation prediction, which affects compression quality and efficiency.

Method used

A method for decoding an encoded video bitstream that involves generating a motion field for blocks using motion vectors indicating reference pictures, and then refining these motion vectors through a decoder-side motion vector refinement process based on bilateral matching to improve prediction accuracy.

Benefits of technology

The proposed solution enhances the accuracy of motion field prediction, leading to improved video encoding efficiency and compression quality, while also reducing computational complexity and memory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516414000001_ABST
    Figure 2025516414000001_ABST
Patent Text Reader

Abstract

A video decoder is provided for decoding a video bitstream encoded in a temporal interpolation prediction (TIP) mode. First and second motion vectors are generated for a block of a current picture, each referring to a reference frame, or a reference picture within those frames. The motion vectors are then adjusted by application of decoder-side motion vector refinement (DMVR) processing based on a bilateral matching process, and the adjusted motion vectors are used to decode the block. The adjustment may more specifically consider adjusted motion vector candidates selected by bilateral matching. The adjustment may be applied to both the block partitioning and sub-block partitioning of the current picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Patents and Related Applications This application is based on U.S. Provisional Patent Application No. 63 / 345,329, filed on May 24, 2022, and U.S. Patent Application No. 17 / 982,071, filed on November 7, 2022, claims the benefit of their priority, and the entire contents thereof are incorporated herein by reference.

[0002] The present disclosure generally relates to video encoding and decoding, and more particularly, to methods and apparatuses for improving the accuracy of motion fields by using temporal interpolation prediction.

Background Art

[0003] Video encoding and decoding are generally widely used with the popularity of connected devices and digital media. AOMedia Video 1 (AV1) is an open video encoding format designed for video transmission over the Internet. Many of the components of the AV1 project were provided from previous research efforts. AV1 is an improvement over existing solutions such as its predecessor codec VP9, but problems regarding interpolation still exist. Therefore, further improvements are needed.

Summary of the Invention

Means for Solving the Problems

[0004] According to an embodiment of the present disclosure, a method for decoding an encoded video bitstream is provided. The method is executed by at least one processor in a video decoder. The method includes receiving an encoded video bitstream including a current picture including at least one block and a syntax element indicating that the at least one block should be predicted in a temporal interpolation prediction (TIP) mode. The method further includes generating a motion field for the at least one block, the motion field including a first motion vector indicating a first reference picture and a second motion vector indicating a second reference picture. The method further includes generating a first adjusted motion vector and a second adjusted motion vector by using the first motion vector and the second motion vector in a decoder-side motion vector refinement (DMVR) process for the at least one block based on a bilateral matching process. The method further includes decoding the at least one block by using the first adjusted motion vector and the second adjusted motion vector.

[0005] According to another embodiment of the present disclosure, a video decoder is provided. The video decoder includes at least one communication module configured to receive a bitstream, at least one non-volatile memory electrically configured to store computer program code, and at least one processor operably connected to the at least one communication module and the non-volatile memory. The at least one processor is configured to operate when instructed by the computer program code. The computer program code includes a received code configured to cause an encoded video bitstream including a current picture including at least one block and a syntax element indicating that the at least one block is to be predicted in a temporal interpolation prediction (TIP) mode to be received by at least one of the at least one processor via the at least one communication module. The computer program code further includes a generation code configured to cause at least one of the at least one processor to generate a motion field including a first motion vector indicating a first reference picture and a second motion vector indicating a second reference picture for the at least one block. The computer program code further includes an adjustment code configured to cause at least one of the at least one processor to generate a first adjusted motion vector and a second adjusted motion vector by using the first motion vector and the second motion vector in a decoder-side motion vector refinement (DMVR) process for the at least one block based on a bilateral matching process. The computer program code further includes a decoding code configured to cause at least one of the at least one processor to decode the at least one block by using the first adjusted motion vector and the second adjusted motion vector.

[0006] According to still other embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The recording medium records instructions executable by at least one processor to execute a method for decoding an encoded video bitstream. The method includes receiving an encoded video bitstream including a current picture including at least one block and a syntax element indicating that at least one block is to be predicted in a temporal interpolation prediction (TIP) mode. The method further includes generating a motion field for at least one block, the motion field including a first motion vector indicating a first reference picture and a second motion vector indicating a second reference picture. The method further includes generating a first adjusted motion vector and a second adjusted motion vector by using the first motion vector and the second motion vector in a decoder-side motion vector refinement (DMVR) process for at least one block based on a bilateral matching process. The method further includes decoding at least one block by using the first adjusted motion vector and the second adjusted motion vector.

[0007] Further aspects are described in part in the following description, may become apparent in part from the description, or may be realized by practice of the presented embodiments of the present disclosure.

[0008] The features, aspects, and advantages of specific exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, in which like reference numerals represent like elements.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

DETAILED DESCRIPTION OF THE INVENTION

[0010] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numbers in different drawings may identify the same or similar elements.

[0011] The foregoing disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the implementation forms strictly to the disclosed forms. Modifications and variations are possible in light of the above disclosure or may be obtained from the practice of the implementation forms. Further, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Further, in the flowcharts and operation descriptions provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed (at least partially) simultaneously, and the order of one or more operations may be switched.

[0012] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the embodiments. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0013] Even if a particular combination of features is recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible implementation forms. In fact, many of these features are not specifically recited in the claims and / or may be combined in ways not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementation forms includes each dependent claim in combination with all other claims in the claim set.

[0014] Elements, operations, or instructions used in this specification should not be construed as important or essential unless explicitly described as such. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar words are used. Also, as used in this specification, terms such as "has," "have," "having," "include," "including," etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0015] Currently, with the expansion of media accessibility via the Internet, video encoding has become more important to reduce network load. Disclosed are a method and an apparatus for video encoding.

[0016] FIG. 1 shows an example of an AV1 quadtree 100 according to an exemplary embodiment. In the quadtree 100 for the image 110, a portion 115 of the image 110 (referred to as a superblock in VP9 / AV1 terminology) is expanded into a ten-way structure 120, and the superblock 115 is divided according to various partitioning patterns (e.g., 125a, 125b, 125c) that can each be processed. The partitioning pattern using rectangular partitioning may not be further subdivided, but the partitioning pattern 125c consists of only square patterns and can be divided in the same way as the superblock 115, resulting in recursive partitioning.

[0017] This section or block of the process may also be referred to as a Coding Tree Unit (CTU), and a group of pixels or pixel data units collectively represented by a CTU may be referred to as a Coding Tree Block (CTB). Note that a single CTU can represent multiple CTBs, and each CTB represents different information components (e.g., a CTB for luminance information and multiple CTBs for different color components such as the "red", "green", and "blue" components).

[0018] AV1 increases the possible maximum size of the starting superblock 115, for example, to 128×128 pixels, compared to the 64×64 pixel superblock in VP9. Also, the ten-way structure 120 includes 4-to-1 and 1-to-4 rectangular partitioning patterns 125a, 125b that did not exist in VP9. Further, AV1 adds additional flexibility in the use of partitions less than the 8×8 pixel level in the sense that 2×2 chroma inter prediction becomes possible in certain cases.

[0019] In High Efficiency Video Coding (HEVC), in order to adapt to various local characteristics, a coding tree unit can be divided into coding units (CUs) using a quadtree structure represented as a coding tree. The determination of whether to encode a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction may be made at the CU level. Each CU may be further divided into one, two, or four prediction units (PUs) according to the prediction unit (PU) partitioning type. Within one PU, the same prediction process may be applied, and related information may be sent to the decoder based on the PU. After obtaining a residual block by applying a prediction process based on the PU partitioning type, the CU may be divided into transform units (TUs) according to another quadtree structure such as the coding tree of the CU. The HEVC structure has multiple partitioning concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can be square, but a PU can be square or rectangular for an inter-predicted block. In HEVC, one coding block may be further divided into four square sub-blocks, and a transform may be performed on each sub-block, i.e., TU. Each TU may be recursively divided into smaller TUs called residual quadtree (RQT) (using quadtree partitioning). At the picture boundary, HEVC may adopt an implicit quadtree partitioning so that the block can maintain quadtree partitioning until its size fits within the picture boundary.

[0020] FIG. 2 shows an example of the division of a CTU 220 using a quadtree and binary tree (QTBT) structure 210 according to an exemplary embodiment. The QTBT structure 210 includes both quadtree nodes and binary tree nodes. In FIG. 2, solid lines indicate branches and leaves resulting from divisions at quadtree nodes such as node 211a and the corresponding block divisions, and dotted lines indicate branches and leaves resulting from divisions at binary tree nodes such as node 211b and the corresponding block divisions.

[0021] The splitting at a binary tree node splits the corresponding block into two sub-blocks of equal size. For each splitting (i.e., non-leaf) binary tree node (e.g., node 211b), a flag or other mark may be used to indicate which splitting type (i.e., horizontal or vertical) is used. For example, 0 indicates a horizontal split and 1 indicates a vertical split. Since the splitting at a quadtree node (e.g., node 211a) splits the corresponding block into four sub-blocks of equal size both horizontally and vertically, the flag indicating the splitting type may be omitted.

[0022] In addition, the QTBT scheme supports the flexibility for the luminance and chrominance to have separate QTBT structures. In the case of P slices and B slices, the luminance CTB and chrominance CTB within one CTU may share the same QTBT structure. However, in the case of I slices, the luminance CTB may be split into CUs by a QTBT structure, and the chrominance CTB may be split into chrominance CUs by another QTBT structure. This means that the CU of an I slice may contain a coding block of the luminance component, or coding blocks of two chrominance components, and the CU of a P slice or B slice may contain coding blocks of all three color components.

[0023] In HEVC, in order to reduce the memory access for motion compensation, the inter-prediction of small blocks is restricted. Therefore, dual prediction is not supported for 4×8 and 8×4 blocks, and inter-prediction is not supported for 4×4 blocks. In the QTBT implemented in certain embodiments, these restrictions are removed.

[0024] In HEVC, a CTU may be divided into CUs by using a quadtree represented as a coding tree so as to conform to various local characteristics. The determination of whether to encode a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction may be made at the CU level. Each CU may be further divided into one, two, or four PUs according to the PU partition type. Within one PU, the same prediction process may be applied, and related information may be transmitted to the decoder based on the PU. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU may be divided into transform units (TUs) according to another quadtree structure similar to the coding tree of the CU. Therefore, the HEVC structure may have the concept of multiple partitions including CUs, PUs, and TUs.

[0025] According to the embodiment shown in FIG. 2, the QTBT structure 210 eliminates the concept of multiple partitions, that is, the separation of the concepts of CUs, PUs, and TUs, and supports further flexibility in the CU partition shape. In the QTBT block structure, a CU may have either a square or rectangular shape. As shown in FIG. 2, a coding tree unit (CTU) 220 may first be divided according to the quadtree node 211a of the QTBT structure 210. The branches of the quadtree node 211a may be further divided according to a binary tree node (e.g., nodes 211b and 211c) or another quadtree node (e.g., node 211d). There may be two types of binary tree divisions: symmetric horizontal division and symmetric vertical division. A binary tree leaf node may be designated as a coding unit (CU), and its segmentation may be used for prediction and conversion processing without further division. This may mean that CUs, PUs, and TUs have the same block size in the QTBT coding block structure.

[0026] In certain embodiments, a CU may contain CBs of different color components (e.g., in the case of P slices and B slices in a 4:2:0 chroma format, one CU may contain one luminance coding block (CB) and two chroma CBs), or instead may contain CBs of a single component (e.g., in the case of I slices, one CU may contain either one luminance CB or two chroma CBs).

[0027] The following parameters are defined for the QTBT partitioning scheme.

[0028] - CTU size: The size of the root node of the quadtree, the same concept as in HEVC

[0029] - MinQTSize: The minimum allowable quadtree leaf node size

[0030] - MaxBTSize: The maximum allowable binary tree root node size

[0031] - MaxBTDepth: The maximum allowable binary tree depth

[0032] - MinBTSize: The minimum allowable binary tree leaf node size

[0033] In an exemplary implementation of the QTBT partitioning structure, the CTU220 size may be set as a 128×128 luminance sample with two corresponding 64×64 blocks of chroma samples, MinQTSize may be set as 16×16, MaxBTSize may be set as 64×64, MinBTSize (both width and height) may be set as 4×4, and MaxBTDepth may be set as 4.

[0034] In such an implementation form, the quadtree splitting is applied to the CTU 220 represented by the quadtree root node 211a to generate quadtree leaf nodes 211b, 211c, 211d, and 211e. The quadtree leaf nodes 211b, 211c, 211d, and 211e can have a size ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the leaf quadtree node size is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), it will not be further split by a binary tree. Otherwise, the leaf quadtree node can be further split by the QTBT splitting structure 210. Therefore, the quadtree leaf node 211b may be treated as the root node of a binary tree with a binary tree depth of 0.

[0035] When the binary tree depth reaches MaxBTDepth (i.e., 4), no further splitting is considered. When the binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal splitting is considered. Similarly, when the binary tree node has a height equal to MinBTSize, no further vertical splitting is considered.

[0036] When the splitting is completed, the final leaf nodes of the QTBT splitting structure 210 (e.g., leaf node 211f) may be further processed by prediction and transformation processing. In a specific implementation form, the maximum CTU size is 256×256 luminance samples.

[0037] FIG. 3 shows an example of a block splitting structure using a ternary tree such as a VVC multi-type tree (MTT) structure according to an exemplary embodiment. Adding the use of a ternary tree having a flag or mark similar to that used for a binary tree node to the splitting structure enables both vertical center-side ternary tree splitting and horizontal center-side ternary tree splitting in addition to the splitting enabled by the above QTBT splitting structure. Ternary tree splitting complements quadtree and binary tree splitting. Ternary tree splitting can capture an object located at the center of a block split by a quadtree or binary tree splitting. The width and height of the ternary tree splitting can each be a power of 2, eliminating the need for additional conversion.

[0038] Theoretically, the complexity of tree traversal is T^D, where T represents the number of splitting types and D represents the depth of the tree. Thus, for reasons of complexity reduction, the tree may be a two-level tree (D = 2).

[0039] FIG. 4 shows an exemplary operation for deriving a spatial motion vector predictor according to an exemplary embodiment. The spatial motion vector predictor (SVMP) may itself take the form of a motion vector or may otherwise include a motion vector. The SVMP may be derived from blocks in the vicinity of the current block 410. More specifically, the SVMP may be derived from spatial neighboring blocks 420 that are adjacent to or near the current block 410 on the upper side and the left side. For example, in FIG. 4, a block is a spatial neighboring block 420 if it is in the three rows of blocks directly above the current block 410, or if it is in the three columns of blocks directly to the left of the current block 410, or if it is to the immediate left or right of the row immediately adjacent to the current block 410. The spatial neighboring blocks 420 may be of a regular size smaller than the current block 410. For example, in FIG. 4, the current block 410 is a 32×32 block and each spatial neighboring block 420 is an 8×8 block.

[0040] The spatial neighboring blocks 420 may be examined to find one or more motion vectors (MVs) associated with the same reference frame index as the current block. The spatial neighboring blocks may be examined for the luminance blocks, for example, according to a set of blocks shown in FIG. 4 labeled according to the order of examination. That is, (1) check the upper adjacent row from left to right. (2) Check the left adjacent column from top to bottom. (3) Check the upper right neighboring block. (4) Check the neighboring blocks of the upper left block. (5) Check the first non-adjacent upper row from left to right. (6) Check the first non-adjacent left column from top to bottom. (7) Check the second non-adjacent upper row from left to right. (8) Check the second non-adjacent left column from top to bottom.

[0041] The "adjacent" spatial MV predictor candidates derived from the "adjacent" blocks (i.e., the blocks of block sets 1-3) may be placed in the MV predictor list before the temporal MV predictor candidates of the temporal motion vector predictor (TMVP) further described herein, and the "non-adjacent" spatial MV predictor candidates derived from the "non-adjacent" blocks (external blocks, also known as the blocks of block sets 4-8) are placed in the MV predictor list after the temporal MV predictor candidates.

[0042] In one embodiment, each SMVP candidate has the same reference picture as the current block. For example, assume that the current block 410 has a single reference picture. If the MV candidate also has the same single reference picture as the reference picture of the current block, this MV candidate may be placed in the MV predictor list. Similarly, if the MV candidate has multiple reference pictures and one of the reference pictures is the same as the reference picture of the current block, this MV candidate may be placed in the MV predictor list. However, if the current block 410 has multiple reference pictures, the MV candidate may be placed in the MV predictor list only when it has the same corresponding reference picture for each of those reference pictures of the current block 410.

[0043] FIG. 5 shows an exemplary operation of a set of temporal motion vector predictors (TMVPs) according to an exemplary embodiment. The TMVP may be derived using the arranged blocks in the reference frame. To generate the TMVP, first, one or more MVs of one or more reference frames may be stored together with the reference index associated with each reference frame. Thereafter, for each 8×8 block of the current frame, the MV of the reference frame through which its trajectory passes through the 8×8 block may be identified using the reference frame index and stored in a temporary MV buffer. In inter prediction using a single reference frame, the MV may be stored in 8×8 units to perform the temporal motion vector prediction of the future frame, regardless of whether the reference frame is a “forward” or “backward” reference frame (i.e., after or before the current frame in a series of frames, respectively). In the case of composite inter prediction, the MVs of the “forward” reference frame may be stored in 8×8 units to perform the temporal motion vector prediction of the future frame.

[0044] An exemplary embodiment of the process for generating TMVP may follow the following operations. In this example, the reference motion vector 550 (also labeled as MVref) of the initial reference frame 510 points from the initial reference frame 510 to a subsequent reference frame 540 which is itself a reference frame of the initial reference frame 510. By doing so, it passes through the 8×8 block 570 (shaded with gray dots) of the current frame 520. MVref 550 may be stored in a temporary MV buffer associated with this current block 570. During the motion projection process for deriving the temporal MV predictor 500, subsequent reference frames (e.g., frames 530 and 540) may be scanned in a predetermined order. For example, using the frame labels defined by the AV1 standard, the scan order may be LAST_FRAME, BWDREF_FRAME, ALTREF_FRAME, ALTREF2_FRAME, and LAST2_FRAME. In one embodiment, the MV from a higher-indexed reference frame (in the scan order) is assigned by a lower-indexed reference frame (in the scan order) and does not replace a previously identified MV.

[0045] Finally, given a predetermined block coordinate, the relevant MV stored in the temporary MV buffer is identified to derive a temporal MV predictor 560 (also labeled as MV0) that points from the current block 570 to an adjacent reference frame 530, and can be projected onto the current block 570.

[0046] FIG. 6 shows an exemplary set of predetermined block positions 600 for deriving a temporal motion predictor for a 16×16 block according to an exemplary embodiment. Up to seven blocks can be checked for valid temporal MV predictors. In FIG. 6, the blocks are labeled B0 - B6. As described with reference to FIG. 4, the temporal MV predictor candidates may be checked after the adjacent spatial MV predictor candidates but before the non-adjacent spatial MV predictor candidates and may be placed in the first MVP list. Then, to derive the motion vector predictor (MVP), all spatial and temporal MVP candidates may be accumulated, and each candidate may be assigned a weight determined during the scan of the spatial and temporal neighboring blocks. Based on the associated weights, the candidates may be classified and ranked, and up to four candidates may be identified and placed in the second MVP list. This second list of MVPs, also referred to as the dynamic reference list (DRL), may be further used in the dynamic MV prediction mode.

[0047] If the DRL is not full, additional search may be performed and the resulting additional MVP candidates may be used to fill the DRL. The additional MVP candidates may include, for example, global MV, zero MV, composite MV without scaling, etc. Then, the adjacent SMVP candidates, TMVP candidates, and non-adjacent SMVP candidates within the DRL may be reordered again. Both AV1 and AVM enable reordering based on, for example, the weights of each candidate. The weights of the candidates may be predefined according to the overlapping region between the current block and the candidate block.

[0048] FIG. 7 shows an exemplary operation of generating a new MV candidate via a single inter-prediction block. If the reference frame of the neighboring block is different from that of the current block but the MVs are in the same direction, a temporal scaling algorithm may be utilized to scale the MV to its reference frame in order to form the MVP of the motion vector of the current block. In the example of FIG. 7, the motion vector 740 (also labeled as mv1 in FIG. 7) from the neighboring block 750 of the current block 710 in the current frame 701 points to the arrayed neighboring block 760 in the reference frame 703. The motion vector 740 may be utilized using temporal scaling to derive the MVP of the motion vector 730 (also labeled as mv0 in FIG. 7) of the current block 710 that points to the arrayed current block 720 in another reference frame 702.

[0049] FIG. 8 shows an exemplary operation of generating a new MV candidate via a composite prediction block. In the example of FIG. 8, the composite MVs 860, 870 point from respective different neighboring blocks 820, 830 of the current block 810 in the current frame 802 to the reference frames 803 and 801. The reference frames 803 and 801 of the composite MVs 860, 870 (also labeled as mv2 and mv3 in FIG. 8) may be the same as that of the current block 810. Composite inter-prediction may derive the MVP of the composite MVs 840, 850 (also labeled as mv0 and mv1 in FIG. 8) of the current block 710, which may be determined as in FIG. 7.

[0050] FIG. 9 shows an exemplary operation of updating a motion vector candidate bank 920. This bank 920 was first proposed in CWG-B023 and is hereby incorporated by reference in its entirety.

[0051] The bank update process may be based on the super block 910. That is, after each super block (e.g., super block 910a) is encoded, a set of first candidate MVs (e.g., candidates such as the first 64) used by each coding block in the super block may be added to the bank 920. Pruning may also be performed during the update.

[0052] After the super block reference MV candidate scan is completed, if there is an open slot in the candidate list, the codec may refer to the MV candidate bank 920 (in the buffer where the reference frame type matches) for additional MV candidates. Moving in reverse direction from the end to the beginning of the buffer, if the MV in the bank buffer does not yet exist in the list, it may be added to the candidate list. More specifically, each buffer may correspond to a unique reference frame type covering single and composite inter modes for a single or a pair of reference frames respectively. All buffers may be of the same size. When a new MV is added to a full buffer, existing MVs may be pushed out to make space for the new MV.

[0053] The coding block may refer to the MV candidate bank 920 in addition to what is obtained in the AV1 reference MV list generation to collect reference MV candidates. After encoding the super block, the MV bank may be updated with the MVs used by the coding blocks of the super block.

[0054] AV1 enables splitting a frame into tiles, and each tile contains multiple super blocks. Each tile may be processed in parallel on different processors. Regarding the candidate bank, each tile may have an independent MV candidate bank utilized by all the super blocks within the tile. At the start of encoding each tile, the corresponding bank is emptied. Then, during encoding each super block within that tile, the MVs from the bank may be used as MV reference candidates. After encoding each super block, the bank may be updated as described above.

[0055] Specific embodiments of bank update and reference processing for bank update and reference are described later in this specification.

[0056] FIG. 10 is a flowchart showing the process of constructing a motion vector prediction list for any video input according to an exemplary embodiment. Adjacent SMVP, TMVP, and non-adjacent SMVP candidates may be generated in S1010, S1020, and S1030, respectively, by the processes described above with reference to FIGS. 4 and 5, for example. Next, the candidates may be classified in S1040 by the process described above with reference to FIG. 6, for example, or sorted otherwise. Further MVP candidates may be derived in S1050 by the processes described above with reference to FIGS. 7 and 8, for example. Optionally, additional MVP candidates may be determined by additional search in S1060 by the process described above with reference to FIG. 6, for example, or obtained from the reference bank in S1070 by the process described above with reference to FIG. 9, for example.

[0057] FIG. 11 shows an exemplary operation of the composite inter prediction mode according to an exemplary embodiment.

[0058] The composite inter mode may create a prediction of a block by combining hypotheses from multiple different reference frames. In the example of FIG. 11, for example, block 1111 of current frame 1110 is predicted by motion vectors 1130a, 1130b (also labeled mv0 and mv1 in FIG. 11) of neighboring reference frames 1120a, 1120b. The neighboring reference frames 1120a, 1120b may be very close (i.e., the frames immediately before and after the current frame 1110 in the sequence), but this is not a requirement. The motion information components of each block (e.g., motion vectors 1130a, 1130b) may be transmitted in the bitstream as overhead.

[0059] However, although motion vectors can typically be adequately predicted using spatial or temporal neighboring motion vectors, or predictors from past motion vectors, the memory capacity used for motion information can still be quite significant for many contents and applications.

[0060] FIG. 12 shows an exemplary operation of the Temporal Interpolation Prediction (TIP) mode according to an exemplary embodiment.

[0061] In the example of FIG. 12, the information in reference frames 1220a, 1220b is combined using simple interpolation processing and projected to the same time instance as the current frame 1210. Multiple TIP modes may be supported. In one TIP mode, the interpolated frame or “TIP frame” 1210’ may be used as an additional reference frame. The coding block of the current frame 1210 may directly refer to the TIP frame 1210’, and utilize information from two different references with only the overhead cost of a single inter-prediction mode. In another TIP mode, the TIP frame 1210’ may be directly assigned as the output of the decoding process for the current frame 1210 while skipping any other conventional encoding steps. This mode can provide significant encoding and simplification advantages, particularly for low bitrate applications.

[0062] There are existing techniques for interpolating frames between two reference frames, such as Frame Rate Up Conversion (FRUC), but achieving a good trade-off between complexity and compression quality can be an important constraint when designing new encoding tools. The method disclosed above is simple and re-uses the motion information already available in the reference frames without the need to perform additional motion searches. Simulation results show that this simple method can achieve good quality with a low-complexity implementation.

[0063] In the example of FIG. 12, the TIP mode operation starts by generating a TIP frame 1210' corresponding to the current frame 1210. Then, the TIP frame 1210' may be used as an additional reference frame for the current frame 1210 or directly assigned as the reconstructed output of the decoder for the current frame 1210. On the decoder side, blocks encoded in the TIP mode may be generated on the fly, thereby eliminating the need to create the entire TIP frame 1210 at the decoder and saving decoding time and processing. This is also compatible with the one-pass decoding pipeline in the decoder and is suitable for hardware implementation.

[0064] The frame-level TIP mode may be indicated using syntax elements. Examples of the modes indicated by the value of the tip_frame_mode parameter are shown in the following table.

[0065] [Table 1]

[0066] A simple interpolation method for interpolating an intermediate frame between two frames is disclosed, and the motion vectors from available references can be fully reused. The same motion vectors may also be used for the temporal motion vector predictor (TMVP) processing after slight modification. This processing may include three operations. 1. Create a rough motion vector field for the TIP frame by projection of the modified TMVP field. 2. Adjust the rough motion vector field by using hole filling and smoothing operations. 3. Generate the TIP frame using the adjusted motion vector field. On the decoder side, blocks encoded in the TIP mode can be generated on the fly without creating the entire TIP frame.

[0067] However, it should be noted that other suitable interpolation methods can be substituted in combination with other features described in this disclosure, which is within the scope of this disclosure.

[0068] FIG. 13 shows an exemplary operation of decoder-side motion vector adjustment based on bilateral matching according to an exemplary embodiment. Multipurpose video coding (VVC) may distribute previously decoded pictures to two reference picture lists 1320a, 1320b. These previously decoded pictures may be used as references for predicting the current picture 1310. In the example of FIG. 13, reference pictures prior to the current picture 1310 are assigned to the “past” reference picture list 1320a according to the display order, and reference pictures after the current picture 1320 may be assigned to the “future” reference picture list 1320b. The corresponding reference picture indexes of each list (not shown) indicate which picture in each list is used to predict the current block 1311 of the current picture 1310. In the case of bidirectional prediction, two predicted blocks 1321a and 1321b predicted using the respective MVs 1331a, 1331b of the past reference picture list 1320a and the future reference picture list 1320b may be combined to obtain a single prediction signal.

[0069] When motion information is encoded in integrated mode, the reference picture index, and the MV of neighboring blocks may be applied directly to the current block 1311. However, this may not accurately predict the current block 1311.

[0070] The decoder-side motion vector adjustment (DMVR) algorithm may be used to improve the accuracy of blocks encoded in integrated mode by involving only decoder-side information. When the DMVR algorithm is applied to blocks 1311, 1321a, and 1321b, the MVs 1331a, 1331b derived from integrated mode may be set as “initial” MVs for DMVR.

[0071] Next, the DMVR may further adjust the initial MVs 1331a and 1331b by block matching. In both reference pictures, candidate blocks surrounding the blocks 1321a and 1321b indicated by the initial MVs may be searched for and a bilateral match may be performed. The best-matching blocks 1323a and 1323b may be used to generate the final prediction signal, and the new MVs 1333a and 1333b indicating these new prediction blocks 1323a and 1323b may be set as the "adjusted" MVs corresponding to the initial MVs 1331a and 1331b, respectively. Many block matching methods suitable for DMVR, such as template matching, methods based on bidirectional template matching, and methods based on bilateral matching adopted in VVC, have been studied.

[0072] In the DMVR based on bilateral matching, the block pair 1321a and 1321b indicated by the initial MV may be defined as the initial block pair. The distortion cost of the initial block pair 1321a and 1321b may be calculated as the initial cost. The blocks surrounding the initial block pair 1321a and 1321b may be used as the DMVR candidate block pair. Each block pair may include one prediction block from a reference picture in the past reference picture list 1320a and one prediction block from a reference picture in the future reference picture list 1320b.

[0073] The distortion costs for DMVR candidate block pairs may be measured and compared. Since the DMVR candidate block pair with the lowest distortion cost comprises the two most similar blocks between reference pictures, it can be assumed that this block pair (i.e., blocks 1323a, 1323b) is the best predictor of the current block 1311. Accordingly, the final dual prediction signal can be generated using the block pair 1323a, 1323b. The corresponding MVs 1333a, 1333b may be shown as adjusted MVs. If all DMVR candidate block pairs have a distortion cost greater than that of the initial block pair 1321a, 1321b, the initial blocks 1321a, 1321b may be used for dual prediction, and the adjusted MVs 1333a, 1333b may be set equal to the initial MVs 1331a, 1331b.

[0074] To simplify the calculation of the distortion cost, the sum of absolute differences (SAD) may be used as the distortion metric, and only the luminance distortion in the DMVR search process may be considered. It should be noted that SAD is evaluated between even rows of the candidate block pair, which can further reduce the computational complexity.

[0075] In the example of FIG. 13, the dotted blocks (1321a, 1321b) within each reference picture indicate the initial block pair. The gray blocks (1323a, 1323b) indicate the best match block pair that may be the block pair with the lowest SAD cost compared to other DMVR candidate block pairs and the initial block pair 1321a, 1321b. The initial MVs 1331a, 1331b may be adjusted to generate the adjusted MVs 1333a, 1333b, and the final dual prediction signal may be generated using the best match block pair 1323a, 1323b. It should be noted that the initial MVs 1331a, 1331b can be derived from the integrated mode, thereby supporting fractional sample MV accuracy of up to 1 / 16, so there is no need to indicate all sample positions.

[0076] Since the difference between the adjusted MV and the corresponding initial MV (shown as ΔMV1335a and 1335b in FIG. 13) can be an integer or a fraction, the adjusted MV may indicate a fractional pixel position. In this case, the intermediate search block and the final prediction block may be generated by DMVR interpolation processing.

[0077] In some embodiments, DMVR based on block-level bilateral matching may be performed on top of the motion field generated by TMVP. Here, with reference to the concepts described above in this specification, an example of such processing will be described.

[0078] The process can start from the motion field generated as part of the TIP for each 8×8 block. The motion field is a representation of 3D motion projected onto a 2D space such as an image, and is typically defined by one or more motion vectors each describing the motion of a corresponding point. Here, the motion field may include two motion vectors (MV0 and MV1) indicating two reference pictures. The motion vectors (MV0 and MV1) may be used as the starting point of the DMVR process. More specifically, corresponding predictors within the reference pictures indicated by the motion vectors may be generated. In this operation, the input may be filtered using a filter such as interpolation, bilinear, etc. Then, candidate predictors surrounding the motion vectors may be generated. These predictors may be searched through a predetermined search range N which is an integer value corresponding to the number of luminance samples. The search accuracy is defined as K and may be a fractional value from 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 up to the number of luminance samples (up to the highest supported MV accuracy). In the next operation, bilateral matching may be performed among all candidate predictors, and the position of the predictor with the lowest distortion cost may be determined to be the adjusted position of this 8×8 block. The distortion cost may be SAD, SATD, SSE, subsampled SAD, mean removed SAD, etc., but is not limited thereto.

[0079] After obtaining the adjusted positions (adjusted motion vectors) of each 8×8 block, TIP processing can be performed. More specifically, the TIP frame may be generated using the DMVR-adjusted motion vector field. The generated frame may be used as a reference for prediction or directly used as a prediction.

[0080] On the decoder side, when a block is encoded as TIP or by the TIP mode, the TIP predictor and DMVR adjustment may be performed on-the-fly without generating the entire frame. In some embodiments, the DMVR based on sub-block level bilateral matching may be performed on top of the motion field generated by TMVP. More specifically, for each 8×8 TIP block set, additional splitting may be performed. Such additional splitting may result in four 4×4 sub-blocks from each 8×8 block. Each sub-block may perform a DMVR search based on bilateral matching to obtain an adjusted motion field for TIP. Further, the DMVR based on sub-block level bilateral matching may be performed on top of the motion field generated by TMVP and the optical flow adjustment. More specifically, for each 8×8 TIP block, additional splitting may be performed. For example, each 8×8 TIP block may be split into four 4×4 sub-blocks, the optical flow adjustment is first applied to adjust the motion vectors, and then the DMVR search based on bilateral matching is additionally applied to adjust the motion field of TIP.

[0081] In some embodiments, DMVR based on sub-block level bilateral matching may be performed on top of the motion field and visual flow adjustment generated by TMVP. For example, for each 8×8 TIP block, an additional splitting operation may generate four 4×4 sub-blocks, and DMVR adjustment based on bilateral matching is applied to adjust the motion vectors, and then visual flow adjustment is further applied to adjust the motion filed for TIP. In some embodiments, multi-stage DMVR may be used to adjust the TIP motion field generated by TMVP. For example, the first block level DMVR may be used to adjust the generated initial motion field. The adjusted MV may be used as the starting point for the second stage. In the second stage, sub-block level DMVR may be executed to further adjust the motion field. Such additional stages are within the scope of the present disclosure.

[0082] In other embodiments, the TIP motion field may use explicitly signaled MV differences and / or corrections. Starting from any level, such as the level of a group of coding blocks, a coding block, or a sub-block, one or more motion vector differences (MVDs) may be signaled to the bitstream. The bitstream may be parsed by a decoder and used as a correction to the TIP motion field. When coding a block in TIP mode, the corresponding motion field of the block may be generated using a method based on TMVP. Next, the parsed MVD may be added to the motion field, and if the block is 8×8 or smaller, the MV of the block may be corrected by the parsed MVD. If the block is larger than 8×8, each MV of each 8×8 sub-block may be added to the parsed MVD.

[0083] In some embodiments, when TIP is applied using two reference pictures for motion compensation, the MVD may be signaled to correct the motion field associated with the selected reference picture. For example, the MVD may be signaled for the future reference picture list, but not for the past reference picture list, or vice versa. The selection of which reference picture requires further MVD signaling may be further signaled or implicitly derived.

[0084] FIG. 14 shows an exemplary usage scenario of the integration mode with motion vector difference (MMVD) according to an exemplary embodiment. The integration mode may typically be used with implicitly derived motion information to predict samples generated by the current coding unit (CU). The integration mode with motion vector difference may use a flag to signal that MMVD is used for the CU. The MMVD flag may be sent after the skip flag is sent. In MMVD, after an integration candidate is selected, it may be further adjusted by the signaled MVD information. Further information may include an integration candidate flag, an index specifying the magnitude of the motion, and an index for indicating the direction of the motion. In the MMVD mode, one of the first two candidates in the integration list may be selected for use as the MV basis. The integration candidate flag may signal which candidate to use.

[0085] This operation may specify the magnitude information of the motion and use a distance index indicating a predefined offset from the starting point. The offset may be added to either the horizontal or vertical component of the starting MV. An exemplary relationship between the distance index and the predefined offset is specified in Table 2.

[0086] [Table 2]

[0087] The direction index may represent the direction of the MVD with respect to the starting point. The direction index may represent one of four directions, as shown in Table 3.

[0088]

Table 3

[0089] The meaning of the sign of the MVD may vary according to the information of the starting MV. If the starting MV is a single predicted MV or a bi-predicted MV where both reference picture lists point to the same side of the current picture (i.e., both reference picture order counts (POCs) are greater than the POC of the current picture or both are less than the POC of the current picture), the signs in Table 3 can specify the signs of the MV offsets added to the starting MV. If the starting MV is a bi-predicted MV where the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), and the difference in POC in the first reference picture list is greater than 1 second, the signs in Table 3 can specify the signs of the MV offsets added to the first list MV component of the starting MV, and the signs of the second list MV can have opposite values. Otherwise, if the difference in POC in the second list is greater than the difference in POC in the first list, the signs in Table 3 can specify the signs of the MV offsets added to the second list MV component of the starting MV, and the signs of the first list MV can have opposite values.

[0090] The MVD may be scaled according to the difference in POC in each direction. If the differences in POC in both lists are the same, the scaling may be omitted. Otherwise, if the difference in POC in one list is greater than the difference in POC in the other list, the MVD of the list with the smaller difference in POC may be scaled. When the starting MV is singly predicted, the MVD may be added to the available MV.

[0091] In addition to one-way prediction and bi-directional prediction mode MVD signaling, a symmetric MVD mode for bi-directional MVD signaling may also be applied. In the symmetric MVD mode, motion information including the reference picture indexes of both reference picture lists and the MVD of the future reference picture list is not signaled but derived.

[0092] In certain embodiments, the decoding process of the symmetric MVD mode may be as follows.

[0093] At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 may be derived as follows. If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the closest reference picture in the past reference picture list L0 and the closest reference picture in the future reference picture list L1 form a pair of in front of and behind the reference picture, or a pair of behind and in front of the reference picture, BiDirPredFlag is set to 1 and the reference pictures of L0 and L1 are both short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.

[0094] At the CU level, when the CU is bi-predicted coded and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used may be explicitly signaled. If the symmetric mode flag is true, mvp_l0_flag, mvp_l1_flag, and MVD0 may be explicitly signaled and other signals may be omitted. The reference indexes for L0 and L1 may be set equal to the pair of reference pictures respectively, and MVD1 may be set equal to (-MVD0).

[0095] In some embodiments, for each coding block within an inter - frame, when the mode of the current block is an inter - coding mode rather than a skip mode, another flag may be signaled to indicate whether a single - reference mode or a composite - reference mode is used for the current block. A predicted block may be generated by one motion vector in the single - reference mode, or may be generated by weighted - averaging two predicted blocks derived from two motion vectors in the composite - reference mode.

[0096] In the case of the single - reference mode, according to the syntax of an exemplary implementation, the following specific modes may be signaled.

[0097] Use one of the motion - vector predictors (MVPs) in the list indicated by the NEARMV - DRL (Dynamic Reference List) index.

[0098] Use one of the motion - vector predictors (MVPs) in the list signaled by the NEWMV - DRL index as a reference and apply a delta to the MVP.

[0099] Use a motion vector based on the GLOBALMV - frame - level global motion parameter.

[0100] In the case of the composite - reference mode, according to the syntax of an exemplary implementation, the following specific modes may be signaled.

[0101] Use one of the motion - vector predictors (MVPs) in the list signaled by the NEAR_NEARMV - DRL index.

[0102] Use one of the motion - vector predictors (MVPs) in the list signaled by the NEAR_NEWMV - DRL index as a reference and transmit the delta MV of the second MV.

[0103] Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEARMV-DRL index as a reference and transmit the delta MV of the first MV.

[0104] Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEWMV-DRL index as a reference and transmit the delta MVs of both MVs.

[0105] Use the MV from each reference based on the GLOBAL_GLOBALMV-frame level global motion parameter.

[0106] In some embodiments, this operation can enable a motion vector accuracy (or precision) of 1 / 8 pixel. In an exemplary implementation, the following syntax may be used to signal the motion vector difference of L0 or L1.

[0107] mv_joint specifies which components of the motion vector difference are non-zero.

[0108] 0 indicates that there is no non-zero MVD in either the horizontal or vertical direction.

[0109] 1 indicates that there is a non-zero MVD only along the horizontal direction.

[0110] 2 indicates that there is a non-zero MVD only in the vertical direction.

[0111] 3 indicates that there are non-zero MVDs along both the horizontal and vertical directions.

[0112] mv_sign specifies whether the motion vector difference is positive or negative.

[0113] mv_class specifies the class of the motion vector difference. As shown in Table 4, a higher class may indicate that the motion vector difference is larger.

[0114]

Table 4

[0115] mv_bit specifies the integer part of the offset between the motion vector difference and the starting size for each MV class.

[0116] mv_fr specifies the first two fractional bits of the motion vector difference.

[0117] mv_hp specifies the third fractional bit of the motion vector difference.

[0118] In the case of the NEW_NEARMV and NEAR_NEWMV modes, the accuracy of the MVD may depend on the associated class and the size of the MVD. For example, the fractional MVD may be allowed only if the size of the MVD is 1 pixel or less. Further, if the value of the associated MV class is MV_CLASS_1 or more, only one MVD value may be allowed, and the MVD value for each MV class is derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5).

[0119] Table 5 shows the MVD values allowed for each MV class according to the above embodiment.

[0120]

Table 5

[0121] In addition, if the current block is encoded as the NEW_NEARMV or NEAR_NEWMV mode, one context may be used to signal mv_joint or mv_class. Otherwise, another context may be used to signal mv_joint or mv_class.

[0122] To indicate whether the MVDs for two reference lists are signaled together, a new inter-coding mode named JOINT_NEWMV may be applied. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs of L0 and L1 may be signaled together. More specifically, only one MVD named joint_mvd may be signaled and sent to the decoder, and the delta MVs of L0 and L1 may be derived from joint_mvd.

[0123] The JOINT_NEWMV mode may be signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No additional context needs to be added. When the JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD may be scaled for L0 or L1 based on the POC distance. Let the POC distance between L0 and the current frame be td0, and the POC distance between L1 and the current frame be td1. If td0 is greater than or equal to td1, joint_mvd may be directly used for L0, and the mvd of L1 may be derived from joint_mvd based on Equation (1).

Number

[0124] Otherwise, if td1 is greater than or equal to td0, joint_mvd may be directly used for L1, and the mvd of L0 may be derived from joint_mvd based on Equation (2).

Number

[0125] (When td0 and td1 are equal, since derived_mvd = joint_mvd according to any of the above equations, joint_mvd may be directly used as the derived MVD for both L0 and L1. In that case, it is obvious that no scaling is performed.)

[0126] Here, an inter-coding mode named the AMVDMV mode can be made available in the single-reference case. In the AMVDMV mode, the adaptive MVD (AMVD) resolution is applied to the signal MVD.)

[0127] To indicate whether AMVD is applied to the joint MVD coding mode, a flag (here denoted as amvd_flag) may be added under the JOINT_NEWMV mode. This may be called joint AMVD coding. In joint AMVD coding, the MVDs of two reference frames may be signaled together, and the accuracy of the MVD may be implicitly determined by the size of the MVD. Otherwise, the MVDs of two (or three or more) reference frames may be signaled together, and MVD coding may be applied.)

[0128] The adaptive motion vector resolution (AMVR) first proposed in CWG-C012 is incorporated herein in its entirety and supports seven MV accuracy values (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8). For each prediction block, the adaptive motion vector (AVM) coder may search all the supported accuracy values and signal the best accuracy to the decoder.)

[0129] To shorten the execution time of the coder, two accuracy sets may be supported. Each accuracy set may include four predetermined accuracies. The accuracy sets may be adaptively selected at the frame level based on the maximum accuracy value of the frame. Similar to the standard AV1, the maximum accuracy may be signaled in the frame header. The following table summarizes the accuracy values supported according to the maximum accuracy at the frame level.)

[0130]

Table 6

[0131] The AOMedia AVM repository related to AV1 provides a frame-level flag to indicate whether the MV of a frame includes sub-pixel accuracy. In certain embodiments, the AMVR may be enabled only when the value of the cur_frame_force_integer_mv flag is 0. If the accuracy of a block is lower than the maximum accuracy, the motion model and interpolation filter may not be signaled and may remain inactive. If the accuracy of a block is lower than the maximum accuracy, the applicable motion model may be inferred as a translational motion model, and the applicable interpolation filter may be inferred as a "regular" filter. If the accuracy of a block is either 4 pixels or 8 pixels, the inter-intra mode may remain not signaled and may be inferred as 0.

[0132] FIG. 15 is a diagram of exemplary components of a device or system 1500 that may implement embodiments of the systems and / or methods described herein. The exemplary system 1500 may be one of various systems such as a personal computer, a mobile device, a computer cluster, a server, an embedded device, an ASIC, a microcontroller, or any other device capable of executing code. A bus 1510 connects the exemplary system 1500 to each other so that all components can communicate with each other. The bus 1510 connects a processor 1520, a memory 1530, a storage component 1540, an input component 1550, an output component 1560, and an interface component.

[0133] Processor 1520 may be a single processor, a processor having multiple processors internally, a cluster of (two or more) processors, and / or distributed processing. The processor executes instructions stored in both memory 1530 and storage component 1540. Processor 1520 operates as a computing device that executes operations to modify a shared Unreal Engine Derived Data Cache. Memory 1530 is high-speed storage and can enable retrieval to any memory device by using cache memory that can be closely associated with one or more CPUs. Storage component 1540 may be one of any long-term storage such as HDD, SSD, magnetic tape, or any other long-term storage format.

[0134] Input component 1550 may be any file type or signal from a user interface component such as an input capture device such as a camera, a handheld controller, a gamepad, a keyboard, a mouse, or a motion capture device. Output component 1560 outputs the processed information to communication interface 1570. The communication interface may be another communication device such as a speaker or a screen, and can display information to a user or another observer such as another computing system.

[0135] The foregoing disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the implementation forms strictly to the disclosed forms. Modifications and variations are possible in light of the above disclosure or may be obtained from the practice of the implementation forms.

[0136] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of integration of technical details. Further, one or more of the components described above may be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-transitory storage medium having computer-readable program instructions for causing a processor to perform operations.

[0137] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0138] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0139] The computer-readable program code / instructions for performing the operations may be in any combination of source code or object code written in one or more programming languages, including assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk or C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing the aspect or operation.

[0140] These computer-readable program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, implement the operations specified in the blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that includes instructions for causing a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium comprises a manufactured article including instructions for implementing the manner of operation specified in the blocks of the flowchart and / or block diagram.

[0141] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device such that the instructions which execute on the computer, other programmable apparatus, or other device implement the operations specified in the blocks of the flowchart and / or block diagram, thereby generating computer-implemented processing.

[0142] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of a module that includes one or more executable instructions for performing the specified logical operation. The methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to those shown in the figures. In some alternative embodiments, the operations shown in the blocks may be performed in an order different from that shown in the figures. For example, two blocks shown in succession may actually be performed simultaneously or substantially simultaneously, or the blocks may be performed in the reverse order depending on the relevant functionality, in some cases. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks of the block diagrams and / or flowcharts, can be implemented by a system based on dedicated hardware that performs the specified operation or action, or a combination of dedicated hardware and computer instructions.

[0143] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting of the embodiments. Accordingly, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

Description of Reference Numerals

[0144] 100 Ternary tree 110 Image 115 Superblock 120 Decenary structure 125a, 125b, 125c segmentation pattern 210 QTBT structure 211a, 211b, 211c, 211d, 211e, 211f nodes 220 CTU 310 Vertical 320 Horizontal 410, 570, 710, 720, 810, 1311 current block 420 Spatial neighborhood block 500, 560 Temporal MV predictor 510 Initial reference frame 520, 701, 802, 1110, 1210 current frame 530, 540, 702, 703, 801, 803, 1120a, 1120b, 1220a, 1220b reference frames 550 Reference motion vector 560MV 600 Block position 730, 740, 1130a, 1130b motion vectors 750, 760, 820, 830 neighborhood blocks 840, 850, 860, 870 composite MV 910, 910a superblock 920 MV candidate bank 1111 block 1210’ TIP frame 1310 current picture 1320a, 1320b, L0, L1 reference picture list 1321a, 1321b initial block pair 1323a, 1323b best matching block 1331a, 1331b, 1335a, 1335b initial MV 1333a, 1333b adjusted MV 1500 System 1510 Bus 1520 Processor 1530 Memory 1540 Storage component 1550 Input component 1560 Output component 1570 Communication interface

Claims

Claim 1 A method for decoding an encoded video bitstream, the method being executed by at least one processor in a video decoder, the method comprising: receiving an encoded video bitstream including a current picture including at least one block and a syntax element indicating that the at least one block is to be predicted in a temporal interpolation prediction (TIP) mode; generating a motion field for the at least one block, the motion field including a first motion vector indicating a first reference picture and a second motion vector indicating a second reference picture; generating a first adjusted motion vector and a second adjusted motion vector using the first motion vector and the second motion vector in a decoder-side motion vector refinement (DMVR) process for the at least one block based on a bilateral matching process; decoding the at least one block using the first adjusted motion vector and the second adjusted motion vector; A method comprising the steps of: Claim 2 The bilateral matching process determines a first set of candidate motion vectors surrounding the first motion vector and a second set of candidate motion vectors surrounding the second motion vector, the first adjusted motion vector is a candidate motion vector derived from the first set of candidate motion vectors and having the lowest distortion cost with respect to the first motion vector, the second adjusted motion vector is a candidate motion vector derived from the second set of candidate motion vectors and having the lowest distortion cost with respect to the second motion vector, The method according to claim 1. Claim 3 The method according to claim 2, wherein at least one candidate motion vector is determined based on a search range N, and N is an integer value corresponding to the number of luminance samples. Claim 4 The method according to claim 2, wherein at least one candidate motion vector is determined based on a search accuracy K. Claim 5 The method according to claim 4, wherein the search accuracy K is a fractional value of the number of luminance samples. Claim 6 The method according to claim 4, wherein the search accuracy K is an integer value corresponding to the number of luminance samples. Claim 7 The step of generating the first adjusted motion vector and the second adjusted motion vector comprises: The step of dividing the at least one block into a plurality of sub - blocks; The step of performing the DMVR process on each of the plurality of sub - blocks; The method according to claim 1, comprising the above.

8. The step of generating the first adjusted motion vector and the second adjusted motion vector: For each of the plurality of sub - blocks, the step of performing a visual flow adjustment process on the sub - block before the DMVR process is performed on the sub - block The method according to claim 7, further comprising the above.

9. The step of generating the first adjusted motion vector and the second adjusted motion vector: For each of the plurality of sub - blocks, the step of performing a visual flow adjustment process on the sub - block after the DMVR process is performed on the sub - block The method according to claim 7, further comprising the above.

10. The step of generating the first adjusted motion vector and the second adjusted motion vector: The step of performing the DMVR process on the at least one block; After the DMVR process is performed on the at least one block, the step of dividing the at least one block into a plurality of sub - blocks; The step of performing the DMVR process on each of the plurality of sub - blocks; The method according to claim 1, comprising the above.

11. A video decoder, comprising: At least one communication module configured to receive a bitstream; At least one non - volatile memory electrically configured to store computer program code; At least one processor operably connected to the at least one communication module and the at least one non - volatile memory, the at least one processor being configured to operate as commanded by the computer program code, wherein the computer program code is: A reception code configured to cause a coded video bitstream including a current picture including at least one block and a syntax element indicating that the at least one block is to be predicted in a temporal interpolation prediction (TIP) mode to be received by at least one of the at least one processor via the at least one communication module; For the at least one block, generation code configured to cause at least one of the at least one processor to generate a motion field including a first motion vector indicating a first reference picture and a second motion vector indicating a second reference picture; Adjustment code configured to cause at least one of the at least one processor to generate a first adjusted motion vector and a second adjusted motion vector using the first motion vector and the second motion vector in decoder-side motion vector refinement (DMVR) processing for the at least one block based on bilateral matching processing; Decoding code configured to cause at least one of the at least one processor to decode the at least one block using the first adjusted motion vector and the second adjusted motion vector; A video decoder, comprising: A video decoder. **Claim 12** The adjustment code is further configured to cause at least one of the at least one processor to determine, by the bilateral matching processing, a first set of candidate motion vectors surrounding the first motion vector and a second set of candidate motion vectors surrounding the second motion vector; The first adjusted motion vector is a candidate motion vector derived from the first set of candidate motion vectors and having the lowest distortion cost with respect to the first motion vector; The second adjusted motion vector is a candidate motion vector derived from the second set of candidate motion vectors and having the lowest distortion cost with respect to the second motion vector. The video decoder according to claim 11. **Claim 13** The video decoder according to claim 12, wherein at least one candidate motion vector is determined based on a search range N, and N is an integer value corresponding to the number of luminance samples. **Claim 14** The video decoder according to claim 12, wherein at least one candidate motion vector is determined based on a search accuracy K. **Claim 15** The video decoder according to claim 14, wherein the search accuracy K is either a fractional value of the number of luminance samples or an integer value corresponding to the number of luminance samples. **Claim 16** The adjustment code causes at least one of the at least one processor to dividing the at least one block into a plurality of sub-blocks, executing the DMVR process on each of the plurality of sub-blocks The video decoder according to claim 11, further configured as described above. **Claim 17** The adjustment code causes at least one of the at least one processor to for each of the plurality of sub-blocks, perform a visual flow adjustment process on the sub-block before the DMVR process is executed on the sub-block The video decoder according to claim 16, further configured as described above. **Claim 18** The adjustment code causes at least one of the at least one processor to for each of the plurality of sub-blocks, perform a visual flow adjustment process on the sub-block after the DMVR process is executed on the sub-block The video decoder according to claim 16, further configured as described above. **Claim 19** The adjustment code causes at least one of the at least one processor to execute the DMVR process on the at least one block, after the DMVR process is executed on the at least one block, divide the at least one block into a plurality of sub-blocks, execute the DMVR process on each of the plurality of sub-blocks The video decoder according to claim 11, further configured as described above. **Claim 20** A non-transitory computer-readable recording medium storing instructions executable by at least one processor to perform a method for decoding an encoded video bitstream, the method comprising: receiving an encoded video bitstream including a current picture including at least one block and a syntax element indicating that the at least one block is to be predicted in a temporal interpolation prediction (TIP) mode; generating a motion field for the at least one block, the motion field including a first motion vector indicating a first reference picture and a second motion vector indicating a second reference picture; generating a first adjusted motion vector and a second adjusted motion vector using the first motion vector and the second motion vector in a decoder-side motion vector refinement (DMVR) process for the at least one block based on a bilateral matching process; A step of decoding the at least one block using the first adjusted motion vector and the second adjusted motion vector; A non-transitory computer-readable recording medium including the above.

Citation Information

Patent Citations

  • Hardware Friendly Constrained Motion Vector Refinement

    US20190238883A1

  • Signaling of slice types in video pictures headers

    WO2021129866A1