Method and apparatus for asymmetric blending of predictions of segmented pictures
Patent Information
- Application Number
- JP2024525693
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2022-11-10
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video encoding and decoding methods, such as AV1, face challenges with interpolation, particularly in handling heterogeneous content where sharp edges and textures are present, leading to suboptimal blending and prediction accuracy.
The proposed method introduces asymmetric blending by using different thresholds on either side of the segmentation boundary, allowing for customized blending masks to be applied to predicted pixels in video encoding and decoding. This approach enhances the prediction accuracy by adapting to the local characteristics of the image.
Asymmetric blending improves the prediction accuracy and visual quality of video encoding and decoding, especially in regions with sharp transitions or complex textures, by providing a more nuanced and adaptive blending strategy.
Smart Images

Figure 00000030_0000 
Figure 00000030_0001 
Figure 00000030_0002
Abstract
Description
Technical Field
[0001] Related Patents and Related Applications This application claims the benefit of priority based on U.S. Provisional Patent Application No. 63 / 345,329, filed on May 24, 2022, U.S. Provisional Patent Application No. 63 / 346,614, filed on May 27, 2022, and U.S. Patent Application No. 17 / 983,017, filed on November 8, 2022, the contents of each of which are hereby incorporated by reference in their entirety.
[0002] This disclosure generally relates to video encoding and decoding, and more particularly, to methods and apparatuses for applying asymmetric blending to predicted partition blocks of a video bitstream encoding.
Background Art
[0003] Video encoding and decoding are generally widely used with the rapid increase of connected devices and digital media. AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. Many of the components of the AV1 project were provided from previous research efforts. AV1 is an improvement over existing solutions such as its predecessor codec VP9, but problems regarding interpolation still exist. Therefore, further improvements are needed.
Summary of the Invention
Means for Solving the Problems
[0004] According to certain embodiments of the present disclosure, a method for predicting a picture area in a decoding process is provided. The method is performed by at least one processor of a decoding device. The method includes receiving an input image including a picture area divided into at least a first part and a second part. The method further includes identifying a segmentation boundary that divides the picture area into the first part and the second part. The method further includes determining a first threshold corresponding to the first part and a second threshold corresponding to the second part. The method further includes applying a first blending mask based on the first threshold to predicted pixels for the first part to generate a first blend area, and applying a second blending mask based on the second threshold to predicted pixels for the second part to generate a second blend area, wherein the first blending mask is different from the second blending mask. The method further includes reconstructing the input image including a prediction for the picture area including the first part and the second part modified by the first blend area and the second blend area.
[0005] According to other embodiments of the present disclosure, a decoding device is provided. The encoding device includes at least one communication module configured to receive a signal, at least one non-volatile memory electrically configured to store computer program code, and at least one processor operably connected to the at least one communication module and the at least one non-volatile memory. The at least one processor is configured to operate as instructed by the computer program code. The computer program code includes input code configured to cause at least one of the at least one processor to receive an input image including a picture area divided into at least a first portion and a second portion through the at least one communication module. The computer program code further includes division code configured to cause at least one of the at least one processor to identify a division boundary that divides the picture area into the first portion and the second portion. The computer program code further includes threshold code configured to cause at least one of the at least one processor to set a first threshold corresponding to the first portion and a second threshold corresponding to the second portion. The computer program code further includes blending code configured to cause at least one of the at least one processor to apply a first blending mask based on the first threshold to predicted pixels for the first portion to generate a first blend area and apply a second blending mask based on the second threshold to predicted pixels for the second portion to generate a second blend area, wherein the first blending mask is different from the second blending mask. The computer program code further includes reconstruction code configured to cause at least one of the at least one processor to reconstruct an input image including predictions for the picture area including the first portion and the second portion modified by the first blend area and the second blend area.
[0006] According to still other embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The recording medium records instructions executable by at least one processor to perform a method for predicting a picture area in a decoding process. The method includes receiving an input image including a picture area divided into at least a first part and a second part. The method further includes identifying a segmentation boundary that divides the picture area into the first part and the second part. The method further includes determining a first threshold corresponding to the first part and a second threshold corresponding to the second part. The method further includes applying a first blending mask based on the first threshold to predicted pixels for the first part to generate a first blend area, and applying a second blending mask based on the second threshold to predicted pixels for the second part to generate a second blend area, wherein the first blending mask is different from the second blending mask. The method further includes reconstructing the input image including a prediction for the picture area including the first part and the second part modified by the first blend area and the second blend area.
[0007] Additional aspects will be described in part in the following description, become apparent in part from the description, or can be realized by practice of the presented embodiments of the disclosure.
[0008] Features, aspects, and advantages of specific exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings, in which like reference numerals represent like elements.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Best Mode for Carrying Out the Invention
[0010] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numbers in different drawings may identify the same or similar elements.
[0011] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementation forms to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure or can be obtained from the practice of the implementation forms. Additionally, one or more features or components of one embodiment may be incorporated into or combined with other embodiments (or one or more features of other embodiments). In addition, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed (at least partially) simultaneously, and the order of one or more operations may be rearranged.
[0012] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation form. Therefore, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0013] Certain combinations of features are recited in the claims and / or disclosed herein, but these combinations do not limit the disclosure of possible implementation forms. In fact, many of these features can be combined in ways not specifically recited in the claims and / or disclosed herein. Each of the dependent claims listed below can depend directly on only one claim, but the disclosure of possible implementation forms includes each dependent claim combined with all other claims in the claim set.
[0014] Elements, acts, or instructions used herein should not be construed as important or essential unless expressly described as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used synonymously with "one or more". When only one item is intended, the term "one" or similar language is used. Also, as used herein, the terms "has", "have", "having", "include", "including", etc. are intended to be open-ended terms. Further, the phrase "based on" shall mean "at least partially based on" unless otherwise specified. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.
[0015] With the current prevalence of media accessibility via the Internet, video encoding has become more important to reduce network load. Methods and apparatuses for video encoding and decoding are disclosed.
[0016] In encoding and decoding, a blending mask or wedge weighting mask can use symmetric blending and the weighting thresholds between segmentation boundaries are equal. This is not suitable for all content types. For example, if a portion of a predicted image is homogeneous while another portion depicts an object, the blending of the homogeneous portion can be sharper than the portion containing the object. An improvement here is desirable.
[0017] In the disclosed method and apparatus, instead of a given symmetric (i.e., one threshold) blend design, the design may have different blend thresholds around the segmentation boundary, e.g., two given thresholds. The blending mask or wedge weighting mask may be calculated either beforehand or on-the-fly based on these two thresholds. The resulting asymmetric blending design can be used, for example, to complement the geometric partitioning mode (GPM) in Versatile Video Coding (VVC) and subsequent codecs, as well as wedge-based prediction in AV1, AV2, and subsequent codecs.
[0018] FIG. 1 shows an exemplary example of an AV1 partition tree 100 according to an exemplary embodiment. In the partition tree 100 of image 110, a portion 115 of image 110 (referred to as a superblock in VP9 / AV1 terminology) is expanded into a 10-way structure 120 and the superblock 115 is segmented according to various partition patterns (e.g., 125a, 125b, 125c) that can each be processed. Partition patterns that use rectangular partitions may not be further subdivided, but partition pattern 125c consists only of square patterns, and the square patterns themselves can be partitioned in the same way as superblock 115, resulting in recursive partitioning.
[0019] A partition or block of this process may also be referred to as a coding tree unit (CTU), and a group of pixels or pixel data units collectively represented by a CTU may be referred to as a coding tree block (CTB). Note that a single CTU may represent multiple CTBs, where each CTB represents a different component of information (e.g., a CTB for luminance information, as well as multiple CTBs for different color components such as "red", "green", and "blue" factors).
[0020] Compared with the 64×64 pixel superblock in VP9, AV1 increases the maximum possible size of the starting superblock 115 to, for example, 128×128 pixels. Also, the 10-way structure 120 includes rectangular partition patterns 125a and 125b of 4:1 and 1:4 that did not exist in VP9. Additionally, AV1 adds further flexibility in the use of partitions below the 8×8 pixel level in the sense that 2×2 chroma inter prediction becomes possible when certain conditions are met.
[0021] In High Efficiency Video Coding (HEVC), in order to adapt to various local characteristics, coding tree units can be divided into coding units (CUs) using a quadtree structure called a coding tree. A decision on whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction can be made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the PU partition type. Within one PU, the same prediction process can be applied, and the relevant information can be sent to the decoder for each PU. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure such as the coding tree for the CU. The HEVC structure has multiple partition concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can be square-shaped, while a PU can be square-shaped or rectangular-shaped for an inter-predicted block. In HEVC, one coding block can be further divided into four square sub-blocks, and the transform can be performed for each sub-block, i.e., TU. Each TU can be recursively divided into smaller TUs (using quadtree partitioning), which is called a residual quadtree (RQT). At the picture boundary, HEVC can adopt implicit quadtree partitioning so that the block can maintain quadtree partitioning until its size conforms to the picture boundary.
[0022] Figure 2 shows an exemplary example of the division of CTU 220 using a quadtree + binary tree (QTBT) structure 210 according to an exemplary embodiment. The QTBT structure 210 includes both quadtree nodes and binary tree nodes. In Figure 2, solid lines indicate branches and leaves resulting from the division at quadtree nodes such as node 211a, as well as the corresponding block divisions, and dotted lines indicate branches and leaves resulting from the division at binary tree nodes such as node 211b, as well as the corresponding block divisions.
[0023] The division at a binary tree node divides the corresponding block into two sub-blocks of equal size. For each division (i.e., non-leaf) binary tree node (e.g., node 211b), a flag or other mark can be used to indicate which division type (i.e., horizontal or vertical) is used. For example, 0 indicates a horizontal division and 1 indicates a vertical division. The division at a quadtree node (e.g., node 211a) divides the corresponding block into four sub-blocks of equal size in both the horizontal and vertical directions, and thus the flag for indicating the division type can be omitted.
[0024] Also, the QTBT method supports flexibility such that luma and chroma have separate QTBT structures. In the case of P slices and B slices, the luma CTB and chroma CTB in one CTU can share the same QTBT structure. However, in the case of I slices, the luma CTB can be divided into CUs by the QTBT structure, and the chroma CTB can be divided into chroma CUs by a different QTBT structure. This means that a CU in an I slice can include a coding block of the luma component or coding blocks of two chroma components, and a CU in a P slice or B slice can include coding blocks of all three color components.
[0025] In HEVC, in order to reduce memory access for motion compensation, inter prediction for small blocks is restricted such that dual prediction is not supported for 4×8 and 8×4 blocks and inter prediction is not supported for 4×4 blocks. In the QTBT implemented in certain embodiments, these restrictions are removed.
[0026] In HEVC, in order to adapt to various local characteristics, a CTU can be divided into CUs by using a quadtree shown as a coding tree. A decision on whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction can be made at the CU level. Each CU can be further divided into one, two, or four PUs according to the PU partition type. Within one PU, the same prediction process can be applied, and relevant information can be sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU. Therefore, the HEVC structure can have multiple partition concepts including CUs, PUs, and TUs.
[0027] According to the embodiment shown in FIG. 2, the QTBT structure 210 removes the concept of multiple partition types, i.e., removes the separation of the concepts of CU, PU, and TU, and supports more flexibility for the CU partition shape. In the QTBT block structure, the CU can have either a square or rectangular shape. As shown in FIG. 2, the coding tree unit (CTU) 220 can be first partitioned according to the quadtree node 211a of the QTBT structure 210. The branches of the quadtree node 211a can be further partitioned according to binary tree nodes (e.g., nodes 211b and 211c) or other quadtree nodes (e.g., node 211d). There can be two types of binary tree partitions: symmetric horizontal partition and symmetric vertical partition. The binary tree leaf nodes can be designated as coding units (CUs), and their segmentation can be used for prediction and transformation processing without any further partitioning. This means that the CUs, PUs, and TUs can have the same block size in the QTBT coding block structure.
[0028] In certain embodiments, a CU can include coding blocks (CBs) of different color components (e.g., in the case of P slices and B slices in 4:2:0 chroma format, one CU can also include one luma CB and two chroma CBs), or alternatively, can include a single-component CB (e.g., in the case of I slices, one CU can include either one luma CB or two chroma CBs).
[0029] The following parameters are defined for the QTBT partitioning method. - CTU size: The size of the root node of the quadtree, which is the same concept as in HEVC - MinQTSize: The minimum allowable quadtree leaf node size - MaxBTSize: The maximum allowable binary tree root node size - MaxBTDepth: The maximum allowable binary tree depth - MinBTSize: The minimum allowable binary tree leaf node size
[0030] In an exemplary implementation of the QTBT partitioning structure, the CTU220 size can be set as a 128×128 luma sample with two corresponding 64×64 blocks of chroma samples, MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for both width and height) is set to 4×4, and MaxBTDepth is set to 4.
[0031] In such an implementation, to generate the quadtree leaf nodes 211b, 211c, 211d, and 211e, the quadtree partitioning is applied to the CTU211 represented by the quadtree root node 220a. The quadtree leaf nodes 211b, 211c, 211d, and 211e can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). When the leaf quadtree node size is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), it is not further divided by a binary tree. Otherwise, the leaf quadtree node can be further partitioned by the QTBT partitioning structure 210. Thus, the quadtree leaf node 211b can also be treated as the root node for a binary tree with a binary tree depth of 0.
[0032] When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting is not considered. When the binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal splitting is not considered. Similarly, when the binary tree node has a height equal to MinBTSize, further vertical splitting is not considered.
[0033] Once the partitioning is complete, the final leaf nodes (e.g., leaf node 211f) of the QTBT partitioning structure 210 can be further processed by prediction and transformation processing. In a particular implementation, the maximum CTU size is 256×256 luma samples.
[0034] FIG. 3 shows an exemplary example of a block partitioning structure using a ternary tree such as a VVC multi-type tree (MTT) structure according to an exemplary embodiment. Adding the use of a ternary tree to the partitioning structure, along with flags or marks similar to those used in binary tree nodes, enables both vertical 310 and horizontal 320 center-side ternary tree partitions in addition to the partitions enabled by the QTBT partitioning structure above. The ternary tree partition complements the quadtree and binary tree partitions. The ternary tree partition can capture an object located at the center of a block that would otherwise be split by a quadtree or binary tree partition. The width and height of the ternary tree partition can each be a power of two, eliminating the need for additional conversions.
[0035] Theoretically, the complexity of tree traversal is T^D, where T represents the number of split types and D represents the depth of the tree. Therefore, for complexity reduction, the tree can be a two-level tree (D = 2).
[0036] FIG. 4 shows an exemplary operation of deriving a spatial motion vector predictor according to an exemplary embodiment. The spatial motion vector predictor (SVMP) can itself take the form of a motion vector or otherwise include a motion vector. The SVMP can be derived from blocks in the vicinity of the current block 410. More specifically, the SVMP can be derived from spatial neighboring blocks 420 adjacent to or near the current block 410 on the upper and left sides. For example, in FIG. 4, if a block is in the block three rows directly above the current block 410, or if a block is in the block three columns directly to the left of the current block 410, or if a block is to the immediate left or right of the row immediately adjacent to the upper part of the current block 410, the block is a spatial neighboring block 420. The spatial neighboring blocks 420 can have a regular size smaller than the current block 410. For example, in FIG. 4, the current block 410 is a 32×32 block and each spatial neighboring block 420 is an 8×8 block.
[0037] The spatial neighborhood block 420 can be examined to find one or more motion vectors (MVs) associated with the same reference frame index as the current block. The spatial neighborhood blocks can be examined for luma blocks, for example, according to the block set shown in FIG. 4, labeled according to the order of examination. That is, (1) the upper adjacent row is checked from left to right. (2) The left adjacent column is checked from top to bottom. (3) The upper right neighborhood block is checked. (4) The upper left block neighborhood block is checked. (5) The first upper non-adjacent row is checked from left to right. (6) The first left non-adjacent column is checked from top to bottom. (7) The second non-adjacent row from the top is checked from left to right. (8) The second non-adjacent column from the left is checked from top to bottom.
[0038] Candidates for "adjacent" spatial MV predictors derived from "adjacent" blocks (i.e., blocks of block sets 1-3) can be placed in the MV predictor list before candidates for the temporal MV predictor of the Temporal Motion Vector Predictor (TMVP) further described herein, and candidates for "non-adjacent" spatial MV predictors derived from "non-adjacent" blocks (outer blocks, also known as blocks of block sets 4-8) are placed in the MV predictor list after candidates for the temporal MV predictor.
[0039] In one embodiment, each SMVP candidate has the same reference picture as the current block. For example, assume that the current block 410 has a single reference picture. If an MV candidate also has a single reference picture that is the same as the ref picture of the current block, this MV candidate can be placed in the MV predictor list. Similarly, if an MV candidate has multiple reference pictures and one of the reference pictures is the same as the reference picture of the current block, this MV candidate can be placed in the MV predictor list. However, if the current block 410 has multiple reference pictures, the MV candidate can be placed in the MV predictor list only when the MV candidate has the corresponding reference picture that is the same for each of those reference pictures of the current block 410.
[0040] FIG. 5 shows an exemplary operation of a set of temporal motion vector predictors (TMVPs) according to an exemplary embodiment. The TMVP can be derived using collocated blocks in a reference frame. To generate the TMVP, first, one or more MVs of one or more reference frames can be stored together with reference indices associated with the respective reference frames. Then, for each 8×8 block of the current frame, the MVs of the reference frames through which the trajectory passes the 8×8 block are identified and can be stored in a temporal MV buffer together with the reference frame indices. In the case of inter prediction using a single reference frame, the MVs can be stored in 8×8 units to perform temporal motion vector prediction of future frames regardless of whether the reference frame is a "forward" reference frame or a "backward" reference frame (i.e., whether it is later or earlier in the frame sequence than the current frame, respectively). In the case of composite inter prediction, the MVs of the "forward" reference frame can be stored in 8×8 units to perform temporal motion vector prediction of future frames.
[0041] Exemplary embodiments for the process of generating TMVP may follow the following operations. In this example, the reference motion vector 550 (also labeled as MVref) of the initial reference frame 510 points from the initial reference frame 510 to a subsequent reference frame 540 which itself is a reference frame of the initial reference frame 510. In doing so, it passes through an 8×8 block 570 (shaded with gray dots) of the current frame 520. MVref 550 may be stored in a temporal MV buffer associated with this current block 570. During the motion projection process for deriving the temporal MV predictor 500, subsequent reference frames (e.g., frames 530 and 540) may be scanned in a predefined order. For example, using the frame labels defined by the AV1 standard, the scan order may be LAST_FRAME, BWDREF_FRAME, ALTREF_FRAME, ALTREF2_FRAME, and LAST2_FRAME. In one embodiment, MVs from higher-indexed reference frames (in the scan order) do not replace previously identified MVs assigned by lower-indexed reference frames (in the scan order).
[0042] Finally, given a predetermined block coordinate, the relevant MV stored in the temporal MV buffer is identified and projected onto the current block 570 to derive a temporal MV predictor 560 (also labeled as MV0) that points from the current block 570 to an adjacent reference frame 530.
[0043] FIG. 6 shows an exemplary set of predefined block positions 600 for deriving a temporal motion predictor for a 16×16 block according to an exemplary embodiment. Up to seven blocks can be checked for valid temporal MV predictors. In FIG. 6, the blocks are labeled B0 to B6. As described with reference to FIG. 4, candidates for the temporal MV predictor are checked after candidates for adjacent spatial MV predictors but before candidates for non-adjacent spatial MV predictors and can be put into the first MVP list. Then, for the derivation of the motion vector predictor (MVP), all spatial and temporal MVP candidates can be pooled, and each candidate can be assigned a weight that is determined during the scan of spatial and temporally adjacent blocks. Based on the associated weights, the candidates can be sorted and ranked, and the top four candidates can be identified and put into the second MVP list. This second list of MVPs is also referred to as the dynamic reference list (DRL) and can be further used in the dynamic MV prediction mode.
[0044] If the DRL is not full, additional searches can be performed, and the resulting additional MVP candidates are used to fill the DRL. The additional MVP candidates can include, for example, global MV, zero MV, combined composite MV without scaling, etc. The adjacent SMVP candidates, TMVP candidates, and non-adjacent SMVP candidates in the DRL can then be rearranged again. Both AV1 and AVM can enable rearrangement, for example, based on the weight of each candidate. The weight of a candidate may be predefined according to the overlapping area between the current block and the candidate block.
[0045] FIG. 7 shows an exemplary operation of generating a new MV candidate via a single inter-prediction block. When the reference frame of a neighboring block is different from that of the current block but the MVs are in the same direction, a temporal scaling algorithm can be utilized to scale the MV to its reference frame in order to form an MVP for the motion vector of the current block. In the example of FIG. 7, the motion vector 740 (also labeled mv1 in FIG. 7) from the neighboring block 750 of the current block 710 in the current frame 701 points to the collocated neighboring block 760 in the reference frame 703. The motion vector 740 can be used to derive an MVP for the motion vector 730 (also labeled mv0 in FIG. 7) of the current block 710 that points to the collocated current block 720 in another reference frame 702 using temporal scaling.
[0046] FIG. 8 shows an exemplary operation of generating a new MV candidate via a composite prediction block. In the example of FIG. 8, the composite MVs 860, 870 point to the reference frames 803 and 801 from different neighboring blocks 820, 830 of the current block 810 in the current frame 802, respectively. The reference frames 803 and 801 of the composed MVs 860, 870 (also labeled mv2 and mv3 in FIG. 8) can be the same as those for the current block 810. Composite inter-prediction can derive an MVP for the composed MVs 840, 850 (also labeled mv0 and mv1 in FIG. 8) of the current block 710, which can be determined as in FIG. 7.
[0047] FIG. 9 shows an exemplary operation of updating a motion vector candidate bank 920. This bank 920 was first proposed in CWG-B023 which is hereby incorporated by reference in its entirety.
[0048] The bank update process may be based on superblock 910. That is, after each superblock (e.g., superblock 910a) is coded, a set of first candidate MVs (e.g., the first 64 such candidates) used by each coding block within the superblock may be added to bank 920. Pruning may also be included during the update.
[0049] After the reference MV candidate scan is complete for a superblock, if there are open slots in the candidate list, the codec may refer to the MV candidate bank 920 for additional MV candidates (in the buffer with the matching reference frame type). MVs within the bank buffer may be added to the candidate list from the end of the buffer towards the beginning if they do not already exist in the list. More specifically, each buffer may correspond to a unique reference frame type covering single inter mode and combined inter mode respectively, corresponding to a single reference frame or a pair of reference frames. All buffers may be of the same size. When a new MV is added to a full buffer, existing MVs may be pushed out to make space for the new MV.
[0050] The coding block may refer to the MV candidate bank 920 to collect reference MV candidates in addition to those obtained in the AV1 reference MV list generation. After coding a superblock, the MV bank may be updated with the MVs used by the coding blocks of the superblock.
[0051] AV1 enables splitting frames into tiles, where each tile contains multiple superblocks. Each tile can be processed in parallel on different processors. Regarding the candidate bank, each tile can have an independent MV candidate bank utilized by all the superblocks within the tile. At the start of encoding each tile, the corresponding bank is emptied. Then, during the encoding of each superblock within that tile, MVs from the bank can be used as MV reference candidates. After encoding each superblock, the bank can be updated as described above.
[0052] Specific embodiments of the bank update and reference processes for bank update and reference are described later in this specification.
[0053] Figure 10 is a flowchart showing the process of constructing a motion vector prediction list for any video input according to an exemplary embodiment. Adjacent SMVP candidates, TMVP candidates, and non - adjacent SMVP candidates can be generated in S1010, S1020, and S1030 respectively, by the processes described previously with reference to, for example, FIGS. 4 and 5. Next, in S1040, the candidates can be sorted or otherwise rearranged by the process described previously with reference to, for example, FIG. 6. Additional MVP candidates can be derived in S1050 by the processes discussed previously with reference to, for example, FIGS. 7 and 8. Optionally, additional MVP candidates can be determined by additional searches in S1060 by the processes discussed previously with reference to, for example, FIG. 6, or retrieved from the reference bank in S1070 by the processes discussed previously with reference to, for example, FIG. 9.
[0054] Figure 11 shows an exemplary operation of the composite inter - prediction mode according to an exemplary embodiment.
[0055] The composite inter mode can create a prediction of a block by combining hypotheses from multiple different reference frames. In the example of FIG. 11, for example, block 1111 of the current frame 1110 is predicted by the motion vectors 1130a, 1130b (also labeled mv0, mv1 in FIG. 11) of the neighboring reference frames 1120a, 1120b. The neighboring reference frames 1120a, 1120b can be the most recent neighboring frames (i.e., the frames immediately before and after the current frame 1110 in the sequence), but this is not a requirement. The motion information component for each block (e.g., motion vectors 1130a, 1130b) can be transmitted in the bitstream as overhead.
[0056] However, although motion vectors can usually be well predicted using predictors from spatial and temporal neighborhoods or historical motion vectors, the bytes used for motion information can still be very important for many contents and applications.
[0057] FIG. 12 shows an exemplary operation of the temporal interpolation prediction (TIP) mode according to an exemplary embodiment.
[0058] In the example of FIG. 12, the information in the reference frames 1220a, 1220b is combined using a simple interpolation process and projected to the same time instance as the current frame 1210. Multiple TIP modes can be supported. In one TIP mode, the interpolated frame or "TIP frame" 1210' can be used as an additional reference frame. The coding block of the current frame 1210 can directly refer to the TIP frame 1210' and utilize information coming from two different references with only the overhead cost of a single inter prediction mode. In another TIP mode, the TIP frame 1210' can be directly assigned as the output of the decoding process for the current frame 1210 while skipping any other conventional coding steps. This mode can provide significant coding and simplification benefits, especially for low bitrate applications.
[0059] There are existing techniques for interpolating frames between two reference frames, such as frame rate up-conversion (FRUC), but achieving a good trade-off between complexity and compression quality can be a significant constraint when designing new coding tools. The method disclosed above is simple and reuses the motion information already available in the reference frame without the need to perform additional motion searches. Simulation results show that this simple method can achieve good quality in a low-complexity implementation form.
[0060] In the example of FIG. 12, the TIP mode operation starts by generating a TIP frame 1210' corresponding to the current frame 1210. The TIP frame 1210' can then be used as an additional reference frame for the current frame 1210 or directly assigned as the reconstructed output of the decoder for the current frame 1210. On the decoder side, blocks coded in the TIP mode can be generated on-the-fly to save decoding time and processing without the need to create the entire TIP frame 1210' in the decoder. This is also compatible with the one-pass decoding pipeline in the decoder and is suitable for hardware implementation forms.
[0061] The frame-level TIP mode can be indicated using syntax elements. Examples of the modes indicated by the value of the tip_frame_mode parameter are shown in the following table.
[0062] [Table 1]
[0063] A simple interpolation method for interpolating intermediate frames between two frames is disclosed, which can fully reuse motion vectors from available references. The same motion vectors can also be used for the Temporal Motion Vector Predictor (TMVP) process after minor modifications. This process may include three operations. 1. Create a rough motion vector field for the TIP frame through projection of the modified TMVP field. 2. Refine the rough motion vector field by filling holes and using smoothing operations. 3. Generate the TIP frame using the refined motion vector field. On the decoder side, blocks coded in the TIP mode can be generated on-the-fly without creating the entire TIP frame.
[0064] However, it should be noted that other suitable interpolation methods may be substituted in combination with other features described in this disclosure, and such are within the scope of this disclosure.
[0065] FIG. 13 shows an exemplary operation of bilateral matching-based decoder-side motion vector refinement according to an exemplary embodiment. Versatile Video Coding (VVC) can distribute previously decoded pictures to two reference picture lists 1320a, 1320b. These previously decoded pictures can be used as references for predicting the current picture 1310. In the example of FIG. 13, according to the display order, reference pictures before the current picture 1310 can be assigned to the "past" reference picture list 1320a, and reference pictures after the current picture 1320 can be assigned to the "future" reference picture list 1320b. Corresponding reference picture indices (not shown) of each list indicate which picture in each list is used to predict the current block 1311 of the current picture 1310. In the case of bidirectional prediction, two predicted blocks 1321a and 1321b predicted using respective MVs 1331a, 1331b for the past reference picture list 1320a and the future reference picture list 1320b can be combined to obtain a single prediction signal.
[0066] When motion information is coded by the merge mode, the reference picture index and MV of neighboring blocks can be directly applied to the current block 1311. However, this may not accurately predict the current block 1311.
[0067] The decoder-side motion vector refinement (DMVR) algorithm can be used to increase the accuracy of blocks coded in the merge mode by including only decoder-side information. When the DMVR algorithm is applied to blocks 1311, 1321a, and 1321b, the MVs 1331a, 1331b derived from the merge mode can be set as the "initial" MVs for DMVR.
[0068] DMVR can then further refine the initial MVs 1331a, 1331b by block matching. In both reference pictures, candidate blocks surrounding the blocks 1321a, 1321b pointed to by the initial MVs can be searched for bilateral matching. The best matching blocks 1323a, 1323b can be used to generate the final prediction signal, and the new MVs 1333a, 1333b indicating these new prediction blocks 1323a, 1323b can be set as the "refined" MVs corresponding to the initial MVs 1331a, 1331b respectively. Many block matching methods suitable for DMVR have been studied, such as template matching, methods based on bidirectional template matching, and methods based on bilateral matching adopted in VVC.
[0069] In the two-sided matching-based DMVR, the block pairs 1321a and 1321b pointed to by the initial MV can be defined as the initial block pairs. The distortion costs of the initial block pairs 1321a and 1321b can be calculated as the initial costs. The blocks surrounding the initial block pairs 1321a and 1321b can be used as DMVR candidate block pairs. Each block pair can include one predicted block from a reference picture in the past reference picture list 1320a and one predicted block from a reference picture in the future reference picture list 1320b.
[0070] The distortion costs of the DMVR candidate block pairs can be measured and compared. Since the DMVR candidate block pair with the lowest distortion cost includes the two most similar blocks between the reference pictures, this block pair (i.e., blocks 1323a and 1323b) can be assumed to be the best predictor of the current block 1311. Therefore, the final dual prediction signal can be generated using the block pair 1323a and 1323b. The corresponding MVs 1333a and 1333b can be shown as the refined MVs. If all the DMVR candidate block pairs have a distortion cost greater than that of the initial block pairs 1321a and 1321b, the initial blocks 1321a and 1321b can be used for dual prediction, and the refined MVs 1333a and 1333b can be set equal to the initial MVs 1331a and 1331b.
[0071] To simplify the distortion cost calculation, the sum of absolute differences (SAD) can be used as the distortion metric, and only the luma distortion can be considered in the DMVR search process. Note that to further reduce the computational complexity, the SAD can be evaluated between the even rows of the candidate block pairs.
[0072] In the example of FIG. 13, the dotted blocks (1321a, 1321b) within each reference picture indicate the initial block pairs. The gray blocks (1323a, 1323b) indicate the best matching block pairs, which may be the block pairs having the lowest SAD cost compared to the other DMVR candidate block pairs and the initial block pairs 1321a, 1321b. The initial MVs 1331a, 1331b can be refined to generate the refined MVs 1333a, 1333b, and the final dual prediction signal can be generated using the best matching block pairs 1323a, 1323b. Note that the initial MVs 1331a, 1331b need not refer to full sample positions since it can be derived from the merge mode, thereby supporting up to 1 / 16 fractional sample MV accuracy.
[0073] Since the difference between the refined MV and the corresponding initial MV (shown in FIG. 13 as ΔMV 1335a, 1335b) can be an integer or a fraction, the refined MV can refer to a fractional pixel position. In this case, the intermediate search block and the final prediction block may be generated by the DMVR interpolation process.
[0074] In some embodiments, the block-level bilateral matching-based DMVR can be performed on the motion field generated by TMVP. Next, an example of such a process will be described with reference to the concepts previously described herein.
[0075] The process may start with the motion field being generated as part of the TIP for each 8×8 block. The motion field is a representation of 3D motion projected onto a 2D space such as a picture, and is generally defined by one or more motion vectors each describing the movement of a corresponding point. Here, the motion field may include two motion vectors (MV0 and MV1) pointing to two reference pictures. The motion vectors (MV0 and MV1) may be used as the starting point of the DMVR process. More specifically, corresponding predictors in the reference pictures indicated by the motion vectors may be generated. In this operation, the input can be filtered using filters such as interpolation, bilinear. Then, candidate predictors surrounding the motion vectors may be generated. These predictors may be searched through a predefined search range N which is an integer value corresponding to the number of luma samples. The search accuracy is defined as K, and K can be a fraction from 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 up to the number of luma samples (up to the highest supported MV accuracy). In the next operation, bilateral matching can be performed between all candidate predictors, and the position of the predictor containing the lowest distortion cost can be determined to be the refined position for this 8×8 block. The distortion cost can be, but is not limited to, SAD, SATD, SSE, subsampled SAD, mean removed SAD, etc.
[0076] After the refined positions (refined motion vectors) for each 8×8 block are obtained, the TIP process can be performed. More specifically, a TIP frame can be generated using the DMVR refined motion vector field. The generated frame can be used as a reference for prediction or directly used as a prediction.
[0077] On the decoder side, when a block is coded as TIP or by the TIP mode, the TIP predictor and DMVR refinement can be performed on-the-fly without generating the entire frame. In some embodiments, sub-block level bilateral matching-based DMVR can be performed on the motion field generated by TMVP. More specifically, for each 8×8 TIP block set, additional splitting can be performed. Such additional splitting can result in four 4×4 sub-blocks from each 8×8 block. Each sub-block can perform bilateral matching-based DMVR search to obtain a refined motion field for TIP. Further, sub-block level bilateral matching-based DMVR can be performed in addition to the motion field generated by TMVP and optical flow refinement. More specifically, for each 8×8 TIP block, further splitting can be performed. For example, each 8×8 TIP block can be split into four 4×4 sub-blocks, where optical flow refinement is first applied to refine the motion vectors, and then bilateral matching-based DMVR search is further applied to refine the motion field for TIP.
[0078] In some embodiments, sub-block level bilateral matching based DMVR can be performed in addition to TMVP generated motion fields and optical flow refinement. For example, for each 8×8 TIP block, additional splitting operations can generate four 4×4 sub-blocks, bilateral matching based DMVR refinement is applied to refine the motion vectors, and then optical flow refinement is further applied to refine the motion field for TIP. In some embodiments, multi-stage DMVR can be used to refine the TMVP generated TIP motion field. For example, a first block level DMVR can be used to refine the initially generated motion field. The refined MV can be used as the starting point for the second stage. In the second stage, sub-block level DMVR can be performed to further refine the motion field. Additional such stages are within the scope of the present disclosure.
[0079] In other embodiments, the TIP motion field can use explicitly signaled MV differences and / or corrections. Starting from any level, such as a group of coding blocks, a coding block, or sub-block level, one or more motion vector differences (MVDs) can be signaled in the bitstream. The bitstream can be parsed by the decoder and used as a correction to the TIP motion field. When a block is encoded as TIP mode, the corresponding motion field for the block can be generated using a TMVP based method. Next, if the block is 8×8 or smaller, the parsed MVD can be added to the motion field so that the MV of the block can be corrected by the parsed MVD. If the block is larger than 8×8, each MV of each 8×8 sub-block can be added to the parsed MVD.
[0080] In some embodiments, when TIP is applied using two reference pictures for motion compensation, an MVD may be signaled to correct the motion field associated with the selected reference picture. For example, the MVD may be signaled for the future reference picture list, but may not be signaled for the past reference picture list, or vice versa. The selection of which reference picture requires the additional MVD to be signaled may be further signaled or implicitly derived.
[0081] FIG. 14 shows an exemplary usage case of the merge mode using motion vector difference (MMVD) according to an exemplary embodiment. The merge mode can generally be used with implicitly derived motion information to predict samples generated by the current coding unit (CU). The merge mode using motion vector difference may use a flag to signal that the MMVD is used for the CU. The MMVD flag may be sent after the skip flag is sent. In MMVD, after a merge candidate is selected, it can be further refined by the signaled MVD information. The additional information may include a merge candidate flag, an index for specifying the magnitude of the motion, and an index for indicating the direction of the motion. In the MMVD mode, one of the first two candidates in the merge list can be selected for use as the MV base. The merge candidate flag can signal which candidate should be used.
[0082] This operation may use a distance index that specifies the magnitude of the motion information and indicates a predefined offset from the starting point. The offset can be added to either the horizontal or vertical component of the starting MV. An exemplary relationship between the distance index and the predefined offset is specified in Table 2.
[0083] [Table 2]
[0084] The direction index may represent the direction of the MVD relative to the starting point. The direction index can represent one of four directions, as shown in Table 3.
[0085]
Table 3
[0086] The meaning of the MVD sign may vary according to the information of the starting MV. When the starting MV is a single-predicted MV, or both reference picture lists point to the same side of the current picture (i.e., the picture order count (POC) of both references is greater than the POC of the current picture, or both are less than the POC of the current picture), the sign in Table 3 may specify the sign of the MV offset added to the starting MV. When the starting MV is a bi-predicted MV with two MVs pointing to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture), and the difference in POC in the first reference picture list is greater than the difference in POC in the second reference picture list, the sign in Table 3 may specify the sign of the MV offset added to the first list MV component for the starting MV, and the sign for the second list MV may have the opposite value. Otherwise, if the difference in POC in the second list is greater than that in the first list, the sign in Table 3 may specify the sign of the MV offset added to the second list MV component for the starting MV, and the sign for the first list MV may have the opposite value.
[0087] The MVD may be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, the scaling may be omitted. Otherwise, if the difference in POC in one list is greater than that in the other list, the MVD for the list with the smaller POC difference may be scaled. When the starting MV is single-predicted, the MVD may be added to the available MV.
[0088] In addition to the unidirectional prediction and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional MVD signaling may also be applied. In the symmetric MVD mode, motion information including the reference picture indexes of both reference picture lists and the MVD of the future reference picture list is not signaled but derived.
[0089] In certain implementations, the decoding process of the symmetric MVD mode may be as follows.
[0090] At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 may be derived as follows. If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the closest reference picture in the past reference picture list L0 and the closest reference picture in the future reference picture list L1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, BiDirPredFlag is set to 1 and both the L0 reference picture and the L1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.
[0091] At the CU level, if the CU is bi-predicted coded and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used may be explicitly signaled. When the symmetric mode flag is true, mvp_l0_flag, mvp_l1_flag, and MVD0 may be explicitly signaled and other signals may be omitted. The reference indexes for L0 and L1 may be set equal to the pair of reference pictures respectively, and MVD1 may be set equal to (-MVD0).
[0092] In some embodiments, for each coding block in an inter-frame, when the mode of the current block is an inter-coding mode rather than a skip mode, another flag may be signaled to indicate whether a single-reference mode or a composite-reference mode is used for the current block. A predicted block may be generated by one motion vector in the single-reference mode, or may be generated by weighted-averaging two predicted blocks derived from two motion vectors in the composite-reference mode.
[0093] In the case of the single-reference mode, the following specific modes may be signaled according to the syntax of an exemplary implementation.
[0094] Use one of the motion vector predictors (MVPs) in the list indicated by the NEARMV-DRL (Dynamic Reference List) index.
[0095] Use one of the motion vector predictors (MVPs) in the list signaled by the NEWMV-DRL index as a reference and apply a difference to the MVP.
[0096] Use a motion vector based on global motion parameters at the globalmv-frame level.
[0097] In the case of the composite-reference mode, the following specific modes may be signaled according to the syntax of an exemplary implementation.
[0098] Use one of the motion vector predictors (MVPs) in the list signaled by the NEAR_NEARMV-DRL index.
[0099] Use one of the motion vector predictors (MVPs) in the list signaled by the NEAR_NEWMV-DRL index as a reference and send a ΔMV for the second MV.
[0100] Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEARMV-DRL index as a reference and send a ΔMV for the first MV.
[0101] Use one of the motion vector predictors (MVPs) in the list signaled by the NEW_NEWMV-DRL index as a reference and send a ΔMV for both MVs.
[0102] GLOBAL_GLOBALMV - Use the MVs from each reference based on their frame-level global motion parameters.
[0103] In some embodiments, the operation may enable 1 / 8 pixel motion vector accuracy (or precision), and in an exemplary implementation, the following syntax may be used to signal the motion vector differences in L0 or L1.
[0104] mv_joint specifies which components of the motion vector difference are non-zero.
[0105] 0 indicates that there is no non-zero MVD along either the horizontal or vertical direction.
[0106] 1 indicates that there is a non-zero MVD only along the horizontal direction.
[0107] Figure 2 indicates that there is a non-zero MVD only along the vertical direction.
[0108] Figure 3 indicates that there are non-zero MVDs along both the horizontal and vertical directions.
[0109] mv_sign specifies whether the motion vector difference is positive or negative.
[0110] The mv_class specifies the class of the motion vector difference. As shown in Table 4, a higher class may indicate that the motion vector difference has a larger magnitude.
[0111]
Table 4
[0112] The mv_bit specifies the integer part of the offset between the motion vector difference and the starting magnitude of each MV class.
[0113] The mv_fr specifies the first two fractional bits of the motion vector difference.
[0114] The mv_hp specifies the third fractional bit of the motion vector difference.
[0115] In the case of the NEW_NEARMV and NEAR_NEWMV modes, the accuracy of the MVD may depend on the associated class and the magnitude of the MVD. For example, a fractional MVD may be allowed only if the magnitude of the MVD is 1 pixel or less. Further, when the value of the associated MV class is MV_CLASS_1 or higher, only one MVD value may be allowed, and the MVD values in each MV class are derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5).
[0116] Table 5 shows the allowable MVD values in each MV class according to the above embodiment.
[0117]
Table 5
[0118] Furthermore, if the current block is coded in NEW_NEARMV or NEAR_NEWMV mode, one context may be used to signal mv_joint or mv_class. Otherwise, other contexts may be used to signal mv_joint or mv_class.
[0119] A new inter-coding mode called JOINT_NEWMV may be applied to indicate whether MVDs for two reference lists are jointly signaled. If the inter-prediction mode is equal to the JOINT_NEWMV mode, MVDs for L0 and L1 may be jointly signaled. More specifically, only one MVD named joint_mvd may be signaled and sent to the decoder, and the delta MVs for L0 and L1 may be derived from joint_mvd.
[0120] The JOINT_NEWMV mode may be signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. There is no need to add additional context. When the JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD may be scaled for L0 or L1 based on the POC distance. Let td0 be the POC distance between L0 and the current frame, and td1 be the POC distance between L1 and the current frame. If td0 is greater than or equal to td1, joint_mvd may be directly used for L0, and the mvd for L1 may be derived from joint_mvd based on Equation (1).
Number
[0121] Otherwise, if td1 is greater than or equal to td0, joint_mvd may be directly used for L1, and the mvd for L0 may be derived from joint_mvd based on Equation (2). [Number]
[0122] (If td0 and td1 are equal, according to any of the above equations, derived_mvd = joint_mvd, and thus joint_mvd can be directly used as the derived MVD for both L0 and L1, and in that case, it will be clear that no scaling is performed.)
[0123] Here, an inter - coding mode called the AMVDMV mode can be made available for the single - reference case. In the AMVDMV mode, the adaptive MVD (AMVD) resolution is applied to the signal MVD.
[0124] To indicate whether AMVD is applied to the joint MVD coding mode, a flag (here labeled amvd_flag) can be added under the JOINT_NEWMV mode, which can be referred to as joint AMVD coding. In joint AMVD coding, the MVDs for two reference frames can be jointly signaled, and the accuracy of the MVD can be implicitly determined by the size of the MVD. Otherwise, the MVDs for two (or more than three) reference frames can be signaled together, and MVD coding can be applied.
[0125] The adaptive motion vector resolution (AMVR), first proposed in CWG - C012 which is incorporated herein in its entirety, supports seven MV accuracy values (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8). For each prediction block, the adaptive motion vector (AVM) encoder can explore all the supported accuracy values and signal the best accuracy to the decoder.
[0126] To reduce the encoder execution time, two precision sets can be supported. Each precision set can include four predefined precisions. The precision set can be adaptively selected at the frame level based on the maximum precision value of the frame. Similar to the AV1 standard, the maximum precision can be signaled in the frame header. The following table summarizes the precision values supported according to the frame-level maximum precision.
[0127]
Table 6
[0128] The AOMedia AVM repository related to AV1 provides a frame-level flag indicating whether the MV of the frame includes sub-pel precision. In certain embodiments, the AMVR can be enabled only when the value of the cur_frame_force_integer_mv flag is 0. When the precision of a block is lower than the maximum precision, the motion model and interpolation filter may not be signaled and may remain inactive. When the precision of a block is lower than the maximum precision, the applicable motion model can be inferred as a translational motion model, and the applicable interpolation filter can be inferred as a "normal" filter. When the precision of a block is either 4 pels or 8 pels, the inter-intra mode may not be signaled and can be inferred as 0.
[0129] FIG. 15 shows an exemplary operation of geometric partitioning mode (GPM) prediction according to an exemplary embodiment. This operation focuses on an inter-picture prediction coding unit (CU). When GPM is applied to the current CU 1510, the current CU 1510 can be divided into two parts 1510a and 1510b by a partitioning boundary. The position of the partitioning boundary can be mathematically defined by an angle parameter φ and an offset parameter ρ. These parameters can be quantized and combined with a GPM partition index lookup table. The GPM partition index of the current CU 1510 can be coded in the bitstream. In total, 64 partitioning modes can be used for a CU 1510 with a size of w×h = 2k×2l (for luma samples), where k, l ∈ {3…6}. Since narrow CUs usually do not contain geometrically separated patterns, the application of GPM can be disabled on CUs 1510 having an aspect ratio greater than 4:1 or less than 1:4.
[0130] The two GPM partitions contain individual motion information that can be used to predict corresponding parts in the current CU 1510. One-directional motion compensation prediction (MCP) can be applied to each CU part 1510a, 1510b so that the required memory bandwidth of the MCP in GPM is equal to the memory bandwidth of the normal bidirectional MCP. To simplify motion information coding and reduce possible combinations for GPM, the motion information can be coded using a merge mode. The GPM merge candidate list can be derived from the conventional merge candidate list to ensure that only one-directional motion information is included.
[0131] In the example of FIG. 15, the right part 1510a of the current CU 1510 is predicted by a first motion vector 1520a (also labeled as MV0 in FIG. 15) from a first reference picture 1530a (also labeled as P0 in FIG. 15), and the left part 1510b is predicted by a second motion vector 1520b (also labeled as MV1 in FIG. 15) from a second reference picture 1530b (also labeled as P1 in FIG. 15).
[0132] Once each part of the CU is predicted, a prediction for the complete CU can be generated by a blending process.
[0133] FIG. 16 shows an exemplary operation of blending in GPM prediction according to an exemplary embodiment. The blending mask can take the form of matrices such as matrices 1600a and 1600b for application to each predicted part of the CU. In the example of FIG. 16, matrices 1600a and 1600b each include weights within the value range of 0 to 8. That is, if the first and second matrices 1600a and 1600b are denoted as W0 and W1 respectively, and the matrix with size w×h is denoted as J, then W0 + W1 = 8J. The weights of the blending matrix can depend on the displacement between the sample location and the partition boundary. Since the computational complexity of deriving the blending matrix is extremely low, these matrices can be generated on-the-fly at the decoder side.
[0134] By applying the matrices, a prediction for the complete CU can be determined based on Equation (3).
Equation
[0135] Here, W0 and W1 represent the first and second matrices 1600a and 1600b respectively, P0 and P1 represent the first and second reference pictures 1530a and 1530b respectively, and PG represents the generated prediction.
[0136] Next, the generated prediction can be subtracted from the original signal to generate a residual. The residual can be transformed, quantized, and coded into a bitstream using, for example, a VVC transform, quantization, and entropy coding engine, or other suitable coding engine. On the decoder side, the signal can be reconstructed by adding the residual to the generated prediction. If the residual is negligible, a "skip mode" can be applied, the residual is dropped by the encoder, and the generated prediction is directly used by the decoder as the reconstructed signal.
[0137] FIG. 17 shows an exemplary codebook for wedge-based prediction in a special composite prediction mode according to an exemplary embodiment. Wedge-based prediction can be implemented in AV1 and can be used for both inter-inter and inter-intra connections.
[0138] In composite wedge prediction, the boundaries of moving objects are often difficult to approximate by on-grid block partitioning. Thus, in some embodiments, when it is selected that the coding unit is further partitioned in such a way, a predefined codebook of 16 possible wedge partitions can be used to signal the wedge index in the bitstream. A 16-value shape codebook can be designed that includes partition orientations that are either horizontal, vertical, or diagonal with slopes of ±2 or ±0.5. In the example of FIG. 17, two codebooks 1710 and 1720 are designed for square and rectangular blocks, respectively.
[0139] To reduce spurious high-frequency components that are often generated by directly juxtaposing two predictors, a soft cliff-shaped 2D wedge mask can be used to smooth the edges around the intended partition. For example, m(i,j) can approach 0.5 around the edge and gradually transform to binary weights at both ends.
[0140] The foregoing blending may utilize a threshold θ that defines a blending interval around the segmentation boundary. A mask can be applied within this interval to generate a blend region. The mask can be defined, for example, by Equation (4) using a ramp function according to the weight of each position (x_c, y_c) having a distance d(x_c, y_c) from the segmentation boundary, and the regions can be blended accordingly.
Number
[0141] Using a fixed threshold θ may not be optimal because a fixed blend region width does not always provide the best blend quality for various types of video content. For example, screen video content typically contains strong textures and sharp edges, which indicates a narrow blending area (i.e., a small threshold) to preserve edge information. In the case of camera-captured content, blending is generally required, but the blending area width depends on several factors, such as the actual boundary of the moving object and the motion discriminability between the two segments. In addition, different CU parts may have different threshold requirements.
[0142] FIG. 18 shows an exemplary operation of asymmetric blending generation according to an exemplary embodiment. The embodiments of the asymmetric blending mask described herein can be applied to geometric segmentation mode prediction in VVC, wedge-based prediction in AV1, or any other similar encoding format and / or technique.
[0143] In the example of FIG. 18, a first threshold θ1 and a second threshold θ2 are defined, where θ1 has a valid negative value reflecting the distance in one direction from the segmentation boundary B, and θ2 has a valid positive value reflecting the distance in the other direction from B. Then, the weight of a specific position can be calculated from the threshold, for example, by Equation (5).
Number
[0144] As described in exemplary formula (4), when the displacement d(x_c, y_c) from the position (x_c, y_c) to the section boundary B is θ1 or less, the position is outside the threshold θ1 with respect to B, and a weight of 0 is used. When d(x_c, y_c) is θ2 or more, the position is outside the threshold θ2 with respect to B, and full weighting (e.g., 8 in this example) is used. When d(x_c, y_c) is between θ1 and θ2, a ramp weighting value between 0 and 8 is used.
[0145] Other equations having other suitable weight values and ramp formulas can be determined empirically, qualitatively, or arbitrarily.
[0146] Note that when the absolute values of θ1 and θ2 are equal, the blending operates similarly to symmetric adaptive blending, and when θ1 and θ2 are not equal, asymmetric adaptive blending occurs.
[0147] In certain embodiments, the blending mask can be calculated based on a wedge-based prediction design that uses two thresholds. In these embodiments, the mask weighting near the section boundary B is equal to half the value (e.g., 32) and gradually converts to binary weighting (e.g., 0 and 64) at either extreme. The gradient may be based on a predetermined threshold in such embodiments, which changes the mask such that, for example, the larger the threshold, the less sharp the conversion on the mask.
[0148] Partial selections for corresponding to different blending thresholds may be signaled explicitly. For example, a binary partial selection flag may signal one of two possible assignments, namely, a first assignment where the first side of the segmentation boundary corresponding to the first CU part is assigned a threshold θ1 (which may be a smaller threshold and result in sharper blending), and the second side of the segmentation boundary corresponding to the second CU part is assigned a threshold θ2 (which may be a larger threshold and result in more blunt or softer blending), and a second assignment which is the reverse of the first assignment.
[0149] Partial selections for corresponding to different blend thresholds may also be derived implicitly by a predetermined method. The selection may be made by selecting various angles, offsets, wedge indices, or any other parameters. Such other parameters may include the magnitude and direction of the corresponding motion vectors of each part, the type of prediction mode of each part, or be based on the reconstructed samples in the neighborhood.
[0150] FIG. 19 shows an exemplary operation of adaptive threshold selection for an asymmetric / symmetric blending mask according to an exemplary embodiment. In this embodiment, two thresholds θ1 and θ2 are used to generate a blending mask 1900. The thresholds may be corresponding indices, signaled values, or implicitly derived. These thresholds may be the same or identical values, or different values. The interval 1910 between the thresholds may be adaptively shifted so that θ1 and θ2 may be the same or similar, or either may be larger than the other to any desired extent.
[0151] Thresholds can be signaled separately and can each have their own syntax elements in the bitstream model and context model. Alternatively, the thresholds can be signaled differentially such that θ1 and (θ2 - θ1) are signaled, or θ2 and (θ1 - θ2) are signaled, and then the remaining thresholds can be derived. Alternatively, the thresholds can have a predefined ratio, such as θ1:θ2 = 1:2, such that only θ1 (or θ2) needs to be signaled.
[0152] Thresholds may also be selected from a predetermined list. For example, a list such as {0.5, 1, 2, 4, 8} may be used as possible thresholds. Using the list, the index of the corresponding threshold can be signaled. At the decoder, based on the predefined list and the parsed index, the values of θ1 and θ2 can be obtained. Combinations of the values of thresholds θ1 and θ2 can alternatively be given for selection in a predefined list, such as {(1,1), (1,2), (2,1), (1,4), (4,1),...}, and an index from the predefined list for the selected combination can be signaled. In some examples, each threshold has its own predetermined list. As an example, the predefined list for θ1 can be {0.5, 1, 2, 4, 8}, and the predefined list for θ2 can be {0.25, 0.5, 1, 2, 4}. Individual indices for each threshold can be signaled.
[0153] When one or more predefined lists for thresholds are used, a subset of each predefined list can be used more specifically. Further, a predetermined threshold may be used for each block of all threshold candidates. The subset of each predefined list can be determined by coding information that may exist for both encoding and decoding of the current block. The coding information in the current block can include neighboring reconstructed samples, block size, prediction mode, or any other relevant information for generating a subset of the predefined threshold list.
[0154] In certain embodiments, the best candidate may be selected by template matching. The template may use the left upper surrounding samples of predictors from each reference frame and may be generated based on a predefined threshold. The generated template may be compared with the left upper surrounding samples of the current block. The candidate with the lowest distortion cost may be used in GPM or wedge-based prediction.
[0155] Candidates may also be sorted based on template matching, and the top N candidates following the lowest distortion may be used. The finally selected threshold may depend on the signaled / parsed index. The value of N may be predefined or signaled in the high-level syntax. Note that when N is equal to 1, the index is not used for signaling.
[0156] In certain embodiments, entropy coding of two thresholds may be performed using content derived from the coded information. The coding information may be the selected threshold from neighboring blocks.
[0157] According to the above disclosure, instead of a predetermined symmetric (i.e., one threshold) blending design, the design may have different blending thresholds around the segmentation boundary, e.g., two predetermined thresholds θ1 and θ2 as shown in FIG. 18. The blending mask or wedge weighting mask may be calculated in advance or on-the-fly based on these two thresholds. Based on the threshold definition of a particular codec, the threshold may be defined as a negative value to indicate displacement (as seen in the GPM of VVC) or as a positive value (as seen in the wedge-based prediction for AV1 and AV2).
[0158] The proposed methods may be used separately or combined in any order. Further, each of the methods (or embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0159] FIG. 20 is a schematic diagram of exemplary components of a device or system 2000 in which embodiments of the systems and / or methods described herein may be implemented. Exemplary system 2000 may be one of a variety of systems such as a personal computer, a mobile device, a computer cluster, a server, an embedded device, an ASIC, a microcontroller, or any other device capable of executing code. Bus 2010 connects exemplary system 2000 to each other so that all components can communicate with each other. Bus 2010 connects processor 2020, memory 2030, storage component 2040, input component 2050, output component 2060, and interface component.
[0160] Processor 2020 may be a single processor, a processor having multiple processors internally, a cluster of (two or more) processors, and / or distributed processing. The processor executes instructions stored in both memory 2030 and storage component 2040. Processor 2020 operates as a computing device and performs operations to modify a shared virtual engine-derived data cache. Memory 2030 is high-speed storage, and retrieval to any of the memory devices may be enabled through the use of cache memory that may be closely associated with one or more CPUs. Storage component 2040 may be one of any long-term storage such as an HDD, an SSD, a magnetic tape, or any other long-term storage format.
[0161] The input component 2050 may be any file type or signal from a user interface component such as an input capture device, such as a camera, a handheld controller, a game pad, a keyboard, a mouse, or a motion capture device. The output component 2060 outputs the processed information to the communication interface 2070. The communication interface may be another communication device, such as a speaker or a screen, that can display the information to a user or other observer, such as another computing system.
[0162] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.
[0163] Some embodiments can relate to systems, methods, and / or computer-readable media in any possible technical detail level of integration. Further, one or more of the components described above can be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or can include at least one processor). The computer-readable medium can include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions for causing a processor to perform operations.
[0164] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction-executing device. The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves in which instructions are recorded, and any appropriate combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0165] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area communication network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.
[0166] The computer-readable program code / instructions for performing the operations can be in any combination of one or more program languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object code such as Smalltalk, C++, and source code or object-oriented programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area communication network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to perform aspects or operations or to personalize the electronic circuit.
[0167] These computer-readable program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus implement the operations specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement the manner of operation specified in one or more blocks of the flowchart and / or block diagram.
[0168] These computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices such that the instructions executed on the computer, other programmable apparatus, or other devices implement the operations specified in one or more blocks of the flowchart and / or block diagram, thereby producing a computer implemented process.
[0169] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical operation(s). The methods, computer systems, and computer-readable media can include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in the figure. In some alternative implementations, the operations shown in the blocks can be performed in an order different from that shown in the figure. For example, two blocks shown in succession can actually be executed simultaneously or substantially simultaneously, or the blocks can sometimes be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified operation or action or a combination of dedicated hardware and computer instructions.
[0170] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. It is understood that the actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting of the implementation. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
Description of Reference Numerals
[0171] 100 AV1 Partition Tree, 110 Image, 115 Portion, 120 Way Structure, 125a Partition Pattern, 125b Partition Pattern, 125c Partition Pattern, 210 QTBT Structure, 211a Node, 211b Node, 211c Node, 211d Node, 211e Node, 220 CTU, 310 Vertical, 320 Horizontal, 410 Current Block, 420 Spatial Neighboring Block, 500 Temporal MV Predictor, 510 Initial Reference Frame, 520 Current Frame, 530 Frame, 540 Frame, 550 Reference Motion Vector, 570 8×8 Block, 701 Current Frame, 702 Reference Frame, 703 Reference Frame, 710 Current Block, 720 Current Block, 730 Motion Vector, 740 Motion Vector, 750 Neighboring Block, 760 Collocated Neighboring Block, 801 Reference Frame, 802 Reference Frame, 803 Reference Frame, 810 Current Block, 820 Neighboring Block, 830 Neighboring Block, 840 MV, 850 MV, 860 Composite MV, 870 Composite MV, 910 Super Block, 910a Super Block, 920 Motion Vector Candidate Bank, 1110 Current Frame, 1111 Block, 1120a Neighboring Reference Frame, 1120b Neighboring Reference Frame, 1130a Motion Vector, 1130b Motion Vector, 1210 Current Frame, 1210’ TIP Frame, 1220a Reference Frame, 1220b Reference Frame, 1310 Current Picture, 1311 Current Block, 1320a Reference Picture List, 1320b Reference Picture List, 1321a Predicted Block, 1321b Predicted Block, 1331a MV, 1331b MV, 1323a Best Matching Block, 1323b Best Matching Block, 1333a MV, 1333b MV, 1335a ΔMV, 1335b ΔMV, 1510 CU, 1510a CU Portion, 1510b CU Portion, 1520a First Motion Vector, 1520b Second Motion Vector, 1530a Reference Picture, 1530b Reference Picture, 1600a Matrix, 1600b Matrix, 1710 Codebook, 1720 Codebook, 1900 Blending Mask, 1910 Interval, 2000 System, 2010 Bus, 2020Processor, 2030 Memory, 2040 Memory Component, 2050 Input Component, 2060 Output Component, 2070 Communication Interface
Claims
1. A method for predicting a picture area in a decoding process, performed by at least one processor of a decoding device, said method comprising: receiving an input image comprising a picture area divided into at least a first portion and a second portion; identifying a partition boundary that divides the picture area into the first portion and the second portion; determining a first threshold value corresponding to the first portion and a second threshold value corresponding to the second portion; applying a first blending mask based on the first threshold to predicted pixels of the first portion to generate a first blending region, and applying a second blending mask based on the second threshold to predicted pixels of the second portion to generate a second blending region, wherein the first blending mask is different from the second blending mask; and reconstructing the input image including a prediction for the picture area including the first portion and the second portion modified by the first blending region and the second blending region.
2. the first threshold and the second threshold are each defined relative to the partition boundary; 2. The method of claim 1, wherein applying each of the first blending mask and the second blending mask comprises applying a weight to a predicted pixel at the location in the picture area based on a distance of the location from at least one of the first threshold and the second threshold.
3. 2. The method of claim 1, wherein the value of at least one of the first threshold and the second threshold is based on at least one consideration derived from the input image, the at least one consideration being based on at least one sample of the input image surrounding the picture area.
4. The method of claim 1 , wherein at least one of the first threshold and the second threshold is based on a candidate value having a lowest distortion cost among a plurality of candidate values.
5. 2. The method of claim 1, wherein each of the first threshold and the second threshold has a respective predefined list of a plurality of selectable thresholds, and each of the first threshold and the second threshold is determined based on a respective index in the respective list indicated in a signaled pair of indices.
6. 2. The method of claim 1, wherein the first threshold and the second threshold are determined based on a signaled index corresponding to a threshold combination in a predefined list of a plurality of selectable combinations of thresholds.
7. the values of the first threshold and the second threshold are determined according to a flag; when the flag is at a first logic level, the first threshold is set to a first value and the second threshold is set to a second value; 2. The method of claim 1, wherein if the flag is at a second logic level, the first threshold is set to the second value and the second threshold is set to the first value.
8. The method of claim 1 , wherein the partition boundaries are geometrically defined according to angle and offset parameters.
9. The method of claim 1 , wherein the partition boundaries are defined according to wedge partitions of a predefined set of wedge partitions.
10. when the first threshold and the second threshold are equal in value, the blending is symmetric adaptive blending; When the values of the first threshold and the second threshold are not equal, the blending is asymmetric adaptive blending. The method of claim 1.
11. A decoding device configured to perform the method according to any one of claims 1 to 10.
12. A computer program for causing a computer to execute the method according to any one of claims 1 to 10.