Fast algorithm of symmetric motion vector difference coding and decoding mode
By employing adaptive motion vector prediction technology, this method addresses the video encoding and decoding issues present in existing video encoding and decoding standards, thereby improving the efficiency and effectiveness of video encoding and decoding.
Patent Information
- Application Number
- CN202511421331.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-01
- Filing Date
- 2020-01-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video codec standards still have room for improvement in terms of encoding and decoding efficiency and bandwidth requirements when processing high-resolution video. In particular, in the HEVC/H.265 standard, motion vector prediction and signaling methods suffer from redundancy and high computational complexity.
An affine mode motion vector prediction (MVP) method with adaptive motion vector resolution (AMVR) is adopted. By modeling non-affine inter-frame modes, context encoding and decoding, multi-mode context, and variable control probability updates, the encoding and decoding process of video blocks is optimized, and the encoding and decoding information of adjacent blocks is used for video processing.
It improves the efficiency of video encoding and decoding, reduces computational complexity, enhances encoding and decoding performance, and meets the needs of high-resolution video.
Smart Images

Figure CN120980252A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a continuation of Chinese Patent Application No. 202080011121.X, filed July 27, 2021, which claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 074216, filed January 31, 2019, and International Patent Application No. PCT / CN2019 / 074433, filed February 1, 2019. All of the above applications are incorporated by reference as part of the disclosure of this application for all purposes. TECHNICAL FIELD
[0003] This patent document relates to video processing techniques, devices, and systems. BACKGROUND
[0004] Despite advances in video compression, digital video consumes the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices that can receive and display video increases, the bandwidth demand for digital video usage is expected to continue to grow. SUMMARY
[0005] Devices, systems, and methods related to digital video coding are described, and in particular, motion vector prediction (MVP) derivation and signaling for affine modes with adaptive motion vector resolution (AMVR) are described. The described methods can be applied to existing video coding standards (e.g., High Efficiency Video Coding (HEVC)) and future video coding standards or video codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a method of video processing. The method includes determining that a conversion between a current video block of a video and a coded representation of the current video block is based on a non-affine inter AMVR mode; and performing the conversion based on the determination, and wherein the coded representation of the current video block is based on context-based coding, and wherein a context used to code the current video block is modeled without using affine AMVR mode information of a neighboring block during the conversion.
[0007] In another representative aspect, the disclosed technology can be used to provide a method of video processing. The method includes determining that a conversion between a current video block of a video and a coded representation of the current video block is based on an affine adaptive motion vector resolution (affine AMVR) mode; and performing the conversion based on the determination, and wherein the coded representation of the current video block is based on context-based coding, and wherein a variable controls a speed of two probability updates of a context.
[0008] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes determining, for a conversion between a current video block of a video and a coded representation of the current video block, a usage of a plurality of contexts for the conversion; and conducting the conversion based on the determining, and wherein the plurality of contexts are utilized to code a syntax element indicative of a coarse motion precision.
[0009] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes determining, for a conversion between a current video block of a video and a coded representation of the current video block, a usage of a plurality of contexts for the conversion; and conducting the conversion based on the determining, and wherein the plurality of contexts are utilized to code a syntax element indicative of a coarse motion precision.
[0010] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes determining, for a conversion between a current video block of a video and a coded representation of the current video block, whether to use a symmetric motion vector difference (SMVD) mode based on a currently selected best mode for the conversion; and conducting the conversion based on the determining.
[0011] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes determining, for a conversion between a current video block of a video and a coded representation of the current video block, whether to use an affine SMVD mode based on a currently selected best mode for the conversion; and conducting the conversion based on the determining.
[0012] In another representative aspect, the above method is implemented in the form of a processor-executable code and stored in a computer readable program medium.
[0013] In yet another representative aspect, an apparatus configured or operable to perform the above method is disclosed. The apparatus can include a processor programmed to implement the method.
[0014] In yet another representative aspect, a video decoder apparatus can implement the method as described herein.
[0015] The above described aspects and features of the disclosed technology, as well as other aspects and features, are more fully described in the accompanying drawings, specification, and claims. BRIEF DESCRIPTION OF DRAWINGS
[0016] FIG. 1 An example of constructing a Merge candidate list is shown.
[0017] FIG. 2An example of a location of a spatial candidate is shown.
[0018] FIG. 3 An example of a pair of candidates subject to a redundancy check for a spatial Merge candidate is shown.
[0019] FIG. 4A And FIG. 4B An example of a location of a second prediction unit (PU) based on the size and shape of the current block is shown.
[0020] FIG. 5 An example of motion vector scaling for a temporal Merge candidate is shown.
[0021] FIG. 6 An example of a candidate location for a temporal Merge candidate is shown.
[0022] FIG. 7 An example of generating a combined bi-predictive Merge candidate is shown.
[0023] FIG. 8 An example of constructing a motion vector prediction candidate is shown.
[0024] FIG. 9 An example of motion vector scaling for a spatial motion vector candidate is shown.
[0025] FIG. 10 An example of motion prediction using an alternative temporal motion vector prediction (ATMVP) algorithm for a coding unit (CU) is shown.
[0026] FIG. 11 An example of a coding unit (CU) with sub-blocks and neighboring blocks used by a spatial-temporal motion vector prediction (STMVP) algorithm is shown.
[0027] FIG. 12A And FIG. 12B An example snapshot of a sub-block when using an overlapping block motion compensation (OBMC) algorithm is shown.
[0028] FIG. 13 An example of neighboring samples used to derive parameters for a local illumination compensation (LIC) algorithm is shown.
[0029] FIG. 14 An example of a simplified affine motion model is shown.
[0030] FIG. 15 An example of an affine motion vector field (MVF) per sub-block is shown.
[0031] FIG. 16 An example of motion vector prediction (MVP) for an AF INTER affine motion mode is shown.
[0032] FIG. 17A and 17B Examples of 4-parameter affine model and 6-parameter affine model are shown, respectively.
[0033] FIG. 18A and 18B An example candidate for AF_MERGE affine motion mode is shown.
[0034] FIG. 19 An example of bilateral matching in the pattern matching motion vector derivation (PMMVD) mode, which is based on a special Merge mode of the frame rate up conversion (FRUC) algorithm, is shown.
[0035] FIG. 20 An example of template matching in the FRUC algorithm is shown.
[0036] FIG. 21 An example of unidirectional motion estimation in the FRUC algorithm is shown.
[0037] FIG. 22 An example of optical flow trajectories used by the bi-directional optical flow (BIO) algorithm is shown.
[0038] FIG. 23A and 23B An example snapshot using the bi-directional optical flow (BIO) algorithm without block expansion is shown.
[0039] FIG. 24 An example of the decoder-side motion vector refinement (DMVR) algorithm based on bilateral template matching is shown.
[0040] FIG. 25A-25F A flowchart of an example method for video processing based on implementations of the disclosed technology is shown.
[0041] FIG. 26 is a block diagram of an example of a hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.
[0042] FIG. 27 An example of a symmetric mode is shown.
[0043] FIG. 28 Another block diagram of an example of a hardware platform for implementing the video processing system described in this document is shown. DETAILED DESCRIPTION
[0044] Video coding methods and techniques are ubiquitous in modern technology due to the increasing demand for higher resolution video. Video codecs, which typically comprise electronic circuits or software that compress or decompress digital video, are continually improved to provide higher coding efficiency. Video codecs convert uncompressed video into a compressed format and vice versa. There is a complex relationship between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end delay (latency). The compressed format typically conforms to a standard video compression specification, e.g., the High Efficiency Video Coding (HEVC) standard (also known as H.265 or MPEG-H Part 2) [1], the Versatile Video Coding standard under finalization, or other current and / or future video coding standards.
[0045] Embodiments of the disclosed technology can be applied to existing video coding standards (e.g., HEVC, H.265) and future standards to improve compression performance. Section headings are used in this document to improve readability of the description and do not limit the discussion or embodiments (and / or implementations) to only the corresponding section in any way.
[0046] 1. Example of inter prediction in HEVC / H.265
[0047] Video coding standards have improved significantly over the years and now provide, in part, high coding efficiency and support for higher resolutions. The latest standards such as HEVC and H.265 are based on a hybrid video coding structure where temporal prediction plus transform coding is used.
[0048] 1.1 Example of prediction modes
[0049] Each inter-predicted prediction unit (PU) has motion parameters for one or two reference picture lists. In some embodiments, the motion parameters include a motion vector and a reference picture index. In other embodiments, inter_pred_idc can also be used to signal the use of one of the two reference picture lists. In still other embodiments, the motion vector can be explicitly coded as a delta relative to a predictor.
[0050] When a coding unit (CU) is coded using skip mode, one PU is associated with the CU and there is no significant residual coefficient, no coded motion vector delta or reference picture index. Merge mode is specified whereby motion parameters for the current PU are derived from neighboring PUs, including spatial and temporal candidates. Merge mode can be applied to any inter- predicted PU, not only to skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list usage are explicitly signaled for each PU.
[0051] When signaling indicates that one of the two reference picture lists is to be used, the PU is generated from one sample block. This is called "uni-prediction". Uni-prediction can be used for P slices and B slices [2].
[0052] When signaling indicates that both reference picture lists are to be used, the PU is generated from two sample blocks. This is called "bi-prediction". Bi-prediction can only be used for B slices.
[0053] 1.1.1 Embodiment of constructing Merge mode candidates
[0054] When a PU is predicted using Merge mode, an index pointing to an entry in the merge candidates list is parsed from the bitstream and used to retrieve the motion information. The construction of this list can be summarized in the following steps order:
[0055] Step 1 : Initial candidate derivation
[0056] Step 1.1 : Spatial candidate derivation
[0057] Step 1.2 : Redundancy check of spatial candidates
[0058] Step 1.3 : Temporal candidate derivation
[0059] Step 2 : Additional candidate insertion
[0060] Step 2.1 : Creation of bi-prediction candidates
[0061] Step 2.2 : Insertion of zero motion candidate
[0062] FIG. 1An example of constructing the Merge candidate list based on the above summarized order of steps is shown. For the spatial Merge candidate derivation, up to four Merge candidates are selected among the candidates located in the five different positions. For the temporal Merge candidate derivation, up to one Merge candidate is selected among the two candidates. Since the number of candidates per PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, a binary unary truncated (TU) is used to encode the index of the best Merge candidate. If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list which is identical to the Merge candidate list of the 2Nx2N prediction unit.
[0063] 1.1.2 Construction of spatial Merge candidates
[0064] In the derivation of spatial Merge candidates, up to four Merge candidates are selected among the candidates located in the positions depicted in FIG. 2 The order of derivation is Al, Bl, B0, A0 and B2. Position B2 is only considered if any of the positions Al, Bl, B0, A0 is not available (e.g. because the PU belongs to another slice or tile) or is intra coded. After adding the candidate at position Al, a redundancy check is performed on the addition of the remaining candidates which ensures that candidates with identical motion information are excluded from the list, improving the coding efficiency.
[0065] To reduce the computational complexity, not all possible pairs of candidates are considered in the mentioned redundancy check. Instead, only pairs linked by an arrow in FIG. 3 are considered and only if the corresponding candidate used for the redundancy check has different motion information, the candidate is added to the list. Another source of redundant motion information is the "second PU" associated with a partition different from 2Nx2N. As an example, FIG. 4A and 4B depict the second PU for the case of Nx2N and 2N x N, respectively. When the current PU is partitioned into Nx2N, the candidate at position Al is not considered for the list construction. In some embodiments, adding this candidate can result in two prediction units with identical motion information which is redundant for a coding unit having only one PU. Similarly, when the current PU is partitioned into 2N x N, position Bl is not considered.
[0066] 1.1.3 Construction of temporal Merge candidates
[0067] In this step, only one candidate is added to the list. In particular, in the derivation of the temporal Merge candidate, a scaled motion vector is derived based on a co-located PU belonging to the picture having the smallest Picture Order Count (POC) difference with respect to the current picture within the given reference picture list. The reference picture list used for the derivation of the co-located PU is explicitly signaled in the slice header.
[0068] FIG. 5 An example of the derivation of the scaled motion vector for a temporal Merge candidate (as dashed line) is shown, which is scaled from the motion vector of a co-located PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set equal to zero. For B slices, two motion vectors are obtained and combined to produce a bi-predictive Merge candidate, one for reference picture list 0 (list 0) and the other for reference picture list 1 (list 1).
[0069] As FIG. 6 shown, among the co-located PUs (Y) belonging to the reference frame, the position for the temporal candidate is selected between candidates C0 and C1. If the PU at position C0 is not available, intra coded or outside the current CTU, then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate.
[0070] 1.1.4 Building additional types of Merge candidates
[0071] Besides the spatial-temporal Merge candidate, there are two additional types of Merge candidates: combined bi-predictive Merge candidate and zero Merge candidate. The combined bi-predictive Merge candidate is generated by exploiting the spatial-temporal Merge candidate. The combined bi-predictive Merge candidate is only used for B slices. The combined bi-predictive candidate is generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another candidate. If these two tuples provide different motion hypotheses, then they will form a new bi-predictive candidate.
[0072] FIG. 7An example of the process is shown in which two candidates with mvL0 and refldxL0 or mvL1 and refldxL1 in the original list (710, on the left) are used to create a combined bi-predictive Merge candidate that is added to the final list (720, on the right).
[0073] Zero motion candidates are inserted to fill the remaining entries in the Merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts from zero and is increased each time a new zero motion candidate is added to the list. The number of reference frames used by these candidates is 1 and 2 for uni- and bi-predictive, respectively. In some embodiments, no redundancy check is performed on these candidates.
[0074] 1.1.5 Example of motion estimation region for parallel processing
[0075] To speed up the encoding process, motion estimation can be performed in parallel, whereby the motion vectors of all prediction units within a given region are derived at the same time. Derivation of Merge candidates from spatial neighbors can interfere with parallel processing, as one prediction unit cannot derive motion parameters from neighboring PUs until its associated motion estimation is complete. To mitigate the trade-off between coding efficiency and processing latency, a motion estimation region (MER) can be defined. The size of the MER can be signaled in the picture parameter set (PPS) using the "log2_parallel_merge_level_minus2" syntax element. When an MER is defined, Merge candidates falling into the same region are marked as unavailable and thus are not considered in the list construction as well.
[0076] 1.2 Embodiments of advanced motion vector prediction (AMVP)
[0077] AMVP exploits the spatial-temporal correlation of motion vectors with neighboring PUs for the explicit transmission of motion parameters. The motion vector candidate list is constructed by first checking the availability of left, top, and temporally neighboring PU positions, removing redundant candidates, and adding a zero vector to make the candidate list of constant length. The encoder can then select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to Merge index signaling, a unary code is used to encode the index of the best motion vector candidate. The maximum value to be encoded in this case is 2 (see FIG. 8 ). Details on the derivation process of motion vector prediction candidates are provided in the following section.
[0078] 1.2.1 Example of constructing motion vector prediction candidates
[0079] FIG. 8 The derivation process for motion vector prediction candidates is summarized and can be implemented for each reference picture list with refidx as input.
[0080] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. As FIG. 2 As previously shown, for spatial motion vector candidate derivation, two motion vector candidates are finally derived based on the motion vector of each PU located in five different positions.
[0081] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two different co-located positions. After making the first list of spatial-temporal candidates, duplicate motion vector candidates in the list are removed. If the number of potential candidates is greater than 2, the motion vector candidates whose reference picture index is greater than 1 within the associated reference picture list are removed from the list. If the number of spatial-temporal motion vector candidates is less than 2, additional zero motion vector candidates are added to the list.
[0082] 1.2.2 Building spatial motion vector candidates
[0083] In the derivation of spatial motion vector candidates, at most two candidates are considered among five potential candidates from PUs located as FIG. 2 Previously shown positions, which are the same as those for motion Merge. The derivation order for the left side of the current PU is defined as A0, A1 and scaled A0, scaled A1. The derivation order for the top side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Thus, for each side, there are four cases available as motion vector candidates, where two cases do not need to use spatial scaling and two cases use spatial scaling. The four different cases are summarized as follows.
[0084] • No spatial scaling
[0085] (1) Same reference picture list, and same reference picture index (same POC)
[0086] (2) Different reference picture list, but same reference picture index (same POC)
[0087] • Spatial scaling
[0088] (3) Same reference picture list, but different reference picture index (different POC)
[0089] (4) Different reference picture list, and different reference picture index (different POC)
[0090] First, the no spatial scaling case is checked, followed by the allowed spatial scaling case. Spatial scaling is considered when the POC is different between the reference picture of the neighboring PU and the reference picture of the current PU, regardless of the reference picture list. If all PUs of the left candidate are not available or are intra coded, scaling of the above motion vector is allowed to help parallel derivation of the left and above MV candidates. Otherwise, no spatial scaling is allowed for the above motion vector.
[0091] As shown in the example in FIG. 9 For the spatial scaling case, the motion vector of the neighboring PU is scaled in a similar way as the temporal scaling. One difference is that the reference picture list and index of the current PU are given as input; the actual scaling process is the same as the temporal scaling process.
[0092] 1.2.3 Construction of temporal motion vector candidates
[0093] Except for the reference picture index derivation, all processes for the derivation of temporal Merge candidates are the same as those for the derivation of spatial motion vector candidates (as shown in the example in FIG. 6
[0094] 2. An example of inter prediction method in Joint Exploration Model (JEM)
[0095] In some embodiments, a reference software called Joint Exploration Model (JEM) [3] [4] is used to explore future video coding technologies. In JEM, subblock-based prediction is employed in several coding tools, such as affine prediction, optional temporal motion vector prediction (ATMVP), spatial-temporal motion vector prediction (STMVP), bi-directional optical flow (BIO), frame rate up conversion (FRUC), local adaptive motion vector resolution (LAMVR), overlapped block motion compensation (OBMC), local illumination compensation (LIC), and decoder-side motion vector refinement (DMVR).
[0096] 2.1 An example of sub-CU based motion vector prediction
[0097] In JEM with quadtree plus binary tree (QTBT), each CU can have at most one set of motion parameters for each prediction direction. In some embodiments, two sub-CU level motion vector prediction methods are considered in the encoder by dividing a large CU into sub-CUs and deriving motion information for all sub-CUs of the large CU. The alternative temporal motion vector prediction (ATMVP) method allows each CU to take multiple sets of motion information from multiple blocks smaller than the current CU in the collocated reference picture. In the spatial-temporal motion vector prediction (STMVP) method, the motion vector of a sub-CU is derived recursively by using temporal motion vector prediction and spatial neighboring motion vectors. In some embodiments, and to preserve more accurate motion field for sub-CU motion prediction, motion compression of reference frames can be disabled.
[0098] 2.1.1 Example of Alternative Temporal Motion Vector Prediction (ATMVP)
[0099] In the ATMVP method, the temporal motion vector prediction (TMVP) method is modified by retrieving multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU.
[0100] FIG. 10 An example of the ATMVP motion prediction process for CU 1000 is shown. The ATMVP method predicts the motion vector of sub-CU 1001 within CU 1000 in two steps. The first step is to identify a corresponding block 1051 in reference picture 1050 using a temporal vector. Reference picture 1050 is also referred to as a motion source picture. The second step is to divide current CU 1000 into sub-CUs 1001 and obtain motion vectors from the blocks corresponding to each sub-CU as well as reference indices for each sub-CU.
[0101] In the first step, reference picture 1050 and the corresponding block are determined by the motion information of the spatial neighboring blocks of current CU 1000. To avoid the repeated scanning process of neighboring blocks, the first Merge candidate in the Merge candidate list of current CU 1000 is used. The first available motion vector and its associated reference index are set as the temporal vector and the index of the motion source picture. In this way, the corresponding block, which is sometimes referred to as the collocated block, can be more accurately identified compared to TMVP, where the corresponding block is always in the right bottom or center position relative to the current CU.
[0102] In the second step, the corresponding block of the sub-CU 1051 is identified by adding the temporal vector to the coordinates of the current CU in the motion source picture 1050. For each sub-CU, the motion information of its corresponding block (e.g., the smallest motion grid covering the center sample) is used to derive the motion information of the sub-CU. After identifying the motion information of the corresponding NxN block, it is converted into the motion vector and reference index of the current sub-CU in the same way as TMVP of HEVC, in which motion scaling and other processes are applied. For example, the decoder checks whether the low-delay condition is satisfied (e.g., the POCs of all reference pictures of the current picture are smaller than the POC of the current picture) and possibly uses the motion vector MVx(e.g., the motion vector corresponding to the reference picture list X) for predicting the motion vector MVy(e.g., where X is equal to 0 or 1 and Y is equal to 1-X) of each sub-CU.
[0103] 2.1.2 Example of Spatial-Temporal Motion Vector Prediction (STMVP)
[0104] In the STMVP method, the motion vector of a sub-CU is derived recursively in the raster scan order. FIG. 11 An example of one CU with four sub-blocks and neighboring blocks is shown. Consider an 8x8 CU 1100 including four 4x4 sub-CUs A (1101), B (1102), C (1103), and D (1104). The neighboring 4x4 blocks in the current frame are labeled a (1111), b (1112), c (1113), and d (1114).
[0105] The motion derivation of sub-CU A starts by identifying its two spatial neighbors. The first neighbor is the NxN block above sub-CU A 1101 (block c 1113). If this block c (1113) is not available or is intra coded, other NxN blocks above sub-CU A (1101) are checked (from left to right, starting from block c 1113). The second neighbor is the block to the left of sub-CU A 1101 (block b 1112). If block b (1112) is not available or is intra coded, other blocks to the left of sub-CU A 1101 are checked (from top to bottom, starting from block b 1112). The motion information obtained from the neighboring blocks for each list is scaled to the first reference frame of the given list. Next, the temporal motion vector prediction (TMVP) of sub-block A 1101 is derived by following the same procedure as specified in TMVP in HEVC. The motion information of the co-located block at block D 1104 is retrieved and scaled accordingly. Finally, after extracting and scaling the motion information, all available motion vectors are averaged separately for each reference list. The averaged motion vector is specified as the motion vector of the current sub-CU.
[0106] 2.1.3 Example of sub-CU motion prediction mode signaling
[0107] In some embodiments, sub-CU modes are enabled as additional Merge candidates and no additional syntax elements are needed to signal the modes. Two additional Merge candidates are added to merge the candidate list of each CU to represent the ATMVP mode and the STMVP mode. In other embodiments, up to seven Merge candidates can be used if the sequence parameter set indicates that ATMVP and STMVP are enabled. The coding logic of the additional Merge candidates is the same as that of the Merge candidates in HM, which means that for each CU in a P or B slice, two additional Merge candidates can require two more RD checks. In some embodiments, such as JEM, all the bins of the Merge index are context coded by CABAC (Context-based Adaptive Binary Arithmetic Coding). In other embodiments, such as HEVC, only the first bin is context coded and the remaining bins are context bypass coded.
[0108] 2.2 Example of Adaptive Motion Vector Difference Resolution
[0109] In some embodiments, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of a PU and the predicted motion vector) is signaled in quarter luma sample units. In JEM, local adaptive motion vector resolution (LAMVR) is introduced. In JEM, the MVD can be coded in quarter luma sample, integer luma sample, or four luma sample units. The MVD resolution is controlled at the coding unit (CU) level, and for each CU with at least one non-zero MVD component, the MVD resolution flag is conditionally signaled.
[0110] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or four luma sample MV precision is used.
[0111] When the first MVD resolution flag of a CU is zero, or is not coded for the CU (meaning that all MVDs in the CU are zero), quarter luma sample MV resolution is used for the CU. When a CU uses integer luma sample MV precision or four luma sample MV precision, the M VP s in the AMVP candidate list of the CU are rounded to the corresponding precision.
[0112] In the encoder, CU-level RD check is used to determine which MVD resolution will be used for the CU. In other words, CU-level RD check is performed three times for each MVD resolution. To speed up the encoder, the following encoding scheme is applied in JEM.
[0113] • During the RD check of a CU with normal quarter-luma-sample MVD resolution, the motion information (integer-luma-sample precision) of the current CU is stored. For the same CU with integer-luma-sample and 4-luma-sample MVD resolution, the stored motion information (after rounding) is used as the starting point for further small-range motion vector refinement during the RD check, so that the time-consuming motion estimation process is not repeated three times.
[0114] • The RD check of a CU with 4-luma-sample MVD resolution is conditionally invoked. For a CU, when the RD cost of integer-luma-sample MVD resolution is much larger than that of quarter-luma-sample MVD resolution, the RD check of 4-luma-sample MVD resolution for the CU is skipped.
[0115] 2.3 Example of higher motion vector storage precision
[0116] In HEVC, the motion vector precision is one-quarter pixel (pel) (one-quarter luma sample and one-eighth chroma sample for 4:2:0 video). In JEM, the precision of the internal motion vector storage and Merge candidate is increased to 1 / 16 pel. The higher motion vector precision (1 / 16 pel) is used for motion-compensated inter prediction of CUs coded in skip mode / Merge mode. For CUs coded using normal AMVP mode, integer-pel or quarter-pel motion is used.
[0117] The SHVC up-sampling interpolation filter with the same filter length and normalization factor as the HEVC motion compensation interpolation filter is used as the motion compensation interpolation filter for additional fractional-pel positions. In JEM, the chroma component motion vector precision is 1 / 32 sample, and the additional interpolation filter for 1 / 32-pel fractional positions is derived by using the average of the filter for two neighboring 1 / 16-pel fractional positions.
[0118] 2.4 Example of overlapped block motion compensation (OBMC)
[0119] In JEM, OBMC can be switched at CU level using a syntax element. When OBMC is used in JEM, it is performed for all motion compensation (MC) block boundaries except the right and bottom boundaries of the CU. In addition, it is applied to both luma and chroma components. In JEM, a MC block corresponds to a coding block. When a CU is coded using sub-CU modes (including sub-CU Merge, affine, and FRUC modes), each sub-block of the CU is a MC block. To handle the boundaries of the CU uniformly, in the case where the size of the sub-block is set to 4x4, OBMC is performed at the sub-block level for all MC block boundaries as shown in FIG. 12A and FIG. 12B
[0120] FIG. 12A Sub-blocks at the CU / PU boundary are shown, and the shaded sub-blocks are the locations where OBMC is applied. Similarly, FIG. 12B Sub-PUs in ATMVP mode are shown.
[0121] When OBMC is applied to the current sub-block, in addition to the current MV, the vectors of the four neighboring sub-blocks (if available and not exactly the same as the current motion vector) are also used to derive the prediction block of the current sub-block. The multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal of the current sub-block.
[0122] The prediction block based on the motion vector of the neighboring sub-block is denoted as PN, where N indicates the index of the neighboring top, bottom, left, and right sub-blocks, and the prediction block based on the motion vector of the current sub-block is denoted as PC. When PN is based on the motion information of the neighboring sub-block and that motion information is the same as the motion information of the current sub-block, no OBMC is performed from PN. Otherwise, each sample of PN is added to the same sample in PC, i.e., four rows / columns of PN are added to PC. The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for PN, and the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for PC. The exception is that only two rows / columns of PN are added to PC for small MC blocks (i.e., when the height or width of the coding block is equal to 4 or the CU is coded using sub-CU modes). In this case, the weighting factors {1 / 4, 1 / 8} are used for PN, and the weighting factors {3 / 4, 7 / 8} are used for PC. For PN generated based on the motion vector of the vertical (horizontal) neighboring sub-block, the samples in the same row (column) of PN are added to PC with the same weighting factor.
[0123] In JEM, for CUs with size smaller than or equal to 256 luma samples, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU or not. For CUs with size larger than 256 luma samples or not coded using AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is taken into account during the motion estimation phase. The prediction signal formed by OBMC using the motion information of the top and left neighboring blocks is used to compensate the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.
[0124] 2.5 Example of Local Illumination Compensation (LIC)
[0125] LIC is based on a linear model for illumination changes, using a scaling factor a and an offset b. Also, for each inter mode coded coding unit (CU), LIC is adaptively enabled or disabled.
[0126] When LIC is applied to a CU, the least square error method is employed to derive the parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. FIG. 13 An example of the neighboring samples used to derive the parameters of the IC algorithm is shown. Specifically, and as shown in FIG. 13 sub-sampled (2: 1 sub-sampling) of the CU and the corresponding samples in the reference picture (identified by the motion information of the current CU or sub-CU) are used. The IC parameters are derived and applied to each prediction direction separately.
[0127] When a CU is coded using Merge mode, the LIC flag is copied from the neighboring block in a similar way as the motion information copy in Merge mode; otherwise, the LIC flag is signaled to the CU to indicate whether LIC is applicable or not.
[0128] When LIC is enabled for a picture, additional CU-level RD check is needed to determine whether to apply LIC to a CU or not. When LIC is enabled for a CU, mean-removed sum of absolute difference (MR-SAD) and mean-removed sum of absolute Hadamard-transformed difference (MR-SATD) are used instead of SAD and SATD for integer and fractional pixel motion search respectively.
[0129] To reduce the encoding complexity, the following encoding scheme is applied in JEM.
[0130] LIC is disabled for the whole picture when there is no significant luminance change between the current picture and its reference pictures. To identify this case, at the encoder, the histograms of the current picture and each reference picture of the current picture are computed. If the histogram difference between the current picture and each reference picture of the current picture is less than a given threshold, LIC is disabled for the current picture; otherwise, LIC is enabled for the current picture.
[0131] 2.6 Example of affine motion compensated prediction
[0132] In HEVC, only translational motion model is applied to motion compensated prediction (MCP). However, cameras and objects can have multiple kinds of motion, such as zooming, rotation, perspective motion, and other irregular motions. On the other hand, JEM applies a simplified affine transform motion compensated prediction. FIG. 14 An example of an affine motion field of a block 1400 described by two control point motion vectors V0 and V1 is shown. The motion vector field (MVF) of the block 1400 can be described by the following equation:
[0133]
[0134] As shown in FIG. 14 (v0x, v0y) is the motion vector of the left top corner control point and (v1x, v1y) is the motion vector of the right top corner control point. To simplify the motion compensated prediction, a subblock-based affine transform prediction can be applied. The subblock size M x N is derived as follows:
[0135]
[0136] where MvPre is the motion vector fractional precision (e.g., 1 / 16 in JEM). (v2x, v2y) is the motion vector of the left bottom control point, which is computed according to equation (1). If needed, M and N can be down-adjusted to be the divisors of w and h, respectively.
[0137] FIG. 15 An example of the affine MVF of each subblock of a block 1500 is shown. To derive the motion vector of each M x N subblock, the motion vector of the center sample of each subblock can be computed according to equation (1) and rounded to the motion vector fractional precision (e.g., 1 / 16 in JEM). Then, a motion compensated interpolation filter can be applied to generate the prediction of each subblock with the derived motion vector. After MCP, the high precision motion vector of each subblock is rounded to the same precision as the normal motion vector and saved.
[0138] 2.6.1 Embodiment of AF INTER mode
[0139] In JEM, there are two affine motion modes: AF_INTER mode and AF_MERGE mode. AF_INTER mode can be applied to CUs with both width and height greater than 8. Signaling in the bitstream informs the affine flag at the CU level to indicate whether AF_INTER mode is used. In AF_INTER mode, adjacent blocks are used to construct motion vector pairs {(v0,v1)|v0={v...}. A ,v B ,v c},v1={v D ,v E The candidate list of}}.
[0140] FIG. 16 An example of motion vector prediction (MVP) for block 1600 in AF_INTER mode is shown. FIG. 16 As shown, v0 is selected from the motion vectors of sub-blocks A, B, or C. Motion vectors from neighboring blocks can be scaled according to a reference list. Motion vectors can also be scaled based on the relationship between the Picture Order Count (POC) of the references of neighboring blocks, the POC of the current CU's reference, and the POC of the current CU. The method for selecting v1 from neighboring sub-blocks D and E is similar. If the number of candidates in the candidate list is less than 2, the list is populated by motion vector pairs constructed by repeating each AMVP candidate. When the candidate list is greater than 2, candidates are first categorized based on adjacent motion vectors (e.g., based on the similarity of two motion vectors in a candidate pair). In some implementations, the first two candidates are retained. In some implementations, rate distortion (RD) cost checking is used to determine which motion vector pair candidate is selected as the Control Point Motion Vector Prediction (CPMVP) for the current CU. The index indicating the position of the CPMVP in the candidate list can be signaled in the bitstream. After determining the CPMVP of the current affine CU, affine motion estimation is applied, and the Control Point Motion Vector (CPMV) is found. The difference between the CPMV and the CPMVP is then signaled in the bitstream.
[0141] In AF_INTER mode, when using the 4 / 6 parameter affine mode, 2 / 3 control points are required, and therefore 2 / 3 MVD encoding / decoding is needed for these control points, such as... FIG. 17A and FIG. 17B As shown. In the existing implementation [5], MV can be derived as follows, for example, it predicts mvd1 and mvd2 from mvd0.
[0142]
[0143] In this article, mvd iand mv1 is the predicted motion vector, motion vector difference, and motion vector of the left-top pixel (i = 0), right-top pixel (i = 1), or left-bottom pixel (i = 2), respectively, as shown in FIG. 18B In some embodiments, the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) equals the summation of the two components individually. For example, newMV = mvA + mvB means that the two components of newMV are set to (xA + xB) and (yA + yB), respectively.
[0144] 2.6.2 Example of fast affine ME algorithm in AF INTER mode
[0145] In some embodiments of affine mode, the MVs of 2 or 3 control points need to be determined jointly. Directly joint search of multiple MVs is computationally complex. In an example, a fast affine ME algorithm [6] is proposed and adopted into VTM / BMS.
[0146] For example, the fast affine ME algorithm is described for 4-parameter affine model, and the idea can be extended to 6-parameter affine model:
[0147]
[0148] Using a' to replace (a - 1) enables the motion vector to be rewritten as:
[0149]
[0150] If the motion vectors of two control points (0, 0) and (0, w) are assumed to be known, then according to equation (5), the affine parameters can be derived as:
[0151]
[0152] The motion vector can be rewritten in vector form as:
[0153]
[0154] In this document, P = (x, y) is a pixel position,
[0155]
[0156] And
[0157]
[0158] In some embodiments, and at the encoder, the MVD of AF INTER can be derived iteratively. Let MVi(P) denote the MV derived in the i-th iteration at position P, and let dMV Ci denotes the increment for the MVC update in the i-th iteration. Then in the (i+1)-th iteration,
[0159]
[0160] Pic ref denotes the reference picture and Pic cur denotes the current picture and Q = P + MV i (P). If MSE is used as the matching criterion, the function to be minimized can be written as:
[0161]
[0162] If it is assumed that is small enough, then can be rewritten as an approximation based on a first order Taylor expansion, as:
[0163]
[0164] In this document, If the notation E i+1 (P) = Pic cur (P) - Pic ref (Q), then
[0165]
[0166] The term can be derived by setting the derivative of the error function to zero and then calculating the increments MV for the control points (0,0) and (0,w) according to
[0167]
[0168] In some embodiments, the MVD derivation process can be iterated n times and the final MVD can be calculated as follows:
[0169]
[0170] In the above implementation [5], predicting the increment MV for the predicted control point (0,w) denoted by mvd1 from the increment MV for the control point (0,0) denoted by mvd0 results in being encoded only for mvd1.
[0171] 2.6.3 Embodiments of AF_MERGE mode
[0172] When a CU is applied in AF_MERGE mode, it obtains a first block coded using affine mode from a valid neighboring reconstructed block. FIG. 18A An example of the selection order of the candidate blocks of the current CU 1800 is shown. As shown, the selection order can be from the left (1801), top (1802), top-right (1803), bottom-left (1804) to the top-left (1805) of the current CU 1800. FIG. 18A FIG. 18B Another example of the candidate blocks of the current CU 1800 in the AF_MERGE mode is shown. As shown, if the neighboring bottom-left block 1801 is coded in the affine mode, the motion vectors v2, v3 and v4 of the top-left corner, top-right corner and bottom-left corner of the CU containing the sub-block 1801 are derived. The motion vector v0 of the top-left corner of the current CU 1800 is calculated based on v2, v3 and v4. The motion vector v1 of the top-right of the current CU can be calculated accordingly. FIG. 18B
[0173] After the CPMVs of the current CU v0 and v1 are calculated according to the affine motion model in equation (1), the MVFs of the current CU can be generated. To identify whether the current CU is coded using the AF_MERGE mode, an affine flag can be signaled in the bitstream when there is at least one neighboring block coded in the affine mode.
[0174] 2.7 Example of motion vector derivation with pattern matching (PMMVD)
[0175] The PMMVD mode is a special Merge mode based on a frame rate up conversion (FRUC) method. With this mode, the motion information of a block is not signaled but derived at the decoder side.
[0176] When the Merge flag of a CU is true, a FRUC flag can be signaled to the CU. When the FRUC flag is false, a Merge index can be signaled and the regular Merge mode is used. When the FRUC flag is true, an additional FRUC mode flag can be signaled to indicate which method (e.g., bilateral matching or template matching) will be used to derive the motion information of the block.
[0177] At the encoder side, the decision on whether to use the FRUC Merge mode for a CU is based on the RD cost selection as done for normal Merge candidates. For example, multiple matching modes (e.g., bilateral matching and template matching) for a CU are checked by using the RD cost selection. The matching mode that results in the minimum cost is further compared with other CU modes. If the FRUC matching mode is the most efficient mode, the FRUC flag is set to true for the CU and the matching mode is used.
[0178] In general, the motion derivation process in FRUC Merge mode has two steps. First, a CU-level motion search is performed, followed by a sub-CU-level motion refinement. At the CU level, an initial motion vector is derived for the whole CU based on bilateral matching or template matching. First, a list of MV candidates is generated, and the candidate that results in the smallest matching cost is selected as the starting point for further CU-level refinement. Then, a local search based on bilateral matching or template matching is performed around the starting point. The MV that results in the smallest matching cost is taken as the MV for the whole CU. Subsequently, the motion information is further refined at the sub-CU level, with the derived CU motion vector as the starting point.
[0179] For example, the following derivation process is performed for WxH CU motion information derivation. In the first stage, the MV of the whole WxH CU is derived. In the second stage, the CU is further divided into MxM sub-CUs. The value of M is calculated as in Equation (3), where D is a predefined division depth, which is set to 3 by default in JEM. Then the MV of each sub-CU is derived.
[0180]
[0181] FIG. 19 An example of bilateral matching used in a frame rate up conversion (FRUC) method is shown. Bilateral matching is used to derive the motion information of a current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures (1910, 1911). Under the assumption of continuous motion trajectory, the motion vectors MV0 (1901) and MV1 (1902) pointing to the two reference blocks are proportional to the temporal distance between the current picture and the two reference pictures, e.g., TD0 (1903) and TD1 (1904). In some embodiments, when the current picture (1900) is in between the two reference pictures (1910, 1911) in the temporal domain and the temporal distance from the current picture to the two reference pictures is the same, bilateral matching becomes mirror-based bi-directional MV.
[0182] FIG. 20An example of template matching used in a frame rate up conversion (FRUC) method is shown. Template matching can be used to derive the motion information of a current CU 2000 by finding the closest match between a template (e.g., the top and / or left neighboring block of the current CU) in the current picture and a block (e.g., of the same size as the template) in the reference picture 2010. In addition to the FRUC Merge mode described above, template matching can also be applied to the AMVP mode. In both JEM and HEVC, AMVP has two candidates. Using the template matching method, a new candidate can be derived. If the newly derived candidate by template matching is different from the first existing AMVP candidate, it is inserted at the beginning of the AMVP candidate list, and then the list size is set to 2 (e.g., by removing the second existing AMVP candidate). When applied to the AMVP mode, only CU-level search is applied.
[0183] The set of MV candidates at the CU level can include: (1) the original AMVP candidates if the current CU is in the AMVP mode, (2) all Merge candidates, (3) a number of MVs in the interpolated MV field (described later), and the motion vectors of the top and left neighbors.
[0184] When using bilateral matching, each valid MV of a Merge candidate can be used as input to generate a pair of MVs under the assumption of bilateral matching. For example, one valid MV of a Merge candidate is (MVa, ref a ). Then, its counterpart bilateral MV is found in the other reference list B, ref b , such that ref a and ref b are located on different sides of the current picture in the temporal domain. If such ref b is not available in the reference list B, ref b is determined to be a different reference from ref a , and ref b is the one with the smallest temporal distance to the current picture in list B. After ref b is determined, MVb is derived by scaling MVa based on the temporal distance between the current picture and ref a , ref b .
[0185] In some implementations, four MVs from the interpolated MV field can also be added to the CU level candidate list. More specifically, the interpolated MVs at the locations (0,0), (W / 2,0), (0,H / 2) and (W / 2,H / 2) of the current CU are added. When FRUC is applied to AMVP mode, the original AMVP candidates are also added to the CU level MV candidate set. In some implementations, at the CU level, 15 MVs for AMVP CUs, 13 MVs for Merge CUs can be added to the candidate list.
[0186] The MV candidate set at the sub-CU level includes: (1) the MVs determined from the CU level search, (2) the neighboring MVs of the top, left, left-top and right-top, (3) scaled versions of the collocated MVs from the reference pictures, (4) one or more ATMVP candidates (e.g., up to four), and (5) one or more STMVP candidates (e.g., up to four). The scaled MVs from the reference pictures are derived as follows. The reference pictures in the two lists are traversed. The MVs at the collocated positions of the sub-CUs in the reference pictures are scaled to the reference of the starting CU level MV. The ATMVP and STMVP candidates can be the first four candidates. At the sub-CU level, one or more MVs (e.g., up to seventeen) are added to the candidate list.
[0187] Generation of Interpolated MV Field. Before the frame is coded, an interpolated motion field is generated for the entire picture based on the unilateral ME. The motion field can then be used later as a CU level or sub-CU level MV candidate.
[0188] In some embodiments, the motion field of each reference picture in the two reference lists is traversed at the 4x4 block level. FIG. 21 An example of the unidirectional motion estimation (ME) 2100 in the FRUC method is shown. For each 4x4 block, if the motion associated with the block goes through a 4x4 block in the current picture and the block has not been assigned any interpolated motion, the motion of the reference block is scaled to the current picture according to the temporal distances TD0 and TD1 (in the same way as the MV scaling of TMVP in HEVC), and the scaled motion is assigned to the block in the current frame. If a non-scaled MV is assigned to the 4x4 block, the motion of the block is marked as unavailable in the interpolated motion field.
[0189] Interpolation and Matching Cost. When the motion vector points to a fractional sample position, motion compensation interpolation is needed. To reduce the complexity, both bilateral matching and template matching can use bilinear interpolation instead of the regular 8-tap HEVC interpolation.
[0190] The calculation of the matching cost is slightly different in different steps. When selecting a candidate from the candidate set at the CU level, the matching cost can be the sum of absolute differences (SAD) of bilateral matching or template matching. After determining the initial MV, the matching cost of the bilateral matching of the sub-CU level search is calculated as follows:
[0191]
[0192] where w is a weighting factor. In some embodiments, w is set to 4 empirically, and MVs denote the current MV and the initial MV, respectively. The SAD can still be used as the matching cost of the template matching of the sub-CU level search.
[0193] In the FRUC mode, the MV is derived by using only the luma samples. The derived motion will be used for both luma and chroma in the MC inter prediction. After determining the MV, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma.
[0194] MV refinement is a pattern-based MV search with the criterion of bilateral matching cost or template matching cost. In JEM, two search patterns are supported - unrestricted center-biased diamond search (UCBDS) and adaptive cross search, respectively, for the MV refinement at the CU level and the sub-CU level. For the CU level and sub-CU level MV refinement, the MV is searched directly at quarter luma sample MV precision, and next at eighth luma sample MV refinement. The search range for the MV refinement at the CU step and the sub-CU step is set to be equal to 8 luma samples.
[0195] In the bilateral matching Merge mode, bi-prediction is applied because the motion information of the CU is derived based on the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures. In the template matching Merge mode, the encoder can choose among uni-prediction from list 0, uni-prediction from list 1, or bi-prediction for the CU. The choice can be based on the template matching cost as follows:
[0196] if costBi <= factor * min(costO, costl)
[0197] bi-prediction is used;
[0198] else if costO <= costl
[0199] uni-prediction from list 0 is used;
[0200] Otherwise,
[0201] Unidirectional prediction from List 1 is used;
[0202] where cost0 is the SAD of List 0 template matching, cost1 is the SAD of List 1 template matching, and costBi is the SAD of bi-prediction template matching. For example, when the value of factor is equal to 1.25, it means that the selection process is biased towards bi-prediction. Inter prediction direction selection can be applied to CU level template matching process.
[0203] 2.8 Example of Bi-directional Optical Flow (BIO)
[0204] The Bi-directional Optical Flow (BIO) method is a sample-wise motion refinement performed on top of the block-wise motion compensation used for bi-prediction. In some implementations, the sample level motion refinement does not use signaling.
[0205] Let I (k) be the luma value of the reference k (k = 0, 1) after block motion compensation, and let be denoted as I (k) and I x , respectively. Assuming the optical flow is valid, the motion vector field (v y , v (k) ) is given by the following equations.
[0206]
[0207] Combining this optical flow equation with the Hermite interpolation of the motion trajectory for each sample, the result is a unique third order polynomial for the end-point matching function values I (k) and derivatives at t = 0. The value of this polynomial at t = 0 is the BIO prediction:
[0208]
[0209] FIG. 22An example of optical flow trajectories in the bi-directional optical flow (BIO) method is shown. Therein, τ0 and τ1 denote the distance to the reference frames. The distances τ0 and τ1 are computed based on the POCs of Ref0 and Ref1 : τ0 = POC(current) - POC(Ref0), τ1 = POC(Ref1) - POC(current). If both predictions come from the same temporal direction (either from the past or from the future), the signs are different (e.g., τ0 · τ1 < 0). In this case, BIO is applied if the predictions are not from the same time instance (e.g., τ0 ≠ τ1). Both reference regions have non-zero motion (e.g., MVx0, MVy0, MVx1, MVy1 ≠ 0) and the block motion vector is proportional to the temporal distance (e.g., MVx0 / MVx1 = MVy0 / MVy1 = -τ0 / τ1).
[0210] The motion vector field (v x ,v y ) is determined by minimizing the difference Δ between the values in point A and point B. FIG. 23A and FIG. 23B An example of the intersection of the motion trajectory and the reference frame plane is shown. The model uses only the first linear term of the local Taylor expansion for Δ:
[0211]
[0212] All values in the above equations depend on the sample position denoted as (i', j'). Assuming that the motion is consistent in the local surrounding area, Δ can be minimized within a (2M+1) x (2M+1) square window Ω centered at the current prediction point (i, j), where M equals 2:
[0213]
[0214] For this optimization problem, JEM uses a simplified approach, first minimizing in the vertical direction and then in the horizontal direction. This leads to
[0215]
[0216] wherein,
[0217]
[0218] To avoid division by zero or very small values, regularization parameters r and m are introduced in equation (28) and equation (29).
[0219] r = 500 · 4 d-8 (12)
[0220] m = 700 · 4 d-8(13) where d is the bit depth of the video samples.
[0221] To keep the memory access of BIO the same as regular bi-prediction motion compensation, all the prediction and gradient values s (k) , FIG. 23A An example of the access positions outside the block 2300 is shown. As FIG. 23A shown in equation (9), a (2M+1) x (2M+1) square window Ω centered at the current prediction point on the boundary of the prediction block needs to access positions outside the block. In JEM, the values of s (k) , outside the block are set to be equal to the nearest available value inside the block. For example, this can be implemented as a padding region 2301, as FIG. 23B shown.
[0222] With BIO, it is possible to refine the motion field for each sample. To reduce the computational complexity, a block-based BIO design can be used in JEM. The motion refinement can be computed based on 4x4 blocks. In block-based BIO, the s n values in equation (9) for all samples in a 4x4 block can be aggregated and then the aggregated s n values are used for the derived BIO motion vector offsets for the 4x4 block. More specifically, the following formulas can be used for block-based BIO derivation:
[0223]
[0224] where b k denotes the set of samples belonging to the k-th 4x4 block of the prediction block. The s n,bk values in equations (28) and (29) are replaced by ((s n >>4) to derive the associated motion vector offsets.
[0225] In some cases, the MV regiment of BIO can be unreliable due to noise or irregular motion. Therefore, in BIO, the magnitude of the MV regiment is clipped to a threshold. The threshold is determined based on whether the reference pictures of the current picture all come from one direction. For example, if all the reference pictures of the current picture come from one direction, the threshold is set to 12x2 14 -d ; otherwise, it is set to 12x2 13-d .
[0226] The gradient of BIO and the motion compensated interpolation can be computed simultaneously, which uses the same operation as the HEVC motion compensation process (e.g., 2D separable Finite Impulse Response (FIR)). In some embodiments, the input of the 2D separable FIR is the same reference frame sample as the motion compensation process according to the fractional position (fracX, fracY) of the fractional part of the block motion vector. For horizontal gradient First, the signal is vertically interpolated using BIO filter S corresponding to the fractional position fracY with de-scaling offset d-8. Then, the gradient filter BIO filter G is applied in horizontal direction corresponding to the fractional position fracX with de-scaling offset 18-d. For vertical gradient First, the gradient filter is applied vertically using BIO filter G corresponding to the fractional position fracY with de-scaling offset d-8. Then, the signal shift is performed using BIO filter S in horizontal direction corresponding to the fractional position fracX with de-scaling offset 18-d. The length of the interpolation filter used for gradient computation BIO filter G and signal shift BIO filter F can be shorter (e.g., 6 taps) to maintain reasonable complexity. Table 1 shows example filters that can be used for gradient computation for different fractional positions of the block motion vector in BIO. Table 2 shows example interpolation filters that can be used for prediction signal generation in BIO.
[0227] Table 1 Example filters for gradient computation in BIO
[0228] Fractional Pixel Position Gradient Interpolation Filter (BIOfilterG) 0 {8,-39,-3,46,-17,5} 1 / 16 {8,-32,-13,50,-18,5} 1 / 8 {7,-27,-20,54,-19,5} 3 / 16 {6,-21,-29,57,-18,5} 1 / 4 {4,-17,-36,60,-15,4} 5 / 16 {3,-9,-44,61,-15,4} 3 / 8 {1,-4,-48,61,-13,3} 7 / 16 {0,1,-54,60,-9,2} 1 / 2 {-1,4,-57,57,-4,1}
[0229] Table 2 Example interpolation filters for prediction signal generation in BIO
[0230]
[0231]
[0232] In JEM, BIO can be applied to all bi-predicted blocks when the two predictions come from different reference pictures. When local illumination compensation (LIC) is enabled for a CU, BIO can be disabled.
[0233] In some embodiments, OBMC is applied to a block after the normal MC process. To reduce the computational complexity, BIO can not be applied during the OBMC process. This means that BIO is applied to the MC process of a block when its own MV is used, and BIO is not applied to the MC process when the MV of a neighboring block is used in the OBMC process.
[0234] 2.9 Example of decoder-side motion vector refinement (DMVR)
[0235] In bi-prediction operation, to predict a block region, two prediction blocks formed using motion vector (MV) of list 0 and MV of list 1 are combined to form a single prediction signal. In decoder-side motion vector refinement (DMVR) method, the two motion vectors of bi-prediction are further refined by a bilateral template matching process. The bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral template and the reconstructed samples in the reference picture in order to obtain refined MVs without transmitting additional motion information.
[0236] As shown in FIG. 24 In DMVR, the bilateral template is generated from the initial MV0 of list 0 and MV1 of list 1 as a weighted combination (i.e. average) of the two prediction blocks, respectively. The template matching operation includes computing a cost metric between the generated template and a region of samples in the reference picture (around the initial prediction block). For each of the two reference pictures, the MV that yields the minimum template cost is considered as the updated MV of that list to replace the original template. In JEM, for each list, nine MV candidates are searched. The nine MV candidates include the original MV and 8 surrounding MVs with one luma sample offset in horizontal direction or vertical direction or both relative to the original MV. Finally, as shown in FIG. 24 The two new MVs, i.e. MV0' and MV1', are used to generate the final bi-prediction result. Sum of absolute difference (SAD) is used as the cost metric.
[0237] DMVR is applied to bi-predictive Merge mode, which uses one MV from a past reference picture and another MV from a future reference picture without transmitting additional syntax elements. In JEM, DMVR is not applied when LIC, affine motion, FRUC or sub-CU Merge candidate is enabled for a CU.
[0238] 2.2.9 Example of symmetric motion vector difference
[0239] In [8], symmetric motion vector difference (SMVD) is proposed to code MVD more efficiently.
[0240] First, in slice level, the variables BiDirPredFlag, RefIdxSymL0 and RefIdxSymL1 are derived as follows:
[0241] Search the forward reference picture in reference picture list 0 that is closest to the current picture. If found, set RefIdxSymL0 equal to the reference index of the forward picture.
[0242] Search for the closest backward reference picture in reference picture list 1 to the current picture. If found, set RefIdxSymL1 equal to the reference index of the backward picture.
[0243] If both the forward and backward pictures are found, set BiDirPredFlag equal to 1.
[0244] Otherwise, the following applies:
[0245] Search for the closest backward reference picture in reference picture list 0 to the current reference picture. If found, set RefIdxSymL0 equal to the reference index of the backward picture.
[0246] Search for the closest forward reference picture in reference picture list 1 to the current reference picture. If found, set RefIdxSymL1 equal to the reference index of the forward picture.
[0247] If both the forward and backward pictures are found, set BiDirPredFlag equal to 1. Otherwise, set BiDirPredFlag equal to 0.
[0248] Second, at the CU level, if the prediction direction of the CU is bi-prediction and BiDirPredFlag is equal to 1, explicitly signal a symmetric mode flag that indicates whether the symmetric mode is used or not.
[0249] When the flag is true, only mvp_l0_flag, mvp_l1_flag and MVD0 are explicitly signaled. For list 0 and list 1, the reference indices are set equal to RefIdxSymL0, RefIdxSymL1, respectively. MVD1 is set equal to -MVD0. The final motion vector is given by the following equation.
[0250]
[0251] FIG. 27 An example of the symmetric mode is shown.
[0252] The modification of the coding unit syntax is shown in Table 3.
[0253] Table 3: Modification of the coding unit syntax
[0254]
[0255]
[0256]
[0257] 2.10.1 Symmetric MVD for affine bi-prediction coding
[0258] SMVD of affine mode is proposed, which extends the symmetric MVD mode to affine bi-prediction. When the symmetric MVD mode is applied for affine bi-prediction coding, the MVD of the control points is not signaled but derived. Based on the assumption of linear motion, the MVD of the top-left control point of Listl is derived from ListO. The MVD of other control points of ListO is set to 0
[0259] 2.11 Context-based adaptive binary arithmetic coding (CABAC)
[0260] 2.11.1 CABAC design in HEVC
[0261] 2.11.1.1 Context representation and initialization process in HEVC
[0262] In HEVC, for each context variable, two variables pStateIdx and valMps are initialized.
[0263] From the 8-bit table entry initValue, two 4-bit variables slopeIdx and offsetIdx are derived as follows:
[0264] slopeIdx = initValue » 4
[0265] offsetIdx = initValue & 15 (34)
[0266] The variables m and n used in the initialization of the context variable are derived from slopeIdx and offsetIdx as follows:
[0267] m = slopeIdx * 5 - 45
[0268] n = (offsetIdx « 3) - 16 (35)
[0269] The two values assigned to pStateIdx and valMps for initialization are derived from the quantization parameter of the luma of the slice represented by SliceQpY. Given the variables m and n, the initialization is specified as follows:
[0270]
[0271] 2.11.1.2 State transition process in HEVC
[0272] The inputs of this process are the current pStateIdx, the decoded value binVal and the valMps value of the context variable associated with ctxTable and ctxIdx.
[0273] The output of this process is the update pStateIdx and the update valMps of the context variables associated with ctxIdx.
[0274] Depending on the decoded value binVal, the update of the two variables pStateIdx and valMps associated with ctxIdx is derived as in (37):
[0275]
[0276] 2.11.2 CABAC design in VVC
[0277] The Context-based Adaptive Binary Arithmetic Coder (BAC) in VVC has been changed from that in HEVC in terms of the context update process and the arithmetic coder.
[0278] The following is an overview of the recently adopted proposal (JVET-M0473, CE test 5.1.13).
[0279] Table 4: Overview of CABAC modifications in VVC
[0280]
[0281] 2.11.2.1 Context initialization process in VVC
[0282] In VVC, the two values assigned to pStateIdx0 and pStateIdx1 for initialization are derived from SliceQpY. Given the variables m and n, the initialization is specified as follows:
[0283]
[0284] 2.11.2.2 State transition process in VVC
[0285] The input of this process is the current pStateIdx0 and the current pStateIdx1 and the decoded value binVal.
[0286] The output of this process is the update pStateIdx0 and the update pStateIdx1 of the context variables associated with ctxIdx.
[0287] The variables shift0 (corresponding to variable a in the Overview of CABAC modifications in VVC Table 4) and shift1 (corresponding to variable b in the Overview of CABAC modifications in VVC Table 4) are derived from the shiftIdx values associated with ctxTable and ctxInc.
[0288] shift0 = (shiftIdx » 2) + 2
[0289] shift1 = (shiftldx & 3) + 3 + shift0 (39)
[0290] Depending on the decoded value binVal, the update of the two variables pStateIdx0 and pStateIdx1 associated with ctxldx is derived as follows:
[0291] pStateIdx0 = pStateIdx0 - (pStateIdx0 » shift0) + (1023 * binVal » shift0)
[0292] pStateIdx1 = pStateIdx1 - (pStateIdx1 » shift1) + (16383 * binVal » shift1) (40)
[0293] 3. Drawbacks of existing implementations
[0294] In some existing implementations, when the MV / MV difference (MVD) can be selected from a set of multiple MV / MVD precisions for an affine coded block, it is still not determined how to obtain a more accurate motion vector.
[0295] In other existing implementations, the MV / MVD precision information also plays an important role in determining the overall coding gain applied to AMVR for affine mode, but it is still not determined how to achieve this goal.
[0296] 4. Example method of MV prediction (MVP) for affine mode with AMVR
[0297] Embodiments of the presently disclosed technology overcome the drawbacks of existing implementations, thereby providing video coding with higher coding efficiency. Based on the disclosed technology, derivation and signaling of motion vector prediction for affine mode with adaptive motion vector resolution (AMVR) can enhance existing and future video coding standards, as set forth in the examples described below for various implementations. The examples of the disclosed technology provided below explain general concepts and are not meant to be interpreted as limiting. In the examples, various features described in these examples can be combined unless explicitly indicated to the contrary.
[0298] In some embodiments, the following examples can apply to affine mode or normal mode when AMVR is applied. These examples assume that the precision Prec (i.e., MVs have 1 / (2Prec) precision) is used for coding MVD in AF INTER mode or for coding MVD in normal inter mode. Motion vector prediction (e.g., inherited from neighboring block MVs) and its precision are denoted by MVPred (MVPred X , MVPred Y ) and PredPrec, respectively.
[0299] In the following discussion, SatShift(x, n) is defined as
[0300]
[0301] Shift(x, n) is defined as Shift(x, n) = (x + offset0) » n. In one example, offset0 and / or offset1 is set to (1 « n) » 1 or (1 « (n - 1)). In another example, offset0 and / or offset1 is set to 0. In another example, offset0 = offset1 = ((1 « n) » 1) - 1 or ((1 « (n - 1)) - 1.
[0302] In the following discussion, an operation between two motion vectors means that the operation will be applied to both components of the motion vectors. For example, MV3 = MV1 + MV2 is equivalent to MV3x = MV1x + MV2x and MV3y = MV1y + MV2y. Alternatively, the operation can be applied to only the horizontal or vertical components of the two motion vectors.
[0303] Example 1. The final MV precision can be kept unchanged, i.e., the same as the precision of the motion vector to be stored.
[0304] (a) In one example, the final MV precision can be set to 1 / 16 pixel or 1 / 8 pixel.
[0305] (b) In one example, the signaled MVD can be scaled first, and then added to the MVP to form the final MV of a block.
[0306] Example 2. The MVP, which is derived directly from a neighboring block (e.g., spatially or temporally) or a default MVP, can be modified first, and then added to the signaled MVD to form the final MV of the (current) block.
[0307] (a) Alternatively, whether and how to apply the modification of the MVP can be different for different Prec values.
[0308] (b) In one example, if Prec is greater than 1 (i.e., MVD has fractional precision), the precision of the neighboring MVs is unchanged and no scaling is performed.
[0309] (c) In one example, if Prec is equal to 1 (i.e., MVD has 1-pixel precision), the MV
[0310] predictor (i.e., the MV of the neighboring block) needs to be scaled, e.g., according to Example 4(b) of PCT application PCT / CN2018 / 104723.
[0311] (d) In one example, if Prec is less than 1 (i.e., MVD has 4-pixel precision), the MV
[0312] predictor (i.e., the MV of the neighboring block) needs to be scaled, e.g., according to Example 4(b) of PCT application PCT / CN2018 / 104723.
[0313] Example 3. In one example, if the signaled precision of the MVD is the same as the precision of the stored MV, no scaling is needed after the affine MV is reconstructed, otherwise, the MV is reconstructed using the signaled precision of the MVD and then scaled to the precision of the stored MV.
[0314] Example 4. In one example, the normal inter mode and AF INTER mode can choose implementation based on the different examples above.
[0315] Example 5. In one example, the following semantics can be used to signal the syntax element indicating the MV / MVD precision for affine mode:
[0316] (a) In one example, the syntax element equal to 0, 1 and 2 indicates 1 / 4-pixel, 1 / 16-pixel and 1-pixel MV precision, respectively.
[0317] (b) Alternatively, in affine mode, the syntax element equal to 0, 1 and 2 indicates 1 / 4-pixel, 1-pixel and 1 / 16-pixel MV precision, respectively.
[0318] (c) Alternatively, in affine mode, the syntax element equal to 0, 1 and 2 indicates 1 / 16-pixel, 1 / 4-pixel and 1-pixel MV precision, respectively.
[0319] Example 6. In one example, whether to enable or disable AMVR for affine mode can be signaled in SPS, PPS, VPS, sequence / picture / slice header / tile, etc.
[0320] Example 7. In one example, an indication of allowed MV / MVD precision can be signaled in SPS, PPS, VPS, sequence / picture / slice header / tile, etc.
[0321] (a) An indication of selected MVD precision can be signaled per coding tree unit (CTU) and / or per region.
[0322] (b) The set of allowed MV / MVD precision can depend on the coding mode (e.g., affine or non-affine) of the current block.
[0323] (c) The set of allowed MV / MVD precision can depend on the slice type / time domain layer index / low delay check flag.
[0324] (d) The set of allowed MV / MVD precision can depend on the block size and / or block shape of the current block or neighboring blocks.
[0325] (e) The set of allowed MV / MVD precision can depend on the precision of MVs to be stored in the decoded picture buffer.
[0326] (i) In one example, if the stored MV is in X pixels, the set of allowed MV / MVD precision can have at least X pixels.
[0327] Improvements to affine mode supporting AMVR
[0328] Example 8. The set of allowed MVD precision can be different from picture to picture, slice to slice, or block to block.
[0329] a. In one example, the set of allowed MVD precision can depend on the coding information, e.g., block size, block shape, etc.
[0330] b. The set of allowed MV precision can be predefined, such as {1 / 16, 1 / 4, 1}.
[0331] c. An indication of allowed MV precision can be signaled in SPS / PPS / VPS / sequence header / picture header / slice header / CTU group, etc.
[0332] d. The signaling of selected MV precision from the set of allowed MV precision further depends on the number of allowed MV precision for the block.
[0333] Example 9. Signaling syntax elements to a decoder to indicate the MVD precision used in affine inter mode.
[0334] a. In one example, only a single syntax element is used to indicate the MVD precision applied to affine mode and AMVR mode.
[0335] i. In one example, the same semantics are used, that is, for AMVR and affine modes, the same values of syntax elements are mapped to the same MVD precision.
[0336] ii. Alternatively, the semantics of a single syntax element differs for AMVR and affine modes. In other words, for AMVR and affine modes, the same value of a syntax element can be mapped to different MVD precisions.
[0337] b. In one example, when the affine mode uses the same set of MVD precision as AMVR (e.g., MVD precision is set to {1, 1 / 4, 4} pixels), the MVD precision syntax element in AMVR is reused in the affine mode, that is, only a single syntax element is used.
[0338] i. Alternatively, when encoding the syntax element in the CABAC encoder / decoder...
[0339] During decoding, the same or different context models can be used for AMVR and affine modes.
[0340] ii. Alternatively, this syntax element can have different semantics in AMVR and affine mode. For example, syntax elements equal to 0, 1, and 2 indicate 1 / 4 pixel, 1 / 2, 1 / 3 pixel, 1 / 4 pixel, 1 / 2 ...
[0341] 1-pixel and 4-pixel MV precision, while in affine mode, the syntax elements are equal to 0, 1, and 2.
[0342] The pixels indicate 1 / 4 pixel, 1 / 16 pixel, and 1 pixel MV precision, respectively.
[0343] c. In one example, when the affine mode uses the same number of MVD precisions as AMVR but a different set of MVD precisions (e.g., the MVD precision is set to {1, 1 / 4, 4} pixels for AMVR, while the MVD precision is set to {1 / 16, 1 / 4, 1} pixels for affine mode), the MVD precision syntax element in AMVR is reused in the affine mode; that is, only a single syntax element is used.
[0344] i. Alternatively, when encoding the syntax element in the CABAC encoder / decoder...
[0345] During decoding, the same or different context models can be used for AMVR and affine modes.
[0346] ii. Alternatively, in addition, the syntax element can have different semantics in AMVR and affine mode.
[0347] d. In one example, affine mode uses less MVD precision than AMVR, the MVD precision syntax element in AMVR is reused in affine mode. However, only a subset of the syntax element values are valid for affine mode.
[0348] i. Alternatively, in addition, when the syntax element is coded in a CABAC encoder / decoder
[0349] When decoded, the same or different context model can be used for AMVR and affine mode.
[0350] ii. Alternatively, in addition, the syntax element can have different semantics in AMVR and affine mode.
[0351] e. In one example, affine mode uses more MVD precision than AMVR, the MVD precision syntax element in AMVR is reused in affine mode. However, such syntax element is extended to allow more values in affine mode.
[0352] i. Alternatively, in addition, when the syntax element is coded in a CABAC encoder / decoder
[0353] When decoded, the same or different context model can be used for AMVR and affine mode.
[0354] ii. Alternatively, in addition, the syntax element can have different semantics in AMVR and affine mode.
[0355] f. In one example, a new syntax element is used to code the MVD precision for affine mode, i.e., two different syntax elements are used to code the MVD precision for AMVR and affine mode.
[0356] g. The syntax to indicate the MVD precision for affine mode can be signaled under one or all of the following conditions:
[0357] i. All control points have non-zero MVD.
[0358] ii. At least one control point has non-zero MVD.
[0359] iii. One control point (e.g., the first CPMV) has non-zero MVD
[0360] In this case, when any or all of the above conditions fail, the MVD precision does not need to be signaled.
[0361] h. The syntax elements for indicating the MVD precision for affine mode or AMVR mode can be context coded, and the context depends on the coded information.
[0362] i. In one example, the context can depend on the current
[0363] whether the block is coded using affine mode.
[0364] i. In one example, the context can depend on the block size / block shape / MVD precision of neighboring blocks / temporal layer index / prediction direction, etc.
[0365] j. Whether to enable or disable the usage of multiple MVD precisions for affine mode can be signaled in SPS / PPS / VPS / sequence header / picture header / slice header / CTU group, etc.
[0366] i. In one example, whether to signal the information to enable or disable the usage of multiple MVD precisions for affine mode can depend on other syntax elements. For example, when affine mode is enabled, signal the information to enable or disable the usage of multiple MVs and / or MVPs and / or MVD precisions for affine mode; when affine mode is disabled, do not signal the information to enable or disable the usage of multiple MVs and / or MVPs and / or MVD precisions for affine mode and infer it as 0.
[0367] k. Alternatively, multiple syntax elements can be signaled to indicate the MVs and / or
[0368] MVPs and / or MVD precisions (in the following discussion, they are all referred to as “MVD precisions”) used in affine inter mode.
[0369] i. In one example, the syntax elements for indicating the MVD precisions used in affine inter mode and normal inter mode can be different.
[0370] 1. The number of syntax elements for indicating the MVD precisions used in affine inter mode and normal inter mode can be different.
[0371] 2. The semantics of the syntax elements for indicating the MVD precisions used in affine inter mode and normal inter mode can be different.
[0372] 3. The context models for coding one syntax element to indicate the MVD precisions used in affine inter mode and normal inter mode in arithmetic coding can be different.
[0373] 4. The method of deriving a context model in arithmetic coding to code a syntax element to indicate the MVD precision used in affine inter mode and normal inter mode can be different.
[0374] ii. In one example, a first syntax element (e.g., amvr_flag) can be signaled to indicate whether AMVR is applied in the affine coded block.
[0375] 1. Conditionally signal the first syntax element.
[0376] a. In one example, the signaling of the first syntax element (amvr_flag) is skipped when the current block is coded using a certain mode (e.g., CPR / IBC mode).
[0377] b. In one example, the signaling of the first syntax element (amvr_flag) is skipped when all MVDs (including horizontal and vertical components) of CPMVs are zero.
[0378] c. In one example, the signaling of the first syntax element (amvr_flag) is skipped when MVDs (including horizontal and vertical components) of a selected CPMV are zero.
[0379] i. In one example, the MVD of the selected CPMV is to be coded / decoded
[0380] the MVD of the first CPMV decoded.
[0381] d. In one example, the signaling of the first syntax element (amvr_flag) is skipped when the usage of multiple MVD precisions for the affine coded block is disabled.
[0382] e. In one example, the first syntax element can be signaled under the following conditions:
[0383] i. the usage of multiple MVD precisions for the affine coded block is enabled, and the current block is coded using affine mode;
[0384] ii. alternatively, the usage of multiple MVD precisions for the affine coded block is enabled, the current block is coded using affine mode, and at least one component of the MVD of a CPMV is not equal to 0.
[0385] iii. alternatively, the usage of multiple MVD precisions for the affine coded block is enabled, the current block is coded using affine mode, and at least one component of the MVD of a selected CPMV is not equal to 0.
[0386] 1. In one example, the MVD of the selected CPMV is the MVD of the first CPMV to be coded / decoded.
[0387] 2. When AMVR is not applied to the affine coded block or the first syntax element is not present, a default MV and / or MVD precision is adopted.
[0388] a. In one example, the default precision is 1 / 4-pel.
[0389] b. Alternatively, the default precision is set to the precision used in the motion compensation of the affine coded block.
[0390] 3. For example, if the MVD precision for affine mode is 1 / 4-pel; otherwise the MVD precision for affine mode can be other values.
[0391] a. Alternatively, in addition, an additional MVD precision can be further signaled via a second syntax element.
[0392] iii. In one example, a second syntax element (e.g., amvr_coarse_precision_flag) can be signaled to indicate the MVD precision for affine mode.
[0393] 1. In one example, whether the second syntax element is signaled can depend on the first syntax element. For example, the second syntax element is signaled only when the first syntax element is 1.
[0394] 2. In one example, if the second syntax element is 0, the MVD precision for affine mode is 1-pel; otherwise, the MVD precision for affine mode is 1 / 16-pel.
[0395] 3. In one example, if the second syntax element is 0, the MVD precision for affine mode
[0396] is 1 / 16-pel; otherwise, the MVD precision for affine mode is full-pel.
[0397] iv. In one example, the syntax element used to indicate the MVD precision used in affine inter mode shares the same context model as the syntax element with the same name but used to indicate the MVD precision used in normal inter mode.
[0398] 1. Alternatively, the syntax element used to indicate the MVD precision used in affine inter mode uses a different context model than the syntax element with the same name but used to indicate the MVD precision used in normal inter mode.
[0399] Example 10.Whether or how to apply AMVR on an affine coded block can depend on the reference picture of the current block.
[0400] a. In one example, if the reference picture is the current picture, then AMVR is not applied, i.e., intra block copy is applied in the current block.
[0401] Fast algorithm for AMVR in encoder affine mode
[0402] Let the RD cost (actual RD cost, or SATD / SSE / SAD cost plus bit cost) of the affine mode and AMVP mode be affineCosti and amvpCosti for IMV = i, where i = 0, 1 or 2. In this context, IMV = 0 means ¼-pel MV, and IMV = 1 means integer MV for AMVP mode and 1 / 16-pel MV for affine mode, and IMV = 2 means 4-pel MV for AMVP mode and integer MV for affine mode. Let the RD cost of the Merge mode be mergeCost.
[0403] Example 11. If the best mode of its parent CU is not AF INTER mode or AF MERGE mode, then propose to disable AMVR for the affine mode of the current CU.
[0404] Alternatively, if the best mode of its parent CU is not AF INTER mode, then disable AMVR for the affine mode of the current CU.
[0405] Example 12. If affineCost0> th1* amvpCost0, where th1 is a positive threshold, then propose to disable AMVR for the affine mode.
[0406] a. Alternatively, in addition, if min(affineCost0, amvpCost0) > th2* mergeCost, where th2 is a positive threshold, then disable AMVR for the affine mode.
[0407] b. Alternatively, in addition, if affineCost0> th3* affineCost1, where th3 is a positive threshold, then disable integer MV for the affine mode.
[0408] Example 12. If amvpCost0> th4* affineCost0, where th4 is a positive threshold, then propose to disable AMVR for the AMVP mode.
[0409] a. Alternatively, if min(affineCost0, amvpCost0) > th5 * mergeCost, where th5 is a positive threshold, disable AMVR for AMVP mode.
[0410] Example 13. The 4 / 6-parameter affine model obtained in one MV precision can be used as the candidate starting search point for other MV precisions.
[0411] a. In one example, the 4 / 6-parameter affine model obtained in 1 / 16 MV can be used as the candidate starting search point for other MV precisions.
[0412] b. In one example, the 4 / 6-parameter affine model obtained in 1 / 4 MV can be used as the candidate starting search point for other MV precisions.
[0413] Example 14. If the parent block of the current block does not select affine mode, do not check AMVR for affine mode at the encoder for the current block.
[0414] Example 15. The rate-distortion calculation of the MV precision of the affine-coded block in the current slice / tile / CTU row can be early terminated with the statistics of the usage of different MV precisions of the affine-coded blocks in the previously coded frame / slice / tile / CTU row.
[0415] a. In one example, the percentage of the affine-coded blocks with a certain MV precision is recorded. If the percentage is too low, skip checking the corresponding MV precision.
[0416] b. In one example, the previously coded frames with the same temporal layer are utilized to decide whether to skip a certain MV precision.
[0417] Contexts for coding affine AMVR
[0418] Example 16. For each context used to code the affine AMVR code, a variable (denoted by shiftldx) is proposed to be set to control the speed of two probability updates associated with this context.
[0419] a. In one example, the faster update speed is defined by (shiftldx » 2) + 2.
[0420] b. In one example, the slower update speed is defined by (shiftldx & 3) + 3 + shift0.
[0421] c. In one example, the consistent bitstream should follow the following rule: the derived faster update speed should be in the range of [2, 5], inclusive.
[0422] d. In one example, the consistent bitstream should follow the rule that the derived faster update speed should be in the range of [3, 6], inclusive.
[0423] Example 17. When coding the AMVR mode of a block, the proposal does not allow the affine AMVR mode information of neighboring blocks to be used for context modeling.
[0424] a. In one example, the AMVR mode index of neighboring blocks can be utilized, and the affine AMVR mode information of neighboring blocks is excluded. An example is shown in Table 5 (including Table 5-1 and 5-2), where (xNbL, yNbL) and (xNbA, yNbA) represent the left and above neighboring blocks. In one example, the context index offset ctxInc = (condL && availableL) + (condA && availableA) + ctxSetIdx * 3.
[0425] Table 5-1 ctxInc specification using left and above syntax elements
[0426] Table 5-2 ctxInc specification using left and above syntax elements
[0427] Table 5-1 ctxInc specification using left and above syntax elements
[0428]
[0429] Table 5-2 ctxInc specification using left and above syntax elements
[0430] b. Alternatively, the affine AMVR mode information of neighboring blocks can be further utilized, but using a function instead of directly using. In one example, the function func described in Table 6-1 can return true when the affine coded neighboring block’s amvr_mode[xNbL]
[0431] [yNbL] indicates a certain MV precision (e.g., ¼ pixel MV precision). In one example, the function func described in Table 6-2 can return true when the affine coded neighboring block’s amvr_flag[xNbL]
[0432] [yNbL] indicates a certain MV precision (e.g., ¼ pixel MV precision).
[0433] Table 6-1 ctxInc specification using left and above syntax elements
[0434]
[0435] Table 6-2 ctxInc specification using left and above syntax elements
[0436]
[0437]
[0438] c. Alternatively, the affine AMVR mode information of the neighboring block can be further used for coding the first syntax element (e.g., amvr_flag) of the AMVR mode (applied to normal inter mode). Table 6-3 and 6-4 give some examples.
[0439] Table 6-3 ctxInc specification using left and above syntax elements
[0440]
[0441] Table 6-4 ctxInc specification using left and above syntax elements
[0442]
[0443] d. When the AMVR mode information is represented by multiple syntax elements (e.g., the first and second syntax elements, represented by amvr_flag,
[0444] amvr_coarse_precision_flag), the above syntax amvr_mode can be replaced by any of the multiple syntax elements, and the above method can still be applied.
[0445] Example 18. When coding the affine AMVR mode, it is proposed to use the AMVR mode information of the neighboring block for context coding.
[0446] a. In one example, the AMVR mode information of the neighboring block is used directly. An example is shown in Table 7. Alternatively, in addition, a context index offset ctxInc = (condL && availableL) +
[0447] (condA && availableA) + ctxSetIdx * 3 is allowed.
[0448] Table 7 ctxInc specification using left and above syntax elements
[0449]
[0450] b. Alternatively, the AMVR mode information of the neighboring block is not allowed to be used for context modeling. An example is shown in Table 8.
[0451] Table 8 ctxInc specification using left and above syntax elements
[0452]
[0453] c. Alternatively, the AMVR mode information of neighboring blocks can be further utilized, but using a function instead of directly using. In one example, when the amvr mode [xNbL] [yNbL] of a non- affine coded neighboring block is available, the function func can be used to determine whether to use the SMVD mode or not.
[0454] The function func described in Table 9 can return true when a certain MV precision (such as ¼ pixel MV precision) is indicated.
[0455] Table 9 specifies the ctxInc using the left and above syntax elements
[0456]
[0457]
[0458] d. When the affine AMVR mode information is represented by multiple syntax elements (e.g., the first and second syntax elements, represented by amvr flag, amvr coarse precision flag), the above syntax amvr mode can be replaced by any one of the multiple syntax elements, and the above method can still be applied.
[0459] Fast algorithm for SMVD and affine SMVD
[0460] When checking the SMVD mode, assume the currently selected best mode is CurBestMode, and the MVD precision in AMVR is MvdPrec or the MVD precision in affine AMVR is MvdPrecAff.
[0461] Example 19. Depending on the currently selected best mode (i.e., CurBestMode), the MVD precision in AMVR, the SMVD mode can be skipped.
[0462] a. In one example, if CurBestMode is Merge mode or / and UMVE mode, the SMVD mode can not be checked.
[0463] b. In one example, if CurBestMode is coded without using the SMVD mode, the SMVD mode can not be checked.
[0464] c. In one example, if CurBestMode is affine mode, the SMVD mode can not be checked.
[0465] d. In one example, if CurBestMode is sub-block Merge mode, the SMVD mode can not be checked.
[0466] e. In one example, if CurBestMode is affine SMVD mode, the SMVD mode can not be checked.
[0467] f. In one example, if CurBestMode is affine Merge mode, the SMVD mode can not be checked.
[0468] g. In one example, the above fast methods, i.e., bullets 13.a-13.f, can be applied only for some MVD precisions.
[0469] i. In one example, the above fast methods can be applied only when the MVD precision is greater than or equal to a precision (e.g., integer pixel precision).
[0470] ii. In one example, the above fast methods can be applied only when the MVD precision is greater than a precision (e.g., integer pixel precision).
[0471] iii. In one example, the above fast methods can be applied only when the MVD precision is less than or equal to a precision (e.g., integer pixel precision).
[0472] iv. In one example, the above fast methods can be applied only when the MVD precision is less than a precision (e.g., integer pixel precision).
[0473] Example 20. Depending on the currently selected best mode (i.e., CurBestMode), the MVD precision in affine AMVR, the affine SMVD mode can be skipped.
[0474] a. In one example, if CurBestMode is Merge mode or / and UMVE mode, the affine SMVD mode can not be checked.
[0475] b. In one example, if CurBestMode is not coded using affine SMVD mode, the affine SMVD mode can not be checked.
[0476] c. In one example, if CurBestMode is sub-block Merge mode, the affine SMVD mode can not be checked.
[0477] SMVD mode.
[0478] d. In one example, if CurBestMode is SMVD mode, the affine SMVD mode can not be checked.
[0479] e. In one example, if CurBestMode is affine Merge mode, the affine SMVD mode can not be checked.
[0480] SMVD mode.
[0481] f. In one example, the above fast method, i.e., bullets 20.a to 20.e, can be applied only for some
[0482] MVD precision.
[0483] i. In one example, the above fast method can be applied only when the affine MVD precision is greater than or equal to the precision (e.g., integer pixel precision).
[0484] ii. In one example, the above fast method can be applied only when the affine MVD precision is greater than the precision (e.g., integer pixel precision).
[0485] iii. In one example, the above fast method can be applied only when the affine MVD precision is less than or equal to the precision (e.g., integer pixel precision).
[0486] iv. In one example, the above fast method can be applied only when the affine MVD precision is less than the precision (e.g., integer pixel precision).
[0487] Example 21. The above proposed methods can be applied under certain conditions, such as block size, slice / picture / tile type, or motion information.
[0488] a. In one example, the proposed method is not allowed when the block size contains less than M*H samples (e.g., 16 or 32 or 64 luma samples).
[0489] b. Alternatively, the proposed method is not allowed when the minimum dimension of the width or / and height of the block is less than or not greater than X. In one example, X is set to 8.
[0490] c. Alternatively, the proposed method is not allowed when the minimum dimension of the width or / and height of the block is not less than X.
[0491] In one example, X is set to 8.
[0492] d. Alternatively, the proposed method is not allowed when the width of the block is >th1 or >=th1 and / or the height of the block is >th2 or >=th2. In one example, th1 and / or th2 is set to 8.
[0493] e. Alternatively, the proposed method is not allowed when the width of the block is <th1 or <=th1 and / or the height of the block is <th2 or <=th2. In one example, th1 and / or th2 is set to 8.
[0494] f.Alternatively, whether to enable or disable the above methods and / or which method to apply can depend on the block size, video processing data unit (VPDU), picture type, low delay check flag, coding information of the current block such as reference picture, uni- or bi-prediction, or previously coded blocks.
[0495] Example 22. When IBC (also known as current picture reference (CPR)) is applied or not applied, the AMVR method for affine mode can be performed in different ways.
[0496] a.In one example, if a block is coded by IBC, AMVR for affine mode can not be used.
[0497] b.In one example, if a block is coded by IBC, AMVR for affine mode can be used, but the candidate MV / MVD / MVP precision can be different from the precision for affine coded blocks that are not coded by IBC.
[0498] Example 23. All the term “slice” in the document can be replaced by “tile group” or “tile”.
[0499] Example 24. In VPS / SPS / PPS / slice header / tile group header, a syntax element (e.g., no_amvr_constraint_flag) equal to 1 specifies a requirement on bitstream conformance that both a syntax element (e.g., sps_amvr_enabled_flag) indicating whether AMVR is enabled and a syntax element (e.g., sps_affine_amvr_enabled_flag) indicating whether affine AMVR is enabled should be equal to 0. A syntax element (e.g., no_amvr_constraint_flag) equal to 0 does not impose a constraint.
[0500] Example 25. In VPS / SPS / PPS / slice header / tile group header or other video data unit, a syntax element (e.g., no_affine_amvr_constraint_flag) can be signaled.
[0501] a.In one example, no_affine_amvr_constraint_flag equal to 1 specifies a requirement on bitstream conformance that a syntax element (e.g.,
[0502] sps_affine_amvr_enabled_flag) should be equal to 0. A syntax element (e.g.,
[0503] no_affine_amvr_constraint_flag) is equal to 0 no constraint is applied.
[0504] Example 26. A second syntax element indicating a coarse motion precision (e.g. amvr_coarse_precision_flag) can be coded with multiple contexts.
[0505] a. In one example, two contexts can be utilized.
[0506] b. In one example, the selection of the context can depend on whether the current block is affine coded.
[0507] c. In one example, for the first syntax, only one context can be used to code it, and for the second syntax, also only one context can be used to code it.
[0508] d. In one example, for the first syntax, only one context can be used to code it, and for the second syntax, it can also be bypass coded.
[0509] e. In one example, for the first syntax, it can be bypass coded, and for the second syntax, it can also be bypass coded.
[0510] f. In one example, for all syntax elements related to motion vector precision, they can be bypass coded.
[0511] Example 27. For example, only the first bin of the syntax element amvr_mode is coded using the arithmetic coding context(s). All the following bins of amvr_mode are coded as bypass coding.
[0512] a. The above disclosed method can also be applied to other syntax elements.
[0513] b. For example, only the first bin of the syntax element SE is coded using the arithmetic coding context(s). All the following bins of SE are coded as bypass coding. SE can be
[0514] 1) alf_ctb_flag
[0515] 2) sao_merge_left_flag
[0516] 3) sao_merge_up_flag
[0517] 4) sao_type_idx_luma
[0518] 5) sao type idx chroma
[0519] 6) split cu flag
[0520] 7) split qt flag
[0521] 8) mtt split cu vertical flag
[0522] 9) mtt split cu binary flag
[0523] 10) cu skip flag
[0524] 11) pred mode ibc flag
[0525] 12) pred mode flag
[0526] 13) intra luma ref idx
[0527] 14) intra subpartitions mode flag
[0528] 15) intra subpartition split flag
[0529] 16) intra luma mpm flag
[0530] 17) intra chroma pred mode
[0531] 18) merge flag
[0532] 19) inter pred idc
[0533] 20) inter affine flag
[0534] 21) cu affine type flag
[0535] 22) ref idx l0
[0536] 23) mvp l0 flag
[0537] 24) ref idx l1
[0538] 25) mvp l1 flag
[0539] 26) avmr flag
[0540] 27) amvr_precision_flag
[0541] 28) gbi_idx
[0542] 29) cu_cbf
[0543] 30) cu_sbt_flag
[0544] 31) cu_sbt_quad_flag
[0545] 32) cu_sbt_horizontal_flag
[0546] 33) cu_sbt_pos_flag
[0547] 34) mmvd_flag
[0548] 35) mmvd_merge_flag
[0549] 36) mmvd_distance_idx
[0550] 37) ciip_flag
[0551] 38) ciip_luma_mpm_flag
[0552] 39) merge_subblock_flag
[0553] 40) merge_subblock_idx
[0554] 41) merge_triangle_flag
[0555] 42) merge_triangle_idx0
[0556] 43) merge_triangle_idx1
[0557] 44) merge_idx
[0558] 45) abs_mvd_greater0_flag
[0559] 46) abs_mvd_greater1_flag
[0560] 47) tu_cbf_luma
[0561] 48) tu_cbf_cb
[0562] 49) tu_cbf_cr
[0563] 50) cu_qp_delta_abs
[0564] 51) transform_skip_flag
[0565] 52) tu_mts_idx
[0566] 53) last_sig_coeff_x_prefix
[0567] 54) last_sig_coeff_y_prefix
[0568] 55) coded_sub_block_flag
[0569] 56) sig_coeff_flag
[0570] 57) par_level_flag
[0571] 58) abs_level_gt1_flag
[0572] 59) abs_level_gt3_flag
[0573] c. Alternatively, in addition, if the syntax element SE is a binary value (i.e. it can only equal 0 or 1), it can be context coded.
[0574] i. Alternatively, in addition, if the syntax element SE is a binary value (i.e. it can only equal 0 or 1), it can be bypass coded.
[0575] d. Alternatively, in addition, only one context can be used for coding the first bin.
[0576] Example 28. The precision of the motion vector prediction (MVP) or motion vector difference (MVD) or reconstructed motion vector (MV) can change depending on the signaled motion precision.
[0577] a. In one example, if the original prediction of the MVP is lower than (or not higher than) the target precision, then
[0578] MVP = MVP « s. s is an integer that can depend on the difference between the original precision and the target precision.
[0579] ii. Alternatively, if the original precision of the MVD is lower than (or not higher than) the target precision, then MVD = MVD « s. s is an integer that can depend on the difference between the original precision and the target precision.
[0580] iii. Alternatively, if the original precision of the MV is lower than (or not higher than) the target precision, then MV = MV « s. s is an integer that can depend on the difference between the original precision and the target precision.
[0581] b. In one example, if the original prediction of the MVP is higher than (or not lower than) the target precision, then MVP = Shift(MVP, s). s is an integer that can depend on the difference between the original precision and the target precision.
[0582] i. Alternatively, if the original precision of the MVD is higher than (or not lower than) the target precision, then MVD = Shift(MVD, s). s is an integer that can depend on the difference between the original precision and the target precision.
[0583] ii. Alternatively, if the original precision of the MV is higher than (or not lower than) the target precision, then MV = Shift(MV, s). s is an integer that can depend on the difference between the original precision and the target precision.
[0584] c. In one example, if the original prediction of the MVP is higher than (or not lower than) the target precision, then MVP = SatShift(MVP, s). s is an integer that can depend on the difference between the original precision and the target precision.
[0585] i. Alternatively, if the original precision of the MVD is higher than (or not lower than) the target precision, then MVD = SatShift(MVD, s). s is an integer that can depend on the difference between the original precision and the target precision.
[0586] ii. Alternatively, if the original precision of the MV is higher than (or not lower than) the target precision, then MV = SatShift(MV, s). s is an integer that can depend on the difference between the original precision and the target precision.
[0587] d. When the current block is not coded using affine mode, the above disclosed methods can be applied.
[0588] e. When the current block is coded using affine mode, the above disclosed methods can be applied.
[0589] 5. Embodiments
[0590] The highlighted part shows the modified specification.
[0591] 5.1 Embodiment 1: Indication of usage of affine AMVR mode
[0592] The indication can be signaled in SPS / PPS / VPS / APS / sequence header / picture header / slice header, etc. This section introduces the signaling in SPS.
[0593] 5.1.1 SPS syntax table
[0594]
[0595]
[0596] The alternative SPS syntax table is given as follows:
[0597]
[0598]
[0599] Semantics:
[0600] sps_affine_amvr_enabled_flag equal to 1 specifies that adaptive motion vector difference resolution is used in the motion vector coding of affine inter mode. amvr_enabled_flag equal to 0 specifies that adaptive motion vector difference resolution is not used in the motion vector coding of affine inter mode.
[0601] 5.2 Parsing process of AMVR mode information
[0602] The syntax of affine AMVR mode information can reuse that of AMVR mode information (applied to normal inter mode). Alternatively, different syntax elements can be used.
[0603] The affine AMVR mode information can be conditionally signaled. The following different embodiments show some examples of the conditions.
[0604] 5.2.1 Embodiment #1: CU syntax table
[0605]
[0606]
[0607]
[0608]
[0609] 5.2.2 Embodiment 2: Alternative CU syntax table design
[0610]
[0611]
[0612]
[0613]
[0614] 5.2.3 Embodiment 3: Third CU syntax table design
[0615]
[0616]
[0617]
[0618]
[0619] 5.2.4 Embodiment 4: Syntax table design with different syntax for AMVR and affine AMVR modes
[0620]
[0621]
[0622]
[0623] • In one example, conditionsA is defined as follows:
[0624] (sps_affine_amvr_enabled_flag && inter_affine_flag == 1 && (MvdCpL0[x0][y0][0][0]!= 0 || MvdCpL0[x0][y0][0][1]!= 0 || MvdCpL1[x0][y0][0][0]!= 0 || MvdCpL1[x0][y0][0][1]!= 0 || MvdCpL0[x0][y0][1][0]!= 0 || MvdCpL0[x0][y0][1][1]!= 0 || MvdCpL1[x0][y0][1][0]!= 0 || MvdCpL1[x0][y0][1][1]!= 0 || MvdCpL0[x0][y0][2][0]!= 0 || MvdCpL0[x0][y0][2][1]!= 0 || MvdCpL1[x0][y0][2][0]!= 0 || MvdCpL1[x0][y0][2][1]!= 0))
[0625] Alternatively, conditionsA is defined as follows:
[0626] (sps_affine_amvr_enabled_flag && inter_affine_flag == 1 && (MvdCpL0[x0][y0][0][0]!= 0 || MvdCpL0[x0][y0][0][1]!= 0 || MvdCpL1[x0][y0][0][0]!= 0 || MvdCpL1[x0][y0][0][1]!= 0))
[0627] Alternatively, conditionsA is defined as follows:
[0628] (sps_affine_amvr_enabled_flag && inter_affine_flag == 1 && (MvdCpLX[x0][y0][0][0]!= 0 || MvdCpLX[x0][y0][0][1]!= 0))
[0629] where X is 0 or 1.
[0630] Alternatively, conditionsA is defined as follows:
[0631] (sps_affine_amvr_enabled_flag && inter_affine_flag == 1)
[0632] • In one example, conditionsB is defined as follows:
[0633] !sps_cpr_enabled_flag ||!(inter_pred_idc[x0][y0] == PRED_L0 && ref_idx_l0[x0][y0] == num_ref_idx_l0_active_minus1)
[0634] Alternatively, conditionsB is defined as follows:
[0635] !sps_cpr_enabled_flag ||!(pred_mode[x0][y0] == CPR)
[0636] Alternatively, conditionsB is defined as follows:
[0637] !sps_ibc_enabled_flag ||!(pred_mode[x0][y0] == IBC)
[0638] Alternatively, conditionsB is defined as follows:
[0639] When AMVR or affine AMVR is coded with different syntax elements, the context modeling and / or contexts applied for affine AMVR for the embodiments in 5.5 can be applied accordingly.
[0640] 5.2.5 Semantics
[0641] Specifies the resolution of the motion vector difference. Array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. amvr_flag[ x0 ][ y0 ] equal to 0 specifies that the resolution of the motion vector difference is 1 / 4 of a luma sample. amvr_flag[ x0 ][ y0 ] equal to 1 specifies that the resolution of the motion vector difference is further specified by amvr_coarse_precision_flag[ x0 ][ y0 ].
[0642] When amvr_flag[ x0 ][ y0 ] is not present, it is inferred as follows:
[0643] — If sps_cpr_enabled_flag is equal to 1, amvr_flag[ x0 ][ y0 ] is inferred to be equal to 1.
[0644] — Otherwise (sps_cpr_enabled_flag is equal to 0), amvr_flag[ x0 ][ y0 ] is inferred to be equal to 0.
[0645] [ x0 ][ y0 ] equal to 1 specifies that the resolution of the motion vector difference is 4 luma samples when inter_affine_flag is equal to 0 and 1 luma sample when inter_affine_flag is equal to 1. Array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.
[0646] When amvr_coarse_precision_flag[ x0 ][ y0 ] is not present, it is inferred to be equal to 0.
[0647] If inter_affine_flag[ x0 ][ y0 ] is equal to 0, the variable MvShift is set equal to ( amvr_flag[ x0 ][ y0 ] + amvr_coarse_precision_flag[ x0 ][ y0 ] ) « 1 and the variables MvdL0[ x0 ][ y0 ][ 0 ], MvdL0[ x0 ][ y0 ][ 1 ], MvdL1[ x0 ][ y0 ][ 0 ], MvdL1[ x0 ][ y0 ][ 1 ] are modified as follows:
[0648] MvdL0[ x0 ][ y0 ][ 0 ] « ( MvShift + 2 ) (7-70)
[0649] MvdL0[ x0 ][ y0 ][ 1 ] « ( MvShift + 2 ) (7-71)
[0650] MvdL1[ x0 ][ y0 ][ 0 ] « ( MvShift + 2 ) (7-72)
[0651] MvdL1[ x0 ][ y0 ][ 1 ] « ( MvShift + 2 ) (7-73)
[0652] If inter affine flag[ x0 ][ y0 ] is equal to 1, the variable MvShift is set equal to ( amvr coarse precisoin flag? ( amvr coarse precisoin flag « 1 ) : ( - ( amvr flag « 1 ) ), and the variables MvdCpL0[ x0 ][ y0 ][ 0 ][ 0 ], MvdCpL0[ x0 ][ y0 ][ 0 ][ 1 ], MvdCpL0[ x0 ][ y0 ][ 1 ][ 0 ], MvdCpL0[ x0 ][ y0 ][ 1 ][ 1 ], MvdCpL0[ x0 ][ y0 ][ 2 ][ 0 ], MvdCpL0[ x0 ][ y0 ][ 2 ][ 1 ] are modified as follows:
[0653] MvdCpL0[ x0 ][ y0 ][ 0 ][ 0 ] « ( MvShift + 2 ) (7-73)
[0654] MvdCpL1[ x0 ][ y0 ][ 0 ][ 1 ] « ( MvShift + 2 ) (7-67)
[0655] MvdCpL0[ x0 ][ y0 ][ 1 ][ 0 ] « ( MvShift + 2 ) (7-66)
[0656] MvdCpL1[ x0 ][ y0 ][ 1 ][ 1 ] « ( MvShift + 2 ) (7-67)
[0657] MvdCpL0[ x0 ][ y0 ][ 2 ][ 0 ] = MvdCpL0[ x0 ][ y0 ][ 2 ][ 0 ] << ( MvShift + 2 ) (7-66)
[0658] MvdCpL1[ x0 ][ y0 ][ 2 ][ 1 ] = MvdCpL1[ x0 ][ y0 ][ 2 ][ 1 ] << ( MvShift + 2 ) (7-67)
[0659] Alternatively, if inter affme flag[ x0 ][ y0 ] is equal to 1, the variable MvShift is set equal to ( affine amvr coarse precisoin flag? ( affine amvr coarse precisoin flag « 1 ) : ( - ( affine amvr flag « 1 ) ) ).
[0660] 5.3 Integerization process of motion vector
[0661] The integerization process is modified when the given rightShift value is equal to 0 (occurring for 1 / 16-pel precision), the integerization offset is set to 0 instead of (1 « (rightShift - 1)).
[0662] For example, the subclause of the integerization process of MV is modified as follows:
[0663] The inputs of this process are:
[0664] - a motion vector mvX,
[0665] - a right shift parameter rightShift used for the integerization,
[0666] - a left shift parameter leftShift used for the resolution increase.
[0667] The output of this process is the integerized motion vector mvX.
[0668] The following applies for the integerization of mvX:
[0669] offset = ( rightShift == 0 )? 0 : ( 1 « ( rightShift - 1 ) ) (8-371)
[0670] mvX[ 0 ] = ( mvX[ 0 ] >= 0? ( mvX[ 0 ] + offset ) » rightShift : - ( ( -mvX[ 0 ] + offset ) » rightShift ) ) « leftShift (8-372)
[0671] mvX[1] = ( mvX[1] >= 0? ( mvX[1] + offset ) » rightShift : - ( ( -mvX[1] + offset ) » rightShift ) ) « leftShift (8-373)
[0672] 5.4 Decoding process
[0673] The rounding process invoked in the affine motion vector derivation process is performed using ( MvShift + 2 ) instead of a fixed input of 2.
[0674] Derivation process of luma affine control point motion vector predictor
[0675] The inputs of the process are:
[0676] - the luma position ( xCb, yCb ) of the top-left sample of the current luma coded block relative to the top-left luma sample of the current picture,
[0677] - two variables cbWidth and cbHeight specifying the width and height of the current luma coded block,
[0678] - the reference index of the current coding unit refldxLX, where X is either 0 or 1,
[0679] - the number of control point motion vectors numCpMv.
[0680] The output of the process is the luma affine control point motion vector predictor mvpCpLX[ cpldx ], where X is either 0 or 1, and cpldx = 0..numCpMv - 1.
[0681] To derive the control point motion vector predictor list, cpMvpListLX, where X is either 0 or 1, the following ordered steps apply:
[0682] The number of control point motion vector predictor candidates in the list numCpMvpCandLX is set equal to 0.
[0683] Both variables availableFlagA and availableFlagB are set to FALSE.
[0684] …
[0685] Invoke the process specified in clause 8.4.2.14 for motion vector scaling, where mvX is set equal to cpMvpLX[ cpIdx ], rightShift is set equal to ( MvShift + 2 ), and leftShift is set equal to ( MvShift + 2 ) as inputs, and the scaled cpMvpLX[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) as output.
[0686] …
[0687] The variable availableFlagA is set equal to TRUE
[0688] Invoke the process specified in clause 8.4.4.5 for derivation of luma affine control point motion vectors from neighboring blocks, where the luma coding block position ( xCb, yCb ), the luma coding block width and height ( cbWidth, cbHeight ), the neighboring luma coding block position ( xNb, yNb ), the neighboring luma coding block width and height ( nbW, nbH ), and the number of control point motion vectors numCpMv are inputs, and the control point motion vector prediction candidates cpMvpLY[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) are outputs.
[0689] Invoke the process specified in clause 8.4.2.14 for motion vector scaling, where mvX is set equal to cpMvpLY[ cpIdx ], rightShift is set equal to ( MvShift + 2 ), and leftShift is set equal to ( MvShift + 2 ) as inputs, and the scaled cpMvpLY[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) as output.
[0690] …
[0691] Invoke the process specified in clause 8.4.4.5 for derivation of luma affine control point motion vectors from neighboring blocks, where the luma coding block position ( xCb, yCb ), the luma coding block width and height ( cbWidth, cbHeight ), the neighboring luma coding block position ( xNb, yNb ), the neighboring luma coding block width and height ( nbW, nbH ), and the number of control point motion vectors numCpMv are inputs, and the control point motion vector prediction candidates cpMvpLX[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) are outputs.
[0692] Invoke the integer-pixel precision process specified in clause 8.4.2.14 with mvX set equal to cpMvpLX[ cpIdx ], rightShift set equal to ( MvShift + 2 ), and leftShift set equal to ( MvShift + 2 ) as inputs and the integer-precise cpMvpLX[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) as output.
[0693] The following assignments are made:
[0694] cpMvpListLX[ numCpMvpCandLX ][ 0 ] = cpMvpLX[ 0 ] (8-618)
[0695] cpMvpListLX[ numCpMvpCandLX ][ 1 ] = cpMvpLX[ 1 ] (8-619)
[0696] cpMvpListLX[ numCpMvpCandLX ][ 2 ] = cpMvpLX[ 2 ] (8-620)
[0697] numCpMvpCandLX = numCpMvpCandLX + 1 (8-621)
[0698] Otherwise, if PredFlagLY[ xNbBk ][ yNbBk ] ( with Y =! X ) is equal to 1 and DiffPicOrderCnt( RefPicListY[ RefIdxLY[ xNbBk ][ yNbBk ] ], RefPicListX[ refIdxLX ] ) is equal to 0, the following applies:
[0699] The variable availableFlagB is set to TRUE
[0700] Invoke the derivation process for luma affine control point motion vectors from neighbouring blocks specified in clause 8.4.4.5 with luma coding block position ( xCb, yCb ), luma coding block width and height ( cbWidth, cbHeight ), neighbouring luma coding block position ( xNb, yNb ), neighbouring luma coding block width and height ( nbW, nbH ), and number of control point motion vectors numCpMv as inputs and control point motion vector prediction candidates cpMvpLY[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) as outputs.
[0701] Invoke the integer-pixel precision process specified in clause 8.4.2.14 with mvX set equal to cpMvpx[ cpIdx ], rightShift set equal to ( MvShift + 2 ), and leftShift set equal to ( MvShift + 2 ) as inputs and the integer-pixel precision cpMvpx[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) as output.
[0702] Make the following assignments:
[0703] cpMvpxListLX[ numCpMvpCandLX ][ 0 ] = cpMvpx[ 0 ] (8-621)
[0704] cpMvpxListLX[ numCpMvpCandLX ][ 1 ] = cpMvpx[ 1 ] (8-622)
[0705] cpMvpxListLX[ numCpMvpCandLX ][ 2 ] = cpMvpx[ 2 ] (8-623)
[0706] numCpMvpCandLX = numCpMvpCandLX + 1 (8-624)
[0707] When numCpMvpCandLX is less than 2, apply the following:
[0708] Invoke the derivation process for constructed affine control point motion vector prediction candidates specified in clause 8.4.4.8 with luma coding block position ( xCb, yCb ), luma coding block width cbWidth, luma coding block height cbHeight, and reference index refldxLX of the current coding unit as inputs and availability flag availableConsFlagLX and cpMvpx[ cpIdx ] ( where cpIdx = 0..numCpMv - 1 ) as outputs.
[0709] When availableConsFlagLX is equal to 1 and numCpMvpCandLX is equal to 0, make the following assignments:
[0710] cpMvpxListLX[ numCpMvpCandLX ][ 0 ] = cpMvpx[ 0 ] (8-626)
[0711] cpMvpListLX[ numCpMvpCandLX ][ 1 ] = cpMvpLX[ 1 ] (8-627)
[0712] cpMvpListLX[ numCpMvpCandLX ][ 2 ] = cpMvpLX[ 2 ] (8-628)
[0713] numCpMvpCandLX = numCpMvpCandLX + 1 (8-629)
[0714] For cpIdx = 0..numCpMv - 1 apply the following:
[0715] When numCpMvpCandLX is less than 2 and availableFlagLX[ cpIdx ] is equal to 1, the following assignment is made:
[0716] cpMvpListLX[ numCpMvpCandLX ][ 0 ] = cpMvpLX[ cpIdx ] (8-630)
[0717] cpMvpListLX[ numCpMvpCandLX ][ 1 ] = cpMvpLX[ cpIdx ] (8-631)
[0718] cpMvpListLX[ numCpMvpCandLX ][ 2 ] = cpMvpLX[ cpIdx ] (8-632)
[0719] numCpMvpCandLX = numCpMvpCandLX + 1 (8-633)
[0720] When numCpMvpCandLX is less than 2, apply the following:
[0721] Invoke the derivation process of temporal luma motion vector prediction specified in clause 8.4.2.11 with luma coding block position ( xCb, yCb ), luma coding block width cbWidth, luma coding block height cbHeight, and refldxLX as inputs and with availableFlagLXCol and temporal motion vector predictor mvLXCol as outputs.
[0722] When availableFlagLXCol is equal to 1, apply the following:
[0723] Invoke the rounding process specified in clause 8.4.2.14 for the motion vector, where mvX is set equal to mvLXCol, rightShift is set equal to (MvShift + 2), and leftShift is set equal to (MvShift + 2) as inputs and the rounded mvLXCol as output.
[0724] The following assignments are made:
[0725] cpMvpListLX[ numCpMvpCandLX ][ 0 ] = mvLXCol (8-634)
[0726] cpMvpListLX[ numCpMvpCandLX ][ 1 ] = mvLXCol (8-635)
[0727] cpMvpListLX[ numCpMvpCandLX ][ 2 ] = mvLXCol (8-636)
[0728] numCpMvpCandLX = numCpMvpCandLX + 1 (8-637)
[0729] When numCpMvpCandLX is less than 2, repeat the following steps until numCpMvpCandLX is equal to 2, where mvZero[ 0 ] and mvZero[ 1 ] are both equal to 0:
[0730] cpMvpListLX[ numCpMvpCandLX ][ 0 ] = mvZero (8-638)
[0731] cpMvpListLX[ numCpMvpCandLX ][ 1 ] = mvZero (8-639)
[0732] cpMvpListLX[ numCpMvpCandLX ][ 2 ] = mvZero (8-640)
[0733] numCpMvpCandLX = numCpMvpCandLX + 1 (8-641) The affine control point motion vector predictor cpMvpLX, where X is 0 or 1, is derived as follows:
[0734] cpMvpLX = cpMvpListLX[ mvp_lX_flag[ xCb ][ yCb ] ] (8-642)
[0735] Derivation process of constructed affine control point motion vector prediction candidate
[0736] The inputs of the process are:
[0737] - the luma position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture,
[0738] - two variables cbWidth and cbHeight specifying the width and height of the current luma coding block,
[0739] - the reference index refldxLX of the current prediction unit partition, where X is either 0 or 1,
[0740] The outputs of the process are:
[0741] - an availability flag availableConsFlagLX of the constructed affine control point motion vector prediction candidate, where X is either 0 or 1,
[0742] - availability flags availableFlagLX[cpldx], where cpldx = 0..2 and X is either 0 or 1,
[0743] - the constructed affine control point motion vector prediction candidates cpMvLX[cpldx], where cpldx = 0..numCpMv - 1 and X is either 0 or 1.
[0744] The first (top-left) control point motion vector cpMvLX[0] and the availability flag availableFlagLX[0] are derived in the following ordered steps:
[0745] Set the sample positions (xNbB2, yNbB2), (xNbB3, yNbB3) and (xNbA2, yNbA2) equal to (xCb - 1, yCb - 1), (xCb, yCb - 1) and (xCb - 1, yCb), respectively.
[0746] Set the availability flag availableFlagLX[0] equal to 0 and both components of cpMvLX[0] equal to 0.
[0747] Apply the following steps for (xNbTL, yNbTL), where TL is replaced by B2, B3 and A2:
[0748] Call the availability derivation process for a coding block specified in the clause with the luma coding block position (xCb, yCb), the luma coding block width cbWidth, the luma coding block height cbHeight, the luma position (xNbY, yNbY) set equal to (xNbTL, yNbTL) as inputs and the output assigned to the coding block availability flag availableTL.
[0749] When availableTL is equal to TRUE and availableFlagLX[ 0 ] is equal to 0, the following steps apply:
[0750] If PredFlagLX[ xNbTL ][ yNbTL ] is equal to 1, and DiffPicOrderCnt( RefPicListX[ RefIdxLX[ xNbTL ][ yNbTL ] ], RefPicListX[ refIdxLX ] ) is equal to 0, and the reference picture corresponding to RefIdxLX[ xNbTL ][ yNbTL ] is not the current picture, availableFlagLX[ 0 ] is set equal to 1 and the following assignments are made:
[0751] cpMvLX[ 0 ] = MvLX[ xNbTL ][ yNbTL ] (8-643)
[0752] Otherwise, when PredFlagLY[ xNbTL ][ yNbTL ] (with Y =!X) is equal to 1, and DiffPicOrderCnt( RefPicListY[ RefIdxLY[ xNbTL ][ yNbTL ] ], RefPicListX[ refIdxLX ] ) is equal to 0, and the reference picture corresponding to RefIdxLY[ xNbTL ][ yNbTL ] is not the current picture, availableFlagLX[ 0 ] is set equal to 1 and the following assignments are made:
[0753] cpMvLX[ 0 ] = MvLY[ xNbTL ][ yNbTL ] (8-644)
[0754] When availableFlagLX[ 0 ] is equal to 1, the process for rounding a motion vector specified in clause 8.4.2.14 is invoked with mvX set equal to cpMvLX[ 0 ], rightShift set equal to ( MvShift + 2 ), and leftShift set equal to ( MvShift + 2 ) as inputs and the rounded cpMvLX[ 0 ] as output.
[0755] The second (right-top) control point motion vector cpMvLX[ 1 ] and the availability flag availableFlagLX[ 1 ] are derived in the following ordered steps:
[0756] The sample positions (xNbB1, yNbB1) and (xNbB0, yNbB0) are set equal to (xCb + cbWidth - 1, yCb - 1) and (xCb + cbWidth, yCb - 1), respectively.
[0757] The availability flag availableFlagLX[1] is set equal to 0 and both components of cpMvLX[1] are set equal to 0.
[0758] The following steps are applied for (xNbTR, yNbTR), where TR is replaced by B1 and B0:
[0759] The availability derivation process for a coding block as specified in clause 6.4.X is invoked with the luma coding block position (xCb, yCb), the luma coding block width cbWidth, the luma coding block height cbHeight, the luma position (xNbY, yNbY) set equal to (xNbTR, yNbTR) as input and the output assigned to the coding block availability flag availableTR.
[0760] When availableTR is equal to TRUE and availableFlagLX[1] is equal to 0, the following steps are applied:
[0761] If PredFlagLX[xNbTR][yNbTR] is equal to 1 and DiffPicOrderCnt(RefPicListX[RefIdxLX[xNbTR][yNbTR]], RefPicListX[refIdxLX]) is equal to 0 and the reference picture corresponding to RefIdxLX[xNbTR][yNbTR] is not the current picture, availableFlagLX[1] is set equal to 1 and the following assignments are made:
[0762] cpMvLX[ 1 ] = MvLX[ xNbTR ][ yNbTR ] (8-645)
[0763] Otherwise, when PredFlagLY[xNbTR][yNbTR] (with Y =!X) is equal to 1 and DiffPicOrderCnt(RefPicListY[RefIdxLY[xNbTR][yNbTR]], RefPicListX[refIdxLX]) is equal to 0 and the reference picture corresponding to RefIdxLY[xNbTR][yNbTR] is not the current picture, availableFlagLX[1] is set equal to 1 and the following assignments are made:
[0764] cpMvLX[ 1 ] = MvLY[ xNbTR ][ yNbTR ] (8-646)
[0765] When availableFlagLX[ 1 ] is equal to 1, the rounding process for motion vectors specified in clause 8.4.2.14 is invoked with mvX set equal to cpMvLX[ 1 ], rightShift set equal to ( MvShift + 2 ), and leftShift set equal to ( MvShift + 2 ) as inputs and the rounded cpMvLX[ 1 ] as output.
[0766] The third (left-bottom) control point motion vector cpMvLX[ 2 ] and the availability flag availableFlagLX[ 2 ] are derived in the following ordered steps:
[0767] The sample positions ( xNbAl, yNbAl ) and ( xNbAo, yNbAo ) are set equal to ( xCb - 1, yCb + cbHeight - 1 ) and ( xCb - 1, yCb + cbHeight ), respectively.
[0768] The availability flag availableFlagLX[ 2 ] is set equal to 0 and both components of cpMvLX[ 2 ] are set equal to 0.
[0769] The following steps are applied for ( xNbBL, yNbBL ), where BL is replaced by Al and Ao:
[0770] The availability derivation process for a coding block specified in clause 6.4.X is invoked with the luma coding block position ( xCb, yCb ), the luma coding block width cbWidth, the luma coding block height cbHeight, the luma position ( xNbY, yNbY ) set equal to ( xNbBL, yNbBL ) as inputs and the output assigned to the coding block availability flag availableBL.
[0771] When availableBL is equal to TRUE and availableFlagLX[ 2 ] is equal to 0, the following steps are applied:
[0772] If PredFlagLX[ xNbBL ][ yNbBL ] is equal to 1, and DiffPicOrderCnt( RefPicListX[ RefIdxLX[ xNbBL ][ yNbBL ] ], RefPicListX[ refIdxLX ] ) is equal to 0, and the reference picture corresponding to RefIdxLY[ xNbBL ][ yNbBL ] is not the current picture, availableFlagLX[ 2 ] is set equal to 1 and the following assignment is made:
[0773] cpMvLX[ 2 ] = MvLX[ xNbBL ][ yNbBL ] (8-647)
[0774] Otherwise, when PredFlagLY[ xNbBL ][ yNbBL ] (with Y =! X) is equal to 1, and DiffPicOrderCnt( RefPicListY[ RefIdxLY[ xNbBL ][ yNbBL ] ], RefPicListX[ refIdxLX ] ) is equal to 0, and the reference picture corresponding to RefIdxLY[ xNbBL ][ yNbBL ] is not the current picture, availableFlagLX[ 2 ] is set equal to 1 and the following assignment is made:
[0775] cpMvLX[ 2 ] = MvLY[ xNbBL ][ yNbBL ] (8-648)
[0776] When availableFlagLX[ 2 ] is equal to 1, the process for rounding a motion vector specified in clause 8.4.2.14 is invoked with mvX set equal to cpMvLX[ 2 ], rightShift set equal to ( MvShift + 2 ), and leftShift set equal to ( MvShift + 2 ) as inputs and the rounded cpMvLX[ 2 ] as output.
[0777] 5.5 Context modelling
[0778] The binary number coded using the context assigns ctxInc to the syntax element:
[0779]
[0780] The binary number coded using the context assigns ctxInc to the syntax element:
[0781] In one example, the context increment offset ctxInc = ( condL && availableL ) + ( condA && availableA ) + ctxSetldx * 3.
[0782] Alternatively, ctxInc = ((condL && availableL) || (condA && availableA)) + ctxSetldx * 3.
[0783] ctxInc = (condL && availableL) + M * (condA && availableA) + ctxSetldx * 3. (e.g., M = 2)
[0784] ctxInc = M * (condL && availableL) + (condA && availableA) + ctxSetldx * 3. (e.g., M = 2)
[0785]
[0786] initValue value for ctxldx of amvr_flag:
[0787] Different contexts are used when the current block is affine or non-affine.
[0788]
[0789] Alternatively,
[0790]
[0791] Alternatively, the same context can be used when the current block is affine or non-affine
[0792]
[0793] Alternatively, amvr_flag is bypass coded.
[0794] initValue value for ctxldx of amvr_coarse_precision_flag:
[0795] Different contexts are used when the current block is affine or non-affine.
[0796]
[0797] Alternatively,
[0798]
[0799] Alternatively, the same context can be used when the current block is affine or non-affine.
[0800]
[0801] Alternatively, the amvr_coarse_precision_flag is bypass coded.
[0802] The above examples can be incorporated in the context of the methods described below, e.g., methods 2510-2540, which can be implemented at a video decoder or a video encoder.
[0803] FIG. 25A-25D A flowchart of an example method for video processing is shown. FIG. 25A The illustrated method 2510 includes, at step 2512, determining that a conversion between a current video block of a video and a coded representation of the current video block is based on a non-affine inter AMVR mode. The method 2510 further includes, at step 2514, conducting the conversion based on the determination. In some implementations, the coded representation of the current video block is based on context-based coding, and wherein a context used to code the current video block is modeled without using affine AMVR mode information of neighboring blocks during the conversion.
[0804] As FIG. 25B The illustrated method 2520 includes, at step 2522, determining that a conversion between a current video block of a video and a coded representation of the current video block is based on an affine adaptive motion vector resolution (affine AMVR) mode. The method 2520 further includes, at step 2524, conducting the conversion based on the determination. In some implementations, the coded representation of the current video block is based on context-based coding, and wherein a variable controls a speed of two probability updates of a context.
[0805] As FIG. 25C The illustrated method 2530 includes, at step 2532, determining that a conversion between a current video block of a video and a coded representation of the current video block is based on an affine AMVR mode. The method 2530 further includes, at step 2534, conducting the conversion based on the determination. In some implementations, the coded representation of the current video block is based on context-based coding, and wherein a context used to code the current video block is modeled using coding information of neighboring blocks that uses AMVR mode of both an affine inter mode and a normal inter mode during the conversion.
[0806] As FIG. 25D The illustrated method 2540 includes, at step 2542, for a conversion between a current video block of a video and a coded representation of the current video block, determining a usage of a plurality of contexts for the conversion. The method 2540 further includes, at step 2544, conducting the conversion based on the determination. In some implementations, the plurality of contexts is used to code a syntax element that indicates a coarse motion precision.
[0807] As FIG. 25E Method 2550 as illustrated in FIG. 25B includes, at 2552, determining, for a conversion between a current video block of a video and a coded representation of the current video block, whether to use a symmetric motion vector difference (SMVD) mode based on a currently selected best mode for the conversion. Method 2554 further includes conducting the conversion based on the determination.
[0808] As FIG. 25F illustrated in FIG. 25B includes, at 2562, determining, for a conversion between a current video block of a video and a coded representation of the current video block, whether to use an affine SMVD mode based on a currently selected best mode for the conversion. Method 2564 further includes conducting the conversion based on the determination.
[0809] 6. Example implementations of the disclosed technology
[0810] FIG. 26 is a block diagram of a video processing device 2600. Device 2600 can be used to implement one or more methods described herein. Device 2600 can be implemented in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 2600 can include one or more processors 2602, one or more memories 2604, and video processing hardware 2606. Processor(s) 2602 can be configured to implement one or more methods described in this document, including but not limited to method 2500. Memory(ies) 2604 can be used for storing data and code used for implementing the methods and techniques described herein. Video processing hardware 2606 can be used to implement, in hardware circuitry, some of the techniques described in this document.
[0811] FIG. 28 is another example of a block diagram of a video processing system that can implement the disclosed technology. FIG. 28 is a block diagram illustrating an example video processing system 2800 that can implement the various techniques disclosed herein. Various implementations can include some or all of the components of the system 2800. The system 2800 can include an input 2802 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be received in a compressed or encoded format. The input 2802 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0812] The system 2800 can include a coding component 2804, which can implement various coding or encoding methods described in this document. The coding component 2804 can reduce the average bitrate of a video from an input 2802 to an output of the coding component 2804 to generate a coded representation of the video. Thus, the coding techniques are sometimes called video compression or video transcoding techniques. The output of the coding component 2804 can be stored, or transmitted via a communication connected, as shown by component 2806. The stored or transmitted bitstream (or coded) representation of the video received at the input 2802 can be used by component 2808 for generating pixel values or displayable video sent to a display interface 2810. The process of generating user-viewable video from a bitstream representation is sometimes called video decompression. Also, while certain video processing operations are referred to as “coding” operations or tools, it should be understood that the coding tools or operations are used at an encoder, and corresponding decoding tools or operations to reverse the coding results would be performed by a decoder.
[0813] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or Displayport, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be implemented in various electronic devices such as mobile phones, laptops, smart phones, or other devices capable of performing digital data processing and / or video display.
[0814] In some embodiments, the video coding methods can be implemented using an apparatus implemented on a hardware platform as described with respect to FIG. 28. FIG. 26
[0815] Various techniques and embodiments can be described using the following clause-based format. These clauses can be implemented as preferred features of some embodiments.
[0816] A first set of clauses uses some of the techniques described in the previous section, including, for example, items 16-18, 23, and 26 in the previous section.
[0817] 1. A method of video processing, comprising:
[0818] determining that a conversion between a current video block of a video and a coded representation of the current video block is based on a non-affine inter (AMVR) mode; and
[0819] performing the conversion based on the determination, and
[0820] wherein the coded representation of the current video block is based on context-based coding, and wherein a context used for coding the current video block is modeled without using affine AMVR mode information of neighboring blocks during the conversion.
[0821] 2. The method of clause 1, wherein a non-affine inter AMVR mode index of the coded neighboring block is utilized.
[0822] 3. The method of clause 2, wherein a context index used for coding a non-affine inter AMVR mode index of the current video block depends on at least one of: the AMVR mode index, an affine mode flag, or availability of left and top neighboring blocks, the non-affine inter AMVR mode index indicating a certain MVD precision.
[0823] 4. The method of clause 3, wherein an offset value is derived for the left and top neighboring blocks, respectively, and the context index is derived as a sum of the offset values of the left and top neighboring blocks.
[0824] 5. The method of clause 4, wherein at least one of the offset values is derived as 1 if a corresponding neighboring block is available and not coded in affine mode and the AMVR mode index of the neighboring block is not equal to 0, and otherwise is derived as 0.
[0825] 6. The method of clause 1, wherein the affine AMVR mode information of the neighboring blocks is not used directly, but indirectly in a function of the affine AMVR mode information.
[0826] 7. The method of clause 6, wherein the function returns true if amvr mode[xNbL][yNbL] or amvr flag[xNbL][yNbL] of the neighboring block covering (xNbL, yNbL) indicates a certain MVD precision.
[0827] 8. The method of clause 1, wherein a first syntax element of a non-affine inter AMVR mode of the current video block is coded with affine AMVR mode information of the neighboring blocks.
[0828] 9. The method of clause 1, wherein in case the AMVR mode of the current video block is represented by multiple syntax elements, any one of the multiple syntax elements is coded with AMVR mode information of the neighboring blocks coded in both non-affine inter mode and affine inter mode.
[0829] 10. The method of clause 1, wherein, in a case where the AMVR mode of the current video block is represented by a plurality of syntax elements, any of the plurality of syntax elements is not coded with AMVR mode information of the neighboring block coded in affine inter mode.
[0830] 11. The method of clause 1, wherein, in a case where the AMVR mode of the current video block is represented by a plurality of syntax elements, any of the plurality of syntax elements is not coded directly but indirectly with AMVR mode information of the neighboring block coded in affine inter mode.
[0831] 12. The method of any of clauses 9-11, wherein the AMVR mode of the current video block includes an affine inter AMVR mode and a non-affine inter AMVR mode.
[0832] 13. A method of video processing, comprising:
[0833] determining that a conversion between a current video block of a video and a coded representation of the current video block is based on an affine adaptive motion vector resolution (affine AMVR) mode; and
[0834] performing the conversion based on the determination, and
[0835] wherein the coded representation of the current video block is based on context-based coding, and wherein two probability update speeds of a context are controlled by a variable.
[0836] 14. The method of clause 13, wherein the two probability update speeds include a faster update speed defined by (shiftldx » 2) + 2, shiftldx indicating the variable.
[0837] 15. The method of clause 13, wherein the two probability update speeds include a slower update speed defined by (shiftldx & 3) + 3 + shiftO, shiftO defined by (shiftldx » 2) + 2, and shiftldx indicating the variable.
[0838] 16. The method of clause 14, wherein the faster update speed is between 2 and 5.
[0839] 17. A method of video processing, comprising:
[0840] determining that a conversion between a current video block of a video and a coded representation of the current video block is based on an affine AMVR mode; and
[0841] performing the conversion based on the determination, and
[0842] wherein the coded representation of the current video block is based on context-based coding, and wherein a context used for coding the current video block is modeled using coded information of neighboring blocks, the coded information of the neighboring blocks using AMVR modes of both affine inter mode and normal inter mode during the conversion.
[0843] 18. The method of clause 17, wherein the AMVR mode information of the neighboring blocks is used directly for the context-based coding.
[0844] 19. The method of clause 17, wherein AMVR mode information of the neighboring blocks coded in normal inter mode is not allowed to be used for the context-based coding.
[0845] 20. The method of clause 17, wherein AMVR mode information of the neighboring blocks coded in normal inter mode is not used directly, but indirectly in a function of the AMVR mode information.
[0846] 21. The method of clause 17, wherein the AMVR mode information of the neighboring blocks is utilized to code a first syntax element of an affine AMVR mode of the current video block.
[0847] 22. The method of clause 17, wherein in case the affine AMVR mode is represented by a plurality of syntax elements, any of the plurality of syntax elements is coded using AMVR mode information of the neighboring blocks coded in both affine inter mode and normal inter mode.
[0848] 23. The method of clause 17, wherein in case the affine AMVR mode is represented by a plurality of syntax elements, any of the plurality of syntax elements is not coded using AMVR mode information of the neighboring blocks coded in normal inter mode.
[0849] 24. The method of clause 17, wherein in case the affine AMVR mode is represented by a plurality of syntax elements, any of the plurality of syntax elements is not coded directly using, but indirectly using AMVR mode information of the neighboring blocks coded in normal inter mode.
[0850] 25. A video processing method, comprising:
[0851] for a conversion between a current video block of a video and a coded representation of the current video block, determining usage of a plurality of contexts for the conversion; and
[0852] performing the conversion based on the determining, and
[0853] wherein the plurality of contexts are utilized to code a syntax element indicative of a coarse motion precision.
[0854] 26. The method of clause 25, wherein the plurality of contexts correspond exactly to two contexts.
[0855] 27. The method of clause 26, wherein the plurality of contexts are selected based on whether the current video block is affine coded.
[0856] 28. The method of clause 25, wherein another syntax element is used to indicate that the conversion is based on an affine AMVR mode.
[0857] 29. The method of clause 28, wherein the other syntax element is coded using only a first context and the syntax element is coded using only a second context.
[0858] 30. The method of clause 28, wherein the other syntax element is coded using only a first context and the syntax element is bypass coded.
[0859] 31. The method of clause 28, wherein the other syntax element is bypass coded and the syntax element is bypass coded.
[0860] 32. The method of clause 25, wherein all syntax elements related to motion vector precision are bypass coded.
[0861] 33. The method of clause 25 or 26, wherein for normal inter modes, the syntax element indicates a selection from a set comprising 1-pel or 4-pel precision.
[0862] 34. The method of clause 25 or 26, wherein for affine modes, the syntax element indicates a selection between at least 1 / 16-pel or 1-pel precision.
[0863] 35. The method of clause 34, wherein the syntax element is equal to 0, the affine mode is 1 / 6-pel precision.
[0864] 36. The method of clause 34, wherein the syntax element is equal to 1, the affine mode is 1-pel precision.
[0865] 37. The method of any of clauses 1 to 36, wherein performing the conversion comprises generating the coded representation from the current video block.
[0866] 38. The method of any of clauses 1-36, wherein performing the conversion comprises generating the current video block from the coded representation.
[0867] 39. An apparatus in a video system comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of one or more of clauses 1-38.
[0868] 40. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for implementing the method of one or more of clauses 1-38.
[0869] A second set of clauses uses some of the techniques described in the previous section, including, for example, items 19-21 and 23 in the previous section.
[0870] 1. A video processing method comprising:
[0871] for a conversion between a current video block of a video and a coded representation of the current video block, determining whether to use a symmetric motion vector difference (SMVD) mode based on a currently selected best mode for the conversion; and
[0872] performing the conversion based on the determination.
[0873] 2. The method of clause 1, wherein the SMVD mode is used without explicitly signaling at least one reference index of a reference list.
[0874] 3. The method of clause 2, wherein the reference index is derived based on a recursive picture order count (POC) calculation.
[0875] 4. The method of clause 1, wherein the determination disables use of the SMVD mode in a case that the currently selected best mode is a Merge mode or a UMVE mode.
[0876] 5. The method of clause 4, wherein the UMVE mode applies a motion vector offset to refine a motion candidate derived from a Merge candidate list.
[0877] 6. The method of clause 1, wherein the determination disables use of the SMVD mode in a case that the currently selected best mode does not use the SMVD mode for coding.
[0878] 7. The method of clause 1, wherein the determination disables use of the SMVD mode in a case that the currently selected best mode is an affine mode.
[0879] 8. The method of clause 1, wherein, in a case that the currently selected best mode is subblock Merge mode, the determining disables use of the SMVD mode.
[0880] 9. The method of clause 1, wherein, in a case that the currently selected best mode is affine SMVD mode, the determining disables use of the SMVD mode.
[0881] 10. The method of clause 1, wherein, in a case that the currently selected best mode is affine Merge mode, the determining disables use of the SMVD mode.
[0882] 11. The method of any of clauses 2-10, wherein the determining is applied only when MVD (motion vector difference) precision is greater than or equal to precision.
[0883] 12. The method of any of clauses 2-10, wherein the determining is applied only when MVD precision is greater than precision.
[0884] 13. The method of any of clauses 2-10, wherein the determining is applied only when MVD precision is less than or equal to precision.
[0885] 14. The method of any of clauses 2-10, wherein the determining is applied only when MVD precision is less than precision.
[0886] 15. A method of video processing, comprising:
[0887] for a conversion between a current video block of a video and a coded representation of the current video block, determining whether to use affine SMVD mode based on a currently selected best mode for the conversion; and
[0888] performing the conversion based on the determining.
[0889] 16. The method of clause 15, wherein, in a case that the currently selected best mode is Merge mode or UMVE mode, the determining disables use of the affine SMVD mode.
[0890] 17. The method of clause 15, wherein, in a case that the currently selected best mode does not use affine SMVD mode coded, the determining disables use of the affine SMVD mode.
[0891] 18. The method of clause 15, wherein, in a case that the currently selected best mode is subblock Merge mode, the determining disables use of the affine SMVD mode.
[0892] 19. The method of clause 15, wherein, in a case where the currently selected best mode is an SMVD mode, the determining disables use of the affine SMVD mode.
[0893] 20. The method of clause 15, wherein, in a case where the currently selected best mode is an affine Merge mode, the determining disables use of the affine SMVD mode.
[0894] 21. The method of any of clauses 16-20, wherein the determining is applied only when an affine MVD precision is greater than or equal to a precision.
[0895] 22. The method of any of clauses 16-20, wherein the determining is applied only when an affine MVD precision is greater than a precision.
[0896] 23. The method of any of clauses 16-20, wherein the determining is applied only when an affine MVD precision is less than or equal to a precision.
[0897] 24. The method of any of clauses 16-20, wherein the determining is applied only when an affine MVD precision is less than a precision.
[0898] 25. The method of clause 1 or 15, further comprising:
[0899] determining a characteristic of the current video block, the characteristic comprising one or more of a block size of the video block, a type of slice or picture or tile related to the video block, or motion information related to the video block, and wherein the determining is also based on the characteristic.
[0900] 26. The method of any of clauses 1-25, wherein the performing the conversion comprises generating the coded representation from the current block.
[0901] 27. The method of any of clauses 1-25, wherein the performing the conversion comprises generating the current block from the coded representation.
[0902] 28. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of one or more of clauses 1-27.
[0903] 29. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for implementing the method of one or more of clauses 1-27.
[0904] From the foregoing, it will be appreciated that specific embodiments of the disclosed technology have been described herein for purposes of illustration, but well-known structures and functions have not been described in detail to avoid obscuring the subject matter of the disclosed technology. Thus, the disclosed technology is not limited to the specific embodiments described herein, but includes any and all implementations falling within the scope of the appended claims.
[0905] Implementations of the subject matter and the functional operations described in this patent document can be implemented in various systems, methods, apparatuses, and computer program products, including those disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by, or to control the operation of, data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of one or more of them, or a combination of one or more of them. The term “data processing apparatus” or “data processing device” encompasses all apparatus, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0906] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0907] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and the apparatus can be implemented as special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0908] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0909] It is intended that the description and examples contained herein serve only as exemplifications of the application. FIG. one The examples are merely illustrative and should not be construed as limiting.
[0910] While this patent document contains many specifics, these should not be construed as limitations on the scope of any invention or patentable aspect, but rather as descriptors of features that can be part of particular embodiments. In this patent document, certain features are described in the context of separate embodiments. It should be understood that these features can be combined with each other in any manner within a single embodiment. Conversely, various features are described in the context of a single embodiment. It should be understood that these features can be combined with each other in any manner within a single embodiment. Further, although the above features can be described as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a sub-combination or variation of a sub-combination.
[0911] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such an order, nor that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0912] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: For the conversion between the current video block and the bitstream of the current video block, it is determined whether to use the Symmetric Motion Vector Difference (SMVD) mode based on the currently selected optimal mode for the conversion; as well as The conversion is performed based on the determination.
2. The method according to claim 1, wherein, The SMVD mode is used when at least one reference index of the signaling notification reference list is not explicitly included.
3. The method according to claim 2, wherein, The reference index is derived based on recursive image sequence counting (POC) calculation.
4. The method according to claim 1, wherein, If the currently selected optimal mode is Merge mode or UMVE mode, then the use of the SMVD mode is disabled.
5. The method according to claim 4, wherein, The UMVE mode applies motion vector offsets to refine the motion candidates derived from the Merge candidate list.
6. The method according to claim 1, wherein, If the currently selected optimal mode does not use the SMVD mode encoding / decoding, the determination disables the use of the SMVD mode.
7. The method according to claim 1, wherein, If the currently selected optimal mode is affine mode, then the use of the SMVD mode is disabled.
8. The method according to claim 1, wherein, If the currently selected optimal mode is the sub-block Merge mode, then the determination disables the use of the SMVD mode.
9. The method according to claim 1, wherein, If the currently selected optimal mode is the affine SMVD mode, then the use of the SMVD mode is disabled.
10. The method according to claim 1, wherein, If the currently selected optimal mode is the affine Merge mode, then the use of the SMVD mode is disabled.
11. The method according to claim 2, wherein, The determination is applied only when the accuracy of the motion vector difference (MVD) is less than the first accuracy.
12. The method according to any one of claims 1 to 11, wherein, The conversion includes generating the bitstream from the current video block.
13. The method according to any one of claims 1 to 11, wherein, The conversion includes generating the current video block from the bitstream.
14. A video processing method, comprising: For the conversion between the current video block and the bitstream of the current video block, it is determined whether to use the affine SMVD mode based on the currently selected best mode for the conversion; as well as The conversion is performed based on the determination.
15. The method according to claim 14, wherein, If the currently selected optimal mode does not use affine SMVD mode encoding / decoding, the determination disables the use of the affine SMVD mode.
16. The method of claim 14, wherein, If the currently selected optimal mode is the sub-block Merge mode, then the use of the affine SMVD mode is disabled.
17. The method of claim 14, wherein, If the currently selected optimal mode is the SMVD mode, then the use of the affine SMVD mode is disabled.
18. The method according to claim 14, wherein, If the currently selected optimal mode is the affine Merge mode, then the use of the affine SMVD mode is disabled.
19. The method of claim 14, wherein, The determination is applied only when the affine MVD accuracy is less than the second accuracy.
20. The method of claim 14, further comprising: The characteristics of the current video block are determined, including one or more of the following: the block size of the video block, the type of strip, image, or slice of the video block, or motion information related to the video block, and wherein the determination is also based on the characteristics.
21. The method according to any one of claims 14 to 20, wherein, The conversion includes generating the bitstream from the current video block.
22. The method according to any one of claims 14 to 20, wherein, The conversion includes generating the current video block from the bitstream.
23. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, When executed by the processor, the instructions cause the processor to perform the method according to any one of claims 1 to 13 or claims 14 to 20.
24. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method of any one of claims 1 to 13 or 14 to 20.