Merge with Motion Vector Difference (MVD) based on geometric partitioning
By using MMVD mode and affine motion compensation prediction technology in video encoding and decoding, the problem of inefficient encoding and decoding in complex motion situations in the prior art is solved, and more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202080008274.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-10
- Filing Date
- 2020-01-10
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-01-10
AI Technical Summary
Existing video encoding and decoding standards are inefficient when dealing with complex motion situations, making it difficult to effectively use affine motion models for efficient encoding and decoding.
The Merge mode (MMVD) based on motion vector difference is adopted, and the video blocks are divided into multiple partitions, and the selection and signaling notification of Merge candidates are refined using the MMVD mode, combined with affine motion compensation prediction technology, the encoding and decoding efficiency is improved.
Improves the efficiency and quality of video encoding and decoding, especially in handling complex motion situations, reducing the need for bitstreams.
Smart Images

Figure CN113273207B_ABST
Abstract
Description
[0001] This application is an application entering the Chinese national phase based on International Patent Application No. PCT / CN2020 / 071444 filed on January 10, 2020. The entire disclosure of which is incorporated by reference as part of the disclosure of this application. Technical Field
[0002] This document deals with video and image encoding and decoding. Background Art
[0003] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of networked user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0004] This document discloses video codec tools that, in one example aspect, improve the signaling of motion vectors for video and image codecs.
[0005] In one aspect, a method for video processing is disclosed, comprising: making a decision about applying a Merge with Motion Vector Difference (MMVD) mode to a current block of a video based on a set of MMVD side information, wherein the current block is partitioned into at least two partitions; and performing conversion between the current block of the video and a bitstream representation of the video using the MMVD mode, wherein, in the MMVD mode, at least one Merge candidate selected for at least one partition is refined based on the set of MMVD side information.
[0006] In one aspect, a method for video processing is disclosed, comprising: making a decision to apply a Merge with Motion Vector Difference (MMVD) mode to a current block of a video, wherein the current block is partitioned into at least two partitions; and performing conversion between the current block of the video and a bitstream representation of the video using the MMVD mode, wherein, in the MMVD mode, at least one Merge candidate selected for at least one partition is refined, and a set of MMVD information associated with the refinement of the at least one Merge candidate is signaled.
[0007] In one aspect, an apparatus in a video system is disclosed, the apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of any of the above examples.
[0008] In one aspect, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product comprising program code for executing the method in any one of the above examples.
[0009] In yet another example aspect, the above method may be implemented by a video encoder device or a video decoder device including a processor.
[0010] In yet another example aspect, the methods may be embodied in the form of processor-executable instructions and stored on a computer-readable program medium.
[0011] These and other aspects are described further throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 An example of a simplified affine motion model is shown.
[0013] Figure 2 An example of the affine motion vector field (MVF) of each sub-block is shown.
[0014] Figure 3A and Figure 3B Examples of a 4-parameter affine model and a 6-parameter affine model are shown respectively.
[0015] Figure 4 An example of a motion vector predictor (MVP) of AF_INTER is shown.
[0016] Figure 5A and Figure 5B An example of a candidate for AF_MERGE is shown.
[0017] Figure 6 Shows examples of candidate positions for the affine merge mode.
[0018] Figure 7 An example of a distance index and distance offset mapping is shown.
[0019] Figure 8 An example of the Ultimate Motion Vector Expression (UMVE) search process is shown.
[0020] Figure 9 An example of a UMVE search point is shown.
[0021] Figure 10 is a flow chart of an example method for video processing.
[0022] Figure 11 is a flow chart of another example method for video processing.
[0023] Figure 12 An example of a hardware platform for implementing the techniques described in this document is shown. DETAILED DESCRIPTION
[0024] This document provides various techniques that can be used by a decoder of a video bitstream to improve the quality of decompressed or decoded digital video. In addition, a video encoder can also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0025] For the sake of clarity, section headings are used in this document and do not limit the embodiments and techniques to the corresponding sections. Thus, the embodiments of one section can be combined with the embodiments of other sections.
[0026] 1. Overview
[0027] This patent document relates to video coding and decoding technology. Specifically, it relates to motion compensation in video coding and decoding. The patent application may be applied to existing video coding and decoding standards, such as HEVC, or to standards to be finalized (e.g., Versatile Video Coding (VVC)). The patent application may also be applied to future video coding and decoding standards or video codecs.
[0028] 2. Introductory Notes
[0029] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the Versatile Video Codec (VVC) standard, with the goal of a 50% bitrate reduction compared to HEVC.
[0030] 2.1 Affine Motion Compensated Prediction
[0031] In HEVC, only the translational motion model is used for motion compensation prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In JEM, a simplified affine transformation motion compensation prediction is applied. Figure 1 As shown, the affine motion field of a block is described by two control point motion vectors.
[0032] The motion vector field (MVF) of a block is described by the following equation:
[0033]
[0034] Where (v 0x ,v 0y ) is the motion vector of the upper left control point, and (v 1x ,v 1y ) is the motion vector of the upper right control point.
[0035] To further simplify motion compensated prediction, sub-block based affine transformation prediction is applied. The sub-block size M×N is derived from Equation 2, where MvPre is the motion vector fractional precision (1 / 16 in JEM), (v 2x ,v 2y ) is the motion vector of the lower left control point calculated according to Equation 1.
[0036]
[0037] After being derived from Equation 2, M and N should be adjusted downward, if necessary, to be divisors of w and h, respectively.
[0038] In order to derive the motion vector of each M×N sub-block, as Figure 2 As shown, the motion vector of the center sample of each sub-block is calculated according to Equation 1 and rounded to a fractional accuracy of 1 / 16.
[0039] After MCP, the high-precision motion vector of each sub-block is rounded and saved to the same precision as the normal motion vector.
[0040] 2.1.1 AF_INTER mode
[0041] In JEM, there are two affine motion modes: AF_INTER mode and AF_MERGE mode. For CUs with width and height greater than 8, AF_INTER mode can be applied. A CU-level affine flag is signaled in the bitstream to indicate whether AF_INTER mode is used. In this mode, neighboring blocks are used to construct a pair of motion vectors {(v0,v1)|v0={v A ,v B ,v c},v1={v D ,v E}} candidate list. Figure 4As shown, v0 is selected from the motion vector of block A, block B or block C. The motion vector from the neighboring block is scaled according to the reference list and the relationship between the POC of the reference of the neighboring block, the POC of the reference of the current block and the POC of the current CU. And the method of selecting v1 from neighboring blocks D and E is similar. If the number of candidate lists is less than 2, the list is filled with motion vector pairs composed by copying each AMVP candidate. When the candidate list is greater than 2, the candidates are first sorted according to the consistency of the neighboring motion vectors (the similarity of the two motion vectors in a pair of candidates), and only the first two candidates are retained. The RD cost is used to determine which motion vector pair candidate is selected as the control point motion vector prediction (CPMVP) of the current CU. And the index used to indicate the position of the CPMVP in the candidate list is signaled in the bitstream. After determining the CPMVP of the current affine CU, affine motion estimation is applied and the control point motion vector (CPMV) is found. Then, the difference between CPMV and CPMVP is signaled in the bitstream.
[0042] Figure 3A An example of a 4-parameter affine model is shown. Figure 3B An example of a 6-parameter affine model is shown.
[0043] In AF_INTER mode, when using 4 / 6 parameter affine mode, 2 / 3 control points are required, and therefore 2 / 3 MVDs need to be encoded and decoded for these control points, such as Figure 3A In the example, it is proposed to derive MV as follows, for example, predicting mvd1 and mvd2 from mvd0.
[0044]
[0045]
[0046]
[0047] in mvd i and mv1 are the predicted motion vector, motion vector difference and motion vector of the upper left pixel (i=0), upper right pixel (i=1) or lower left pixel (i=2), respectively. Figure 3B Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the separate summation of the two components, i.e., newMV = mvA + mvB, where the two components of newMV are set to (xA + xB) and (yA + yB), respectively.
[0048] 2.1.2 Fast Affine ME Algorithm in AF_INTER Mode
[0049] In affine mode, the MVs of two or three control points need to be jointly determined. Direct joint search of multiple MVs is computationally complex. A fast affine ME algorithm is proposed and applied to VTM / BMS.
[0050] The fast affine ME algorithm is described for a 4-parameter affine model, and the idea can be extended to a 6-parameter affine model.
[0051]
[0052]
[0053] Replacing (a-1) with a', the motion vector can be rewritten as:
[0054]
[0055] Assuming that the motion vectors of the two control points (0,0) and (0,w) are known, we can derive the affine parameters from equation (5),
[0056]
[0057] The motion vector can be rewritten in vector form as:
[0058]
[0059] in
[0060]
[0061]
[0062] P = (x, y) is the pixel position.
[0063] At the encoder, the MVD of AF_INTER is derived iteratively. i (P) is the MV derived in the i-th iteration at position P, and dMV C i It is expressed as MV in the i-th iteration C Then, in the (i+1)th iteration,
[0064]
[0065] Pic ref Indicated as a reference picture, and Pic cur Represents the current picture, and represents Q=P+MVi (P). Assuming we use MSE as the matching criterion, we need to minimize:
[0066]
[0067] Assumptions Small enough, we can rewrite it approximately using the first-order Taylor expansion as follows:
[0068]
[0069] in, Indicates E i+1 (P)=Pic cur (P)-Pic ref (Q),
[0070]
[0071] We can deduce this by setting the derivative of the error function to zero. Then, according to To calculate the incremental MV of the control points (0,0) and (0,w):
[0072]
[0073]
[0074]
[0075]
[0076] Assuming that such an MVD derivation process is iterated n times, the final MVD is calculated as follows:
[0077]
[0078]
[0079]
[0080]
[0081] As an example, for example, the delta MV of the control point (0, 0) represented by mvd0 is predicted from the delta MV of the control point (0, 0) represented by mvd1. Now, only mvd1 is actually encoded.
[0082] 2.1.3 AF_MERGE Mode
[0083] When applying CU in AF_MERGE mode, it obtains the first block encoded and decoded in affine mode from the valid adjacent reconstructed blocks. And the selection order of candidate blocks is from left, top, top right, bottom left to top left, such as Figure 5A If the adjacent lower left block A is encoded and decoded in affine mode, such as Figure 5B As shown, the motion vectors v2, v3, and v4 of the upper left, upper right, and lower left corners of the CU containing block A are derived. The motion vector v0 of the upper left corner of the current CU is calculated based on v2, v3, and v4. Next, the motion vector v1 of the upper right corner of the current CU is calculated.
[0084] After deriving the CPMV v0 and v1 of the current CU, the MVF of the current CU is generated according to the simplified affine motion model equation 1. In order to identify whether the current CU is coded or decoded in AF_MERGE mode, the affine flag is signaled in the bitstream when there is at least one neighboring block coded or decoded in affine mode.
[0085] In the example planned for adoption into VTM 3.0, the affine merge candidate list is constructed using the following steps:
[0086] 1) Insert inherited affine candidates
[0087] An inherited affine candidate is one that is derived from the affine motion model of its valid neighboring affine codec blocks. In a common basis, such as Figure 6 As shown, the scanning order of candidate positions is: A1, B1, B0, A0 and B2.
[0088] After a candidate is derived, a full pruning process is performed to check whether the same candidate has been inserted into the list. If there is an identical candidate, the derived candidate is discarded.
[0089] 2) Insert the constructed affine candidate
[0090] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (set to 5 in this paper), the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to a candidate constructed by combining the neighboring motion information of each control point.
[0091] The motion information of the control point is first obtained from Figure 5B The CPk (k=1, 2, 3, 4) is derived from the specified spatial and temporal neighbors shown. CPk (k=1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial locations of the predicted CPk (k=1, 2, 3); T is the temporal location of the predicted CP4.
[0092] The coordinates of CP1, CP2, CP3, and CP4 are (0,0), (W,0), (H,0), and (W,H), respectively, where W and H are the width and height of the current block.
[0093] Figure 6 Examples showing candidate positions for the affine merge mode
[0094] The motion information for each control point is obtained according to the following priority order:
[0095] For CP1, the priority is B2->B3->A2. If B2 is available, B2 is used. Otherwise, if B2 is not available, B3 is used. If both B2 and B3 are unavailable, A2 is used. If all three candidates are unavailable, motion information for CP1 cannot be obtained.
[0096] For CP2, the checking priority is B1->B0.
[0097] For CP3, the checking priority is A1->A0.
[0098] For CP4, use T.
[0099] Second, affine merge candidates are constructed using combinations of control points.
[0100] Constructing a 6-parameter affine candidate requires motion information from three control points. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} are converted into a 6-parameter motion model represented by the top left, top right, and bottom left control points.
[0101] Constructing a 4-parameter affine candidate requires motion information for two control points. These two control points can be selected from one of the following six combinations ({CP1, CP4}, {CP2, CP3}, {CP1, CP2}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4}). The combination {CP1, CP4}, {CP2, CP3}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4} will be converted into a 4-parameter motion model represented by the top left and top right control points.
[0102] The combinations of constructed affine candidates are inserted into the candidate list in the following order: {CP1,CP2,CP3}, {CP1,CP2,CP4}, {CP1,CP3,CP4}, {CP2,CP3,CP4}, {CP1,CP2}, {CP1,CP3}, {CP2,CP3}, {CP1,CP4}, {CP2,CP4}, {CP3,CP4}.
[0103] For a combined reference list X (X is 0 or 1), the reference index with the highest usage in the control point is selected as the reference index of list X, and motion vectors pointing to different reference pictures will be scaled.
[0104] After a candidate is derived, a full pruning process is performed to check whether the same candidate has been inserted into the list. If there is an identical candidate, the derived candidate will be discarded.
[0105] 3) Fill with zero motion vectors
[0106] If the number of candidates in the affine merge candidate list is less than 5, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.
[0107] 2.2 Affine Merge Mode with Prediction Offset
[0108] In this example, UMVE is extended to the affine merge mode, which we will refer to as the UMVE affine mode. The proposed method selects the first available affine merge candidate as the base predictor. It then applies a motion vector offset to the motion vector value of each control point from the base predictor. If no affine merge candidate is available, the proposed method is not used.
[0109] The inter prediction directions of the selected basic predictor and the reference index for each direction are used without change.
[0110] In the current implementation, the affine model of the current block is assumed to be a 4-parameter model, and only 2 control points need to be derived. Therefore, only the first 2 control points of the basic prediction amount will be used as the control point prediction amount.
[0111] For each control point, the zero_MVD flag is used to indicate whether the control point of the current block has the same MV value as the corresponding control point prediction. If the zero_MVD flag is true, no further signaling is required for the control point. Otherwise, the distance index and offset direction index are signaled for the control point.
[0112] A distance offset table of size 5 is used, as shown in the following table. The distance index is signaled to indicate which distance offset to use. The mapping between distance index and distance offset value is as follows: Figure 7 shown.
[0113] Table - Distance Offset Table
[0114] Distance DX 0 1 2 3 4 Distance Offset 1 / 2 pixel 1 pixel 2 pixels 4 pixels 8 pixels
[0115] The direction index can represent four directions as shown below, where only the x or y direction may have MV differences, but not both directions.
[0116] Offset direction IDX 00 01 10 11 x-dir-factor +1 –1 0 0 y-dir-factor 0 0 +1 –1
[0117] If inter prediction is unidirectional, the signaled distance offset is applied to the offset direction of the prediction quantity for each control point. The result will be the MV value for each control point.
[0118] For example, when the basic predictor is unidirectional and the motion vector value of the control point is MVP (v px ,v py ). When the distance offset and direction index are signaled, the motion vector of the corresponding control point of the current block will be calculated as follows.
[0119] MV(v x ,v y )=MVP(v px ,v py )+MV(x-dir-factor*distance-offset,y-dir-factor*distance-offset)
[0120] If inter prediction is bidirectional, the signaled distance offset is applied in the signaled offset direction of the L0 motion vector of the control point predictor; and the same distance offset with the opposite direction is applied to the L1 motion vector of the control point predictor. The result will be the MV value of each control point in each inter prediction direction.
[0121] For example, when the basic predictor is unidirectional and the motion vector value of the control point on L0 is MVP L0 (v 0px ,v 0py ), and the motion vector of the control point on L1 is MVP L1 (v 1px ,v 1py ). When the distance offset and direction index are signaled, the motion vector of the corresponding control point of the current block will be calculated as follows.
[0122] MV L0 (v 0x ,v 0y )=MVP L0 (v 0px ,v0py )+MV(x-dir-factor*distance-offset,y-dir-factor*distance-offset);
[0123] MV L1 (v 0x ,v 0y )=MVP L1 (v 0px ,v 0py )+MV(-x-dir-factor*distance-offset,-y-dir-factor*distance-offset).
[0124] 2.3 Final Motion Vector Expression
[0125] In an example, a final motion vector expression (UMVE) is proposed. The UMVE is used for Skip or Merge mode with the proposed motion vector expression method.
[0126] UMVE reuses the same Merge candidates as those included in the regular Merge candidate list in VVC. Among the Merge candidates, basic candidates can be selected and further extended by the proposed motion vector expression method.
[0127] UMVE provides a new method for representing motion vector difference (MVD), in which the starting point, motion magnitude and motion direction are used to represent the MVD.
[0128] Figure 8 An example of the UMVE search process is shown.
[0129] Figure 9 An example of a UMVE search point is shown.
[0130] The proposed technique uses the Merge candidate list as is, but for the extension of UMVE, only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered.
[0131] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.
[0132] Table 1. Basic candidate IDX
[0133] Basic Candidate IDX 0 1 2 3 Nth MVP First MVP Second MVP Third MVP Fourth MVP
[0134] If the number of basic candidates is equal to 1, the basic candidate IDX is not signaled.
[0135] The distance index is the motion amplitude information. The distance index indicates the predefined distance from the starting point information. The predefined distances are as follows:
[0136] Table 2a. Distance IDX
[0137]
[0138] During the entropy encoding and decoding process, the distance IDX is binarized into bits (bins) using a truncated unary code:
[0139] Table 2b: Distance IDX binarization
[0140]
[0141] In arithmetic coding, the first bit is coded using a probability context, while subsequent bits are coded using an equal probability model (also known as bypass coding).
[0142] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown below.
[0143] Table 3. Direction IDX
[0144] Direction IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –
[0145] The UMVE flag is signaled immediately after the Skip and Merge flags are sent. If the Skip and Merge flags are true, the UMVE flag is parsed. If the UMVE flag is 1, the UMVE syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is 1, this is AFFINE mode. However, if it is not 1, the Skip / Merge index is parsed for VTM Skip / Merge mode.
[0146] No additional line buffering is required for UMVE candidates, as software skip / merge candidates are used directly as base candidates. Using the input UMVE index, the MV complement is determined before motion compensation. No long line buffer is required for this purpose.
[0147] Under the current general test conditions, the first or second Merge candidate in the Merge candidate list may be selected as the basic candidate.
[0148] UMVE is called Merge with MVD (MMVD).
[0149] 2.4 Generalized Bidirectional Prediction
[0150] In traditional bidirectional prediction, the predictions from L0 and L1 are averaged to generate the final prediction using equal weight 0.5. The prediction generation formula is shown in Equation (3).
[0151] P TraditionalBiPred =(P L0 +P L1 +RoundingOffset)>>shiftNum (1)
[0152] In equation (3), the final prediction quantity of the traditional bidirectional prediction is P TraditionalBiPred , P L0 and P L1 are the predictions from L0 and L1 respectively, and RoundingOffset and shiftNum are used to normalize the final prediction.
[0153] Generalized Bi-prediction (GBI) is proposed to allow different weights to be applied to the prediction quantities from L0 and L1. The prediction quantity is generated as shown in equation (4).
[0154] P GBi =((1-w1)*P L0 +w1*P L1 +RoundingOffset GBi )>>shiftNum GBi (2)
[0155] In equation (4), P GBi is the final prediction of GBi, (1-w1) and w1 are the selected GBI weights applied to the predictions of L0 and L1 respectively. GBi and shiftNum GBi It is used to normalize the final prediction in GBi.
[0156] The supported weights for w1 are {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}. One equal-weight set and four unequal-weight sets are supported. For the equal-weight case, the process for generating the final prediction is exactly the same as in the traditional bidirectional prediction mode. For the true bidirectional prediction case under random access (RA) conditions, the number of candidate weight sets is reduced to three.
[0157] For Advanced Motion Vector Prediction (AMVP) mode, if the CU is bi-predictive, the weight selection in the GBI is explicitly signaled at the CU level. For Merge mode, the weight selection is inherited from the Merge candidate. In this proposal, GBI supports weighted averaging of DMVR-generated templates and the final prediction of BMS-1.0.
[0158] 2.5 Adaptive Motion Vector Difference Resolution
[0159] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the PU's motion vector and the predicted motion vector) is signaled in units of one-quarter luma samples. In VTM-3.0, Locally Adaptive Motion Vector Resolution (LAMVR) was introduced. In JEM, MVD can be encoded and decoded in units of one-quarter luma samples, integer luma samples, or four luma samples. The MVD resolution is controlled at the codec unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.
[0160] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or quad luma sample MV precision is used.
[0161] When the first MVD resolution flag of a CU is zero or is not coded for the CU (meaning all MVDs in the CU are zero), a quarter luma sample MV resolution is used for the CU. When the CU uses integer luma sample MV precision or four luma sample MV precision, the MVP in the CU's AMVP candidate list is rounded to the corresponding precision.
[0162] In arithmetic coding, the first MVD resolution flag is coded with one of three probability contexts: C0, C1, or C2, while the second MVD resolution flag is coded with a fourth probability context C3. The probability context Cx of the first MVD resolution flag is derived as (L represents the left neighboring block and A represents the top neighboring block):
[0163] If L is available, is inter-coded, and its first MVD resolution flag is not equal to 0, then xL is set equal to 1; otherwise, xL is set equal to 0.
[0164] If A is available, is inter-coded, and its first MVD resolution flag is not equal to 0, then xA is set equal to 1; otherwise, xA is set equal to 0.
[0165] x is set equal to xL+xA.
[0166] In the encoder, CU-level RD check is used to determine which MVD resolution to use for the CU. That is, for each MVD resolution, three CU-level RD checks are performed. In order to speed up the encoder, the following encoding scheme is applied in JEM:
[0167] During RD check of a CU with normal quarter luma sample MVD resolution, the motion information of the current CU (integer luma sample accuracy) is stored. The stored motion information (after rounding) is used as a starting point for further small-scale motion vector refinement during RD check of the same CU with integer luma sample and 4 luma sample MVD resolution, so that the time-consuming motion estimation process is not repeated three times.
[0168] ○ Conditionally call RD check for CUs with 4 luma sample MVD resolution. For a CU, when the RD cost of integer luma sample MVD resolution is much greater than the RD cost of quarter luma sample MVD resolution, skip the RD check for the CU's 4 luma sample MVD resolution.
[0169] In VTM-3.0, LAMVR is also called Integer Motion Vector (IMV).
[0170] 2.6 Current Image Reference
[0171] Decoder:
[0172] In this method, the currently (partially) decoded picture is considered a reference picture. The current picture is placed at the last position in reference picture list 0. Therefore, for slices that use the current picture as the only reference picture, the slice type is considered a P slice. In this method, the bitstream syntax follows the same syntax structure as that used for inter codecs, and the decoding process is unified with inter codecs. The only significant difference is that the block vector (the motion vector pointing to the current picture) always uses integer pixel resolution.
[0173] The changes with the block level CPR_flag method are:
[0174] When searching for this pattern in the encoder, the width and height of the block are both less than or equal to 16.
[0175] When the luma block vector is an odd integer, chroma interpolation is enabled.
[0176] When the SPS flag is turned on, the adaptive motion vector resolution (AMVP) of CPR mode is enabled. In this case, when using AMVR, the block vector can be switched between 1 pixel integer resolution and 4 pixel integer resolution at the block level.
[0177] Encoder side:
[0178] The encoder performs RD checks on blocks with width or height no greater than 16. For non-Merge mode, a block vector search is first performed using a hash-based search. If no valid candidate is found from the hash search, a local search based on block matching is performed.
[0179] In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 blocks. For larger current block sizes, a hash key match with a reference block occurs when all of its 4×4 blocks match the hash keys in the corresponding reference positions. If multiple reference blocks are found to match the current block with the same hash key, the block vector cost of each candidate is calculated and the block with the lowest cost is selected.
[0180] In the block matching search, the search range is set to 64 pixels to the left and top of the current block, and the search range is limited to within the current CTU.
[0181] 2.7 Merge List Design in an Example
[0182] VVC supports three different Merge list construction processes:
[0183] 1) Sub-block Merge candidate list: It includes ATMVP and affine merge candidates. Affine mode and ATMVP mode share a merge list construction process. Here, ATMVP and affine merge candidates can be added in sequence. The sub-block merge list size is signaled in the slice header and has a maximum value of 5.
[0184] 2) One-way prediction TPM Merge list: For triangle prediction mode, two partitions share a single merge list construction process, even though they can select their own merge candidate indices. When building the merge list, the block's spatial neighbors and two temporal blocks are checked. The motion information derived from the spatial neighbors and temporal blocks is referred to herein as regular motion candidates. These regular motion candidates are further used to derive multiple TPM candidates. Note that the transform is performed at the whole block level, even though the two partitions can use different motion vectors to generate their own prediction blocks.
[0185] In some embodiments, the one-way predictive TPM Merge list size is fixed to 5.
[0186] 3) Rule Merge list: For the remaining codec blocks, a common Merge list construction process is used. Here, spatial / temporal / HMVP, paired bi-prediction Merge candidates, and zero motion candidates can be inserted in sequence. The regular Merge list size is signaled in the slice header and has a maximum value of 6.
[0187] Sub-block Merge candidate list
[0188] It is suggested to put all sub-block related motion candidates into a separate Merge list in addition to the regular Merge list for non-sub-block Merge candidates.
[0189] The sub-block related motion candidates are put into a separate Merge list, which is named "sub-block Merge candidate list".
[0190] In one example, the sub-block Merge candidate list includes affine Merge candidates, and ATMVP candidates and / or sub-block-based STMVP candidates.
[0191] In this paper, the ATMVP Merge candidate in the normal Merge list is moved to the first position of the affine Merge list, so that all Merge candidates in the new list (i.e., the sub-block based Merge candidate list) are based on the sub-block codec tool.
[0192] Construct the affine merge candidate list by following the steps below:
[0193] 1) Insert inherited affine candidates
[0194] An inherited affine candidate is one that is derived from the affine motion model of its valid neighboring affine codec block. Up to two inherited affine candidates are derived from the affine motion model of the neighboring block and inserted into the candidate list. For the left predictor, the scan order is {A0, A1}; for the top predictor, the scan order is {B0, B1, B2}.
[0195] 2) Insert the constructed affine candidate
[0196] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (set to 5), the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to a candidate constructed by combining the neighboring motion information of each control point.
[0197] The motion information of the control point is first obtained from Figure 7 The CPk (k=1, 2, 3, 4) is derived from the specified spatial and temporal neighbors shown. CPk (k=1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial locations of the predicted CPk (k=1, 2, 3); T is the temporal location of the predicted CP4.
[0198] The coordinates of CP1, CP2, CP3, and CP4 are (0,0), (W,0), (H,0), and (W,H), respectively, where W and H are the width and height of the current block.
[0199] The motion information for each control point is obtained according to the following priority order:
[0200] For CP1, the priority is B2->B3->A2. If B2 is available, B2 is used. Otherwise, if B2 is not available, B3 is used. If both B2 and B3 are unavailable, A2 is used. If all three candidates are unavailable, motion information for CP1 cannot be obtained.
[0201] For CP2, the checking priority is B1->B0.
[0202] For CP3, the checking priority is A1->A0.
[0203] For CP4, use T.
[0204] Second, affine merge candidates are constructed using combinations of control points.
[0205] Constructing a 6-parameter affine candidate requires motion information from three control points. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} are converted into a 6-parameter motion model represented by the top-left, top-right, and bottom-left control points.
[0206] Constructing a 4-parameter affine candidate requires motion information for two control points. These two control points can be selected from one of the following two combinations ({CP1, CP2}, {CP1, CP3}). These two combinations will be converted into a 4-parameter motion model represented by the top left and top right control points.
[0207] The combinations of constructed affine candidates are inserted into the candidate list in the following order:
[0208] {CP1,CP2,CP3}, {CP1,CP2,CP4}, {CP1,CP3,CP4}, {CP2,CP3,CP4}, {CP1,CP2}, {CP1,CP3}.
[0209] Only when the CPs have the same reference index, the available combination of motion information of the CPs is added to the affine Merge list.
[0210] 3) Fill with zero motion vectors
[0211] If the number of candidates in the affine merge candidate list is less than 5, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.
[0212] 2.8 MMVD with Affine Merge Candidates in Examples
[0213] For example, the MMVD concept is applied to affine merge candidates (called affine merge with prediction offset). It is an extension of MVD (or "distance" or "offset") and is signaled after the affine merge candidate (called) is signaled. All CPMVs are added to the MVD to obtain a new CPMV. The distance table is specified as
[0214] Distance from IDX 0 1 2 3 4 Distance-Offset 1 / 2-pixel 1-pixel 2-pixel 4-pixel 8-pixel
[0215] In some embodiments, an offset mirroring method based on POC distance is used for bi-directional prediction. When the base candidate is bi-directionally predicted, the offset applied to L0 is signaled, and the offset on L1 depends on the temporal position of the reference pictures on list 0 and list 1.
[0216] If both reference pictures are on the same temporal side of the current picture, the same distance offset and the same offset direction are applied to the CPMV of L0 and L1.
[0217] When the two reference pictures are on different sides of the current picture, the CPMV of L1 will apply the distance offset in opposite offset directions.
[0218] 3. Examples of Problems Solved by Disclosed Embodiments
[0219] There are some potential problems in the design of MMVD:
[0220] The encoding / decoding / parsing process of UMVE information can be inefficient because it uses truncated unary binarization to encode and decode distance (MVD accuracy) information and uses a fixed length to bypass encoding and decoding direction indices. This is based on the assumption that 1 / 4 pixel accuracy is the highest percentage. However, this is not true for all types of sequences.
[0221] Possible distance designs may not be efficient.
[0222] 4. Examples of Techniques Implemented by Various Embodiments
[0223] The following list should be considered as an example to explain the general concept. These inventions should not be interpreted narrowly. In addition, these technologies can be combined in any way.
[0224] Parsing of distance indices (e.g., MVD precision indices)
[0225] 1. The distance index (DI) used in UMVE is proposed to be binarized without truncated unary codes.
[0226] a. In one example, DI can be binarized using a fixed length code, an Exponential-Golomb code, a truncated Exponential-Golomb code, a Rice code, or any other code.
[0227] 2. The distance index may be signaled using more than one syntax element.
[0228] 3. It is proposed to classify the set of allowed distances into multiple subsets, for example, K subsets (K is greater than 1). The subset index is first signaled (first syntax), followed by the distance index within the subset (second syntax).
[0229] a. For example, first signal mmvd_distance_subset_idx, then mmvd_distance_idx_in_subset.
[0230] i. In one example, mmvd_distance_idx_in_subset may be binarized with a unary code, a truncated unary code, a fixed length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0231] (i) Specifically, if there are only two possible distances in the subset, mmvd_distance_idx_in_subset may be binarized as a flag.
[0232] (ii) Specifically, if there is only one possible distance in the subset, mmvd_distance_idx_in_subset is not signaled.
[0233] (iii) Specifically, if mmvd_distance_idx_in_subset is binarized into a truncated code, the maximum value is set to the number of possible distances in the subset minus 1.
[0234] b. In one example, there are two subsets (eg, K=2).
[0235] i. In one example, one subset includes all fractional MVD precisions (e.g., 1 / 4 pixel, 1 / 2 pixel). Another subset includes all integer MVD precisions (e.g., 1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, 32 pixels).
[0236] ii. In one example, one subset may have only one distance (eg, 1 / 2 pixel) and another subset have all remaining distances.
[0237] c. In one example, there are three subsets (eg, K=3).
[0238] i. In one example, the first subset includes fractional MVD precisities (e.g., 1 / 4 pixel, 1 / 2 pixel); the second subset includes integer MVD precisities less than 4 pixels (e.g., 1 pixel, 2 pixels); and the third subset includes all other MVD precisities (e.g., 4 pixels, 8 pixels, 16 pixels, 32 pixels).
[0239] d. In one example, there are K subsets, and the number of K is set equal to the MVD accuracy allowed in LAMVR.
[0240] i. Optionally, in addition, the signaling of the subset index can be reused as performed for LAMVR (e.g., reusing the way to derive the context offset index; reusing the context, etc.)
[0241] ii. The distance within a subset can be determined by the associated LAMVR index (e.g., AMVR_mode in the specification).
[0242] e. In one example, how to define subsets and / or how many subsets there are can be predefined or dynamically adaptively changed.
[0243] f. In one example, the first syntax may be encoded or decoded using a truncated unary code, a fixed length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0244] g. In one example, the second grammar may be encoded and decoded using a truncated unary code, a fixed length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0245] h. In one example, the subset index (e.g., the first syntax) may not be explicitly encoded in the bitstream. Alternatively, the subset index may be dynamically derived, for example, based on encoding information (e.g., block dimensions) of the current block and / or previously encoded blocks.
[0246] i. In one example, the distance index within the subset may not be explicitly encoded in the bitstream (eg, the second syntax).
[0247] i. In one example, when the subset has only one distance, no further signaling of the distance index is required.
[0248] ii. Optionally, furthermore, the second syntax may be dynamically derived, for example based on codec information (eg, block dimensions) of the current block and / or previously coded blocks.
[0249] j. In one example, a first resolution bit is signaled to indicate whether DI is less than a predetermined number T. Optionally, a first resolution bit is signaled to indicate whether the distance is less than a predetermined number.
[0250] i. In one example, two syntax elements are used to represent the distance index: mmvd_resolution_flag is signaled first, followed by mmvd_distance_idx_in_subset.
[0251] ii. In one example, three syntax elements are used to represent the distance index: mmvd_resolution_flag is signaled first, followed by mmvd_short_distance_idx_in_subset when mmvd_resolution_flag is equal to 0, and followed by mmvd_long_distance_idx_in_subset when mmvd_resolution_flag is equal to 1.
[0252] iii. In one example, the distance index number T corresponds to a 1-pixel distance. For example, in Table 2a defined in VTM-3.0, T=2.
[0253] iv. In one example, the distance index number T corresponds to a 1 / 2 pixel distance. For example, in Table 2a defined in VTM-3.0, T=1.
[0254] v. In one example, the distance index number T corresponds to a distance of W pixels. For example, in Table 2a defined in VTM-3.0, T=3 corresponds to a distance of 2 pixels.
[0255] vi. In one example, if DI is less than T, the first resolution bit is equal to 0. Alternatively, if DI is less than T, the first resolution bit is equal to 1.
[0256] vii. In one example, if DI is less than T, a code of a short distance index is further signaled after the first resolution bit to indicate the value of DI.
[0257] (i) In one example, DI is signaled. DI can be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0258] a. When DI is binarized into a truncated code (such as a truncated unary code), the maximum codec value is T-1.
[0259] (ii) In one example, S=T-1-DI is signaled. T-1-DI can be binarized using a unary code, a truncated unary code, a fixed length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0260] a. When T-1-DI is binarized into a truncated code (such as a truncated unary code), the maximum codec value is T-1.
[0261] b. After S is parsed, DI is reconstructed as DI=TS-1.
[0262] viii. In one example, if DI is not less than T, a code of a long distance index is further signaled after the first resolution bit to indicate the value of DI.
[0263] (i) In one example, B=DI-T is signaled. DI-T can be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0264] a. When DI-T is binarized into a truncated code (such as a truncated unary code), the maximum codec value is DMax-T, where DMax is the maximum allowed distance, such as 7 in VTM-3.0.
[0265] b. After B is parsed, DI is reconstructed as DI=B+T.
[0266] (ii) In one example, B'=DMax-DI is signaled, where DMax is the maximum allowed distance, such as 7 in VTM-3.0. DMax-DI can be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0267] a. When DMax-DI is binarized into a truncated code (such as a truncated unary code), the maximum codec value is DMax-T, where DMax is the maximum allowed distance, such as 7 in VTM-3.0.
[0268] b. After B' is resolved, DI is reconstructed as DI = DMax - B'.
[0269] k. In one example, a first resolution bit is signaled to indicate whether DI is greater than a predetermined number T. Optionally, a first resolution bit is signaled to indicate whether the distance is greater than a predetermined number.
[0270] i. In one example, the distance index number T corresponds to a 1-pixel distance. For example, in Table 2a defined in VTM-3.0, T=2.
[0271] ii. In one example, the distance index number T corresponds to a 1 / 2 pixel distance. For example, in Table 2a defined in VTM-3.0, T=1.
[0272] iii. In one example, the distance index number T corresponds to a distance of W pixels. For example, in Table 2a defined in VTM-3.0, T=3 corresponds to a distance of 2 pixels.
[0273] iv. In one example, if DI is greater than T, the first resolution bit is equal to 0. Alternatively, if DI is greater than T, the first resolution bit is equal to 1.
[0274] v. In one example, if DI is not greater than T, a code of a short distance index is further signaled after the first resolution bit to indicate the value of DI.
[0275] (i) In one example, DI is signaled. DI can be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0276] a. When DI is binarized into a truncated code (such as a truncated unary code), the maximum codec value is T.
[0277] (ii) In one example, S=T-DI is signaled. T-DI can be binarized with a unary code, a truncated unary code, a fixed length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0278] a. When T-DI is binarized into a truncated code (such as a truncated unary code), the maximum codec value is T.
[0279] b. After S is parsed, DI is reconstructed as DI=TS.
[0280] vi. In one example, if DI is greater than T, a code of a long distance index is further signaled after the first resolution bit to indicate the value of DI.
[0281] (i) In one example, B=DI-1-T is signaled. DI-1-T may be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0282] a. When DI-1-T is binarized into a truncated code (such as a truncated unary code), the maximum codec value is DMax-1-T, where DMax is the maximum allowed distance, such as 7 in VTM-3.0.
[0283] b. After B is parsed, DI is reconstructed as DI=B+T+1.
[0284] (ii) In one example, B'=DMax-DI is signaled, where DMax is the maximum allowed distance, such as 7 in VTM-3.0. DMax-DI can be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential-Golomb code, a truncated exponential-Golomb code, a Rice code, or any other code.
[0285] a. When DMax-DI is binarized into a truncated code (such as a truncated unary code), the maximum codec value is DMax-1-T, where DMax is the maximum allowed distance, such as 7 in VTM-3.0.
[0286] b. After B' is resolved, DI is reconstructed as DI = DMax - B'.
[0287] 1. Several possible binarization methods for distance indices are: (It should be noted that two binarization methods should be considered the same if the process of changing all "1"s to "0"s in one method and changing all "0"s to "1"s in another method produces the same codeword as the other method.)
[0288]
[0289] 4. It is proposed to use one or more probabilistic contexts to encode and decode the first grammar
[0290] a. In one example, the first syntax is the first resolution bit mentioned above.
[0291] b. In one example, which probability context to use is derived from the first resolution bits of neighboring blocks.
[0292] c. In one example, which probability context to use is derived from the LAMVR values (eg, AMVR_mode values) of neighboring blocks.
[0293] 5. It is proposed to use one or more probability contexts to encode and decode the second grammar.
[0294] a. In one example, the second syntax is the short distance index mentioned above.
[0295] i. In one example, the first bit used to encode the short distance index is encoded with the probability context, and the other bits are bypass-encoded.
[0296] ii. In one example, the first N bits used to encode the short distance index are encoded with the probability context, and the other bits are bypass-encoded.
[0297] iii. In one example, all bits used to encode the short-distance index are encoded with the probability context.
[0298] iv. In one example, different bits may have different probability contexts.
[0299] v. In one example, several bits share a single probability context.
[0300] (i) In one example, the bits are consecutive.
[0301] vi. In one example, which probability context to use is derived from the short distance index of the neighboring blocks.
[0302] b. In one example, the second syntax is the long distance index mentioned above.
[0303] i. In one example, the first bit used to encode the long-range index is encoded with the probability context, and the other bits are bypass-encoded.
[0304] ii. In one example, the first N bits used to encode the long-distance index are encoded with the probability context, and the other bits are bypass-encoded.
[0305] iii. In one example, all bits used to encode and decode the long-range index are encoded and decoded with the probability context.
[0306] iv. In one example, different bits may have different probability contexts.
[0307] v. In one example, several bits share a single probability context.
[0308] (i) In one example, the bits are consecutive.
[0309] vi. In one example, which probability context to use is derived from the long-range indices of neighboring blocks.
[0310] Interaction with LAMVR
[0311] 6. It is proposed to encode and decode the first syntax (eg, first resolution bits) according to a probability model for encoding and decoding LAMVR information.
[0312] a. In one example, the first resolution bits are encoded and decoded in the same manner as the first MVD resolution flag (eg, shared context, or same context index derivation method, but LAMVR information of neighboring blocks is replaced by MMVD information).
[0313] i. In one example, which probability context to use to encode the first resolution bits is derived from the LAMVR information of neighboring blocks.
[0314] (i) In one example, which probability context to use to encode the first resolution bit is derived from the first MVD resolution flag of the neighboring block.
[0315] b. Optionally, the first MVD resolution flag is encoded and used as the first resolution bit when the distance index is encoded.
[0316] c. In one example, which probability model to use to encode and decode the first resolution bits may depend on the encoded LAMVR information.
[0317] i. For example, which probability model to use to encode the first resolution bits may depend on the MV resolution of the neighboring blocks.
[0318] 7. It is proposed to encode and decode the first bit of the short-distance index using a probabilistic context.
[0319] a. In one example, the first bit for encoding the short distance index is encoded in the same manner as encoding the second MVD resolution flag (e.g., shared context, or the same context index derivation method, but the LAMVR information of the neighboring block is replaced by MMVD information).
[0320] b. Optionally, the second MVD resolution flag is encoded and used as the first bit for encoding the short distance index when the distance index is encoded.
[0321] c. In one example, which probability model is used to encode the first bit for encoding the short distance index may depend on the encoded LAMVR information.
[0322] i. For example, which probability model is used to encode the first bit for encoding the short-distance index may depend on the MV resolution of the neighboring blocks.
[0323] 8. It is proposed to encode and decode the first bit of the long-distance index using a probabilistic context.
[0324] a. In one example, the first bit used to encode the long distance index is encoded in the same way as the second MVD resolution flag is encoded (e.g., shared context, or the same context index derivation method, but the LAMVR information of the neighboring block is replaced by MMVD information).
[0325] b. Optionally, the second MVD resolution flag is encoded and used as the first bit for encoding the long distance index when the distance index is encoded.
[0326] c. In one example, which probability model is used to encode the first bit for encoding the long-distance index may depend on the LAMVR information of the codec.
[0327] i. For example, which probability model is used to encode the first bit for encoding the long-range index may depend on the MV resolution of the neighboring blocks.
[0328] 9. For LAMVR mode, in arithmetic coding, the first MVD resolution flag is coded using one of three probability contexts (C0, C1, or C2); while the second MVD resolution flag is coded using a fourth probability context (C3). The following describes an example of deriving the probability context used to codec the distance index.
[0329] a. The probability context Cx of the first resolution bit is derived as (L represents the left neighboring block and A represents the top neighboring block):
[0330] If L is available, is inter-coded, and its first MVD resolution flag is not equal to 0, then xL is set equal to 1; otherwise, xL is set equal to 0.
[0331] If A is available, is inter-coded, and its first MVD resolution flag is not equal to 0, then xA is set equal to 1; otherwise, xA is set equal to 0.
[0332] X is set equal to xL+xA.
[0333] b. The probability context of the first bit of the codec long distance index is C3.
[0334] c. The probability context of the first bit of the short-range index for encoding and decoding is C3.
[0335] 10. It is proposed that when the MMVD mode is applied, the LAMVR MVD resolution is signaled.
[0336] a. It is proposed that when encoding and decoding the side information of the MMVD mode, the syntax used for signaling the LAMVR MVD resolution is reused.
[0337] b. When the signaled LAMVR MVD resolution is 1 / 4 pixel, a short-range index is signaled to indicate the MMVD distance in the first subset. For example, the short-range index can be 0 or 1 to represent MMVD distances of 1 / 4 pixel or 1 / 2 pixel respectively.
[0338] c. When the signaled LAMVR MVD resolution is 1 pixel, a medium-range index is signaled to indicate the MMVD distance in the second subset. For example, the medium-range index can be 0 or 1 to represent MMVD distances of 1 pixel or 2 pixels respectively.
[0339] d. When the signaled LAMVR MVD resolution is 4 pixels in the third subset, a long-range index is signaled to indicate the MMVD distance. For example, the medium-range index can be X to represent an MMVD distance of (4 << X) pixels.
[0340] e. In the following disclosure, the subset distance index can refer to the short-range index, the medium-range index, or the long-range index.
[0341] i. In one example, the subset distance index can be binary-coded using unary code, truncated unary code, fixed-length code, exponential-Golomb code, truncated exponential-Golomb code, Rice code, or any other code.
[0342] (i) Specifically, if there are only two possible distances in the subset, the subset distance index can be binary-coded as a flag.
[0343] (ii) Specifically, if there is only one possible distance in the subset, the subset distance index is not signaled.
[0344] (iii) Specifically, if the subset distance index is binary-coded as a truncated code, the maximum value is set to the number of possible distances in the subset minus 1.
[0345] ii. In one example, the first bit for encoding the subset distance index is encoded and decoded using probability context, and the other bits are bypassed for encoding and decoding.
[0346] iii. In one example, probability context is used to encode and decode the first N bits for encoding and decoding the subset distance index, and the other bits are bypass encoded and decoded.
[0347] iv. In one example, probability context is used to encode and decode all bits for encoding and decoding the subset distance index.
[0348] v. In one example, different bits can have different probability contexts.
[0349] vi. In one example, several bits share a single probability context.
[0350] (i) In one example, these bits are consecutive.
[0351] vii. It is proposed that a distance cannot appear in two different distance subsets.
[0352] viii. In one example, more distances can be signaled in the short distance subset.
[0353] (i) For example, the distances signaled in the short distance subset must be in sub-pixels, not integer pixels. For example, 5 / 4 pixels, 3 / 2 pixels, 7 / 4 pixels can be in the short distance subset, but 3 pixels cannot be in the short distance subset.
[0354] ix. In one example, more distances can be signaled in the medium distance subset.
[0355] (i) For example, the distances signaled in the medium distance subset must be integer pixels, but not in the form of 4N, where N is an integer. For example, 3 pixels, 5 pixels can be in the medium distance subset, but 24 pixels cannot be in the medium distance subset.
[0356] x. In one example, more distances can be signaled in the long distance subset.
[0357] (i) For example, the distances signaled in the long distance subset must be integer pixels in the form of 4N, where N is an integer. For example, 4 pixels, 8 pixels, 16 pixels or 24 pixels can be in the long distance subset.
[0358] 11. It is proposed that the variable for storing the MV resolution of the current block can be determined by the UMVE distance.
[0359] a. In one example, if the UMVE distance < T1 or <= T1, the MV resolution of the current block is set to 1 / 4 pixel.
[0360] b. In one example, if the UMVE distance < T1 or <= T1, the first and second MVD resolution flags of the current block are set to 0.
[0361] c. In one example, if the UMVE distance > T1 or >= T1, the MV resolution of the current block is set to 1 pixel.
[0362] d. In one example, if the UMVE distance > T1 or >= T1, the first MVD resolution flag of the current block is set to 1, and the second MVD resolution flag of the current block is set to 0.
[0363] e. In one example, if the UMVE distance > T2 or >= T2, the MV resolution of the current block is set to 4 pixels.
[0364] f. In one example, if the UMVE distance > T2 or >= T2, the first and second MVD resolution flags of the current block are set to 1.
[0365] g. In one example, if the UMVE distance > T1 or >= T1 and the UMVE distance < T2 or <= T2, the MV resolution of the current block is set to 1 pixel.
[0366] h. In one example, if the UMVE distance > T1 or >= T1 and the UMVE distance < T2 or <= T2, the first MVD resolution flag of the current block is set to 1, and the second MVD resolution flag of the current block is set to 0.
[0367] i. T1 and T2 can be any numbers. For example, T1 = 1 pixel, and T2 = 4 pixels.
[0368] 12. It is proposed that the variable for storing the MV resolution of the current block can be determined by the UMVE distance index.
[0369] a. In one example, if the UMVE distance index < T1 or <= T1, the MV resolution of the current block is set to 1 / 4 pixel.
[0370] b. In one example, if the UMVE distance index < T1 or <= T1, the first and second MVD resolution flags of the current block are set to 0.
[0371] c. In one example, if the UMVE distance index > T1 or >= T1, the MV resolution of the current block is set to 1 pixel.
[0372] d. In one example, if the UMVE distance index > T1 or >= T1, the first MVD resolution flag of the current block is set to 1, and the second MVD resolution flag of the current block is set to 0.
[0373] e. In one example, if the UMVE distance index > T2 or >= T2, the MV resolution of the current block is set to 4 pixels.
[0374] f. In one example, if the UMVE distance index > T2 or >= T2, the first and second MVD resolution flags of the current block are set to 1.
[0375] g. In one example, if the UMVE distance > T1 or >= T1 and the UMVE distance index < T2 or <= T2, the MV resolution of the current block is set to 1 pixel.
[0376] h. In one example, if the UMVE distance index > T1 or >= T1 and the UMVE distance index < T2 or <= T2, the first MVD resolution flag of the current block is set to 1, and the second MVD resolution flag of the current block is set to 0.
[0377] i. T1 and T2 can be any numbers. For example, T1 = 2, and T2 = 3, or T1 = 2, and T2 = 4;
[0378] 13. The variable for storing the MV resolution of the UMVE coded block can be used for coding and decoding subsequent blocks coded in the LAMVR mode.
[0379] a. Optionally, the variable for storing the MV resolution of the UMVE coded block can be used for coding and decoding subsequent blocks coded in the UMVE mode.
[0380] b. Optionally, the MV precision of the LAMVR coded block can be used for coding and decoding subsequent UMVE coded blocks.
[0381] 14. The above items can also be applied to the coding and decoding direction index.
[0382] Mapping between distance indices and distances
[0383] 15. It is proposed that the relationship between the distance index (DI) and the distance is not an exponential relationship as in VTM-3.0. (Distance = 1 / 4 pixel × 2 DI )
[0384] a. In one example, the mapping can be piecewise.
[0385] i. For example, when T0 <= DI < T1, distance = f1(DI), when T1 <= DI < T2, distance = f2(DI), … when Tn-1 <= DI < Tn, distance = fn(DI).
[0386] (i) For example, when DI < T1, distance = 1 / 4 pixel × 2 DI; When T1 <= DI < T2, distance = a × DI + b; when DI >= T2, distance = c × 2 DI . In one example, T1 = 4, a = 1, b = -1, T2 = 6, c = 1 / 8.
[0387] 16. It is proposed that the size of the distance table can be greater than 8, such as 9, 10, 12, 16.
[0388] 17. It is proposed that distances less than 1 / 4 pixel can be included in the distance table, such as 1 / 8 pixel, 1 / 16 pixel or 3 / 8 pixel.
[0389] 18. It is proposed that distances in non-2 X pixel form can be included in the distance table, such as 3 pixels, 5 pixels, 6 pixels, etc.
[0390] 19. It is proposed that the distance table can be different for different directions.
[0391] a. Accordingly, for different directions, the process of parsing the distance index can be different.
[0392] b. In one example, the four directions with direction indices 0, 1, 2, and 3 have different distance tables.
[0393] c. In one example, the two x-directions with direction indices 0 and 1 have the same distance table.
[0394] d. In one example, the two y-directions with direction indices 2 and 3 have the same distance table.
[0395] e. In one example, the x-direction and the y-direction can have two different distance tables.
[0396] i. Accordingly, for the x-direction and the y-direction, the process of parsing the distance index can be different.
[0397] ii. In one example, the distance table for the y-direction can have fewer possible distances than the distance table for the x-direction.
[0398] iii. In one example, the shortest distance in the distance table for the y-direction can be shorter than the shortest distance in the distance table for the x-direction.
[0399] iv. In one example, the longest distance in the distance table for the y-direction can be shorter than the longest distance in the distance table for the x-direction.
[0400] 20. It is proposed that different distance tables can be used for different block widths and / or heights.
[0401] a. In one example, when the direction is along the x-axis, different distance tables can be used for different block widths.
[0402] b. In one example, when the direction is along the y-axis, different distance tables can be used for different block heights.
[0403] 21. It is proposed that different distance tables can be used when the POC distances are different. The POC difference is calculated as |POC of the current picture – POC of the reference picture|.
[0404] 22. Different distance tables are proposed to be used for different basic candidates.
[0405] 23. It is proposed that the ratio of two distances (MVD precision) with consecutive indices is not fixed to 2.
[0406] a. In one example, the ratio of two distances (MVD precisions) with consecutive indices is fixed to M (eg, M=4).
[0407] b. In one example, the increment (rather than the ratio) of two distances (MVD precision) with consecutive indices can be fixed for all indices. Alternatively, the increment of two distances (MVD precision) with consecutive indices can be different for different indices.
[0408] c. In one example, the ratio of two distances (MVD precision) with consecutive indices may be different for different indices.
[0409] i. In one example, a distance set such as {1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, 32 pixels, 48 pixels, 64 pixels} may be used.
[0410] ii. In one example, a distance set such as {1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, 32 pixels, 64 pixels, 96 pixels} may be used.
[0411] iii. In one example, a distance set such as {1 pixel, 2 pixels, 3 pixels, 4 pixels, 5 pixels, 16 pixels, 32 pixels} may be used.
[0412] 24. Signaling of MMVD side information can be accomplished in the following ways:
[0413] a. When the current block is in inter mode and non-Merge mode (which may include, for example, non-skip, non-subblock, non-triangle, and non-MHIntra), the MMVD flag may be signaled first, followed by the subset index of the distance, the distance index within the subset, and the direction index. Here, MMVD is considered a different mode from the Merge mode.
[0414] b. Optionally, when the current block is in Merge mode, the MMVD flag may be further signaled, followed by the distance subset index, the distance index within the subset, and the direction index. Here, MMVD is considered a special Merge mode.
[0415] 25. The direction of the MMVD and the distance of the MMVD can be jointly signaled.
[0416] a. In one example, whether and how to signal the MMVD distance may depend on the MMVD direction.
[0417] b. In one example, whether and how to signal the MMVD direction may depend on the MMVD distance.
[0418] c. In one example, one or more syntax elements are used to signal a joint codeword. The MMVD distance and MMVD direction can be derived from the codeword. For example, the codeword is equal to the MMVD distance index + the MMVD direction index * 7. In another example, an MMVD codeword table is designed. Each codeword corresponds to a unique combination of MMVD distance and MMVD direction.
[0419] 26. Some exemplary UMVE distance tables are listed below:
[0420] a. The distance table size is 9:
[0421]
[0422]
[0423] b. The distance table size is 10:
[0424]
[0425] a. The distance table size is 12:
[0426]
[0427]
[0428] 27. It is proposed that a granular signaling method can be used to signal the MMVD distance. The distance is first signaled by an index with a coarse granularity, and then one or more indices with a finer granularity.
[0429] a. For example, the first index F1 represents the distance in the ordered set M1; the second index F2 represents the distance in the ordered set M2. The final distance is calculated as, for example, M1[F1]+M2[F2].
[0430] b. For example, the first index F1 represents the distance in the ordered set M1; the second index F2 represents the distance in the ordered set M2; and so on, up to the nth index F n representing the distance in the ordered set M n . The final distance is calculated as M1[F1] + M2[F2] +... + M n [F n .
[0431] c. For example, the signaling or binarization of F k can depend on the F k-1 signaled.
[0432] i. In one example, when F k-1 is not the maximum index pointing to M k-1 , the entry in M k [F k must be less than M k-1 [F k-1 + 1] - M k-1 [F k-1 . 1 < k <= n.
[0433] d. For example, the signaling or binarization of F k can depend on the F S signaled for all 1 <= s < k.
[0434] i. In one example, when 1 < k <= n, the entry in M k [F k must be less than M S [F S + 1] - M S [F S .
[0435] e. In one example, if F k is the maximum index pointing to M k , then F k+1 is no longer signaled, and the final distance is calculated as M1[F1] + M2[F2] +... + M k [F k , where 1 <= k <= n.
[0436] f. In one example, the entry in M k [F k can depend on the F k-1 signaled.
[0437] g. In one example, M k [F kThe entries in [] may depend on the signaling notified F for all 1 <= s < k S .
[0438] h. For example, n = 2. M1 = {1 / 4 pixel, 1 pixel, 4 pixels, 8 pixels, 16 pixels, 32 pixels, 64 pixels, 128 pixels},
[0439] i. When F1 = 0 (M1[F1] = 1 / 4 pixel), M2 = {0 pixel, 1 / 4 pixel};
[0440] ii. When F1 = 1 (M1[F1] = 1 pixel), M2 = {0 pixel, 1 pixel, 2 pixels};
[0441] iii. When F1 = 2 (M1[F1] = 4 pixels), M2 = {0 pixel, 1 pixel, 2 pixels, 3 pixels};
[0442] iv. When F1 = 3 (M1[F1] = 8 pixels), M2 = {0 pixel, 2 pixels, 4 pixels, 6 pixels};
[0443] v. When F1 = 4 (M1[F1] = 16 pixels), M2 = {0 pixel, 4 pixels, 8 pixels, 12 pixels};
[0444] vi. When F1 = 5 (M1[F1] = 32 pixels), M2 = {0 pixel, 8 pixels, 16 pixels, 24 pixels};
[0445] vii. When F1 = 6 (M1[F1] = 32 pixels), M2 = {0 pixel, 16 pixels, 32 pixels, 48 pixels};
[0446] Strip / Picture Level Control
[0447] 28. It is proposed how to signal MMVD side information (e.g., MMVD distance) and / or how to interpret the signaled MMVD side information (e.g., distance distance index), which may depend on the information signaled or inferred at a level higher than the CU level (e.g., sequence level, or picture level or slice level, or slice group level, such as in VPS / SPS / PPS / slice header / picture header / slice group header).
[0448] a. In one example, the code table index is signaled or inferred at a higher level. The specific code table is determined by the table index. The distance index can be signaled by the method disclosed in items 1 - 26. Then, the distance is derived by querying the entry with the signaled distance index in the specific code table.
[0449] b. In one example, parameter X is signaled or inferred at a higher level. The distance index can be signaled using the methods disclosed in items 1 - 26. Then, distance D’ is derived by looking up the entry with the signaled distance index in the code table. Then, the final distance D is calculated as D = f(D’, X). f can be any function. For example, f(D’, X) = D’ << X or f(D’, X) = D’ * X, or f(D’, X) = D’ + X, or f(D’, X) = D’ shifted right by X (with or without rounding).
[0450] c. In one example, the effective MV resolution is signaled or inferred at a higher level. Only the MMVD distances with the effective MV resolution can be signaled.
[0451] i. For example, the signaling method of MMVD information at the CU level can depend on the effective MV resolution signaled at a higher level.
[0452] (i) For example, the signaling method of MMVD distance resolution information at the CU level can depend on the effective MV resolution signaled at a higher level.
[0453] (ii) For example, the number of distance subsets can depend on the effective MV resolution signaled at a higher level.
[0454] (iii) For example, the meaning of each subset may depend on the effective MV resolution signaled at a higher level.
[0455] ii. For example, the minimum MV resolution (such as 1 / 4 pixel or 1 pixel or 4 pixels) is signaled.
[0456] (i) For example, when the minimum MV resolution is 1 / 4 pixel, the distance index is signaled as described in items 1 - 26.
[0457] (ii) For example, when the minimum MV resolution is 1 pixel, the flag (such as the first resolution bit in LAMVR) for signaling whether the distance resolution is 1 / 4 pixel is not signaled. Only the medium distance index and the long distance index disclosed in item 10 can be signaled after the LAMVR information.
[0458] (iii) For example, when the minimum MV resolution is 4 pixels, the flag (such as the first resolution bit in LAMVR) for signaling whether the distance resolution is 1 / 4 pixel is not signaled; and the flag (such as the second resolution bit in LAMVR) for signaling whether the distance resolution is 1 pixel is not signaled. Only the long distance index disclosed in item 10 can be signaled after the LAMVR information.
[0459] (iv) For example, when the minimum MV resolution is 1 pixel, the range resolution is signaled in the same way as when the minimum MV resolution is 1 / 4 pixel. However, the meaning of the range subset may be different.
[0460] For example, the short distance subset represented by the short distance index is redefined as a very long distance subset. For example, the two distances that can be signaled in this very long distance subset are 64 pixels and 128 pixels.
[0461] 29A. It is proposed that the encoder can determine whether a slice / slice / picture / sequence / CTU group / block group is screen content by checking the ratio of blocks with one or more similar or identical blocks within the same slice / slice / picture / sequence / CTU group / block group.
[0462] a. In one example, if the ratio is greater than a threshold, it is considered as screen content.
[0463] b. In one example, if the ratio is greater than a first threshold and less than a second threshold, it is considered as screen content.
[0464] c. In one example, a slice / slice / picture / sequence / CTU group / block group can be partitioned into M×N non-overlapping blocks. For each M×N block, the encoder checks whether there is another (or more) M×N block that is similar or identical to it. For example, M×N is equal to 4×4.
[0465] d. In one example, only some blocks are examined when calculating the ratio, for example, only blocks in even rows and even columns are examined.
[0466] e. In one example, a key value, such as a cyclic redundancy check (CRC) code, may be generated for each M×N block, and the key values of two blocks may be compared to check whether the two blocks are identical.
[0467] i. In one example, the key value may be generated using only some of the color components of the block. For example, the key value may be generated using only the luminance component.
[0468] ii. In one example, only some pixels of a block may be used to generate the key value, for example, only the even-numbered rows of the block may be used.
[0469] f. In one example, SAD / SATD / SSE or mean-removed SAD / SATD / SSE may be used to measure the similarity of two blocks.
[0470] i. In one example, SAD / SATD / SSE or mean-removed SAD / SATD / SSE may be calculated for only some pixels, for example, only for even-numbered rows.
[0471] Affine MMVD
[0472] 29B. It is proposed that the use indication of affine MMVD can be signaled only when the Merge index of the sub-block Merge list is greater than K (where K=0 or 1).
[0473] a. Optionally, when there are separate lists for the affine merge list and other merge lists (such as the ATMVP list), the use of the affine MMVD can be signaled only when the affine mode is enabled. In addition, the use of the affine MMVD can be signaled only when the affine mode is enabled and there is more than one basic affine candidate.
[0474] 30. It is proposed that in addition to the affine mode, the MMVD method can also be applied to other sub-block based codec tools, such as the ATMVP mode. In one example, if the current CU applies ATMVP and the MMVD on / off flag is set to 1, then MMVD is applied to ATMVP.
[0475] In one example, one set of MMVD side information may be applied to all sub-blocks, in which case one set of MMVD side information is signaled. Alternatively, different sub-blocks may select different sets, in which case multiple sets of MMVD side information may be signaled.
[0476] b. In one embodiment, the MV of each sub-block is added to the signaled MVD (also called offset or distance).
[0477] c. In one embodiment, when the sub-block Merge candidate is an ATMVP Merge candidate, the method of signaling the MMVD information is the same as the method when the sub-block Merge candidate is an affine Merge candidate.
[0478] d. In one embodiment, when the sub-block Merge candidate is an ATMVP Merge candidate, an offset mirror method based on POC distance is used for bidirectional prediction to add MVD to the MV of each sub-block.
[0479] 31. It is proposed that when the sub-block Merge candidate is an affine Merge candidate, the MV of each sub-block is added to the signaled MVD (also called offset or distance).
[0480] 32. It is proposed that the MMVD signaling method disclosed in item 1-28 can also be applied to signal the MVD used by the affine MMVD mode.
[0481] a. In one embodiment, the LAMVR information used to signal the MMVD information of the affine MMVD mode may be different from the LAMVR information used to signal the MMVD information of the non-affine MMVD mode.
[0482] i. For example, the LAMVR information used to signal the MMVD information of the affine MMVD mode is also used to signal the MV accuracy used in the affine inter mode; but the LAMVR information used to signal the MMVD information of the non-affine MMVD mode is used to signal the MV accuracy used in the non-affine inter mode.
[0483] 33. It is proposed that the MVD information of the sub-block Merge candidate in MMVD mode should be signaled in the same way as the MVD information of the regular Merge candidate in MMVD mode.
[0484] a. For example, they share the same distance table;
[0485] b. For example, they share the same mapping between distance indices and distances.
[0486] c. For example, they share the same direction.
[0487] d. For example, they share the same binarization method.
[0488] e. For example, they share the same arithmetic coding / decoding context.
[0489] 34. It is proposed that the MMVD side information signaling can depend on the codec mode, such as affine or normal merge or triangle merge mode or ATMVP mode.
[0490] 35. It is proposed that the predetermined MMVD edge information may depend on the encoding / decoding mode, such as affine or normal merge or triangle merge mode or ATMVP mode.
[0491] 36. It is proposed that the predetermined MMVD side information may depend on the color subsampling method (eg, 4:2:0, 4:2:2, 4:4:4:4) and / or the color component.
[0492] Triangle MMVD
[0493] 37. It is proposed that MMVD can be applied to triangle prediction mode.
[0494] a. After the TPM Merge candidate is signaled, the MMD information is signaled. The signaled TPM Merge candidate is considered the basic Merge candidate.
[0495] i. For example, using the same signaling method as the rule Merge MMVD to signal MMVD information;
[0496] ii. For example, the MMVD information is signaled using the same signaling method as that of the affine Merge or other types of sub-block Merge MMVD;
[0497] iii. For example, different from the regular Merge, affine Merge or other types of sub-block Merge MMVD signaling method to signaling MMVD information;
[0498] b. In one example, the MV of each triangular partition is added to the signaled MVD;
[0499] c. In one example, the MV of one triangular partition is added to the signaled MVD, and the MV of another triangular partition is added to f(signaled MVD), where f is any function.
[0500] i. In one example, f depends on the reference picture POCs or reference indices of the two triangle partitions.
[0501] ii. In one example, if the reference picture of one triangle partition precedes the current picture in display order, and the reference picture of another triangle partition follows the current picture in display order, then f(MVD)=-MVD.
[0502] 38. It is proposed that the MMVD signaling method disclosed in item 1-28 can also be applied to signaling MVD used in triangular MMVD mode.
[0503] a. In one embodiment, the LAMVR information used to signal the MMVD information of the affine MMVD mode may be different from the LAMVR information used to signal the MMVD information of the non-affine MMVD mode.
[0504] i. For example, the LAMVR information used to signal the MMVD information of the affine MMVD mode is also used to signal the MV accuracy used in the affine inter-frame mode; but the LAMVR information used to signal the MMVD information of the non-affine MMVD mode is used to signal the MV accuracy used in the non-affine inter-frame mode.
[0505] 39. For all of the above items, the MMVD side information may include, for example, offset tables (distances) and direction information.
[0506] 5. Example Embodiments
[0507] This section shows some embodiments of improved MMVD designs.
[0508] 5.1 Example #1 (MMVD distance index encoding and decoding)
[0509] In one embodiment, the first resolution bit is encoded to encode the MMVD distance. For example, it can be encoded using the same probability context as the first flag of the MV resolution.
[0510] – If the resolution bit is 0, the following flag is encoded or decoded. For example, it can be encoded or decoded with another probability context to indicate the short distance index. If the flag is 0, the index is 0; if the flag is 1, the index is 1.
[0511] Otherwise (resolution bit is 0), the long distance index L is encoded or decoded as a truncated unary code with a maximum value of MaxDI-2, where MaxDI is the maximum possible distance index, which is equal to 7 in this embodiment. After parsing out L, the distance index is reconstructed as L+2. In the exemplary C-type embodiment:
[0512] DI=2;
[0513]
[0514] The first bit of the long-distance index is encoded with the probability context, and the other bits are bypassed. In the C-type embodiment:
[0515] DI=2;
[0516]
[0517] Examples of proposed grammatical changes are highlighted, and deleted sections are marked with strikethrough.
[0518]
[0519]
[0520] In one example, mmvd_distance_subset_idx represents a resolution index as described above, and mmvd_distance_idx_in_subset represents a short distance or long distance index according to the resolution index. Truncated unary may be used to encode mmvd_distance_idx_in_subset.
[0521] In a random access test under common test conditions, this embodiment can achieve an average coding gain of 0.15% and a gain of 0.34% for a UHD sequence (class A1).
[0522]
[0523] 5.2 Example #2 (MMVD side information encoding and decoding)
[0524] MMVD is considered as a separate mode, not a Merge mode. Therefore, only when the Merge flag is 0 can the MMVD flag be further encoded or decoded.
[0525]
[0526]
[0527]
[0528] In one embodiment, the MMVD information is signaled as:
[0529]
[0530] mmvd_distance_idx_in_subset[x0][y0] is binarized to a truncated unary code. If amvr_mode[x0][y0] < 2, the maximum value of the truncated unary code is 1; otherwise (amvr_mode[x0][y0] is equal to 2), the maximum value is set to 3.
[0531] mmvd_distance_idx[x0][y0] is set equal to mmvd_distance_idx_in_subset[x0][y0]+2*amvr_mode[x0][y0].
[0532] Which probability context mmvd_distance_idx_in_subset[x0][y0] uses depends on amvr_mode[x0][y0].
[0533] 5.3 Example #3 (MMVD Stripe Level Control)
[0534] In the slice header, the syntax element mmvd_integer_flag is signaled.
[0535] The syntax changes are described below, with new additions highlighted in italics.
[0536] 7.3.2.1 Sequence Parameter Set RBSP Syntax
[0537]
[0538] 7.3.3.1 General Strip Header Syntax
[0539]
[0540] 7.4.3.1 Sequence Parameter Set RBSP Semantics
[0541] sps_fracmmvd_enabled_flag equal to 1 specifies that slice_fracmmvd_flag is present in the slice header syntax for B slices and P slices. sps_fracmmvd_enabled_flag equal to 0 specifies that slice_fracmmvd_flag is not present in the slice header syntax for B slices and P slices.
[0542] 7.4.4.1 General Strip Header Semantics
[0543] slice_fracmmvd_flag specifies the distance table used to derive MmvdDistance[x0][y0]. When not present, the value of slice_fracmmvd_flag is inferred to be 1.
[0544]
[0545]
[0546]
[0547] In one embodiment, the MMVD information is signaled as:
[0548]
[0549] mmvd_distance_idx_in_subset[x0][y0] is binarized to a truncated unary code. If amvr_mode[x0][y0] < 2, the maximum value of the truncated unary code is 1; otherwise (amvr_mode[x0][y0] is equal to 2), the maximum value is set to 3.
[0550] mmvd_distance_idx[x0][y0] is set equal to mmvd_distance_idx_in_subset[x0][y0] + 2*amvr_mode[x0][y0]. In one example, the probability context used by mmvd_distance_idx_in_subset[x0][y0] depends on amvr_mode[x0][y0].
[0551] The array index x0, y0 specifies the position (x0, y0) of the top left corner luma sample of the considered coding block relative to the top left corner luma sample of the picture. mmvd_distance_idx[x0][y0] and MmvdDistance[x0][y0] are as follows:
[0552] Table 7-9 - Specification of mmvdDistance[x0][y0] based on mmvd_distance_idx[x0][y0] when slice_fracmmvd_flag is equal to 1.
[0553] mmvd_distance_idx[x0][y0] MmvdDistance[x0][y0] 0 1 1 2 2 4 3 8 4 16 5 32 6 64 7 128
[0554] Table 7-9 - Specification of mmvdDistance[x0][y0] based on mmvd_distance_idx[x0][y0] when slice_fracmmvd_flag is equal to 0.
[0555] mmvd_distance_idx[x0][y0] MmvdDistance[x0][y0] 0 4 1 8 2 16 3 32 4 64 5 128 6 256 7 512
[0556] When mmvd_integer_flag is equal to 1, mmvd_distance=mmvd_distance<<2.
[0557] Figure 10 is a flow chart of an example method 1000 for video processing. The method 1000 includes: making (1002) a decision to apply a Merge with Motion Vector Difference (MMVD) mode to a current block of a video based on a set of MMVD side information, wherein the current block is partitioned into at least two partitions; and performing (1004) a conversion between the current block of the video and a bitstream representation of the video using the MMVD mode, wherein, in the MMVD mode, at least one Merge candidate selected for at least one partition is refined based on the set of MMVD side information.
[0558] Figure 11 is a flow chart of an example method 1100 for video processing. The method 1100 includes making (1102) a decision to apply a Merge with Motion Vector Difference (MMVD) mode to a current block of a video, wherein the current block is partitioned into at least two partitions; and performing (1104) a conversion between the current block of the video and a bitstream representation of the video using the MMVD mode, wherein, in the MMVD mode, at least one Merge candidate selected for at least one partition is refined, and a set of MMVD information associated with the refinement of the at least one Merge candidate is signaled.
[0559] With reference to methods 1000, 1100, some examples of motion vector signaling are described in Section 4 of this document, and the aforementioned methods may include the features and steps described below.
[0560] In one aspect, a method for video processing is disclosed, comprising: making a decision about applying a Merge with Motion Vector Difference (MMVD) mode to a current block of a video based on a set of MMVD side information, wherein the current block is partitioned into at least two partitions; and performing conversion between the current block of the video and a bitstream representation of the video using the MMVD mode, wherein, in the MMVD mode, at least one Merge candidate selected for at least one partition is refined based on the set of MMVD side information.
[0561] In another aspect, a method for video processing is disclosed, comprising: making a decision to apply a Merge with Motion Vector Difference (MMVD) mode to a current block of a video, wherein the current block is partitioned into at least two partitions; and performing conversion between the current block of the video and a bitstream representation of the video using the MMVD mode, wherein, in the MMVD mode, at least one Merge candidate selected for at least one partition is refined, and a set of MMVD information associated with the refinement of the at least one Merge candidate is signaled.
[0562] In one example, at least one of the at least two partitions is not rectangular.
[0563] In one example, the at least two partitions are two triangle partitions encoded using a triangle prediction mode (TPM).
[0564] In one example, the at least two partitions are two partitions coded in a geometric prediction mode (GEO).
[0565] In an example, the set of MMVD side information is signaled for at least one Merge candidate in the same manner as for at least one of a regular Merge candidate, an affine Merge candidate, and a sub-block based Merge candidate.
[0566] In an example, the set of MMVD side information is signaled for at least one Merge candidate in a different manner than for at least one of a regular Merge candidate, an affine Merge candidate, and a sub-block based Merge candidate.
[0567] In an example, at least one Merge candidate is used as a basic Merge candidate and is signaled before the set of MMVD side information.
[0568] In an example, the set of MMVD side information comprises a motion vector difference (MVD), and the MVD is added to a motion vector (MV) of at least one of the two partitions.
[0569] In an example, the set of MMVD side information includes a motion vector difference (MVD); the MVD is added to the motion vector (MV) of one of the two partitions, and f(MVD) is added to the MV of the other of the two partitions, f(MVD) representing a function of the MVD included in the set of MMVD side information.
[0570] In an example, the function depends on the picture order counts (POCs) or reference indices of the reference pictures of the two partitions.
[0571] In an example, if a reference picture of one partition and a reference picture of another partition are respectively located on both sides of a current picture to which a current block belongs in a display order, f(MVD)=-MVD.
[0572] In an example, the set of MMVD side information is indicated in a set of Locally Adaptive Motion Vector Resolution (LAMVR) information.
[0573] In an example, a set of MMVD side information for the affine MMVD mode and a set of MMVD side information for the non-affine MMVD mode are signaled in different sets of LAMVR information.
[0574] In an example, the LAMVR information indicating the set of MMVD side information for the affine MMVD mode further indicates the MV precision used in the affine inter mode, and the LAMVR information indicating the set of MMVD side information for the non-affine MMVD mode further indicates the MV precision used in the non-affine inter mode.
[0575] In an example, the set of MMVD edge information further includes at least one of an offset distance table and a direction information table of Merge candidates.
[0576] In one example, the converting includes encoding the current block into a bitstream representation of the video and decoding the current block from the bitstream representation of the video.
[0577] In one aspect, an apparatus in a video system is disclosed, the apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of any of the above examples.
[0578] In one aspect, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product comprising program code for executing the method in any one of the above examples.
[0579] Figure 12is a block diagram of a video processing device 1200. The device 1200 can be used to implement one or more methods described herein. The device 1200 can be included in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 1200 may include one or more processors 1202, one or more memories 1204, and video processing hardware 1206. The processor(s) 1202 can be configured to implement one or more methods described in this document. The memory(s) 1204 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 1206 can be used to implement some of the techniques described in this document in hardware circuits and can be partially or completely part of the processor 1202 (e.g., a graphics processor core GPU or other signal processing circuitry).
[0580] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits that are collocated or distributed at different locations within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on a transformed and coded error residual value, and may also be encoded using bits in the header and other fields in the bitstream.
[0581] It should be appreciated that several techniques have been disclosed that will benefit video encoder and decoder embodiments incorporated into video processing devices such as smartphones, laptops, desktops, and similar devices by allowing the use of virtual motion candidates constructed based on the various rules disclosed in this document.
[0582] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that implements a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an operating environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0583] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one or more computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0584] The processes and logic flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0585] By way of example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices. However, a computer need not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0586] Although this patent document contains many details, these should not be construed as limitations on any subject matter or the scope of the claims, but rather as descriptions of features unique to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, alone or in any suitable subcombination. Furthermore, although the features described above may be described as functioning in certain combinations, and even initially claimed to be so protected, in some cases one or more features in the combination may be deleted from the claimed combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0587] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that these operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0588] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: determining that a first mode applies to a current block of the video; determining motion information for each of at least two partitions of the current block, wherein at least one of the at least two partitions is non-square and non-rectangular, wherein the motion information of each partition is determined based on a base candidate index included in mode information of the first mode indicated in a bitstream of a video; Refining motion information of at least one of the at least two partitions, wherein the motion information is refined based on distance information and direction information, wherein the distance information and the direction information are included in the pattern information, the distance information indicates an offset of the refined motion information relative to the motion information, and the direction information indicates a direction of the offset; as well as performing conversion between the current block and the bitstream using the refined motion information of the at least one partition, wherein the motion vector of the motion information of the first partition of the at least two partitions is added to the offset, and The motion vector of the motion information of the second partition among the at least two partitions is added to f(a), where a is equal to the offset.
2. The method according to claim 1, wherein The at least two partitions are two triangle partitions coded in a triangle prediction mode.
3. The method according to claim 1, wherein The at least two partitions are two partitions coded in a geometric prediction mode.
4. The method according to claim 1, wherein f(a)=-a.
5. The method according to claim 1, wherein f(a) depends on the picture order count (POC) or reference index of the reference pictures of the at least two partitions.
6. The method according to claim 1, wherein The signaling notification method of the distance information and the direction information of the current block is the same as the signaling notification method of the block encoded and decoded in the affine mode or the regular Merge mode.
7. The method according to claim 1, wherein Determining motion information for each of the at least two partitions includes: constructing a candidate list for the current block; and The motion information of each partition is determined based on the basic candidate index and the candidate list corresponding to each partition.
8. The method according to claim 1, wherein The mode information is indicated in a set of local adaptive motion vector resolution information.
9. The method according to claim 1, wherein The converting includes encoding the current block into the bitstream.
10. The method according to claim 1, wherein The converting includes decoding the current block from the bitstream.
11. An apparatus for processing video data, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: determining that a first mode applies to a current block of the video; determining motion information for each of at least two partitions of the current block, wherein at least one of the at least two partitions is non-square and non-rectangular; wherein, The motion information of each partition is determined based on a basic candidate index included in the mode information of the first mode indicated in the bitstream of the video; Refining motion information of at least one of the at least two partitions, wherein the motion information is refined based on distance information and direction information, wherein the distance information and the direction information are included in the pattern information, the distance information indicates an offset of the refined motion information relative to the motion information, and the direction information indicates a direction of the offset; as well as performing conversion between the current block and the bitstream using the refined motion information of the at least one partition, wherein the motion vector of the motion information of the first partition of the at least two partitions is added to the offset, and The motion vector of the motion information of the second partition among the at least two partitions is added to f(a), where a is equal to the offset. 12 . The apparatus of claim 11 , wherein the at least two partitions are two triangle partitions encoded using a triangle prediction mode.
13. The apparatus according to claim 11, wherein the at least two partitions are two partitions coded in a geometric prediction mode. The apparatus according to claim 11 , wherein f(a)=−a.
15. A non-transitory computer-readable storage medium storing instructions that cause a processor to: determining that a first mode applies to a current block of the video; determining motion information for each of at least two partitions of the current block, wherein at least one of the at least two partitions is non-square and non-rectangular; wherein, The motion information of each partition is determined based on a basic candidate index included in the mode information of the first mode indicated in the bitstream of the video; Refining motion information of at least one of the at least two partitions, wherein the motion information is refined based on distance information and direction information, wherein the distance information and the direction information are included in the pattern information, the distance information indicates an offset of the refined motion information relative to the motion information, and the direction information indicates a direction of the offset; as well as performing conversion between the current block and the bitstream using the refined motion information of the at least one partition, wherein the motion vector of the motion information of the first partition of the at least two partitions is added to the offset, and The motion vector of the motion information of the second partition among the at least two partitions is added to f(a), where a is equal to the offset.
16. A method for storing a bitstream, the method comprising: determining that a first mode applies to a current block of the video; determining motion information for each of at least two partitions of the current block, wherein at least one of the at least two partitions is non-square and non-rectangular; wherein the motion information of each partition is determined based on a base candidate index included in mode information of the first mode indicated in a bitstream of a video; Refining motion information of at least one of the at least two partitions, wherein the motion information is refined based on distance information and direction information, wherein the distance information and the direction information are included in the pattern information, the distance information indicates an offset of the refined motion information relative to the motion information, and the direction information indicates a direction of the offset; generating the bitstream using the refined motion information of the at least one partition; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the motion vector of the motion information of the first partition of the at least two partitions is added to the offset, and The motion vector of the motion information of the second partition among the at least two partitions is added to f(a), where a is equal to the offset.
17. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of any one of claims 5-10, 16.
18. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for executing the method of any one of claims 2-10, 16.