Use of the Converted Single Prediction Candidate

The patent addresses the challenge of high bandwidth usage in digital video transmission by providing video processing techniques that optimize motion vector determination and affine models, leading to efficient encoding and decoding with reduced bandwidth and line buffer requirements.

JP7683069B2Active Publication Date: 2025-05-26DOUYIN VISION CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024039982
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-24
Filing Date
2024-03-14
Publication Date
2025-05-26
Estimated Expiration
2040-03-06

AI Technical Summary

Technical Problem

The increasing demand for digital video transmission leads to high bandwidth usage, and existing video coding technologies face challenges in efficiently encoding and decoding video data while minimizing bandwidth and line buffer requirements.

Method used

The patent describes various video processing techniques that can be employed by video encoders and decoders, including methods for determining motion vectors, affine models, and applicable coding tools based on block sizes and positions, to optimize the conversion between video blocks and bitstream representations.

Benefits of technology

These techniques enhance the efficiency of video encoding and decoding, reducing bandwidth and line buffer requirements, and improving the overall performance of video processing operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683069000033
    Figure 0007683069000033
  • Figure 0007683069000034
    Figure 0007683069000034
  • Figure 0007683069000035
    Figure 0007683069000035
Patent Text Reader

Abstract

To provide a technique for implementing a video processing technique.SOLUTION: In an example implementation, a method of video processing includes steps of determining a modified set of motion vectors for a conversion between a current block of video and a bitstream representation of the video, and performing the conversion on the basis of the modified set of motion vectors. Where the current block satisfies a condition, the modified set of motion vectors is a modified version of the set of motion vectors associated with the current block.SELECTED DRAWING: Figure 34
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This patent document is related to image and video coding and decoding.

Background Art

[0002] Digital video occupies the largest bandwidth usage on the Internet and other digital communication networks. As the number of user devices capable of receiving and displaying video increases, the bandwidth demand for digital video utilization is expected to continue to grow.

Summary of the Invention

[0003] This patent document discloses various video processing techniques that can be used by video encoders and decoders during encoding and decoding operations.

[0004] In one exemplary aspect, a method of video processing is disclosed. The method includes determining that a first motion vector of a sub-block of a current block and a second motion vector, which is a representative motion vector for the current block, comply with size constraints for conversion between the current block of the video and the bitstream representation of the video using an affine coding tool. The method also includes performing the conversion based on the determination.

[0005] In another exemplary aspect, a method of video processing is disclosed. The method includes determining an affine model having six parameters for conversion between the current block of the video and the bitstream representation of the video. The affine model is inherited from affine coding information of adjacent blocks of the current block. The method also includes performing the conversion based on the affine model.

[0006] In another exemplary aspect, a method of video processing is disclosed. The method includes determining whether dual-prediction coding technology is applicable to a block for conversion between a block of video and a bitstream representation of the video, based on the size of the block having a width W and a height H, where W and H are positive integers. The method also includes performing the conversion based on the determination.

[0007] In another exemplary aspect, a method of video processing is disclosed. The method includes determining whether a coding tree splitting process is applicable to a block for conversion between a block of video and a bitstream representation of the video, based on the size of a sub-block that is a child coding unit of the block according to the coding tree splitting process. The sub-block has a width W and a height H, where W and H are positive integers. The method also includes performing the conversion according to the determination.

[0008] In another exemplary aspect, a method of video processing is disclosed. The method includes determining whether an index of a coding unit level weighted bi-prediction (BCW) coding mode is derived for conversion between a current block of video and a bitstream representation of the video, based on rules regarding the position of the current block. In the BCW coding mode, a set of weights including a plurality of weights is used to generate a dual-prediction value of the current block. The method also includes performing the conversion according to the determination.

[0009] In another exemplary aspect, a method of video processing is disclosed. The method includes determining an intra prediction mode of a current block independently of intra prediction modes of adjacent blocks for conversion between the current block of video coded using combined inter and intra prediction (CIIP) coding technology and a bitstream representation of the video. The CIIP coding technology uses intermediate inter prediction values and intermediate intra prediction values to derive a final prediction value for the current block. The method also includes performing the conversion based on the determination.

[0010] In another exemplary aspect, a method of video processing is disclosed. The method includes determining an intra prediction mode of the current block according to a first intra prediction mode of a first adjacent block and a second intra prediction mode of a second adjacent block for conversion between the current block of video coded using combined inter and intra prediction (CIIP) coding technology and a bitstream representation of the video. The first adjacent block is coded using intra prediction coding technology, and the second adjacent block is coded using CIIP coding technology. The first intra prediction mode is given a different priority than the second intra prediction mode. The CIIP coding technology uses intermediate inter prediction values and intermediate intra prediction values to derive a final prediction value for the current block. The method also includes performing the conversion based on the determination.

[0011] In another exemplary aspect, a method of video processing is disclosed. The method includes determining whether a combined inter and intra prediction (CIIP) process is applicable to a color component of the current block based on a size of the current block for conversion between the current block of video and a bitstream representation of the video. The CIIP coding technology uses intermediate inter prediction values and intermediate intra prediction values to derive a final prediction value for the current block. The method also includes performing the conversion based on the determination.

[0012] In another exemplary aspect, a method of video processing is disclosed. The method includes determining, based on characteristics of a current block of video, whether inter and intra combined prediction (CIIP) coding techniques should be applied to the current block for conversion between the current block of video and a bitstream representation of the video. The CIIP coding techniques use an intermediate inter prediction value and an intermediate intra prediction value to derive a final prediction value for the current block. The method also includes performing the conversion based on the determination.

[0013] In another exemplary aspect, a method of video processing is disclosed. The method includes determining, based on whether the current block of video is coded by inter and intra combined prediction (CIIP) coding techniques, whether coding tools should be disabled for the current block for conversion between the current block of video and a bitstream representation of the video. The coding tools include at least one of bidirectional optical flow (BDOF), overlapped block motion compensation (OBMC), or decoder-side motion vector refinement process (DMVR). The method also includes performing the conversion based on the determination.

[0014] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a first prediction P1 used for a motion vector for spatial motion prediction and a second prediction P2 used for a motion vector for temporal motion prediction for conversion between a block of video and a bitstream representation of the video. P1 and / or P2 are fractional and neither P1 nor P2 is conveyed in the bitstream representation. The method also includes performing the conversion based on the determination.

[0015] In another exemplary aspect, a method of video processing is disclosed. The method includes determining motion vectors (Mvx, MVy) by prediction (Px, Py) for conversion between a block of video and a bitstream representation of the video. Px is associated with MVx and Py is associated with MVy. MVx and MVy are stored as integers, each having N bits, where MinX ≦ MVx ≦ MaxX and MinY ≦ MVy ≦ MaxY, and MinX, MaxX, MinY, and MaxY are real numbers. The method also includes performing the conversion based on the determination.

[0016] In another exemplary aspect, a method of video processing is disclosed. The method includes determining whether a shared merge list is applicable to a current block of video for conversion between the current block of video and a bitstream representation of the video, according to the coding mode of the current block. The method also includes performing the conversion based on the determination.

[0017] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a second block of dimensions (W + N - 1) × (H + N - 1) for motion compensation during conversion for conversion between a current block of video having a size of W × H and a bitstream representation of the video. The second block is determined based on a reference block of dimensions (W + N - 1 - PW) × (H + N - 1 - PH). N represents a filter size, and W, H, N, PW, and PH are non-negative integers. Neither PW nor PH is equal to 0. The method also includes performing the conversion based on the determination.

[0018] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a second block of dimension (W+N-1)×(H+N-1) for motion compensation during conversion for conversion between a current block of a video having a size of W×H and a bitstream representation of the video. W and H are non-negative integers, and N is a non-negative integer based on a filter size. During conversion, refined motion vectors are determined based on multi-point search according to a motion vector refinement operation on the original motion vectors, and the pixel length boundaries of the reference blocks are determined by repeating one or more non-boundary pixels. The method also includes performing the conversion based on the determination.

[0019] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a predicted value at a position within a block for conversion between the block of a video coded using inter-intra composite prediction (CIIP) coding technology and a bitstream representation of the video, based on a weighted sum of an inter-predicted value and an intra-predicted value at that position. The weighted sum is based on adding an offset to an initial sum obtained based on the inter-predicted value and the intra-predicted value, and the offset is added before a right shift operation performed to determine the weighted sum. The method also includes performing the conversion based on the determination.

[0020] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a manner in which coding information of a current block is represented in a bitstream representation, at least in part, based on whether a condition related to the size of the current block is satisfied, for conversion between the current block of a video and the bitstream representation of the video. The method also includes performing the conversion based on the determination.

[0021] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a modified set of motion vectors for conversion between a current block of video and a bitstream representation of the video, and performing the conversion based on the modified set of motion vectors. The modified set of motion vectors is a modified version of the set of motion vectors associated with the current block when the current block satisfies a condition.

[0022] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a unidirectional motion vector from a bidirectional motion vector when a block size condition is satisfied for conversion between a current block of video and a bitstream representation of the video. The unidirectional motion vector is then used as a merge candidate for the conversion. The method also includes performing the conversion based on the determination.

[0023] In another exemplary aspect, a method of video processing is disclosed. The method includes determining that motion candidates for a current block of video are restricted to be in a unidirectional prediction direction based on the size of the current block for conversion between the current block of video and a bitstream representation of the video, and performing the conversion based on the determination.

[0024] In another exemplary aspect, a method of video processing is disclosed. The method includes determining a size limit between a representative motion vector of a current video block being affine coded and motion vectors of sub-blocks of the current video block, and performing a conversion between the bitstream representation and pixel values of the current video block or sub-block by using the size limit.

[0025] In another exemplary aspect, other methods of video processing are disclosed. The method includes determining, for a current video block to be affine-coded, one or more sub-blocks of the current video block, where each sub-block has a size of M×N pixels, with M and N being multiples of 2 or 4; aligning the motion vectors of the sub-blocks with a size limit; and conditionally performing, based on a trigger, a conversion between a bitstream representation and pixel values of the current video block by using the size limit.

[0026] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes determining that a current video block satisfies a size condition and, based on that determination, performing a conversion between a bitstream representation and pixel values of the current video block by excluding a bi-prediction coding mode for the current video block.

[0027] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes determining that a current video block satisfies a size condition and, based on that determination, performing a conversion between a bitstream representation and pixel values of the current video block, where an inter-prediction mode is signaled in the bitstream according to a size limit.

[0028] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes determining that a current video block satisfies a size condition and, based on that determination, performing a conversion between a bitstream representation and pixel values of the current video block, where generation of a merge candidate list during the conversion depends on the size condition.

[0029] In yet another example aspect, other methods of video processing are disclosed. The method includes determining that a child coding unit of a current video block satisfies a size condition, and based on that determination, performing a conversion between a bitstream representation and pixel values of the current video block, wherein the coding tree splitting process used to generate the child coding unit depends on the size condition.

[0030] In yet another example aspect, other methods of video processing are disclosed. The method includes determining a weight index for a generalized bi-prediction (GBi) process for a current video block based on the position of the current video block, and performing a conversion between the current video block and its bitstream representation using the weight index to implement the GBi process.

[0031] In yet another example aspect, other methods of video processing are disclosed. The method includes determining that the current video block is coded as an intra-inter prediction (IIP) coding block, and performing a conversion between the current video block and its bitstream representation using a simplification rule that determines an intra prediction mode or a most probable mode (MPM) for the current video block.

[0032] In yet another example aspect, other methods of video processing are disclosed. The method includes determining that the current video block satisfies a simplification criterion, and performing a conversion between the current video block and the bitstream representation by disabling the use of an inter-intra prediction mode for the conversion or by disabling additional coding tools used for the conversion.

[0033] In yet another example aspect, other methods of video processing are disclosed. The method includes performing a conversion between a current video block and a bitstream representation for the current video block using an encoding process based on motion vectors, where (a) precision P1 is used to store spatial motion prediction results, precision P2 is used to store temporal motion prediction results during the conversion process, P1 and P2 are fractions, or (b) precision Px is used to store x motion vectors, precision Py is used to store y motion vectors, and Px and Py are fractions.

[0034] In yet another example aspect, other methods of video processing are disclosed. The method includes fetching a block of (W2+N-1-PW)×(H2+N-1-PH) where W1, W2, H1, H2, and PW and PH are integers, pixel padding the fetched block, performing boundary pixel repetition on the pixel-padded block, and obtaining pixel values of small sub-blocks to interpolate a small sub-block of size W1×H1 within a large sub-block of size W2×H2 of the current video block, and performing a conversion between the current video block and a bitstream representation of the current video block using the interpolated pixel values of the small sub-block.

[0035] In another example aspect, other methods of video processing are disclosed. The method includes fetching (W+N-1-PW)×(H+N-1-PH) reference pixels during the conversion of a current video block of size W×H and a bitstream representation of the current video block, and performing a motion compensation operation by padding reference pixels larger than the fetched reference pixels during the motion compensation operation, and performing a conversion between the current video block and a bitstream representation of the current video block using the result of the motion compensation operation, where W, H, N, PW, and PH are integers.

[0036] In yet another example aspect, other methods of video processing are disclosed. The method includes determining, based on the size of a current video block, that dual prediction or single prediction of the current video block is not permitted, and based on that determination, performing a conversion between a bitstream representation and pixel values of the current video block by disabling the dual prediction or single prediction mode.

[0037] In yet another example aspect, other methods of video processing are disclosed. The method includes determining, based on the size of a current video block, that dual prediction or single prediction of the current video block is not permitted, and based on that determination, performing a conversion between a bitstream representation and pixel values of the current video block by disabling the dual prediction or single prediction mode.

[0038] In yet another example aspect, a video encoder device is disclosed. The video encoder has a processor configured to implement the above method.

[0039] In yet another example aspect, a video decoder device is disclosed. The video decoder has a processor configured to implement the above method.

[0040] In yet another example aspect, a computer-readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0041] These and other features are described throughout this document.

Brief Description of the Drawings

[0042]

Figure 1

Figure 2A

Figure 2B

Figure 3

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

[0043] Section headings are used in this document to facilitate understanding, and do not limit the applicability of the technologies and embodiments disclosed in each section to that section only.

[0044] [1. Summary] This patent document is related to video / image coding technology. Specifically, it is related to reducing the bandwidth and line buffer of some coding tools in video / image coding. It may be applied to existing video coding standards such as HEVC, or standards to be completed (Versatile Video Coding). It may also be applicable to future video / image coding standards or video / image codecs.

[0045] [2. Background] Video coding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video and H264 / MPEG-4 AVC (Advanced Video Coding) as well as the H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure, using temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, the JVET (Joint Video Exploration Team) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been introduced by JVET and placed in the reference software named JEM (Joint Exploration Model). In April 2018, the JVET (Joint Video Expert Team) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was formed to work on the VVC standard aiming for a 50% bitrate reduction compared to HEVC.

[0046] [Inter Prediction in HEVC / VVC] [Interpolation Filter]< In HEVC, the luma subsample is generated by an 8-tap interpolation filter, and the chroma subsample is generated by a 4-tap interpolation filter.

[0047] The filter is separable in two dimensions. Samples are first filtered horizontally and then vertically.

[0048] [Sub-Block Based Prediction Techniques] Sub-block based prediction was first introduced into the video coding standard by HEVC Annex I (3D-HEVC). According to sub-block based prediction, a block such as a Coding Unit (CU) or a Prediction Unit (PU) is divided into several non-overlapping sub-blocks. Different sub-blocks may be assigned different motion information such as a reference index or a Motion Vector (MV), and Motion Compensation (MC) is performed individually for each sub-block. Figure 1 shows the concept of sub-block based prediction.

[0049] To explore future video coding technologies beyond HEVC, the JVET (Joint Video Exploration Team) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been introduced by the JVET and placed in the reference software named JEM (Joint Exploration Model).

[0050] In JEM, sub-block based prediction is adopted in several coding tools such as affine prediction, Alternative Temporal Motion Vector Prediction (ATMVP), Spatial-Temporal Motion Vector Prediction (STMVP), Bi-directional Optical flow (BIO), and Frame-Rate Up Conversion (FRUC). Affine prediction is also adopted in VVC.

[0051] [2.3 Affine Prediction] In HEVC, only the translational motion model is applied for motion compensation prediction (MCP). On the other hand, in the real world, there are many types of motions, such as zoom-in / out, rotation, projective motion, and other irregular motions. In VVC, a simplified affine transform motion compensation prediction is applied. As shown in FIGS. 2A to 2C, the affine motion field of a block is represented by two control point motion vectors (in the 4-parameter affine model) or three control point motion vectors (in the 6-parameter affine model).

[0052] The motion vector field (MVF) of a block is represented by the following equations according to the 4-parameter affine model (the 4 parameters are defined as variables a, b, e, and f) in Equation (1) and the 6-parameter affine model (the 6 parameters are defined as variables a, b, c, d, e, and f) in Equation (2), respectively:

Equation

[0053] Here, (mv h 0 , mv v 0 ) is the motion vector of the control point at the upper left corner, (mv h 1 , mv v 1 ) is the motion vector of the control point at the upper right corner, (mv h 2 , mv v 2is the motion vector of the bottom-left control point. All three of these motion vectors are called Control Point Motion Vectors (CPMV). (x, y) represents the coordinates of the representative point with respect to the top-left sample within the current block. The CP motion vector may be signaled (as in the affine AMVP mode) or derived on-the-fly (as in the affine merge mode). w and h are the width and height of the current block. In practice, the division is implemented by a right shift with a rounding operation. In VTM, the representative point is defined to be the center position of the sub-block. For example, if the coordinates of the top-left corner of the sub-block with respect to the top-left sample within the current block are (xs, ys), the coordinates of the representative point are defined to be (xs + 2, ys + 2).

[0054] In a design without partitioning, equations (1) and (2) are:

Number

[0055] For the 4-parameter affine model shown in equation (1):

Number

[0056] For the 6-parameter affine model shown in equation (2):

Number

[0057] Finally,

Number

[0058] Here, S represents the calculation accuracy. For example, in VVC, S = 7. In VVC, for the MC of a sub-block by the top-left sample at (xs, ys), the MV used is calculated by Equation (6) using x = xs + 2 and y = ys + 2.

[0059] To derive the motion vector of each 4×4 sub-block, as shown in FIG. 3, the motion vector of the central sample of each sub-block is calculated according to Equation (1) or Equation (2) and rounded to 1 / 16 fractional accuracy. Then, a motion compensation interpolation filter is applied to generate a prediction for each sub-block based on the derived motion vector.

[0060] The affine model can be inherited from spatial adjacent affine coding blocks such as the left, top, top-right, bottom-left, and top-left adjacent blocks as shown in FIG. 4A. For example, when the bottom-left adjacent block A in FIG. 4 is coded in affine mode as represented by A0 in FIG. 4B, the control point (CP) motion vectors mv 0 N , mv 1 N and mv 2 N of the top-left corner, top-right corner, and bottom-left corner of the adjacent CU / PU containing block A are fetched. Then, the bottom-left / top-right / bottom-left motion vectors mv 0 C , mv 1 C and mv 2 C on the current CU / PU are mv 0 N , mv 1 N and mv 2 NIt is calculated based on. It should be noted that in VTM-2.0, when the current block is affine-coded, for a sub-block (e.g., a 4×4 block in VTM), LT saves mv0 and RT saves mv1. When the current block is coded by a 6-parameter affine model, LB saves mv2, and when it is not (coded by a 4-parameter affine model), LB saves mv2'. Other sub-blocks save the MVs used for MC.

[0061] It should be noted that when a CU is coded in the affine merge mode, e.g., in the AF_MERGE mode, it obtains the first block coded in the affine mode from a valid adjacent reconstructed block. And the selection order of candidate blocks is from left to top, upper right, lower left, upper left as shown in Figure 4A.

[0062] The derived CP MV mv of the current block 0 C , mv 1 C and mv 2 C can be used as the CP MV in the affine merge mode. Alternatively, they can be used as the MVP for the affine inter mode in VVC. It should be noted that for the merge mode, when the current block is coded in the affine mode, after deriving the CP MV of the current block, the current block may be further divided into multiple sub-blocks, and each block derives its motion information based on the derived CP MV of the current block.

[0063] [2.4 Example embodiments in JVET] Different from VTM where only one affine spatial adjacent block can be used to derive the affine motion of a block, in some embodiments, separate lists of affine candidates are constructed for the AF_MERGE mode.

[0064] 1) Insert the inherited affine candidates into the candidate list The inherited affine candidates mean that candidates are derived from valid adjacent reconstruction blocks coded in the affine mode. As shown in Figure 5, the scanning order of the candidate blocks is A 1 , B 1 , B 0 , A 0 , and B 2 . When a block is selected (e.g., A 1 ), a two - step procedure is applied.

[0065] 1.a First, use the motion vectors of the three corners of the CU covering the block to derive two / three control points of the current block.

[0066] 1.b Based on the control points of the current block, derive the sub - block motion of each sub - block within the current block.

[0067] 2) Insert the constructed affine candidates When the number of candidates in the affine merge candidate list is less than MaxNumAffineCand, the constructed affine candidates are inserted into the candidate list.

[0068] The constructed affine candidates mean that candidates are constructed by combining the adjacent motion information of each control point.

[0069] The motion information of the control points is first derived from the specified spatial neighborhood and temporal neighborhood shown in Figure 5. CPk (k = 1, 2, 3, 4) represents the k - th control point. A 0 , A 1 , A 2 , B 0 , B 1 , B 2 and B 3 are the spatial positions for predicting CPk (k = 1, 2, 3), and T is the temporal position for predicting CP4.

[0070] The coordinates of CP1, CP2, CP3, and CP4 are (0,0), (W,0), (H,0), and (W,H), respectively, where W and H are the width and height of the current block.

[0071] The movement information of each control point is obtained according to the following priority order.

[0072] 2.a For CP1, the check priority is B 2 ->B 3 ->A 2 That is. B 2 is used if it is available. If not, if B 3 is available, B 3 is used. If both B2 and B3 are unavailable, A 2 is used. If all three candidates are unavailable, the movement information of CP1 cannot be obtained.

[0073] 2.b For CP2, the check priority is B 1 ->B 0 That is.

[0074] 2.c For CP3, the check priority is A 1 ->A 0 That is.

[0075] 2.d For CP4, T is used.

[0076] Second, combinations of control points are used to construct a motion model.

[0077] The motion vectors of three control points are required to calculate the transformation parameters in a six-parameter affine model. The three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}). For example, using the CP1, CP2, and CP3 control points, a six-parameter affine motion model represented as Affine(CP1, CP2, CP3) is constructed.

[0078] The motion vectors of two control points are required to calculate the transformation parameters in a four-parameter affine model. The two control points can be selected from one of the following six combinations ({CP1, CP4}, {CP2, CP3}, {CP1, CP2}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4}). For example, using the CP1 and CP2 control points, a four-parameter affine motion model represented as Affine(CP1, CP2) is constructed.

[0079] The constructed combinations of affine candidates are inserted into the candidate list in the following order: {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}, {CP2, CP3}, {CP1, CP4}, {CP2, CP4}, {CP3, CP4}.

[0080] 3) Insert zero motion vector When the number of candidates in the affine merge candidate list is less than MaxNumAffineCand, zero motion vectors are inserted into the candidate list until the list is full.

[0081] [2.5 Affine merge candidate list] [2.5.1 Affine merge mode] In the affine merge mode of VTM-2.0.1, only the first available affine neighborhood can be used to derive the motion information of the affine merge mode. In some embodiments, the candidate list for the affine merge mode is constructed by searching for valid affine neighborhoods and combining the adjacent motion information of each control point.

[0082] The affine merge candidate list is constructed as the following steps.

[0083] 1) Insert inherited affine candidates The inherited affine candidates mean that the candidates are derived from the affine motion model of its valid adjacent affine coding block. On a general basis, as shown in FIG. 5, the scanning order of the candidate positions is A 1 , B 1 , B 0 , A 0 , and B 2 .

[0084] After the candidates are derived, a full pruning process is executed to check whether the same candidate is inserted into the list. If the same candidate exists, the derived candidate is discarded.

[0085] 2) Insert constructed affine candidates When the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (set to 5 in this application), the constructed affine candidates are inserted into the candidate list. The constructed affine candidates mean that the candidates are constructed by combining the adjacent motion information of each control point.

[0086] The motion information of the control points is first derived from the specified spatial neighborhood and temporal neighborhood. CPk (k = 1, 2, 3, 4) represents the k-th control point. A 0 , A 1 , A 2 , B 0 , B 1 , B 2 and B3 is the spatial position for predicting CPk (k = 1, 2, 3), and T is the temporal position for predicting CP4.

[0087] The coordinates of CP1, CP2, CP3, and CP4 are (0, 0), (W, 0), (H, 0), and (W, H), respectively, where W and H are the width and height of the current block.

[0088] The movement information of each control point is obtained according to the following priority order.

[0089] For CP1, the check priority is B 2 ->B 3 ->A 2 is. B 2 is used if it is available. Otherwise, if B 3 is available, B 3 is used. If both B2 and B3 are unavailable, A 2 is used. If all three candidates are unavailable, the movement information of CP1 cannot be obtained.

[0090] For CP2, the check priority is B 1 ->B 0 is.

[0091] For CP3, the check priority is A 1 ->A 0 is.

[0092] For CP4, T is used.

[0093] Second, combinations of control points are used to form affine merge candidates.

[0094] The motion information of three control points is required to constitute a six-parameter affine candidate. The three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}). The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4} are converted into a six-parameter motion model represented by the upper left, upper right, and lower left control points.

[0095] The motion vectors of two control points are required to constitute a four-parameter affine candidate. The two control points can be selected from one of the following six combinations ({CP1, CP4}, {CP2, CP3}, {CP1, CP2}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4}). The combinations {CP1, CP4}, {CP2, CP3}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4} are converted into a four-parameter motion model represented by the upper left and upper right control points.

[0096] The combinations of the constituted affine candidates are inserted into the candidate list in the following order: {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}, {CP2, CP3}, {CP1, CP4}, {CP2, CP4}, {CP3, CP4}.

[0097] For the reference list X of the combination (X is 0 or 1), the reference index with the maximum utilization rate among the control points is selected as the reference index of the list X, and the motion vector indicating the differential reference picture is scaled.

[0098] After the candidate is derived, a complete pruning process is executed to check whether the same candidate is inserted into the list. If the same candidate exists, the derived candidate is discarded.

[0099] 3) Padding with Zero Motion Vectors If the number of candidates in the affine merge candidate list is less than 5, zero motion vectors with zero reference indices are inserted into the candidate list until the list is full.

[0100] [2.5.2 Example Affine Merge Mode] In some embodiments, the affine merge mode can be simplified as follows.

[0101] 1) The pruning process for inherited affine candidates is simplified by comparing coding units covering adjacent positions instead of comparing the affine candidates derived in VTM-2.0.1. Up to two inherited affine candidates are inserted into the affine merge list. The pruning process for the constructed affine candidates is removed entirely.

[0102] 2) The MV scaling operation in the constructed affine candidates is removed. If the reference indices of the control points are different, the constructed motion model is discarded.

[0103] 3) The number of constructed affine candidates is reduced from 10 to 6.

[0104] 4) In some embodiments, other merge candidates by sub-block prediction such as ATMVP are also put into the affine merge candidate list. In that case, the affine merge candidate list may be renamed by some other name such as a sub-block merge candidate list.

[0105] [2.6 Example Control Point MV Offsets for Affine Merge Mode] New affine merge candidates are generated based on the CPMV offset of the first affine merge candidate. When the first affine merge candidate enables a 4-parameter affine model, two CPMVs for each new affine merge candidate are derived by offsetting the two CPMVs of the first affine merge candidate; otherwise (when a 6-parameter affine model is enabled), three CPMVs for each new affine merge candidate are derived by offsetting the three CPMVs of the first affine merge candidate. In single prediction, the CPMV offset is applied to the CPMN of the first candidate. In dual prediction with lists 0 and 1 in the same direction, the CPMN offset is applied to the first candidate as follows: MV new(L0),i =MV old(L0) +MV offset(i) Equation (8) MV new(L1),i =MV old(L1) +MV offset(i) Equation (9)

[0106] In dual prediction with lists 0 and 1 in the reverse direction, the CPMV offset is applied to the first candidate as follows: MV new(L0),i =MV old(L0) +MV offset(i) Equation (10) MV new(L1),i =MV old(L1) -MV offset(i) Equation (11)

[0107] Various offset directions with various offset magnitudes can be used to generate new affine merge candidates. Two implementations were tested.

[0108] (1) Sixteen new affine merge candidates with eight different offset directions at two different offset magnitudes were the following offset pairs: Set of offsets = {(4, 0), (0, 4), (-4, 0), (0, -4), (-4, -4), (4, -4), (4, 4), (-4, 4), (8, 0), (0, 8), (-8, 0), (0, -8), (-8, -8), (8, -8), (8, 8), (-8, 8)} As shown by to , it is generated.

[0109] The affine merge list is increased up to 20 for this design. The total number of possible affine merge candidates is 31.

[0110] (2) Four affine merge candidates with four different offset directions for one offset size are from the following set of offsets: Set of offsets = {(4, 0), (0, 4), (-4, 0), (0, -4)} As shown by to , it is generated.

[0111] The affine merge list is kept at 5 as VTM2.0.1 does. Four temporal reconstruction affine merge candidates are removed so as not to change the number of possible affine merge candidates (e.g., 15 in total). Let the coordinates of CPMV1, CPMV2, CPMV3, and CPMV4 be (0, 0), (W, 0), (H, 0), and (W, H), respectively. Note that CPMV4 is the point derived from the temporal MV as shown in Figure 6. The removed candidates are the following four temporally related reconstructed affine merge candidates: {CP2, CP3, CP4}, {CP1, CP4}, {CP2, CP4}, {CP3, CP4}.

[0112] [2.7 Bandwidth Issues of Affine Motion Compensation] Since the current block is divided into 4×4 sub - blocks for the luma component and 2×2 sub - blocks for the two chroma components to perform motion compensation, the total bandwidth requirement is much higher than non - sub - block inter - prediction. To address the bandwidth issue, several approaches are proposed.

[0113] [2.7.1 Example 1] The 4×4 block is used as the sub-block size of the CU coded with one-directional affine, while the 8×4 / 4×8 block is used as the sub-block size of the CU coded with bi-directional prediction affine coding.

[0114] [2.7.2 Example 2] For the affine mode, the sub-block motion vectors of the affine CU are constrained to be within a predefined motion vector field. If the motion vector of the first (top left) sub-block is (v 0x ,v 0y ) and the second sub-block is (v ix ,v iy ), then the values of v ix and v iy show the following constraints: v ix ∈[v 0x -H,v 0x +H] Equation (12) v iy ∈[v 0y -V,v 0y +V] Equation (13)

[0115] If the motion vector of any sub-block exceeds the predefined motion vector field, the motion vector is clipped. An example of the concept of the constrained sub-block motion vector is given in FIG. 6.

[0116] If the memory is read per CU instead of per sub-block, the values of H and V are selected such that the worst memory bandwidth of the affine CU does not exceed that of the normal inter-MC of an 8×8 bi-prediction block. Note that the values of H and V can be adapted to the CU size and single prediction or bi-prediction.

[0117] [2.7.3 Example 3] To reduce the memory bandwidth requirement in affine prediction, each 8×8 block within a block is regarded as a basic unit. The MVs of all four 4×4 sub-blocks within an 8×8 block are constrained such that the maximum difference between the integer parts of the four 4×4 sub-blocks is not greater than 1 pixel. Thereby, the bandwidth is (8 + 7 + 1)×(8 + 7 + 1) / (8×8) = 4 samples / pixel.

[0118] In some cases, after the MVs of all sub-blocks within the current block are calculated by the affine model, the MVs of the sub-blocks containing control points are first replaced with the corresponding control point MVs. This means that the MVs of the top-left, top-right, and bottom-left sub-blocks are replaced by the top-left, top-right, and bottom-left control point MVs respectively. Then, for each 8×8 block within the current block, the MVs of all four 4×4 sub-blocks are clipped to ensure that the maximum difference between the integer parts of those four MVs is not greater than 1 pixel. Here, it should be noted that the sub-blocks containing control points (top-left, top-right, and bottom-left sub-blocks) use the corresponding control point MVs for the MV clipping process. During the clipping process, the MV of the top-right control point remains unchanged.

[0119] The clipping process applied to each 8×8 block is described as follows.

[0120] 1. The minimum and maximum values of the MV components, MVminx, MVminy, MVmaxx, MVmaxy, are first determined for each 8×8 block as follows: a) Obtain the minimum MV component among the four 4×4 sub-block MVs MVminx = min(MVx0, MVx1, MVx2, MVx3) MVminy = min(MVy0, MVy1, MVy2, MVy3) b) Use the integer parts of MVminx and MVminy as the minimum MV component MVminx = MVminx >> MV_precision << MV_precision MVminy = MVminy >> MV_precision << MV_precision c) The maximum MV component is calculated as follows: MVmaxx = MVminx + (2 << MV_precision) - 1 MVmaxy = MVminy + (2 << MV_precision) - 1 d) If the upper - right control point is in the current 8×8 block If MV1x > MVmaxx MVminx = (MV1x >> MV_precision << MV_precision) - (1 << MV_precision) MVmaxx = MVminx + (2 << MV_precision) - 1 If MV1y > MVmaxy, MVminy = (MV1y >> MV_precision << MV_precision) - (1 << MV_precision) MVmaxy = MVminy + (2 << MV_precision) - 1

[0121] 2. The MV components of each 4×4 block within this 8×8 block are clipped as follows: MVxi = max(MVminx, min(MVmaxx, MVxi)) MVyi = max(MVminy, min(MVmaxy, MVyi)) Here, (MVxi, MVyi) is the MV of the i - th sub - block within one 8×8 block, where i is 0, 1, 2, 3, (MV1x, MV1y) is the MV of the upper - right control point, and MV_precision is equal to 4 corresponding to 1 / 16 motion vector fractional precision. Since the difference between the integer parts of MVminx and MVmaxx (MVminy and MVmaxy) is 1 pixel, the maximum difference between the integer parts of the four 4×4 sub - block MVs is not greater than 1 pixel.

[0122] A similar method may also be used in some embodiments to handle the planar mode.

[0123] [2.7.4 Example 4] In some embodiments, a restriction to the affine mode for worst - case bandwidth reduction. To ensure that the worst - case bandwidth of the affine block is not worse than that of the INTER_4×8 / INTER_8×4 block or the INTER_9×9 block, the motion vector difference between the affine control points is used to determine whether the sub - block size of the affine block is 4×4 or 8×8.

[0124] <General Affine Restrictions for Worst - Case Bandwidth Reduction> Memory bandwidth reduction for the affine mode is controlled by restricting the motion vector difference (also referred to as the control point difference) between the affine control points. Generally, when the control point difference satisfies the following restrictions, the affine motion is using 4×4 sub - blocks (i.e., 4×4 affine mode). Otherwise, it is using 8×8 sub - blocks (8×8 affine mode). The restrictions for the 6 - parameter and 4 - parameter models are given as follows.

[0125] To derive the constraints for different block sizes (w×h), the motion vector difference of the control points is: Norm(v 1x -v 0x )=(v 1x -v 0x )×(128 / w) Norm(v 1y -v 0y )=(v 1y -v 0y )×(128 / w) Norm(v 2x -v 0x )=(v 2x -v 0x )×(128 / h) Norm(v 2y -v 0y )=(v 2y -v 0y )×(128 / h) Equation (14) is normalized as.

[0126] In the 4-parameter affine model, (v 2x -v 0x ) and (v 2y -v 0y ) are set as follows: (v 2x -v 0x ) = -(v 1y -v 0y ) (v 2y -v 0y ) = -(v 1x -v 0x ) Equation (15)

[0127] Therefore, the norms of (v 2x -v 0x ) and (v 2y -v 0y ) are: Norm(v 2x -v 0x ) = -Norm(v 1y -v 0y ) Norm(v 2y -v 0y ) = Norm(v 1x -v 0x ) Equation (16) is given by.

[0128] The limit to ensure the worst-case bandwidth is to achieve INTER_4×8 or INTER_8×4: |Norm(v 1x -v 0x ) + Norm(v 2x -V 0x ) + 128| + |Norm(v 1y -v 0y ) + Norm(v 2y -v 0y ) + 128| + |Norm(v 1x -v 0x ) - Norm(v 2x -v 0x )| + |Norm(v 1y -v 0y)-Norm(v 2y -v 0y )| <128×3.25 Equation (17) Here, the left side of Equation (17) represents the shrink or span level of the sub-affine block, while (3.25) indicates a pixel shift of 3.25.

[0129] The limitation to ensure the worst-case bandwidth is to achieve INTER_9×9: (4×Norm(v 1x -v 0x ) > -4×pel && +4×Norm(v 1x -v 0x ) < pel) && (4×Norm(v 1y -v 0y ) > -pel && 4×Norm(v 1y -v 0y ) < pel) && (4×Norm(v 2x -v 0x ) > -pel && 4×Norm(v 2x -v 0x ) < pel) && (4×Norm(v 2y -v 0y ) > -4×pel && 4×Norm(v 2y -v 0y ) < pel) && ((4×Norm(v 1x -v 0x ) + 4×Norm(v 2y -v 0y ) > -4×pel) && (4×Norm(v 1x -v 0x ) + 4×Norm(v 2x -v 0x ) < pel)) && ((4×Norm(v 1y -v 0y ) + 4×Norm(v 2x -v 0x ) > -4×pel) && (4×Norm(v 1y -v 0y) + 4×Norm(v 2y - v 0yx ) < pel)) Equation (18) Here, pel = 128×16 (128 and 16 are the normalization coefficient and the motion vector accuracy, respectively).

[0130] [2.8 Generalized Dual-Prediction Improvement] Some embodiments improve the trade-off between the gain and complexity for GBi and are adopted in BMS2.1. GBi is also called Bi-prediction with CU-level Weight (BCW). BMS2.1 GBi applies unequal weights to the predictors from L0 and L1 in the dual-prediction mode. In the inter-prediction mode, multiple weight pairs including the equal weight pair (1 / 2, 1 / 2) are evaluated based on Rate-Distortion Optimization (RDO), and the GBi index of the selected weight pair is notified to the decoder. In the merge mode, the GBi index is inherited from the adjacent CU. In BMS2.1 GBi, the predictor generation in the dual-prediction mode is shown in Equation (19): P GBi = (w 0 × P L0 + w 1 × P L1 + RoundingOffset GBi ) >> shiftNum GBi Equation (19)

[0131] Here, P GBi is the final predictor of GBi. w 0 and w 1 are the selected GBi weight pair and are applied to list 0 (L0) and list 1 (L1), respectively. RoundingOffset GBi and shiftNum GBi are used to normalize the final predictor in GBi. The supported w 1is set to {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}, and these five weights correspond to one pair of equal weights and four pairs of unequal weights. The blending gain, e.g., w 1 and w 0 sum is fixed to 1.0. Thus, the corresponding w 0 weight set is {5 / 4, 5 / 8, 1 / 2, 3 / 8, -1 / 4}. The weight pair selection is at the CU level.

[0132] For non-low-delay pictures, the weight set size is reduced from 5 to 3. At this time, the w 1 weight set is {3 / 8, 1 / 2, 5 / 8}, and the w 0 weight set is {5 / 8, 1 / 2, 3 / 8}. The reduction of the weight set size for non-low-delay pictures is applied to BMS2.1 GBi and all GBi tests in this contribution.

[0133] In some embodiments, the following changes are applied on top of the existing GBi design in BMS2.1 to further improve GBi performance.

[0134] [2.8.1 GBi Encoder Bug Fix] To reduce the GBi encoding time, in the current encoder design, the encoder stores the unidirectional motion vectors estimated from the GBi weights equal to 4 / 8 and reuses them for the unidirectional search of other GBi weights. This fast encoding method is applicable to both the translational motion model and the affine motion model. In VTM2.0, the 6-parameter affine model was adopted together with the 4-parameter affine model. The BMS2.1 encoder does not distinguish between the 4-parameter affine model and the 6-parameter affine model when the GBi weight is equal to 4 / 8 and it holds the unidirectional affine MV. As a result, the 4-parameter affine MV can be overwritten by the 6-parameter affine MV after encoding with the GBi weight 4 / 8. The stored 6-parameter affine MV may be used for the 4-parameter affine ME for other GBi weights, or the stored 4-parameter affine MV may be used for the 6-parameter affine ME. The proposed GBi encoder bug fix is to separate the 4-parameter and 6-parameter affine MV storage. The encoder stores these affine MVs based on the affine model type when the GBi weight is equal to 4 / 8 and reuses the corresponding affine MVs based on the affine model type of other GBi weights.

[0135] [2.8.2 CU Size Constraints for GBi] In this method, GBi is disabled for small CUs. In the inter prediction mode, bi-prediction is used and GBi is disabled without any signaling when the CU area is smaller than 128 luma samples.

[0136] [2.8.3 Merge Mode by GBi] According to the merge mode, the GBi index is not signaled. Instead, it is inherited from the adjacent block it is merged with. When the TMVP candidate is selected, GBi is turned off in this block.

[0137] [2.8.4 Affine Prediction by GBi] GBi is available when the current block is coded by affine prediction. For the affine inter mode, the GBi index is signaled. For the affine merge mode, the GBi index is inherited from the adjacent block it is merged with. When the configured affine model is selected, GBi is turned off for this block.

[0138] [2.9 Example Inter-Intra Prediction Mode (IIP)] According to the inter-intra prediction mode, also called inter and intra combined prediction (CIIP), the multi-hypothesis prediction combines one intra prediction and one merge indexing prediction. Such blocks are treated as special inter-coding blocks. In the merge CU, one flag is signaled for the merge mode to select the intra mode from the intra candidate list if the flag is true. For the luma component, the intra candidate list is derived from four intra prediction modes including the DC, planar, horizontal, and vertical modes, and the size of the intra candidate list can be 3 or 4 depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is removed from the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra list mode. One intra prediction mode selected by the intra mode index and one merge indexing prediction selected by the merge index are combined using a weighted average. For the chroma component, DM is always applied regardless of the extra signaling.

[0139] The weights for combining predictions are explained as follows. Equal weights are applied when the DC or planar mode is selected, or when the CB width or height is less than 4. For a CB where the CB width and height are 4 or more, when the horizontal / vertical mode is selected, one CB is first divided into 4 equal-area regions in the vertical / horizontal direction. Let i range from 1 to 4, and assuming (w_intra 1 , w_inter 1 ) = (6, 2), (w_intra 1 , w_inter 2 ) = (5, 3), (w_intra 3 , w_inter 3 ) = (3, 5), and (w_intra 4 , w_inter 4 ) = (2, 6), each weight set represented as (w_intra i , w_inter i ) will be applied to the corresponding region. (w_intra 1 , w_inter 1 ) is for the region closest to the reference sample, and (w_intra 4 , w_inter 4 ) is for the region farthest from the reference sample. Then, the combined prediction can be calculated by summing the two weighted predictions and right-shifting by 3 bits. Further, the intra prediction mode for the predictor's intra hypothesis can be saved for reference to the next adjacent CU.

[0140] Assume that the intra and inter prediction values are PIntra and PInter, and the weight coefficients are w_intra and w_inter respectively. The prediction value at position (x, y) is calculated as (PIntra(x, y) × w_intra(x, y) + PInter(x, y) × w_inter(x, y)) >> N. Here, w_intra(x, y) + w_inter(x, y) = 2 N .

[0141] <Signaling of Intra Prediction Mode in IIP Coding Block> When inter-intra mode is used, one of the four permitted intra prediction modes, DC, planar, horizontal, and vertical, is selected and signaled. The three Most Probable Mode(s) (MPM) are formed from the left and upper adjacent blocks. The intra prediction mode of an intra-coded adjacent block or an IIP-coded adjacent block is treated as one of the MPMs. When the intra prediction mode is not one of the four permitted intra prediction modes, it will be rounded to the vertical mode or the horizontal mode according to the angular difference. The adjacent block should be on the same CTU line as the current block.

[0142] Assume that the width and height of the current block are W and H. When W > 2×H or H > 2×W, only one of the three MPMs is available for use in inter-intra mode. Otherwise, all four valid intra prediction modes are available for use in inter-intra mode.

[0143] It should be noted that the intra prediction mode in inter-intra mode cannot be used to predict the intra prediction mode in a normal intra-coded block.

[0144] Inter-intra prediction is available only when W×H >= 64.

[0145] [2.10 Example Triangular Prediction Mode] The concept of Triangular Prediction Mode (TPM) is to introduce a new triangular partition for motion compensated prediction. As shown in FIGS. 7A - 7B, it divides the CU into two triangular prediction units either in the diagonal or anti - diagonal direction. Each triangular prediction unit within the CU is inter - predicted using its own single prediction motion vector and reference frame index derived from a single prediction candidate list. The adaptive weighting process is performed on the diagonal side after predicting the triangular prediction units. Then, the transform and quantization processes are applied to the entire CU. This mode is known to be applicable only to the skip and merge modes.

[0146] [2.10.1 Single Prediction Candidate List for TPM] The single prediction candidate list consists of five single prediction motion vector candidates. It is derived from seven adjacent blocks including five spatially adjacent blocks (1 to 5) and two temporally co - located blocks (6 to 7) as shown in FIG. 8. The motion vectors of the seven adjacent blocks are collected and placed in the single prediction candidate list according to the order of the single prediction motion vectors, the L0 motion vector of the bi - prediction motion vector, the L1 motion vector of the bi - prediction motion vector, and the averaged motion vector of the L0 and L1 motion vectors of the bi - prediction motion vector. If the number of candidates is less than 5, a zero motion vector is added to the list. The motion candidates added to this list are called TPM motion candidates.

[0147] More specifically, the following steps are included.

[0148] 1) Regardless of any pruning operation, obtain motion candidates from A 1 , B 1 , B 0 , A 0 , B 2 , Col and Col2 (corresponding to blocks 1 - 7 in FIG. 8).

[0149] 2) Set the variable numCurrMergeCand = 0.

[0150] 3) A 1 、B 1 、B 0 、A 0 、B 2 Each motion candidate is derived from Col and Col2. When numCurrMergeCand is less than 5, if the motion candidate is a uni-prediction (from either list 0 or list 1), then it increments numCurrMergeCand by 1 and is added to the merge list. Such an added motion candidate is called an "originally uni-predicted candidate". Complete pruning is applied.

[0151] 4) A 1 、B 1 、B 0 、A 0 、B 2 Each motion candidate is derived from Col and Col2. When numCurrMergeCand is less than 5, if the motion candidate is a bi-prediction, then the motion information from list 0 is added to the merge list (i.e., changed to be a uni-prediction from list 0), and numCurrMergeCand is incremented by only 1. Such an added candidate is called a "Truncated List0-predicted candidate". Complete pruning is applied.

[0152] 5) A 1 、B 1 、B 0 、A 0 、B 2, when each motion candidate is derived from Col and Col2 and numCurrMergeCand is less than 5, if the motion candidate is a dual prediction, the motion information from List 1 is added to the merge list (i.e., changed to be a single prediction from List 1), and numCurrMergeCand is incremented by only 1. Such added candidates are called "Truncated List1-predicted candidates". Complete pruning is applied.

[0153] 6)A 1 , B 1 , B 0 , A 0 , B 2 , when each motion candidate is derived from Col and Col2 and numCurrMergeCand is less than 5, if the motion candidate is a dual prediction, - When the slice quantization parameter (QP) of the List 0 reference picture is smaller than the slice QP of the List 1 reference picture, the motion information of List 1 is first scaled to the List 0 reference picture, and the average of two MVs (one from the original List 0 and the other from the scaled MV of List 1) is added to the merge list. This is the averaged single prediction from the List 0 motion candidate, and numCurrMergeCand is incremented by only 1. - Otherwise, the motion information of List 0 is first scaled to the List 1 reference picture, and the average of two MVs (one from the original List 1 and the other from the scaled MV of List 0) is added to the merge list. This is the averaged single prediction from the List 1 motion candidate, and numCurrMergeCand is incremented by only 1.

[0154] Complete pruning is applied.

[0155] 7) When numCurrMergeCand is less than 5, a zero motion vector candidate is added.

[0156] [2.11 Motion Vector Refinement (DMVR) on the Decoder Side in VVC] For DMVR in VVC, MVD mirroring between List 0 and List 1 is considered as shown in FIG. 13, and bilateral matching is performed to find the best MVD from several MVD candidates, for example, to refine the MV. The MVs of two reference pictures are represented by MVL0 (L0X, L0Y) and MVL1 (L1X, L1Y). The MVD represented by (MvdX, MvdY) for List 0 that can minimize the cost function (e.g., SAD) is defined as the best MVD. For the SAD function, it is defined as the SAD between the reference block of List 0 derived by the motion vector (L0X + MvdX, L0Y + MvdY) in the List 0 reference picture and the reference block of List 1 derived by the motion vector (L1X - MvdX, L1Y - MvdY) in the List 1 reference picture.

[0157] The motion vector refinement process may be repeated twice. In each repetition, at most six MVDs (with integer pel precision) are checked in two steps as shown in FIG. 14. In the first step, MVDs (0, 0), (-1, 0), (1, 0), (0, -1), (0, 1) are checked. In the second step, one of MVDs (-1, -1), (-1, 1), (1, -1), or (1, 1) is selected and may be further checked. Assume that the function Sad(x, y) returns the SAD value of MVD (x, y). The MVD represented by (MvdX, MvdY) checked in the second step is determined as follows: MvdX = -1; MvdY = -1; if (Sad(1, 0) < Sad(-1, 0)) MvdX = 1; if (Sad(0, 1) < Sad(0, -1)) MvdY = 1

[0158] In the first iteration, the starting point is the notified MV, and in the second iteration, the starting point is the notified MV plus the selected best MVD in the first iteration. DMVR is applied only when one reference picture is a preceding picture, the other reference picture is a following picture, and those two reference pictures are at the same picture order count distance from the current picture.

[0159] To further simplify the DMVR process, the following main features can be implemented in some embodiments.

[0160] 1. Early termination when the (0,0) position SAD between list 0 and list 1 is smaller than the threshold.

[0161] 2. Early termination when the SAD between list 0 and list 1 is zero at a certain position.

[0162] 3. DMVR block size: W×N >= 64 && H >= 8, where W and H are the width and height of the block.

[0163] 4. In the case of DMVR with CU size > 16×16, divide the CU into multiple 16×16 sub-blocks. If only the width or height of the CU is greater than 16, it is divided only in the vertical or horizontal direction.

[0164] 5. Reference block size (W + 7)×(H + 7) (for luma).

[0165] 6. 25-point SAD-based integer pixel search (e.g., (+-)2 narrowing search range, single stage).

[0166] 7. Bilinear interpolation-based DMVR.

[0167] 8. Sub - pixel refinement based on the "parametric error surface equation". This procedure is only executed when the minimum SAS cost is not equal to zero and the best MVD is (0, 0) in the last MV refinement iteration.

[0168] 9. Luma / Chroma MC with reference block padding (if necessary).

[0169] 10. Refined MVs used only for MC and TMVP.

[0170] [2.11.1 Use of DMVR] DMVR can be enabled when all of the following conditions are met: - The DMVR enabling flag (e.g., sps_dmvr_enabled_flag) in the SPS is equal to 1. - The TPM flag, the inter - affine flag, and the sub - block merge flag (either for ATMVP or affine merge), and the MMVD flag are all equal to 0. - The merge flag is equal to 1. - The current block is bi - predicted and the Picture Order Count (POC) distance between the current picture and the reference picture in list 1 is equal to the POC distance between the reference picture in list 0 and the current picture. - The height of the current CU is 8 or more. - The number of luma samples (CU width × height) is 64 or more.

[0171] [2.11.2 Sub - pixel refinement based on the "parametric error surface equation"] The method is summarized as follows.

[0172] 1. The parametric error surface fit is calculated only when the center position is the best cost position in a given iteration.

[0173] 2. The cost at the center position and the costs at positions (-1,0), (0,-1), (1,0), and (0,1) from the center are as follows E(x,y)=A(x - x 0 ) 2 +B(y - y 0 ) 2 +C is used to fit the two - dimensional radiation error surface equation of 0 ,y 0 ). Here, (x 0 ,y 0 ) corresponds to the position with the minimum cost, and C corresponds to the minimum cost value. By solving five equations for five unknowns, (x 0 ,y 0 ) is: x 0 =(E(-1,0)-E(1,0)) / (2E(-1,0)+E(1,0)-2E(0,0)) y 0 =(E(0,-1)-E(0,1)) / (2E(0,-1)+E(0,1)-2E(0,0)) and is calculated as. (x 0 ,y 0 ) can be calculated to any required sub - pixel accuracy by adjusting the precision at which the division is performed (e.g., how many bits of the quotient are calculated). For 1 / 16 pel accuracy, only 4 bits of the absolute value of the quotient need to be calculated. This helps in the fast shift - subtract - based implementation of the two divisions required per CU.

[0174] 3. The calculated (x 0 ,y 0 ) is added to the integer - distance refined MV to obtain the sub - pixel accurate refined delta MV.

[0175] [2.11.3 Required Reference Samples in DMVR] For a block of size W×H, assuming that the maximum allowable MVD value is ±offSet (e.g., 2 in VVC) and the filter size is filterSize (e.g., 8 for luma and 4 for chroma in VVC), (W + 2×offSet + filterSize - 1)×(H + 2×offSet + filterSize - 1) reference samples are required. To reduce the memory bandwidth, the central (W + filterSize - 1)×(H + filterSize - 1) reference samples are fetched, and the remaining pixels are generated by repeating the boundaries of the fetched samples. An example of an 8×8 block is shown in FIG. 15, where 15×15 reference samples are fetched and the boundaries of the fetched samples are repeated to generate a 17×17 region.

[0176] During motion vector refinement, bilinear motion compensation is performed using those reference samples. On the other hand, the final motion compensation is also performed using those reference samples.

[0177] [2.12 Bandwidth calculation for different block sizes] Based on the current 8-tap luma interpolation filter and 4-tap chroma interpolation filter, the memory bandwidth of each block unit (4:2:0 color format, one M×N luma block with two M / 2×N / 2 chroma blocks) is shown in Table 1 below.

Table 1

[0178] Similarly, based on the current 8-tap luma interpolation filter and 4-tap chroma interpolation filter, the memory bandwidth of each M×N luma block unit is tabulated in Table 2 below.

Table 2

[0179] Therefore, regardless of the color format, the bandwidth requirements for each block size in descending order are: 4×4Bi > 4×8Bi > 4×16Bi > 4×4Uni > 8×8Bi > 4×32Bi > 4×64Bi > 4×128Bi > 8×16Bi > 4×8Uni > 8×32Bi > ··· That is.

[0180] [2.13 Motion Vector Accuracy Problem in VTM-3.0] In VTM-3.0, the MV accuracy is 1 / 16 luma pixel in storage. When the MV is signaling, the finest accuracy is 1 / 4 luma pixel.

[0181] [3. Examples of Problems Solved by the Disclosed Embodiments] 1. The method for bandwidth control for affine accuracy is not sufficiently clear and should be more flexible.

[0182] 2. In the HEVC design, in the worst case of memory bandwidth requirements, even if a coding unit (CU) can be divided by an asymmetric accuracy mode (for example, one 16×16 divided into PUs of sizes equal to 4×16 and 12×16), it is 8×8 bi-prediction. In VVC, with the new QTBT partition structure, one CU can be set to 4×16 and bi-prediction can be enabled. A bi-predicted 4×16 CU requires a much higher memory bandwidth compared to a bi-predicted 8×8 CU. Methods for handling block sizes that require a higher bandwidth (such as 4×16 or 16×4) are not known.

[0183] 3. New coding tools such as GBi introduce more line buffer problems.

[0184] 4. The inter-intra mode requires more memory and logic to convey the intra prediction mode used in the inter-coded block.

[0185] 1 / 16 luma pixel MV accuracy requires higher memory storage.

[0186] 6. To interpolate four 4×4 blocks within one 8×8 block, it is necessary to fetch (8 + 7 + 1) × (8 + 7 + 1) reference pixels, which requires approximately 14% more pixels compared to non-affine / non-planar mode 8×8 blocks.

[0187] 7. The averaging operations in intra and inter composite prediction should be aligned with other coding tools, such as weighted prediction, local luminance compensation, OBMC, and triangular prediction, and the offset is added before shifting.

[0188] [4. Examples of Embodiments] The techniques disclosed herein can reduce the bandwidth and line buffer required by affine prediction and other new coding tools.

[0189] The following description should be regarded as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, embodiments can be combined in any way.

[0190] In the following discussion, the width and height of the currently CU coded affinely are w and h, respectively. The interpolation filter taps (in motion compensation) are N (e.g., 8, 6, 4, or 2), and it is assumed that the current block size is W×H.

[0191] <Bandwidth Control for Affine Prediction> Example 1: If the motion vector of sub-block SB within an affinely coded block is MV SB (represented as (MVx, MVy)), then MV SB can be within a specific range with respect to the representative motion vector MV’(MV’x, MV’y).

[0192] In some embodiments, MVx >= MV’x - DH0 and MVx <= MV’x + DH1, and MVy >= MV’y - DV0 and MVy <= MV’y + DV1, where MV’ = (MV’x, MV’y). In some embodiments, DH0 may or may not be equal to DH1, and DV0 may or may not be equal to DV1. In some embodiments, DH0 may or may not be equal to DV0, and DH1 may or may not be equal to DV1. In some embodiments, DH0 may not be equal to DH1, and DV0 may not be equal to DV1. In some embodiments, DH0, DH1, DV0, and DV1 may be signaled from the encoder to the decoder, for example, in the VPS / SPS / PPS / slice header / tile group header / tile / CTU / CU / PU. In some embodiments, DH0, DH1, DV0, and DV1 may be specified to be different for different standard profiles / levels / tiers. In some embodiments, DH0, DH1, DV0, and DV1 may depend on the width and height of the current block. In some embodiments, DH0, DH1, DV0, and DV1 may depend on whether the current block is uni-predicted or bi-predicted. In some embodiments, DH1, DH1, DV0, and DV1 may depend on the position of the sub-block SB. In some embodiments, DH0, DH1, DV0, and DV1 may depend on the method of obtaining MV’.

[0193] In some embodiments, MV’ can be one CPMV such as MV0, MV1, or MV2.

[0194] In some embodiments, MV’ can be the MV used for MC for one of the corner sub-blocks, such as MV0’, MV1’, or MV2’ in FIG. 3.

[0195] In some embodiments, MV’ can be the MV derived using the affine model of the current block for any position either inside or outside the current block. For example, it may be derived for the center position of the current block (e.g., x = w / 2 and y = h / 2).

[0196] In some embodiments, MV’ can be the MV used for MC for any sub-block of the current block, such as one of the central sub-blocks (C0, C1, C2, or C3 shown in FIG. 3).

[0197] In some embodiments, MV SB should be clipped to a valid region when it does not satisfy the constraints. In some embodiments, the clipped MV SB is saved in the MV buffer and will subsequently be used to predict the MV of the coded block. In some embodiments, the MV SB before clipping is saved in the MV buffer. SB is saved in the MV buffer.

[0198] In some embodiments, MV SB when it does not satisfy the constraints, the bitstream is considered non-compliant (invalid). In one example, the MV SB may be specified by the standard as something that must or should satisfy the constraints. This constraint should be followed by any compliant encoder; otherwise, the encoder is considered non-compliant.

[0199] In some embodiments, MV SB and MV’ can be represented by signaling MV precision (e.g., 1 / 4 pixel precision). In some embodiments, MV SB and MV’ can be represented by storage MV precision (e.g., 1 / 16 precision). In some embodiments, MV SBAnd MV' may be rounded to a precision different from the signaling or storage precision (e.g., integer precision).

[0200] Example 2: For an affine-coded block, each M×N (e.g., 8×4, 4×8, or 8×8) block within the block is regarded as a basic unit. The MVs of all 4×4 sub-blocks within M×N are constrained such that the maximum difference between the integer parts of the four 4×4 sub-block MVs is not greater than K pixels.

[0201] In some embodiments, whether and how to apply this constraint depends on whether the current block applies dual prediction or single prediction. For example, the constraint is applied only to dual prediction and not to single prediction. As another example, M, N, and K are different for dual prediction and single prediction.

[0202] In some embodiments, M, N, and K may depend on the width and height of the current block.

[0203] In some embodiments, whether to apply the constraint may be notified from the encoder to the decoder, e.g., in the VPS / SPS / PPS / slice header / tile group header / tile / CTU / CU / PU. For example, an on / off flag is notified to indicate whether to apply the constraint. As another example, M, N, and K are notified.

[0204] In some embodiments, M, N, and K may be specified to be different for different standard profiles / levels / tiers.

[0205] Example 3: The width and height of the sub-block may be calculated differently for different affine-coded blocks.

[0206] In some embodiments, the calculation method varies for each affine-coded block by single prediction and dual prediction. In one example, the sub-block size is fixed for blocks by single prediction (e.g., 4×4, 4×8, or 8×4). In other examples, the sub-block size is calculated for blocks by dual prediction. In this case, the sub-block size may vary for each of the two different, dually predicted affine blocks.

[0207] In some embodiments, for dually predicted affine blocks, the width and / or height of the sub-block from reference list 0 and the width and / or height of the sub-block from reference list 1 may be different. In one example, the width and height of the sub-block from reference list 0 are Wsb0 and Hsb0, respectively, and the width and height of the sub-block from reference list 1 are Wsb1 and Hsb1, respectively. In that case, the final width and height of the sub-block for both reference list 0 and reference list 1 are calculated as Max(Wsb0, Wsb1) and Max(Hsb0, Hsb1), respectively.

[0208] In some embodiments, the calculated width and height of the sub-block are applied only to the luma component. For the chroma component, it is always fixed, e.g., a 4×4 chroma sub-block corresponding to an 8×8 luma block in a 4:2:0 color format.

[0209] In some embodiments, MVx - MV’x and MVy - MV’y are calculated to determine the width and height of the sub-block. (MVx, MVy) and (MV’x, MV’y) are defined in Example 1.

[0210] In some embodiments, the MVs involved in the calculation may be represented by a signaling MV accuracy (e.g., 1 / 4 pixel accuracy). In one example, these MVs may be represented by a storage MV accuracy (e.g., 1 / 16 accuracy). As another example, these MVs may be rounded to an accuracy different from the signaling or storage accuracy (e.g., integer accuracy).

[0211] In some embodiments, the threshold used in the calculation to determine the width and height of the sub-block may be notified from the encoder to the decoder, for example, in the VPS / SPS / PPS / slice header / tile group header / tile / CTU / CU / PU.

[0212] In some embodiments, the threshold used in the calculation to determine the width and height of the sub-block may vary for different standard profiles / levels / tiers.

[0213] Example 4: To interpolate a W1×H1 sub-block within one W2×H2 sub-block / block, a (W2+N-1-PW)×(H2+N-1-PH) block is first fetched, and then the pixel padding method (e.g., boundary pixel repetition method) described in Example 6 is applied to generate a larger block. The larger block is then used to interpolate the W1×H1 sub-block. For example, W2 = H2 = 8, W1 = H1 = 4, and PW = PH = 0.

[0214] In some embodiments, the integer part of the MV of any W1×H1 sub-block may be used to fetch the entire W2×H2 sub-block / block, and different boundary pixel repetition methods may be required accordingly. For example, when the maximum difference between the integer parts of the MVs of all W1×H1 sub-blocks is not greater than 1 pixel, the integer part of the MV of the upper-left W1×H1 sub-block is used to fetch the entire W2×H2 sub-block / block. The right and lower boundaries of the reference block are repeated once. As another example, when the maximum difference between the integer parts of the MVs of all W1×H1 sub-blocks is not greater than 1 pixel, the integer part of the MV of the lower-right W1×H1 sub-block is used to fetch the entire W2×H2 sub-block / block. The left and upper boundaries of the reference block are repeated once.

[0215] In some embodiments, the MV of any W1×H1 sub-block may first be modified and then used to fetch the entire W2×H2 sub-block / block, and different boundary pixel repetition methods may be required accordingly. For example, when the maximum difference between the integer parts of the MVs of all W1×H1 sub-blocks is not greater than 2 pixels, the integer part of the MV of the upper-left W1×H1 sub-block may be incremented by only (1,1) (where 1 means a distance of 1 integer pixel), and then used to fetch the entire W2×H2 sub-block / block. In this case, the left, right, upper, and lower boundaries of the reference block are repeated once. As another example, when the maximum difference between the integer parts of the MVs of all W1×H1 sub-blocks is not greater than 2 pixels, the integer part of the MV of the lower-right W1×H1 sub-block may be incremented by only (-1,-1) (where 1 means a distance of 1 integer pixel), and then used to fetch the entire W2×H2 sub-block / block. In this case, the left, right, upper, and lower boundaries of the reference block are repeated once.

[0216] <Bandwidth control for specific block sizes> Example 5: Bi-prediction is not allowed when the w and h of the current block satisfy one or more of the following conditions.

[0217] A. w is equal to T1 and h is equal to T2, or h is equal to T1 and w is equal to T2. In one example, T1 = 4 and T2 = 16.

[0218] B. w is equal to T1 and h is not greater than T2, or h is equal to T1 and w is not greater than T2. In one example, T1 = 4 and T2 = 16.

[0219] C. w is not greater than T1 and h is not greater than T2, or h is not greater than T1 and w is not greater than T2. In one example, T1 = 8 and T2 = 8. In another example, T1 == 8, T2 == 4. In yet another example, T1 == 4 and T2 == 4.

[0220] In some embodiments, dual prediction may be disabled for 4×8 blocks. In some embodiments, dual prediction may be disabled for 8×4 blocks. In some embodiments, dual prediction may be disabled for 4×16 blocks. In some embodiments, dual prediction may be disabled for 16×4 blocks. In some embodiments, dual prediction may be disabled for 4×8 and 8×4 blocks. In some embodiments, dual prediction may be disabled for 4×16 and 16×4 blocks. In some embodiments, dual prediction may be disabled for 4×8 and 16×4 blocks. In some embodiments, dual prediction may be disabled for 4×16 and 8×4 blocks. In some embodiments, dual prediction may be disabled for 4×N blocks, e.g., where N <= 16. In some embodiments, dual prediction may be disabled for N×4 blocks, e.g., where N <= 16. In some embodiments, dual prediction may be disabled for 8×N blocks, e.g., where N <= 16. In some embodiments, dual prediction may be disabled for N×8 blocks, e.g., where N <= 16. In some embodiments, dual prediction may be disabled for 4×8, 8×4, and 4×16 blocks. In some embodiments, dual prediction may be disabled for 4×8, 8×4, and 16×4 blocks. In some embodiments, dual prediction may be disabled for 8×4, 4×16, and 16×4 blocks. In some embodiments, dual prediction may be disabled for 4×8, 8×4, 4×16, and 16×4 blocks.

[0221] In some embodiments, the block sizes disclosed herein may refer to one color component, such as the luma component, and the determination as to whether dual prediction is disabled may be applied to all color components. For example, if dual prediction is disabled according to the block size of the luma component of a block, dual prediction is also disabled for the corresponding blocks of other color components. In some embodiments, the block sizes disclosed herein may refer to one color component, such as the luma component, and the determination as to whether dual prediction is disabled may be applied only to that color component.

[0222] In some embodiments, when dual prediction is disabled for a block and the selected merge candidate is dual predicted, only one MV from reference list 0 or reference list 1 of that merge candidate is assigned to that block.

[0223] In some embodiments, the triangular prediction mode (TPM) is not allowed for a block when dual prediction is disabled for that block.

[0224] In some embodiments, the method of notifying the prediction direction (single prediction from list 0 / 1, dual prediction) may depend on the block size. In one example, an indication of single prediction from list 0 / 1 may be notified when 1) the block width × block height < 62, or 2) the block width × block height = 64 but the width is not equal to the height. As another example, an indication of single prediction or dual prediction from list 0 / 1 may be notified when 1) the block width × block height > 64, or 2) the block width × block height = 64 and the width is equal to the height.

[0225] In some embodiments, both single prediction and dual prediction may be disabled for 4×4 blocks. In some embodiments, it may be disabled for affine-coded blocks. Alternatively, it may be disabled for non-affine-coded blocks. In some embodiments, the indication of quadtree splitting for 8×8 blocks, binary tree splitting for 8×4 or 4×8 blocks, and ternary tree splitting for 4×16 or 16×4 blocks may be skipped. In some embodiments, 4×4 blocks should be coded as intra blocks. In some embodiments, the MV of 4×4 blocks should be at integer precision. For example, the IMV flag for 4×4 blocks should be 1. As another example, the MV of 4×4 blocks should be rounded to integer precision.

[0226] In some embodiments, dual prediction is permitted. However, if the interpolation filter tap is N, instead of fetching (W+N-1)×(H+N-1) reference pixels, only (W+N-1-PW)×(H+N-1-PH) reference pixels are fetched. On the other hand, pixels at the reference block boundaries (top, left, bottom, and right boundaries) are repeated to generate a (W+N-1)×(H+N-1) block as shown in FIG. 9 for use in the final interpolation. In some embodiments, PH is zero and only the left or / and right boundaries are repeated. In some embodiments, PW is zero and only the top or / and bottom boundaries are repeated. In some embodiments, both PW and PH are greater than zero, and first the left or / and right boundaries are repeated, and then the top or / and bottom boundaries are repeated. In some embodiments, both PW and PH are greater than zero, and first the top or / and bottom boundaries are repeated, and then the left or / and right boundaries are repeated. In some embodiments, the left boundary is repeated M1 times and the right boundary is repeated PW-M1 times. In some embodiments, the top boundary is repeated M2 times and the bottom boundary is repeated PH-M2 times. In some embodiments, such boundary pixel repetition methods may be applied to some or all of the reference pixels. In some embodiments, PW and PH may be different for different color components such as Y, Cb, and Cr.

[0227] FIG. 9 shows an example of repeated boundary pixels of a reference block before interpolation.

[0228] Example 6: In some embodiments, (W+N-1-PW)×(H+N-1-PH) reference pixels may be fetched for motion compensation of a W×H block (instead of (W+N-1)×(H+N-1) reference pixels). Samples that are outside the range (W+N-1-PW)×(H+N-1-PH) but within (W+N-1)×(H+N-1) are padded to perform an interpolation process. In one padding method, pixels at the reference block boundaries (top, left, bottom, and right boundaries) are repeated to generate a (W+N-1)×(H+N-1) block as shown in FIG. 11, which is used for the final interpolation.

[0229] In some embodiments, PH is zero and only the left or / and right boundary is repeated.

[0230] In some embodiments, PW is zero and only the top or / and bottom boundary is repeated.

[0231] In some embodiments, both PW and PH are greater than zero. First, the left or / and right boundary is repeated, and then the top or / and bottom boundary is repeated.

[0232] In some embodiments, both PW and PH are greater than zero. First, the top or / and bottom boundary is repeated, and then the left or / and right boundary is repeated.

[0233] In some embodiments, the left boundary is repeated M1 times and the right boundary is repeated PW-M1 times.

[0234] In some embodiments, the top boundary is repeated M2 times and the bottom boundary is repeated PH-M2 times.

[0235] In some embodiments, such boundary pixel repetition methods may be applied to some or all of the reference pixels.

[0236] In some embodiments, PW and PH may be different for different color components such as Y, Cb, and Cr.

[0237] In some embodiments, PW and PH may be different for different block sizes or shapes.

[0238] In some embodiments, PW and PH may be different for single prediction and dual prediction.

[0239] In some embodiments, padding may not be performed in affine mode.

[0240] In some embodiments, samples that are outside the range (W + N - 1 - PW) × (H + N - 1 - PH) but within (W + N - 1) × (H + N - 1) are set to be a single value. In some embodiments, the single value is 1 << (BD - 1), where BD is the bit depth of the sample, for example, 8 or 10. In some embodiments, the single value is notified from the encoder to the decoder in the VPS / SPS / PPS / slice header / tile group header / tile / CTU / CU / PU. In some embodiments, the single value is derived from samples within the range (W + N - 1 - PW) × (H + N - 1 - PH).

[0241] Example 7: Instead of fetching (W + filterSize - 1) × (H + filterSize - 1) samples in DMVR, (W + filterSize - 1 - PW) × (H + filterSize - 1 - PH) reference samples may be fetched, and all other required samples may be generated by repeating the boundaries of the fetched reference samples, where PW >= 0 and PH >= 0.

[0242] In some embodiments, the method proposed in Example 6 may be used to pad samples that are not fetched.

[0243] In some embodiments, in the final motion compensation of DMVR, padding may not be performed again.

[0244] In some embodiments, whether to apply the above method may depend on the block size.

[0245] Example 8: The signaling method of inter_pred_odc may depend on whether w and h satisfy the conditions of Example 5. One example is shown in Table 3 below.

Table 3

[0246] Another example is shown in Table 4 below.

Table 4

[0247] Yet another example is shown in Table 5 below.

Table 5

[0248] Example 9: The merge candidate list construction process may depend on whether w and h satisfy the conditions of Example 4. The following embodiments will describe the case where w and h satisfy the conditions.

[0249] In some embodiments, when one merge candidate uses bi-prediction, only the prediction from reference list 0 is retained, and the merge candidate is treated as a uni-prediction referring to reference list 0.

[0250] In some embodiments, when one merge candidate uses bi-prediction, only the prediction from reference list 1 is retained, and the merge candidate is treated as a uni-prediction referring to reference list 1.

[0251] In some embodiments, when one merge candidate uses dual prediction, the candidate is treated as unavailable. That is, such a merge candidate is removed from the merge list.

[0252] In some embodiments, a merge candidate list construction process for the triangular prediction mode is used instead.

[0253] Example 10: The coding tree splitting process may depend on whether the width and height of the child CU after splitting satisfy the conditions of Example 5.

[0254] In some embodiments, when the width and height of the child CU after splitting satisfy the conditions of Example 5, splitting is not permitted. In some embodiments, the signaling of the coding tree splitting may depend on whether one type of splitting is permitted. In one example, when one type of splitting is not permitted, the codeword representing the splitting is deleted.

[0255] Example 11: The signaling of the skip flag or / and the Intra Block Copy (IBC) flag may depend on whether the width and / or height of the block satisfy certain conditions (e.g., the conditions described in Example 5).

[0256] In some embodiments, the condition is that the luma block contains no more samples than X. For example, X = 16.

[0257] In some embodiments, the condition is that the luma block contains X samples. For example, X = 16.

[0258] In some embodiments, the condition is that both the width and height of the luma block are equal to X. For example, X = 4.

[0259] In some embodiments, when one or some of the above conditions are met, the inter-mode and / or the IBC mode are not permitted for such blocks.

[0260] In some embodiments, when the inter-mode is not permitted for a block, the skip flag may not be notified thereof. Alternatively, further, the skip flag may be presumed to be false.

[0261] In some embodiments, when the inter-mode and the IBC mode are not permitted for a block, the skip flag may not be notified thereof and may be implicitly derived as false (e.g., it is derived that the block is coded in a non-skip mode).

[0262] In some embodiments, when the inter-mode is not permitted for a block but the IBC mode is permitted for that block, the skip flag may still be notified. In some embodiments, the IBC flag may not be notified when the block is coded in skip mode, and the IBC flag may be implicitly derived as true (e.g., it is derived that the block is coded in IBC mode).

[0263] Example 12: The signaling of the prediction mode may depend on whether the width and / or height of the block satisfy certain conditions (e.g., the conditions described in Example 5).

[0264] In some embodiments, the condition is that the luma block contains no more than X samples. For example, X = 16.

[0265] In some embodiments, the condition is that the luma block contains X samples. For example, X = 16.

[0266] In some embodiments, the condition is that both the width and height of the rum block are equal to X. For example, X = 4.

[0267] In some embodiments, if one or some of the above conditions apply, the intermode and / or IBC mode are not permitted for such blocks.

[0268] In some embodiments, the signaling of the indication of a particular mode may be skipped.

[0269] In some embodiments, if the intermode and IBC mode are not permitted for a block, the signaling of the indication of the inter and IBC modes is skipped, and the remaining permitted modes (e.g., whether it is the intramode or the pallet mode) may still be notified.

[0270] In some embodiments, if the intermode and IBC mode are not permitted for a block, the prediction mode may not be notified. Alternatively, further, the prediction mode may be implicitly derived as being the intramode.

[0271] In some embodiments, if the intermode is not permitted for a block, the signaling of the indication of the intermode is skipped, and the remaining permitted modes (e.g., whether it is the intramode or the IBC mode) may still be notified. Alternatively, the remaining permitted modes (e.g., whether it is the intramode or the IBC mode or the pallet mode) may still be notified.

[0272] In some embodiments, if the intermode is not permitted for a block, but the IBC mode and the intramode are permitted for it, the IBC flag may be notified to indicate whether the block is coded in the IBC mode. Alternatively, further, the prediction mode may not be notified.

[0273] Example 13: The signaling of the triangle mode may depend on whether the width and / or height of the block satisfy certain conditions (e.g., the conditions described in FIG. 5).

[0274] In some embodiments, the condition is that the luma block size is one of several specific sizes. For example, the specific sizes may include 4×16 or / and 16×4.

[0275] In some embodiments, when the above conditions are met, the triangle mode may not be permitted, the flag indicating whether the current block is coded in triangle mode may not be signaled, and may be derived as false.

[0276] Example 14: The signaling of the inter-prediction direction may depend on whether the width and / or height of the block satisfy certain conditions (e.g., the conditions described in FIG. 5).

[0277] In some embodiments, the condition is that the luma block size is one of several specific sizes. For example, the specific sizes may include 8×4 or / and 4×8 or / and 4×16 or / and 16×4.

[0278] In some embodiments, when the above conditions are met, the block may simply be single-predicted, the flag indicating whether the current block is dual-predicted may not be signaled, and may be derived as false.

[0279] Example 15: The signaling of the SMVD (symmetric MVD) flag may depend on whether the width and / or height of the block satisfy certain conditions (e.g., the conditions described in Example 5).

[0280] In some embodiments, the condition is that the luma block size is one of several specific sizes. In some embodiments, the condition is defined as whether the block size has no more than 32 samples. In some embodiments, the condition is defined as whether the block size is 4×8 or 8×4. In some embodiments, the condition is defined as whether the block size is 4×4, 4×8 or 8×4. In some embodiments, the specific size may include 8×4 or / and 4×8 or / and 4×16 or / and 16×4.

[0281] In some embodiments, when a particular condition is met, an indication of the use of SMVD (e.g., an SMVD flag) may not be signaled and may be derived as false. For example, the block may be set when singly predicted.

[0282] In some embodiments, when a particular condition is met, an indication of the use of SMVD (e.g., an SMVD flag) may still be signaled, but only the motion information in list 0 or list 1 may be utilized in the motion compensation process.

[0283] Example 16: A motion vector or block vector (such as a motion vector derived in the regular merge mode, ATMVP mode, MMVD merge mode, MMVD skip mode, etc. used for IBC) may be changed depending on whether the width and / or height of the block satisfies a particular condition.

[0284] In some embodiments, the condition is that the luma block is one of several specific sizes. For example, the specific size may include 8×4 or / and 4×8 or / and 4×16 or / and 16×4.

[0285] In some embodiments, when the above conditions are met, the block motion vector or block vector may be changed to a unidirectional motion vector if the derived motion information is bidirectional (e.g., having some offsets and inherited from adjacent blocks). Such a process is called a conversion process, and the final unidirectional motion vector is referred to as the "converted unidirectional" motion vector. In some embodiments, the motion information of reference picture list X (e.g., X is 0 or 1) may be retained, and the motion information of list Y (Y is 1 - X) may be discarded. In some embodiments, the motion information of reference picture list X (e.g., X is 0 or 1) and that of list Y (Y is 1 - X) may be jointly used to derive new motion candidate points for list X. In one example, the motion vector of the new motion candidate may be the averaged motion vector of the two reference picture lists. As another example, the motion information of list Y may first be scaled to list X. Then, the motion vector of the new motion candidate may be the averaged motion vector of the two reference picture lists. In some embodiments, the motion vector in prediction direction X may not be used (e.g., the motion vector in prediction direction X is changed to (0, 0) and the reference index in prediction direction X is changed to -1), and the prediction direction may be changed to 1 - X (X = 0 or 1). In some embodiments, the converted unidirectional motion vector may be used to update the HMVP lookup table. In some embodiments, the derived bidirectional motion information, e.g., the bidirectional MV before being converted to a unidirectional MV, may be used to update the HMVP lookup table. In some embodiments, the converted unidirectional motion vector may be saved and then used for motion prediction of the next coded block, TMVP, deblocking, etc. In some embodiments, the derived bidirectional motion information, e.g., the bidirectional MV before being converted to a unidirectional MV, may be retained and then used for motion prediction of the next coded block, TMVP, deblocking, etc.In some embodiments, the converted unidirectional motion vectors may be used for motion refinement. In some embodiments, the derived bidirectional motion information may be used for motion refinement and / or sample refinement, e.g., by an optical flow method. In some embodiments, the predicted block generated according to the derived bidirectional motion information may be refined first, and then only one predicted block may be utilized to derive the final prediction and / or reconstruction block of one block.

[0286] In some embodiments, when certain conditions are met, the (bi-predicted) motion vectors may be converted to unidirectional motion vectors before being used as basic merge candidates in MMVD.

[0287] In some embodiments, when certain conditions are met (e.g., the block size satisfies the conditions specified in Example 5 above), the (bi-predicted) motion vectors may be converted to unidirectional motion vectors before being inserted into the merge list.

[0288] In some embodiments, the converted unidirectional motion vectors may be from reference list 0 only. In some embodiments, when the current slice / tile group / picture is bi-predicted, the converted unidirectional motion vectors may be from reference list 0 or list 1. In some embodiments, when the current slice / tile group / picture is bi-predicted, the converted unidirectional motion vectors from reference list 0 and list 1 may be interleaved in the merge list and / or the MMVD-based merge candidate list.

[0289] In some embodiments, the method of converting motion information into unidirectional motion vectors may depend on the reference picture. In some embodiments, when all the reference pictures of one video data unit (e.g., tile / tile group) are past pictures in the display order, the motion information in List 1 may be used. In some embodiments, when at least one of the reference pictures of one video data unit (e.g., tile / tile group) is a past picture and at least one is a future picture in the display order, the motion information in List 0 may be used. In some embodiments, the method of converting motion information into unidirectional motion vectors may depend on the low latency check flag.

[0290] In some embodiments, the conversion process may be called immediately before the motion compensation process. In some embodiments, the conversion process may be called immediately after the motion candidate list (e.g., merge list) construction process. In some embodiments, the conversion process may be called before calling the additional MVD process in the MMVD process. That is, the additional MVD process follows a design of single prediction rather than bi-prediction. In some embodiments, the conversion process may be called before calling the sample refinement process in the PROF process. That is, the sample refinement process follows a design of single prediction rather than bi-prediction. In some embodiments, the conversion process may be called before calling the BIO (also known as BDOF) process. That is, in some cases, BIO may be invalidated because it has been converted to single prediction. In some embodiments, the conversion process may be called before calling the DMVR process. That is, in some cases, DMVR may be invalidated because it has been converted to single prediction.

[0291] Example 17: In some embodiments, the method of generating a motion candidate list may depend on the block size, for example, as described in Example 5 above.

[0292] In some embodiments, for a particular block size, motion candidates derived from spatial blocks and / or temporal blocks and / or HMVP and / or other types of motion candidates may be restricted to be uni-predicted.

[0293] In some embodiments, for a particular block size, if one motion candidate derived from spatial blocks and / or temporal blocks and / or HMVP and / or other types of motion candidates is bi-predicted, it may be converted to be uni-predicted when it is first added to the candidate list.

[0294] Example 18: Whether a shared merge list is permitted may depend on the encoding mode.

[0295] In some embodiments, the shared merge list may not be permitted for blocks coded in the regular merge mode, and may be permitted for blocks coded in the ICB mode.

[0296] In some embodiments, when one block split from a parent shared node is coded in the regular merge mode, the update of the HMVP table may be invalidated after encoding / decoding the block.

[0297] Example 19: In the above examples disclosed, the block size / width / height of the luma block may also be changed to the block size / width / height of a chroma block such as Cb, Cr, or G / B / R.

[0298] <Line buffer reduction for GBi mode> Example 20: Whether a GBi weighted index can be inherited from an adjacent block or predicted (including CABC context selection) depends on the position of the current block.

[0299] In some embodiments, the GBi weighted index cannot be inherited or predicted from adjacent blocks that are not in the same coding tree unit (CTU, also known as the largest coding unit LCU) as the current block.

[0300] In some embodiments, the GBi weighted index cannot be inherited or predicted from adjacent blocks that are not in the same CTU line or CTU row as the current block.

[0301] In some embodiments, the GBi weighted index cannot be inherited or predicted from adjacent blocks that are not in the same M×N region as the current block. For example, M = N = 64. In this case, a tile / slice / picture is divided into a plurality of non-overlapping M×N regions.

[0302] In some embodiments, the GBi weighted index cannot be inherited or predicted from adjacent blocks that are not in the same M×N region line or M×N region row as the current block. For example, M = N = 64. CTU lines / rows and region lines / rows are shown in FIG. 10.

[0303] In some embodiments, assuming that the upper left corner (or other position) of the current block is (x, y) and the upper left corner (or other position) of the adjacent block is (x’, y’), it cannot be inherited or predicted from the adjacent block if the following conditions are satisfied: (1) x / M!= x’ / M. For example, M = 128 or 64. (2) y / N!= y’ / N. For example, N = 128 or 64. (3) ((x / M!= x’ / M) && (y / N!= y’ / N)). For example, M = N = 128 or M = N = 64. (4) ((x / M!= x’ / M) || (y / N!= y’ / N)). For example, M = N = 128 or M = N = 64. (5)x >> M != x' >> M. For example, M = 7 or 6. (6)y >> N != y' >> N. For example, N = 7 or 6. (7)((x >> M != x' >> M) && (y >> N != y' >> N)). For example, M = N = 7 or M = N = 6. (8)((x >> M != x' >> M) || (y >> N != y' >> N)). For example, M = N = 7 or M = N = 6.

[0304] In some embodiments, the flag is signaled in the PPS or slice header or tile group header or tile to indicate whether GBi is applicable in the picture / slice / tile group / tile. In some embodiments, whether CBi is used and how GBi is used (e.g., several candidate weights and weight values) may be derived for the picture / slice / tile group / tile. In some embodiments, the derivation may depend on information such as QP, temporal layer, POC distance, etc.

[0305] FIG. 10 shows an example of CTU (region) lines. The shared CTU (region) is on one CUT (region) line, and the non-shared CTU (region) is on other CUT (region) lines.

[0306] <Simplification of Inter-Intra Prediction (IIP)> Example 21: The coding of the intra prediction mode in an IIP-coded block is performed independently of the intra prediction mode of the adjacent IIP-coded blocks.

[0307] In some embodiments, only the intra prediction mode of the intra-coded blocks can be used in the coding of the intra prediction mode of the IIP-coded blocks, for example, during the MPM list construction process.

[0308] In some embodiments, the intra prediction mode in an IIP-coded block is coded without any mode prediction from adjacent blocks.

[0309] Example 22: The intra prediction mode of an IIP-coded block may have a lower priority than that of an intra-coded block when they are both used to code the intra prediction mode of newly IIP-coded blocks.

[0310] In some embodiments, when deriving the MPM of an IIP-coded block, the intra prediction modes of both the IIP-coded block and the intra-coded block are utilized. However, the intra prediction modes from intra-coded adjacent blocks may be inserted into the MPM before those from IIP-coded adjacent blocks.

[0311] In some embodiments, the intra prediction modes from intra-coded adjacent blocks may be inserted into the MPM after those from IIP-coded adjacent blocks.

[0312] Example 23: The intra prediction mode in an IIP-coded block can also be used to predict that of an intra-coded block.

[0313] In some embodiments, the intra prediction mode in an IIP-coded block can be used to derive the MPM for a normal intra-coded block. In some embodiments, the intra prediction mode in an IIP-coded block may have a lower priority than the intra prediction mode in an intra-coded block when they are used to derive the MPM for a normal intra-coded block.

[0314] In some embodiments, the intra prediction mode in an IIP-coded block may be used to predict the intra prediction mode of a normal intra-coded block or an IIP-coded block only if one or more of the following conditions are satisfied: 1. The two blocks are on the same CTU line. 2. The two blocks are in the same CTU. 3. The two blocks are in the same M×N region (e.g., M = N = 64). 4. The two blocks are on the same M×N region line (e.g., M = N = 64).

[0315] Example 24: In some embodiments, the MPM construction process for an IIP-coded block should be the same as that for a normal intra-coded block.

[0316] In the same embodiment, six MPMs are used for an inter-coded block by inter-intra prediction.

[0317] In some embodiments, only a part of the MPMs are used for an IIP-coded block. In some embodiments, the first one is always used. Alternatively, further, neither the MPM flag nor the MPM index needs to be signaled. In some embodiments, the first four MPMs may be utilized. Alternatively, further, the MPM flag need not be signaled, but the MPM index needs to be signaled.

[0318] In some embodiments, each block may select one from the MPM list according to the intra prediction mode included in the MPM list, for example, select the mode having the minimum index compared with a given mode (e.g., planar).

[0319] In some embodiments, each block may select a subset of modes from the MPM list and notify the mode indices within that subset.

[0320] In some embodiments, the context used to code the intra MPM modes is reused to code the intra modes in the IIP-coded blocks. In some embodiments, different contexts used to code the intra MPM modes are used to code the intra modes in the IIP-coded blocks.

[0321] Example 25: In some embodiments, for the angular intra prediction modes other than the horizontal and vertical directions, equal weights are applied to the intra prediction block and the inter prediction block generated for the IIP-coded block.

[0322] Example 26: In some embodiments, for a particular position, zero weight may be applied in the IIP coding process.

[0323] In some embodiments, zero weight may be applied to the intra prediction block used in the IIP coding process.

[0324] In some embodiments, zero weight may be applied to the inter prediction block used in the IIP coding process.

[0325] Example 27: In some embodiments, the intra prediction mode of the IIP-coded block can only be selected as one of the MPMs, regardless of the size of the current block.

[0326] In some embodiments, the MPM flag is not notified and is assumed to be 1, regardless of the size of the current block.

[0327] Example 28: For IIP-coded blocks, the luma prediction chroma mode (LM) mode is used instead of the Derived Mode (DM) mode to perform intra prediction for the chroma component.

[0328] In some embodiments, both DM and LM may be permitted.

[0329] In some embodiments, multiple intra prediction modes may be permitted for the chroma component.

[0330] In some embodiments, whether multiple modes should be permitted for the chroma component may depend on the color format. In one example, for the 4:4:4 color format, the permitted chroma intra prediction modes may be the same as those for the luma component.

[0331] Example 29: Inter-intra prediction may not be permitted in one or more of the following specific cases: A. w == T1 || h == T1, for example, T = 4. B. w > T1 || h > T1, for example, T1 = 64. C. (w == T1 && h == T2) || (w == T2 && h == T1), for example, T1 = 4, T2 = 16.

[0332] Example 30: Inter-intra prediction may not be permitted for blocks using dual prediction.

[0333] In some embodiments, if the selected merge candidate of an IIP-coded block uses dual prediction, it is converted to a single prediction merge candidate. In some embodiments, only the prediction from reference list 0 is retained and the merge candidate is treated as a single prediction referring to reference list 0. In some embodiments, only the prediction from reference list 1 is retained and the merge candidate is treated as a single prediction referring to reference list 1.

[0334] In some embodiments, a restriction is added that the selected merge candidate should be a unidirectional merge candidate. Alternatively, the notified merge index of the IIP-coded block indicates the index of the unidirectional merge candidate (i.e., the bidirectional merge candidates are not counted).

[0335] In some embodiments, the merge candidate list construction process used in the triangular prediction mode may be utilized to derive a motion candidate list for the IIP-coded block.

[0336] Example 31: When inter-intra prediction is applied, some coding tools may not be permitted.

[0337] In some embodiments, bi-directional optical flow (BIO) is not applied to bi-prediction.

[0338] In some embodiments, overlapped block motion compensation (OBMC) is not applied.

[0339] In some embodiments, the motion vector derivation / refinement process on the decoder side is not permitted.

[0340] Example 32: The intra prediction process used in inter-intra prediction may be different from that used in normal intra-coded blocks.

[0341] In some embodiments, adjacent samples may be filtered in various ways. In some embodiments, adjacent samples are not filtered before constructing the intra prediction used in inter-intra prediction.

[0342] In some embodiments, the position-dependent intra prediction sample filtering process is not configured for intra prediction used in inter-intra prediction. In some embodiments, multi-line intra prediction is not permitted in inter-intra prediction. In some embodiments, wide-angle intra prediction is not permitted in inter-intra prediction.

[0343] Example 33: Assume that the intra and inter prediction values in intra and inter composite prediction are PIntra and PInter, and the weight coefficients are w_intra and w_inter, respectively. The predicted value at position (x, y) is calculated as (PIntra(x, y) × w_intra(x, y) + PInter(x, y) × w_inter(x, y) + offset(x, y)) >> N, where w_inter(x, y) + w_inter(x, y) = 2 N and offset(x, y) = 2 (N-1) and in one example, N = 3.

[0344] Example 34: In some embodiments, the MPM flags signaled in a normal intra-coded block and in an IIP-coded block should share the same arithmetic coding context.

[0345] Example 35: In some embodiments, MPM is not required to code the intra prediction mode in an IIP-coded block (assuming the block width and height are w and h).

[0346] In some embodiments, the four modes {PLANAR, DC, VERTICAL, HORIZONTAL} are binary coded as 00, 01, 10, and 11 (any mapping rule such as 00 - PLANAR, 01 - DC, 10 - VERTICAL, 11 - HORIZONTAL can be used).

[0347] In some embodiments, the four modes {planner, DC, vertical, horizontal} are binary coded as 0, 10, 110, and 111 (it can be according to any mapping rule such as 0 - planner, 10 - DC, 110 - vertical, 111 - horizontal).

[0348] In some embodiments, the four modes {planner, DC, vertical, horizontal} are binary coded as 1, 01, 001, and 000 (it can be according to any mapping rule such as 1 - planner, 01 - DC, 001 - vertical, 000 - horizontal).

[0349] In some embodiments, when W > N×H (N is an integer such as 2), only three modes {planner, DC, vertical} can be used. The three modes are binary coded as 1, 01, and 11 (it can be according to any mapping rule such as 1 - planner, 01 - DC, 11 - vertical).

[0350] In some embodiments, when W > N×H (N is an integer such as 2), only three modes {planner, DC, vertical} can be used. The three modes are binary coded as 0, 10, and 00 (it can be according to any mapping rule such as 0 - planner, 10 - DC, 00 - vertical).

[0351] In some embodiments, when H > N×W (N is an integer such as 2), only three modes {planner, DC, horizontal} can be used. The three modes are binary coded as 1, 01, and 11 (it can be according to any mapping rule such as 1 - planner, 01 - DC, 11 - horizontal).

[0352] In some embodiments, when H > N×W (N is an integer such as 2), only three modes {planner, DC, horizontal} can be used. The three modes are binary coded as 0, 10, and 00 (it can be according to any mapping rule such as 0 - planner, 10 - DC, 00 - horizontal).

[0353] Example 36: In some embodiments, only DC and planar modes are used in IP-coded blocks. In some embodiments, one flag is signaled to indicate whether DC or planar is used.

[0354] Example 37: In some embodiments, the IIP is configured differently for different color components.

[0355] In some embodiments, inter-intra prediction is not performed for chroma components (e.g., Cb and Cr).

[0356] In some embodiments, the intra prediction mode for chroma components is different from that for luma components in IIP-coded blocks. In some embodiments, the DC mode is always used for chroma. In some embodiments, the planar mode is always used for chroma. In some embodiments, the LM mode is always used for chroma.

[0357] In some embodiments, how to perform IIP for different color components may depend on the color format (e.g., 4:2:0 or 4:4:4).

[0358] In some embodiments, how to perform IIP for different color components may depend on the block size. For example, inter-intra prediction is not performed for chroma components (e.g., Cb and Cr) when the width or height of the current block is 4 or less.

[0359] <MV Prediction Problem> In the following discussion, the accuracy used for the MVs stored for spatial motion prediction is denoted as P1, and the accuracy used for the MVs stored for temporal motion prediction is denoted as P2.

[0360] Example 38: P1 and P2 may be the same, or they may be different.

[0361] In some embodiments, P1 is a 1 / 16 luma pixel and P2 is a 1 / 4 luma pixel. In some embodiments, P1 is a 1 / 16 luma pixel and P2 is a 1 / 8 luma pixel. In some embodiments, P1 is a 1 / 8 luma pixel and P2 is a 1 / 4 luma pixel. In some embodiments, P1 is a 1 / 8 luma pixel and P2 is a 1 / 8 luma pixel. In some embodiments, P2 is a 1 / 16 luma pixel and P1 is a 1 / 4 luma pixel. In some embodiments, P2 is a 1 / 16 luma pixel and P1 is a 1 / 8 luma pixel. In some embodiments, P2 is a 1 / 8 luma pixel and P1 is a 1 / 4 luma pixel.

[0362] Example 39: P1 and P2 may not be fixed. In some embodiments, P1 / P2 may vary for different standard profiles / levels / tiers. In some embodiments, P1 / P2 may vary for pictures in different temporal layers. In some embodiments, P1 / P2 may vary for pictures of different widths / heights. In some embodiments, P1 / P2 may be signaled in the VPS / SPS / PPS / slice header / tile group header / tile / CTU / CU from the encoder to the decoder.

[0363] Example 40: For MV (MVx, MVy), the accuracies of MVx and MVy may be different and are represented as Px and Py.

[0364] In some embodiments, Px / Py may vary for different standard profiles / levels / tiers. Px / Py may vary for each picture in different temporal layers. In some embodiments, Px may vary for each picture of different widths. In some embodiments, Py may vary for each picture of different heights. In some embodiments, Px / Py may be signaled in the VPS / SPS / PPS / slice header / tile group header / tile / CTU / CU from the encoder to the decoder.

[0365] Example 41: Before storing the MV (MVx, MVy) for temporal motion prediction, it should be modified to an accurate precision.

[0366] In some embodiments, when P1 >= P2, MVx = Shift(MVx, P1 - P2), MVy = Shift(MVy, P1 - P2). In some embodiments, when P1 >= P2, MVx = SignShift(MVx, P1 - P2), MVy = SignShfit(MVy, P1 - P2). In some embodiments, when P1 < P2, MVx = MVx << (P2 - P1), MVy = MVy << (P2 - P1).

[0367] Example 42: Assume that the precision of MV (MVx, MVy) is Px and Pz, and MVx or MVy is stored by an integer having N bits. The range of MV (MVx, MVy) is MinX <= MVx <= MaxX and MinY <= MVy <= MaxY.

[0368] In some embodiments, MinX may be equal to MinY, or it may not be equal to MinY. In some embodiments, MaxX may be equal to MaxY, or it may not be equal to MaxY. In some embodiments, {MinX, MaxX} may depend on Px. In some embodiments, {MinY, MaxY} may depend on Py. In some embodiments, {MinX, MaxX, MinY, MaxY} may depend on N. In some embodiments, {MinX, MaxX, MinY, MaxY} may be different for the MV saved for spatial motion prediction and the MV saved for temporal motion prediction. In some embodiments, {MinX, MaxX, MinY, MaxY} may be different for different standard profiles / levels / tiers. In some embodiments, {MinX, MaxX, MinY, MaxY} may be different for pictures in different temporal layers. In some embodiments, {MinX, MaxX, MinY, MaxY} may be different for pictures of different widths / heights. In some embodiments, {MinX, MaxX, MinY, MaxY} may be signaled in the VPS / SPS / PPS / slice header / tile group header / tile / CTU / CU from the encoder to the decoder. In some embodiments, {MinX, MaxX} may be different for pictures of different widths. In some embodiments, {MinY, MaxY} may be different for pictures of different heights. In some embodiments, MVx is clipped to [MinX, MaxX] before being placed in storage for spatial motion prediction. In some embodiments, MVx is clipped to [MinX, MaxX] before being placed in storage for temporal motion prediction. In some embodiments, MVy is clipped to [MinY, MaxY] before being placed in storage for spatial motion prediction. In some embodiments, MVy is clipped to [MinY, MaxY] before being placed in storage for temporal motion prediction.

[0369] <Line Buffer Reduction for Affine Merge Mode> Example 43: The affine model (derived CPMV or affine parameters) inherited by an affine merge candidate from an adjacent block is always a 6-parameter affine model.

[0370] In some embodiments, when an adjacent block is coded by a 4-parameter affine model, the affine model is still inherited as a 6-parameter affine model.

[0371] In some embodiments, whether a 4-parameter affine model from an adjacent block is inherited as a 6-parameter affine model or a 4-parameter affine model may depend on the position of the current block. In some embodiments, a 4-parameter affine model from an adjacent block is inherited as a 6-parameter affine model when the adjacent block is not in the same coding tree unit (CTU, also known as the largest coding unit LCU) as the current block. In some embodiments, a 4-parameter affine model from an adjacent block is inherited as a 6-parameter affine model when the adjacent block is not in the same CTU line or CTU row as the current block. In some embodiments, a 4-parameter affine model from an adjacent block is inherited as a 6-parameter affine model when the adjacent block is not in the same M×N region as the current block. For example, M = N = 64. In this case, a tile / slice / picture is divided into a plurality of non-overlapping M×N regions. In some embodiments, a 4-parameter affine model from an adjacent block is inherited as a 6-parameter affine model when the adjacent block is not in the same M×N region line or M×N region row as the current block. For example, M = N = 64. CTU lines / rows and region lines / rows are shown in FIG. 10.

[0372] In some embodiments, assuming that the upper left corner (or other position) of the current block is (x, y) and the upper left corner (or other position) of the adjacent block is (x’, y’), the 4-parameter affine model from the adjacent block is inherited as a 6-parameter affine model when the adjacent block satisfies one or more of the following conditions: (a) x / M!= x’ / M. For example, M = 128 or 64. (b) y / N!= y’ / N. For example, N = 128 or 64. (c) ((x / M!= x’ / M) && (y / N!= y’ / N)). For example, M = N = 128 or M = N = 64. (d) ((x / M!= x’ / M) || (y / N!= y’ / N)). For example, M = N = 128 or M = N = 64. (e) x >> M!= x’ >> M. For example, M = 7 or 6. (f) y >> N!= y’ >> N. For example, N = 7 or 6. (g) ((x >> M!= x’ >> M) && (y >> N!= y’ >> N)). For example, M = N = 7 or M = N = 6. (h) ((x >> M!= x’ >> M) || (y >> N!= y’ >> N)). For example, M = N = 7 or M = N = 6.

[0373] [5. Embodiment] The following description shows examples of how the disclosed technology can be implemented within the syntax structure of the current VVC standard. New additions are shown in bold (or underlined), and deletions are shown in italics.

[0374] <Embodiment #1 (Disable 4×4 inter prediction and disable dual prediction for 4×8, 8×4, 4×16, and 16×4 blocks)>

[0375] 7.3.6.6 Coding Unit Syntax

Table 6

[0376] 7.4.7.6 Coding Unit Semantics That the pred_mode_flag is equal to 0 specifies that the current coding unit is coded in the inter prediction mode. That the pred_mode_flag is equal to 1 specifies that the current coding unit is coded in the intra prediction mode. The variable CuPredMode[x][y] is derived as follows when x = x0··x0 + cbWidth - 1 and y = y0··y0 + cbHeight - 1: - When the pred_mode_flag is equal to 0, CuPredMode[x][y] is set equal to MODE_INTER. - Otherwise (the pred_mode_flag is equal to 1), CuPredMode[x][y] is set equal to MODE_INTRA.

[0377] When the pred_mode_flag does not exist, it is assumed to be equal to 1 when decoding an I tile group, and is assumed to be equal to 0 when decoding a P or B tile group. or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 1 when decoding an I tile group, and is assumed to be equal to 0 when decoding a P or B tile group.

[0378] That the pred_mode_ibc_flag is equal to 1 specifies that the current coding unit is coded in the IBC prediction mode. That the pred_mode_ibc_flag is equal to 0 specifies that the current coding unit is not coded in the IBC prediction mode.

[0379] When the pred_mode_ibc_flag does not exist, it is assumed to be equal to the value of sps_ibc_enabled_flag when decoding an I tile group, and is assumed to be equal to 0 when decoding a P or B tile group. or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 and is coded in skip mode equal to the value of sps_ibc_enabled_flag when decoding an I tile group, and is assumed to be equal to 0 when decoding a P or B tile group.

[0380] When pred_mode_ibc_flag is equal to 1, the variable CuPredMode[x][y] is set equal to MODE_IBC for x = x0··x0 + cbWidth - 1 and y = y0··y0 + cbHeight - 1.

[0381] inter_pred_idc[x0][y0] specifies whether list 0, list 1, or bi-prediction is used for the current coding unit according to Table 7-9. The array indices x0, y0 identify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.

Table 7

[0382] If inter_pred_idc[x0][y0] does not exist, it is assumed to be equal to PRED_L0.

[0383] 8.5.2.1 Overview The inputs to this process are: - The luma position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, - The variable cbWidth that specifies the width of the current coding block within the luma sample, - The variable cbHeight that specifies the height of the current coding block within the luma sample are.

[0384] The outputs of this process are: - Luma motion vectors at 1 / 16 fractional sample accuracy mvL0[0][0] and mvL1[0][0], - Reference indices refIdxL0 and refIdxL1, - Prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], - Bi-prediction weight index gbiIdx are.

[0385] Assume that X is 0 or 1, and let the variable LX be RefPicList[X] of the current picture.

[0386] For the derivation of the variables mvL0[0][0] and mvL1[0][0], refIdxL0 and refIdxL1, and predFlagL0[0][0] and predFlagL1[0][0], the following is applied: - When merge_flag[xCb][yCb] is equal to 1, the process for deriving the luma motion vectors for the merge mode specified in Section 8.5.2.2 is called with the luma position (xCb, yCb), the inputs of the variables cbWidth and cbHeight, and the outputs of the luma motion vectors mvL0[0][0], mvL1[0][0], the reference index refIdxL0, refIdxL1, the prediction list usage flag predFlagL0[0][0], predFlagL1[0][0], and the bi-prediction weight index gbiIdx. - Otherwise, the following is applied: - When X is replaced by either 0 or 1 in the variables predFlagLX[0][0], mvLX[0][0] and refIdxLX, in PRED_LX and in the syntax elements ref_idx_lX and mvdLX, the following ordered steps are applied: 1. The variables refIdxLX and predFlagLX[0][0] are derived as follows: · When inter_pred_idc[xCb][yCb] is equal to PRED_LX or PRED_BI, refIdxLX = ref_idx_lX[xCb][yCb] (8-266) predFlagLX[0][0]=1 (8-267) · Otherwise, the variables refIdxLX and predFlagLX[0][0] are, refIdxLX = 1 (8 - 268) predFlagLX[0][0] = 0 (8 - 269) is specified by 2. The variable mvdLX is derived as follows: mvdLX[0] = MvdLX[xCb][yCb][0] (8 - 270) mvdLX[1] = mvdLX[xCb][yCb][1] (8 - 271) 3. When predFlagLX[0][0] is equal to 1, the process of deriving the luma motion vector prediction in Section 8.5.2.8 is called by the luma coding block position (xCb, yCb), the coding block width cbWidth as input, the coding block height cbHeight, and the variables refIdxLX, and the output mvpLX. 4. When predFlagLX[0][0] is equal to 1, the luma motion vector mvLX[0][0] is derived as follows: uLX[0] = (mvpLX[0] + mvdLX[0] + 2 18 ) % 2 18 (8 - 272) mvLX[0][0][0] = (uLX[0] >= 2 17 )? (uLX[0] - 2 18 ): uLX[0] (8 - 273) uLX[1] = (mvpLX[1] + mvdLX[1] + 2 18 ) % 2 18 (8 - 274) mvLX[0][0][1] = (uLX[1] >= 2 17 )? (uLX[1] - 2 18 ): uLX[1] (8 - 275) Note 1 - The values obtained as a result of mvLX[0][0][0] and mvLX[0][0][1] above are always in the range of -2 17 or more and 2 17 -1 or less. - The dual prediction weight index gbiIdx is set equal to gbi_idx[xCb][yCb].

[0387] If all of the following conditions are true, refIdxL1 is set equal to -1, predFlagL1 is set equal to 0, and gbiIdx is set equal to 0: - predFlagL0[0][0] is equal to 1. - predFlagL1[0][0] is equal to 1. - (cbWidth + cbHeight == 8) || (cbWidth + cbHeight == 12) || (cbWidth + cbHeight == 20) (The conditions "cbWidth is equal to 4" and "cbHeight is equal to 4" have been removed.)

[0388] The update process for the history-based motion vector predictor list specified in section 8.5.2.16 is called using the luma motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], and the dual prediction weight index gbiIdx.

[0389] 9.5.3.8 Binaryization process for inter_pred_idc The inputs to this process are the requirements for the binaryization of the syntax element inter_pred_idc, the width cbWidth of the current luma coding block, and the height cbHeight of the current luma coding block. The output of this process is the binaryization of the syntax element. The binaryization of the syntax element inter_pred_idc is clearly stated in Table 9-9.

Table 8

[0390] 9.5.4.2.1 Overview

Table 9

[0391] <Embodiment #2 (Disable 4×4 Inter Prediction)>

[0392] 7.3.6.6 Coding Unit Syntax

Table 10

[0393] 7.4.7.6 Coding Unit Semantics That the pred_mode_flag is equal to 0 specifies that the current coding unit is coded in the inter prediction mode. That the pred_mode_flag is equal to 1 specifies that the current coding unit is coded in the intra prediction mode. The variable CuPredMode[x][y] is derived as follows when x = x0··x0 + cbWidth - 1 and y = y0··y0 + cbHeight - 1: - When the pred_mode_flag is equal to 0, CuPredMode[x][y] is set equal to MODE_INTER. - Otherwise (the pred_mode_flag is equal to 1), CuPredMode[x][y] is set equal to MODE_INTRA.

[0394] When the pred_mode_flag does not exist, it is presumed to be equal to 1 when decoding an I tile group, or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 and presumed to be equal to 0 when decoding a P or B tile group.

[0395] That the pred_mode_ibc_flag is equal to 1 specifies that the current coding unit is coded in the IBC prediction mode. That the pred_mode_ibc_flag is equal to 0 specifies that the current coding unit is not coded in the IBC prediction mode.

[0396] When pred_mode_ibc_flag does not exist, it is assumed to be equal to the value of sps_ibc_enabled_flag when decoding an I tile group, and equal to 0 when decoding a P or B tile group. or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 and is coded in skip mode When pred_mode_ibc_flag is equal to 1, variable CuPredMode[x][y] is set equal to MODE_IBC for x = x0··x0+cbWidth-1 and y = y0··y0+cbHeight-1.

[0397] When pred_mode_ibc_flag is equal to 1, variable CuPredMode[x][y] is set equal to MODE_IBC for x = x0··x0+cbWidth-1 and y = y0··y0+cbHeight-1.

[0398] <Embodiment #3 (Invalidation of Bi-prediction for 4×8, 8×4, 4×16, and 16×4 Blocks)>

[0399] 7.4.7.6 Coding Unit Semantics inter_pred_idc[x0][y0] specifies whether list 0, list 1, or bi-prediction is used for the current coding unit according to Table 7-9. Array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture. [Table 11]

[0400] When inter_pred_idc[x0][y0] does not exist, it is assumed to be equal to PRED_L0.

[0401] 8.5.2.1 Overview The inputs to this process are: - The luma position (xCb,yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, - The variable cbWidth that specifies the width of the current coding block within the luma sample, - The variable cbHeight that specifies the height of the current coding block within the luma sample are as follows.

[0402] The output of this process is: - Luma motion vectors at 1 / 16 fractional sample precision mvL0[0][0] and mvL1[0][0], - Reference indices refIdxL0 and refIdxL1, - Prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], - The bi-prediction weight index gbiIdx are as follows.

[0403] Assuming X is 0 or 1, let the variable LX be RefPicList[X] of the current picture.

[0404] For the derivation of the variables mvL0[0][0] and mvL1[0][0], refIdxL0 and refIdxL1, and predFlagL0[0][0] and predFlagL1[0][0], the following applies: - When merge_flag[xCb][yCb] is equal to 1, the process for deriving the luma motion vector for the merge mode specified in Section 8.5.2.2 is called with the luma position (xCb,yCb), the inputs of the variables cbWidth and cbHeight, and the outputs of the luma motion vectors mvL0[0][0], mvL1[0][0], the reference indices refIdxL0, refIdxL1, the prediction list usage flags predFlagL0[0][0], predFlagL1[0][0], and the bi-prediction weight index gbiIdx. - Otherwise, the following applies: - In the variables predFlagLX[0][0], mvLX[0][0], and refIdxLX, when X is replaced by either 0 or 1 in PRED_LX and in the syntax elements ref_idx_lX and mvdLX, the following ordered steps are applied: 5. The variables refIdxLX and predFlagLX[0][0] are derived as follows: · When inter_pred_idc[xCb][yCb] is equal to PRED_LX or PRED_BI, refIdxLX = ref_idx_lX[xCb][yCb] (8-266) predFlagLX[0][0]=1 (8-267) · Otherwise, the variables refIdxLX and predFlagLX[0][0] are refIdxLX = 1 (8-268) predFlagLX[0][0]=0 (8-269) specified by 6. The variable mvdLX is derived as follows: mvdLX[0]=MvdLX[xCb][yCb][0] (8-270) mvdLX[1]=mvdLX[xCb][yCb][1] (8-271) 3. When predFlagLX[0][0] is equal to 1, the process for deriving the luma motion vector prediction in Section 8.5.2.8 is called with the luma coding block position (xCb, yCb), the coding block width cbWidth as input, the coding block height cbHeight, and the variable refIdxLX, and the output mvpLX. 8. When predFlagLX[0][0] is equal to 1, the luma motion vector mvLX[0][0] is derived as follows: uLX[0]=(mvpLX[0]+mvdLX[0]+2 18 )%2 18 (8-272) mvLX[0][0][0] = (uLX[0] >= 2 17 )? (uLX[0] - 2 18 ): uLX[0] (8 - 273) uLX[1] = (mvpLX[1] + mvdLX[1] + 2 18 ) % 2 18 (8 - 274) mvLX[0][0][1] = (uLX[1] >= 2 17 )? (uLX[1] - 2 18 ): uLX[1] (8 - 275) Note 1 - The values obtained as a result of mvLX[0][0][0] and mvLX[0][0][1] above are always within the range of -2 17 or more and 2 17 -1 or less. - The dual prediction weight index gbiIdx is set equal to gbi_idx[xCb][yCb.

[0405] When all of the following conditions are true, refIdxL1 is set equal to -1, predFlagL1 is set equal to 0, and gbiIdx is set equal to 0: - predFlagL0[0][0] is equal to 1. - predFlagL1[0][0] is equal to 1. - (cbWidth + cbHeight == 8) || (cbWidth + cbHeight == 12) || (cbWidth + cbHeight == 20) ("cbWidth is equal to 4" and "cbHeight is equal to 4" conditions have been removed.)

[0406] The update process of the history-based motion vector predictor list specified in section 8.5.2.16 is called using the luma motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], and the dual prediction weight index gbiIdx.

[0407] 9.5.3.8 Binarization Process for inter_pred_idc The input to this process is the requirement for the binarization of the syntax element inter_pred_idc, the width cbWidth of the current luma coding block, and the height cbHeight of the current luma coding block. The output of this process is the binarization of the syntax element. The binarization of the syntax element inter_pred_idc is clearly described in Table 9-9. [Table 12]

[0408] 9.5.4.2.1 Overview [Table 13]

[0409] <Embodiment #4 (Disable 4×4 inter prediction and disable dual prediction for 4×8 and 8×4 blocks)>

[0410] 7.3.6.6 Coding Unit Syntax [Table 14] TIFF0007683069000022.tif230168

[0411] 7.4.7.6 Coding Unit Semantics That pred_mode_flag is equal to 0 specifies that the current coding unit is coded in the inter prediction mode. That pred_mode_flag is equal to 1 specifies that the current coding unit is coded in the intra prediction mode. The variable CuPredMode[x][y] is derived as follows when x = x0··x0 + cbWidth - 1 and y = y0··y0 + cbHeight - 1: - When pred_mode_flag is equal to 0, CuPredMode[x][y] is set equal to MODE_INTER. - Otherwise (when pred_mode_flag is equal to 1), CuPredMode[x][y] is set equal to MODE_INTRA.

[0412] In the case where pred_mode_flag does not exist, it is assumed to be equal to 1 when decoding an I tile group, or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 and is assumed to be equal to 0 when decoding a P or B tile group.

[0413] That pred_mode_ibc_flag is equal to 1 specifies that the current coding unit is coded in IBC prediction mode. That pred_mode_ibc_flag is equal to 0 specifies that the current coding unit is not coded in IBC prediction mode.

[0414] In the case where pred_mode_ibc_flag does not exist, it is assumed to be or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 and is coded in skip mode equal to the value of sps_ibc_enabled_flag when decoding an I tile group, and is assumed to be equal to 0 when decoding a P or B tile group.

[0415] When pred_mode_ibc_flag is equal to 1, the variable CuPredMode[x][y] is set equal to MODE_IBC when x = x0··x0+cbWidth-1 and y = y0··y0+cbHeight-1.

[0416] inter_pred_idc[x0][y0] specifies whether list 0, list 1, or bi-prediction is used for the current coding unit according to Table 7-9. The array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.

Table 15

[0417] If inter_pred_idc[x0][y0] does not exist, it is assumed to be equal to PRED_L0.

[0418] 8.5.2.1 Overview The input to this process is: - The luma position (xCb, yCb) of the top-left sample of the current luma coding block with respect to the top-left luma sample of the current picture, - The variable cbWidth that specifies the width of the current coding block within the luma sample, - The variable cbHeight that specifies the height of the current coding block within the luma sample That is.

[0419] The output of this process is: - Luma motion vectors at 1 / 16 fractional sample accuracy mvL0[0][0] and mvL1[0][0], - Reference indices refIdxL0 and refIdxL1, - Prediction list utilization flags predFlagL0[0][0] and predFlagL1[0][0], - Biprediction weight index gbiIdx That is.

[0420] Assuming X is 0 or 1, let variable LX be RefPicList[X] of the current picture.

[0421] For the derivation of variables mvL0[0][0] and mvL1[0][0], refIdxL0 and refIdxL1, and predFlagL0[0][0] and predFlagL1[0][0], the following is applied: - When -merge_flag[xCb][yCb] is equal to 1, the process for deriving the luma motion vector for the merge mode specified in Section 8.5.2.2 is called with the luma position (xCb, yCb), the input of the variables cbWidth and cbHeight, and the output of the luma motion vectors mvL0[0][0], mvL1[0][0], the reference indices refIdxL0, refIdxL1, the prediction list usage flags predFlagL0[0][0], predFlagL1[0][0], and the bi-prediction weight index gbiIdx. - Otherwise, the following applies: - When X is replaced by either 0 or 1 in the variables predFlagLX[0][0], mvLX[0][0] and refIdxLX, in PRED_LX and in the syntax elements ref_idx_lX and mvdLX, the following ordered steps apply: 1. The variables refIdxLX and predFlagLX[0][0] are derived as follows: · When inter_pred_idc[xCb][yCb] is equal to PRED_LX or PRED_BI, refIdxLX = ref_idx_lX[xCb][yCb] (8-266) predFlagLX[0][0]=1 (8-267) · Otherwise, the variables refIdxLX and predFlagLX[0][0] are refIdxLX = 1 (8-268) predFlagLX[0][0]=0 (8-269) as specified by. 2. The variable mvdLX is derived as follows: mvdLX[0]=MvdLX[xCb][yCb][0] (8-270) mvdLX[1]=mvdLX[xCb][yCb][1] (8-271) When predFlagLX[0][0] is equal to 1, the process of deriving the luma motion vector prediction in Section 8.5.2.8 is called by the luma coding block position (xCb, yCb), the coding block width cbWidth as input, the coding block height cbHeight, and the variables refIdxLX, and the output of mvpLX. When predFlagLX[0][0] is equal to 1, the luma motion vector mvLX[0][0] is derived as follows: uLX[0]=(mvpLX[0]+mvdLX[0]+2 18 )%2 18 (8-272) mvLX[0][0][0]=(uLX[0]>=2 17 )?(uLX[0]-2 18 ):uLX[0] (8-273) uLX[1]=(mvpLX[1]+mvdLX[1]+2 18 )%2 18 (8-274) mvLX[0][0][1]=(uLX[1]>=2 17 )?(uLX[1]-2 18 ):uLX[1] (8-275) Note 1 - The values obtained as a result of mvLX[0][0][0] and mvLX[0][0][1] above are always within the range of -2 17 above 2 17 -1 or less. - The dual prediction weight index gbiIdx is set equal to gbi_idx[xCb][yCb.

[0422] When all of the following conditions are true, refIdxL1 is set equal to -1, predFlagL1 is set equal to 0, and gbiIdx is set equal to 0: - predFlagL0[0][0] is equal to 1. - predFlagL1[0][0] is equal to 1. - (cbWidth + cbHeight == 8) || (cbWidth + cbHeight == 12) ("The conditions 'cbWidth is equal to 4' and 'cbHeight is equal to 4' were removed.")

[0423] The update process of the history-based motion vector predictor list specified in Section 8.5.2.16 is called using the luma motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], and the dual prediction weight index gbiIdx.

[0424] 9.5.3.8 Binarization Process for inter_pred_idc The input to this process is the requirement for the binarization of the syntax element inter_pred_idc, the width cbWidth of the current luma coding block, and the height cbHeight of the current luma coding block. The output of this process is the binarization of the syntax element. The binarization of the syntax element inter_pred_idc is clearly stated in Table 9-9. [Table 16]

[0425] 9.5.4.2.1 Overview [Table 17]

[0426] <5.5 Embodiment #5 (Disable 4×4 inter prediction, disable dual prediction for 4×8 and 8×4 blocks, and disable the shared merge list for the regular merge mode)>

[0427] 7.3.6.6 Coding Unit Syntax [Table 18] TIFF0007683069000027.tif231167

[0428] 7.4.7.6 Coding Unit Semantics That the pred_mode_flag is equal to 0 specifies that the current coding unit is coded in the inter prediction mode. That the pred_mode_flag is equal to 1 specifies that the current coding unit is coded in the intra prediction mode. The variable CuPredMode[x][y] is derived as follows when x = x0··x0 + cbWidth - 1 and y = y0··y0 + cbHeight - 1: - When the pred_mode_flag is equal to 0, CuPredMode[x][y] is set equal to MODE_INTER. - Otherwise (the pred_mode_flag is equal to 1), CuPredMode[x][y] is set equal to MODE_INTRA.

[0429] When the pred_mode_flag does not exist, it is assumed to be equal to 1 when decoding an I tile group, or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 and is assumed to be equal to 0 when decoding a P or B tile group.

[0430] That the pred_mode_ibc_flag is equal to 1 specifies that the current coding unit is coded in the IBC prediction mode. That the pred_mode_ibc_flag is equal to 0 specifies that the current coding unit is not coded in the IBC prediction mode.

[0431] When the pred_mode_ibc_flag does not exist, it is assumed to be equal to or when decoding a coding unit where cbWidth is equal to 4 and cbHeight is equal to 4 and is coded in skip mode the value of sps_ibc_enabled_flag when decoding an I tile group, and is assumed to be equal to 0 when decoding a P or B tile group.

[0432] When pred_mode_ibc_flag is equal to 1, the variable CuPredMode[x][y] is set equal to MODE_IBC for x = x0··x0 + cbWidth - 1 and y = y0··y0 + cbHeight - 1.

[0433] inter_pred_idc[x0][y0] specifies whether list 0, list 1, or bi-prediction is used for the current coding unit according to Table 7-9. The array indices x0, y0 identify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture. [Table 19]

[0434] If inter_pred_idc[x0][y0] does not exist, it is assumed to be equal to PRED_L0.

[0435] 8.5.2.1 Overview The inputs to this process are: - The luma position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, - The variable cbWidth that specifies the width of the current coding block within the luma samples, - The variable cbHeight that specifies the height of the current coding block within the luma samples That is.

[0436] The outputs of this process are: - Luma motion vectors at 1 / 16 fractional sample accuracy mvL0[0][0] and mvL1[0][0], - Reference indices refIdxL0 and refIdxL1, - Prediction list utilization flags predFlagL0[0][0] and predFlagL1[0][0], - Dual prediction weight index gbiIdx is as follows.

[0437] Assume that X is 0 or 1, and let variable LX be RefPicList[X] of the current picture.

[0438] For the derivation of variables mvL0[0][0] and mvL1[0][0], refIdxL0 and refIdxL1, and predFlagL0[0][0] and predFlagL1[0][0], the following is applied: - When merge_flag[xCb][yCb] is equal to 1, the process for deriving the luma motion vector for the merge mode specified in Section 8.5.2.2 is called with the luma position (xCb, yCb), the inputs of variables cbWidth and cbHeight, and the outputs of the luma motion vectors mvL0[0][0], mvL1[0][0], the reference indices refIdxL0, refIdxL1, the prediction list utilization flags predFlagL0[0][0], predFlagL1[0][0], and the dual prediction weight index gbiIdx. - Otherwise, the following is applied: - When X is replaced by either 0 or 1 in variables predFlagLX[0][0], mvLX[0][0] and refIdxLX, in PRED_LX and in syntax elements ref_idx_lX and mvdLX, the following ordered steps are applied: 5. Variables refIdxLX and predFlagLX[0][0] are derived as follows: · When inter_pred_idc[xCb][yCb] is equal to PRED_LX or PRED_BI, refIdxLX = ref_idx_lX[xCb][yCb] (8-266) predFlagLX[0][0] = 1 (8-267) · Otherwise, the variables refIdxLX and predFlagLX[0][0] are refIdxLX = 1 (8-268) predFlagLX[0][0] = 0 (8-269) specified by 6. The variable mvdLX is derived as follows: mvdLX[0] = MvdLX[xCb][yCb][0] (8-270) mvdLX[1] = mvdLX[xCb][yCb][1] (8-271) 7. When predFlagLX[0][0] is equal to 1, the process of deriving the luma motion vector prediction in Section 8.5.2.8 is called by the luma coding block position (xCb, yCb), the coding block width cbWidth as input, the coding block height cbHeight, and the variables refIdxLX, and the output mvpLX. 8. When predFlagLX[0][0] is equal to 1, the luma motion vector mvLX[0][0] is derived as follows: uLX[0]=(mvpLX[0]+mvdLX[0]+2 18 )%2 18 (8-272) mvLX[0][0][0]=(uLX[0]>=2 17 )?(uLX[0]-2 18 ):uLX[0] (8-273) uLX[1]=(mvpLX[1]+mvdLX[1]+2 18 )%2 18 (8-274) mvLX[0][0][1]=(uLX[1]>=2 17 )?(uLX[1]-2 18 ):uLX[1] (8-275) Note 1 - The values obtained as a result of mvLX[0][0][0] and mvLX[0][0][1] above are always -2 17 or more and 2 17It is within the range of -1 or less. - The dual prediction weight index gbiIdx is set equal to gbi_idx[xCb][yCb].

[0439] When all of the following conditions are true, refIdxL1 is set equal to -1, predFlagL1 is set equal to 0, and gbiIdx is set equal to 0: - predFlagL0[0][0] is equal to 1. - predFlagL1[0][0] is equal to 1. - (cbWidth + cbHeight == 8) || (cbWidth + cbHeight == 12) (The conditions "cbWidth is equal to 4" and "cbHeight is equal to 4" have been removed.)

[0440] The update process of the history-based motion vector predictor list specified in Section 8.5.2.16 is called using the luma motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], and the dual prediction weight index gbiIdx.

[0441] 8.5.2.2 Luma Motion Vector Derivation Process for Merge Mode This process is called only when merge_flag[xCb][yCb] is equal to 1. (xCb,yCb) identifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture.

[0442] The inputs to this process are: - The luma position (xCb,yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, - The variable cbWidth that specifies the width of the current coding block within the luma sample, - The variable cbHeight that specifies the height of the current coding block in the luma sample is as follows.

[0443] The output of this process is: - Luma motion vectors at 1 / 16 fractional sample precision mvL0[0][0] and mvL1[0][0], - Reference indices refIdxL0 and refIdxL1, - Prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], - Dual prediction weight index gbiIdx is as follows.

[0444] The dual prediction weight index gbiIdx is set equal to 0.

[0445] The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpCand are derived as follows:

Number

[0446] 8.5.2.6 Derivation process of merge candidates based on history The inputs to this process are: - The merge candidate list mergeCandList, - The variable isInSmr that indicates whether the current coding unit is within the shared merge candidate region, - The number of available merge candidates in the list numCurrMergeCand is as follows.

[0447] The output of this process is: - The modified merge candidate list mergeCandList, - The modified number of merge candidates in the list numCurrMergeCand is as follows.

[0448] Variable isPrunedA 1 and isPrunedB 1 are both set equal to FALSE.

[0449] The array smrHmvpCandList and the variable smrNumHmvpCand are derived as follows:

Number

[0450] For each candidate in smrHmvpCandList[hMvpIdx] having an index hMVpIdx = 1··smrNumHmvpCand, the following ordered steps are repeated until numCurrMergeCand becomes equal to (MaxNumMergeCand - 1): 1. The variable sameMotion is derived as follows: · For any merge candidate N that is A 1 or B 1 both sameMotion and isPrunedN are set equal to TRUE if all of the following conditions are met: - hMvpIdx is less than or equal to 2. - The candidate smrHmvpCandList[smrNumHmvpCand - hMvpIdx] is equal to the merge candidate N. - isPrunedN is equal to FALSE. · Otherwise, sameMotion is set equal to FALSE. 2. If sameMotion is equal to FALSE, the candidate smrHmvpCandList[smrNumHmvpCand - hMvpIdx] is added to the merge candidate list as follows: mergeCandList[numCurrMergeCand++] = smrHmvpCandList[smrNumHmvpCand - hMvpIdx] (8 - 355)

[0451] 9.5.3.8 Binaryization Process for inter_pred_idc The input to this process is the requirement for the binaryization of the syntax element inter_pred_idc, the width cbWidth of the current luma coding block, and the height cbHeight of the current luma coding block. The output of this process is the binaryization of the syntax element. The binaryization of the syntax element inter_pred_idc is clearly described in Table 9-9. [Table 20]

[0452] 9.5.4.2.1 Overview [Table 21]

[0453] FIG. 11 is a block diagram of a video processing apparatus 1100. The apparatus 1100 may be used to implement one or more of the methods described herein. The apparatus 1100 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1100 may include one or more processors 1102, one or more memories 1104, and video processing hardware 1106. The processor 1102 may be configured to implement one or more of the methods described herein. The memory 1104 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 1106 may be used to implement some of the techniques described herein in a hardware circuit.

[0454] FIG. 12 is a flowchart of an exemplary method 1200 for video processing. The method 1200 includes a step (1202) of determining a size limit between a representative motion vector of a current video block to be affine coded and a motion vector of a sub-block of the current video block, and a step (1204) of performing a conversion between a pixel value and a bitstream representation of the current video block or the sub-block by using the size limit.

[0455] As used herein, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits that are spread at different locations within the bitstream or are at the same location, as defined by the syntax, for example. For example, a macroblock may be encoded using bits in the header and other fields within the bitstream with respect to the transformed and coded error residues.

[0456] As is apparent, the disclosed techniques are useful for implementing embodiments in which the implementation complexity of video processing is reduced by reducing memory requirements or line buffer size requirements. Some of the techniques currently disclosed may be described using the following bullet points.

[0457] 1. A method for video processing, comprising: determining a size limit between a representative motion vector of a current video block to be affine coded and a motion vector of a sub-block of the current video block; and performing a conversion between a bitstream representation and a pixel value of the current video block or the sub-block by using the size limit. A method having the above.

[0458] 2. The step of performing the conversion includes generating the bitstream representation from the pixel values. The method according to item 1.

[0459] 3. The step of performing the conversion includes generating the pixel values from the bitstream representation. The method according to item 1.

[0460] 4. The size limitation includes restricting the values of the motion vectors (MVx, MVy) of the sub-blocks according to MVx >= MV’x - DH0 and MVx <= MV’x + DH1 and MVy >= MV’y - DV0 and MVy <= MV’y + DV1, where MV’ = (MV’x, MV’y), MV’ represents the representative motion vector, and DH0, DH1, DV0, and DV1 represent positive numbers. The method according to any one of items 1 to 3.

[0461] 5. The size limitation is as follows: i. DH0 is equal to DH1, or DV0 is equal to DV1. ii. DH0 is equal to DV0, or DH1 is equal to DV1. iii. DH0 and DH1 are different, or DV0 and DV1 are different. iv. DH0, DH1, DV0, and DV1 are notified in the bitstream representation at the video parameter set level or sequence parameter set level or picture parameter set level or slice header level or tile group header level or tile level or coding tree unit level or coding unit level or prediction unit level. v. DH0, DH1, DV0, and DV1 are functions of the video processing mode. vi. DH0, DH1, DV0, and DV1 depend on the width and height of the current video block. vii. DH0, DH1, DV0, and DV1 depend on whether the current video block is coded using single prediction or dual prediction. viii. DH0, DH1, DV0, and DV1 depend on the positions of the sub-blocks and include at least one of the method described in item 4.

[0462] 6. The representative motion vector corresponds to the control point motion vector of the current video block, and is the method described in any one of items 1 to 5.

[0463] 7. The representative motion vector corresponds to the motion vector of the corner sub-block of the current video block, and is the method described in any one of items 1 to 5.

[0464] 8. The accuracy used for the motion vector of the sub-block and the representative motion vector corresponds to the motion vector signaling accuracy in the bitstream representation, and is the method described in any one of items 1 to 7.

[0465] 9. The accuracy used for the motion vector of the sub-block and the representative motion vector corresponds to the storage accuracy for storing the motion vector, and is the method described in any one of items 1 to 7.

[0466] 10. A method for video processing, comprising: determining, for a current video block to be affine-coded, one or more sub-blocks of the current video block, each sub-block having a size of M×N pixels where M and N are multiples of 2 or 4; said determining step; scaling the motion vector of the sub-block to fit the size limit; conditionally based on a trigger, performing a conversion between the bitstream representation and the pixel values of the current video block by using the size limit; and having the method.

[0467] 11. The step of performing the conversion includes generating the bitstream representation from the pixel values. The method according to clause 10.

[0468] 12. The step of performing the conversion includes generating the pixel values from the bitstream representation. The method according to clause 10.

[0469] 13. The size limitation is such that the maximum difference between the integer parts of the sub-block motion vectors of the current video block is at most K pixels, where K is an integer. The method according to any one of clauses 10 to 12.

[0470] 14. The method is applicable only when the current video block is coded using bi-prediction. The method according to any one of clauses 10 to 13.

[0471] 15. The method is applicable only when the current video block is coded using uni-prediction. The method according to any one of clauses 10 to 13.

[0472] 16. The value of M, N, or K is a function of the uni-prediction or bi-prediction mode of the current video block. The method according to any one of clauses 10 to 13.

[0473] 17. The value of M, N, or K is a function of the height or width of the current video block. The method according to any one of clauses 10 to 13.

[0474] 18. The trigger is included in the bitstream representation at the video parameter set level or the sequence parameter set level or the picture parameter set level or the slice header level or the tile group header level or the tile level or the coding tree unit level or the coding unit level or the prediction unit level, The method according to any one of items 10 to 17.

[0475] 19. The trigger conveys a value of M, N, or K, The method according to item 18.

[0476] 20. The one or more sub-blocks of the current video block are calculated based on the type of affine coding used for the current video block, The method according to any one of items 10 to 19.

[0477] 21. Two different methods are used to calculate sub-blocks for the single prediction and dual prediction affine prediction modes, The method according to item 20.

[0478] 22. When the current video block is a dual-predicted affine block, the width or height of the sub-blocks from different reference lists is different, The method according to item 21.

[0479] 23. The one or more sub-blocks correspond to the luma component, The method according to any one of items 20 to 22.

[0480] 24. The width and height of one of the one or more sub-blocks are determined using the motion vector difference between the motion vector value of the current video block and that of the one of the one or more sub-blocks, The method according to any one of items 10 to 23.

[0481] 25. The step of calculating is based on the pixel accuracy conveyed in the bitstream representation. The method according to any one of clauses 20 to 23.

[0482] 26. A method for video processing, determining that a current video block satisfies a size condition; executing a conversion between the bitstream representation and pixel values of the current video block by excluding a dual-prediction encoding mode for the current video block based on the determination; and a method having the above.

[0483] 27. A method for video processing, determining that a current video block satisfies a size condition; executing a conversion between the bitstream representation and pixel values of the current video block based on the determination; and having, wherein the inter-prediction mode is conveyed in the bitstream representation according to the size condition. A method.

[0484] 28. A method for video processing, determining that a current video block satisfies a size condition; executing a conversion between the bitstream representation and pixel values of the current video block based on the determination; and having, wherein the generation of the merge candidate list during the conversion depends on the size condition. A method.

[0485] 29. A method for video processing, determining that a child coding unit of a current video block satisfies a size condition; executing a conversion between the bitstream representation and pixel values of the current video block based on the determination; having The coding tree splitting process used to generate the sub-coding unit depends on the size condition, method.

[0486] Assuming that 30.w is the width and h is the height, the size condition is as follows (a) w is equal to T1 and h is equal to T2, or h is equal to T1 and w is equal to T2, (b) w is equal to T1 and h is not greater than T2, or h is equal to T1 and w is not greater than T2, (c) w is not greater than T1 and h is not greater than T2, or h is not greater than T1 and w is not greater than T2, is one of The method according to any one of clauses 26 to 29.

[0487] 31. T1 = 8 and T2 = 8, or T1 = 8 and T2 = 4, or T1 = 4 and T2 = 4, or T1 = 4 and T2 = 16, The method according to clause 30.

[0488] 32. The conversion includes generating the bitstream representation from the pixel values of the current video block, or generating the pixel values of the current video block from the bitstream representation, The method according to any one of clauses 26 to 29.

[0489] 33. A method for video processing, determining a weight index of a generalized bi-prediction (GBi) process for a current video block based on the position of the current video block; executing a conversion between the current video block and its bitstream representation using the weight index to implement the GBi process and having a method.

[0490] 34. The conversion includes generating the bitstream representation from the pixel values of the current video block or generating the pixel values of the current video block from the bitstream representation. The method according to item 33.

[0491] 35. The determining step includes, for the current video block at the first position, inheriting or predicting other weight indices of adjacent blocks, and for the current video block at the second position, calculating the GBi without inheriting from the adjacent blocks. The method according to any one of items 33 or 34.

[0492] 36. The second position has the current video block located in a coding tree unit different from the adjacent block. The method according to item 35.

[0493] 37. The second position corresponds to the current video block being in a coding tree unit line or a coding tree unit row different from the adjacent block. The method according to item 35.

[0494] 38. A method for video processing, determining that the current video block is coded as an intra-inter prediction (IIP) coding block, executing a conversion between the current video block and its bitstream representation using a simplification rule for determining the intra prediction mode or the most probable mode (MPM) of the current video block, and a method including the above.

[0495] 39. The conversion includes generating the bitstream representation from the pixel values of the current video block or generating the pixel values of the current video block from the bitstream representation. The method according to clause 38.

[0496] 40. The simplification rule is defined to determine the intra prediction coding mode of the current video block that is intra-inter prediction (IIP) coded so as to be independent of the other intra prediction coding modes of adjacent video blocks. The method according to any one of clauses 38 to 39.

[0497] 41. The intra prediction coding mode is represented in the bitstream representation using coding that is independent of that of adjacent blocks. The method according to any one of clauses 38 to 39.

[0498] 42. The simplification rule is defined to prioritize the selection that favors the coding mode of an intra-coded block over that of an intra-prediction-coded block. The method according to any one of clauses 38 to 40.

[0499] 43. The simplification rule is defined to determine the MPM by inserting the intra prediction mode from an intra-coded adjacent block before inserting the intra prediction mode from an IIP-coded adjacent block. The method according to clause 38.

[0500] 44. The simplification rule is defined to use the same configuration process as that used for other normal intra-coded blocks to determine the MPM. The method according to clause 38.

[0501] 45. A method for video processing, determining that a current video block satisfies a simplification criterion; performing the conversion between the current video block and the bitstream representation by disabling use of an inter-intra prediction mode for the conversion or by disabling additional coding tools used for the conversion; A method having

[0502] 46. The conversion includes generating the bitstream representation from pixel values of the current video block or generating pixel values of the current video block from the bitstream representation, The method according to clause 45.

[0503] 47. The simplification criterion includes that the width or height of the current video block is equal to T1, assuming that T1 is an integer, The method according to any one of clauses 45 to 46.

[0504] 48. The simplification criterion includes that the width or height of the current video block is greater than T1, assuming that T1 is an integer, The method according to any one of clauses 45 to 46.

[0505] 49. The simplification criterion includes that the width or height of the current video block is equal to T1 and the height of the current video block is equal to T2, The method according to any one of clauses 45 to 46.

[0506] 50. The simplification criterion defines that the current video block uses a bi-prediction mode, The method according to any one of clauses 45 to 46.

[0507] 51. The additional coding tool includes bi-directional optical flow (BIO) coding, The method according to any one of clauses 45 to 46.

[0508] 52. The additional coding tool includes an overlapping block motion compensation mode, the method according to any one of clauses 45 to 46.

[0509] 53. A method for video processing, executing a conversion between a current video block and a bitstream representation of the current video block using an encoding process based on motion vectors, (a) Precision P1 is used to store a spatial motion prediction result, precision P2 is used to store a temporal motion prediction result during the conversion process, P1 and P2 are fractions, or (b) Precision Px is used to store an x motion vector, precision Py is used to store a y motion vector, Px and Py are fractions, method.

[0510] 54. P1, P2, Px, and Py are different numbers, the method according to clause 53.

[0511] 55. P1 is 1 / 16 luma pixel, P2 is 1 / 4 luma pixel, or P1 is 1 / 16 luma pixel, P2 is 1 / 8 luma pixel, or P1 is 1 / 8 luma pixel, P2 is 1 / 4 luma pixel, or P1 is 1 / 8 luma pixel, P2 is 1 / 8 luma pixel, or P2 is 1 / 16 luma pixel, P1 is 1 / 4 luma pixel, or P2 is 1 / 16 luma pixel, P1 is 1 / 8 luma pixel, or P2 is 1 / 8 luma pixel, P1 is 1 / 4 luma pixel, the method according to clause 54.

[0512] 56. P1 and P2 are different for different pictures in different temporal layers included in the bitstream representation, the method according to clauses 53 to 54.

[0513] 57. The calculated motion vector is processed through an accuracy correction process before storage as the temporal motion prediction, the method according to clauses 53 to 54.

[0514] 58. The storage includes storing the x motion vector and the y motion vector as N-bit integers, the value range of the x motion vector is [MinX, Max], the value range of the y motion vector is [MinY, MaxY], and their ranges are a. minX is equal to MinY, b. MaxX is equal to MaxY, c. {MinX, MaxX} depends on Px, d. {MinY, MaxY} depends on Py, e. {MinX, MaxX, MinY, MaxY} depends on N, f. {MinX, MaxX, MinY, MaxY} is different for the stored MV for spatial motion prediction and other MVs stored for temporal motion prediction, g. {MinX, MaxX, MinY, MaxY} is different for pictures in different temporal layers, h. {MinX, MaxX, MinY, MaxY} is different for pictures having different widths or heights, i. {MinX, MaxX} is different for pictures having different widths, j. {MinY, MaxY} is different for pictures having different heights, k. MVx is clipped to [MinX, MaxX] before storage for spatial motion prediction, l. MVx is clipped to [MinX, MaxX] before storage for temporal motion prediction, m. MVy is clipped to [MinY, MaxY] before storage for spatial motion prediction, n.MVy is clipped to [MinY, MaxY] before storage for time motion prediction, satisfying one or more of the following, the method according to clauses 53 to 54.

[0515] 59. A method of video processing, assuming that W1, W2, H1, H2, and PW and PH are integers, fetching (W2 + N - 1 - PW) × (H2 + N - 1 - PH) blocks, pixel padding the fetched blocks, performing boundary pixel repetition on the pixel-padded blocks, and obtaining pixel values of small sub-blocks, thereby interpolating the small sub-blocks of size W1 × H1 within a large sub-block of size W2 × H2 of the current video block; performing conversion between the current video block and a bitstream representation of the current video block using the interpolated pixel values of the small sub-blocks; A method having the above steps.

[0516] 60. The conversion includes generating the current video block from the bitstream representation or generating the bitstream representation from the current sub-block. The method according to clause 59.

[0517] 61. W2 = H2 = 8, W1 = H1 = 4, and PW = PH = 0, The method according to any one of clauses 59 to 60.

[0518] 62. A method of video processing, During conversion of a current video block of W × H dimensions and a bitstream representation of the current video block, fetching (W + N - 1 - PW) × (H + N - 1 - PH) reference pixels and performing the motion compensation operation by padding reference pixels larger than the fetched reference pixels during the motion compensation operation; Performing a conversion between the current video block and a bitstream representation of the current video block using the result of the motion compensation operation comprising wherein W, H, N, PW, and PH are integers Method

[0519] 63. The conversion includes generating the current video block from the bitstream representation or generating the bitstream representation from the current sub-block The method according to clause 62

[0520] 64. Padding includes repeating the left or right boundary of the fetched pixels The method according to any one of clauses 62 to 63

[0521] 65. Padding includes repeating the upper or lower boundary of the fetched pixels The method according to any one of clauses 62 to 63

[0522] 66. Padding includes setting pixel values to a constant The method according to any one of clauses 62 to 63

[0523] 67. The rule determines that the same arithmetic coding context used for other intra-coded blocks is used during the conversion The method according to clause 38

[0524] 68. The conversion of the current video block excludes using the MPM of the current video block The method according to clause 38

[0525] 69. The simplification rule determines that only DC and planar modes are used for the bitstream representation of the current video block which is the IIP-coded block The method according to item 38.

[0526] 70. The simplification rule defines different intra prediction modes for the luma component and the chroma component. The method according to item 38.

[0527] 71. The subset of the MPM is used for the current video block to be IIP-coded. The method according to item 44.

[0528] 72. The simplification rule indicates that the MPM is selected based on the intra prediction modes included in the MPM list. The method according to item 38.

[0529] 73. The simplification rule indicates that a subset of the MPM should be selected from the MPM list and notifies the mode index associated with the subset. The method according to item 38.

[0530] 74. The context used for coding the intra MPM mode is used for coding the intra mode of the current video block to be IIP-coded. The method according to item 38.

[0531] 75. Equal weights are used for the intra prediction block and the inter prediction block generated for the current video block which is an IIP-coded block. The method according to item 44.

[0532] 76. Zero weights are used for the positions in the IIP coding process for the current video block. The method according to item 44.

[0533] 77. The zero weights are applied to the intra prediction blocks used in the IIP coding process. The method according to item 76.

[0534] 78. The zero weight is applied to an inter-prediction block used in the IIP coding process, The method according to item 76.

[0535] 79. A method for video processing, determining, based on the size of a current video block, that dual prediction or single prediction of the current video block is not permitted; performing a conversion between a bitstream representation and pixel values of the current video block by disabling a dual prediction or single prediction mode based on the determination; and a method having the above steps. For example, the non-permitted mode is not used for encoding or decoding the current video block. The conversion operation may represent either video coding or compression, or video decoding or decompression.

[0536] 80. The current video block is 4×8, and the determining step includes determining that dual prediction is not permitted, The method according to item 79. Other examples are given in Example 5.

[0537] 81. The current video block is 4×8 or 8×4, and the determining step includes determining that dual prediction is not permitted, The method according to item 79.

[0538] 82. The current video block is 4×N, where N is an integer less than or equal to 16, and the determining step includes determining that dual prediction is not permitted, The method according to item 79.

[0539] 83. The size of the current video block corresponds to the size of a color component or a luma component of the current video block, The method according to any one of clauses 26 to 29 or 79 to 82.

[0540] 84. Disabling the dual prediction or the CA prediction mode is applied to all three components of the current video block, The method according to clause 83.

[0541] 85. Disabling the dual prediction or the CA prediction mode is applied only to the color component for which the size is used as the size of the current video block, The method according to clause 83.

[0542] 86. The conversion is performed by disabling the use of dual prediction, and further, the merge candidates for dual prediction, and then allocating only one motion vector from only one reference list to the current video block. The method according to any one of clauses 79 to 85.

[0543] 87. The current video block is 4×4, and the determining step includes determining that neither dual prediction nor uni - prediction is permitted. The method according to clause 79.

[0544] 88. The current video block is coded as an intra block. The method according to clause 87.

[0545] 89. The current video block is restricted to using integer - pixel motion vectors. The method according to clause 87.

[0546] Further examples and embodiments for clauses 78 to 89 are described in Example 5.

[0547] 90. A method for video processing, Determining video coding conditions for a video block based on the size of the current video block; Performing a conversion between the current video block and the bitstream representation of the current video block based on the video coding conditions A method having

[0548] 91. The video coding conditions define selective signaling of a skip flag or an intra block coding flag in the bitstream representation The method according to clause 90

[0549] 92. The video coding conditions define selective signaling of a prediction mode for the current video block The method according to clause 90 or 91

[0550] 93. The video coding conditions define selective signaling of triangular mode coding of the current video block The method according to any one of clauses 90 to 92

[0551] 94. The video coding conditions define selective signaling of an inter prediction direction for the current video block The method according to any one of clauses 90 to 93

[0552] 95. The video coding conditions define selectively changing a motion vector or a block vector used for intra block copy of the current video block The method according to any one of clauses 90 to 94

[0553] 96. The video coding conditions depend on the height of the pixels of the current video block The method according to any one of clauses 90 to 95

[0554] 97. The video coding conditions depend on the width of the pixels of the current video block The method according to any one of clauses 90 to 96.

[0555] 98. The video coding condition depends on whether the current video block is square-shaped. The method according to any one of clauses 90 to 95.

[0556] Further examples of clauses 90 to 98 are provided in Examples 11 to 16 in "4. Examples of Embodiments" of this specification.

[0557] 99. A video encoder device having a processor configured to execute the method described in one or more of clauses 1 to 98.

[0558] 100. A video decoder device having a processor configured to execute the method described in one or more of clauses 1 to 98.

[0559] 101. A computer-readable medium storing code that causes a processor to execute the method described in any one or more of clauses 1 to 98 when executed by the processor.

[0560] FIG. 16 is a block diagram showing an exemplary video processing system 1600 in which various techniques disclosed herein may be implemented. For various implementations, some or all of the components of system 1600 may be included. System 1600 may include an input unit 1602 that receives video content. The video content may be received in a raw or uncompressed format, for example, 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input unit 1602 may correspond to a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0561] System 1600 may include a coding component 1604 that can implement various coding or encoding methods described herein. The coding component 1604 may reduce the average bit rate of the video from the input unit 1602 to the output unit of the coding component 1604 so as to generate a coded representation of the video. Accordingly, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the coding component 1604 may be stored or transmitted via a connected communication as represented by the component 1606. The stored or communicated bitstream (or coded) representation of the video received at the input unit 1602 may be used by a component 1608 that generates a displayable video that is sent to the pixel values or the display interface 1610. The process of generating a video that a user can view from the bitstream is sometimes referred to as video decompression. Further, while certain video processing operations are referred to as "coding" operations or tools, it will be understood that such coding tools or operations are used in an encoder and the corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.

[0562] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI (registered trademark)) or a Displayport (registered trademark), etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, an IDE interface, etc. The techniques described herein may be embodied in various electronic devices such as a cellular phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.

[0563] FIG. 17 is a flowchart representation of a method 1700 for video processing according to the present disclosure. Method 1700 includes, at operation 1702, determining that a first motion vector of a sub-block of a current block and a second motion vector that is a representative motion vector of the current block comply with size constraints for conversion between the current block of the video and a bitstream representation of the video using an affine coding tool. Method 1700 also includes, at operation 1704, performing the conversion based on the determination.

[0564] In some embodiments, the first motion vector of the sub-block is represented as (MVx, MVy), and the second motion vector is represented as (MV'x, MV'y). The size constraints indicate that MVx >= MV'x - DH0, MVx <= MV'x + DH1, MVy >= MV'y - DV0, and MVy <= MV'y + DV1, where DH0, DH1, DV0, and DV1 are positive numbers. In some embodiments, DH0 = DH1. In some embodiments, DH0 ≠ DH1. In some embodiments, DV0 = DV1. In some embodiments, DV0 ≠ DV1. In some embodiments, DH0 = DV0. In some embodiments, DH0 ≠ HV0. In some embodiments, DH1 = DV1. In some embodiments, DH1 ≠ DV1.

[0565] In some embodiments, at least one of DH0, DH1, DV0, and DV1 is signaled in the bitstream representation at the video parameter set level, sequence parameter set level, picture parameter set level, slice header, tile group header, tile level, coding tree unit level, coding unit level, or prediction unit level. In some embodiments, DH0, DH1, DV0, and DV1 are different for different profiles, levels, or tiers of the transform. In some embodiments, DH0, DH1, DV0, and DV1 are based on the width or height of the current block. In some embodiments, DH0, DH1, DV0, and DV1 are based on the prediction mode of the current block, where the prediction mode is a single prediction mode or a dual prediction mode. In some embodiments, DH0, DH1, DV0, and DV1 are based on the position of sub-blocks within the current block.

[0566] In some embodiments, the second motion vector has the control point motion vector of the current block. In some embodiments, the second motion vector has the motion vector of the second sub-block of the current block. In some embodiments, the second sub-block has the central sub-block of the current block. In some embodiments, the second sub-block has the corner sub-block of the current block. In some embodiments, the second motion vector has a motion vector derived for a position inside or outside the current block, where the position is coded using the same affine model as the current block. In some embodiments, the position has the central position of the current block.

[0567] In some embodiments, the first motion vector is adjusted to satisfy a size constraint. In some embodiments, the bitstream is not valid if the first motion vector cannot satisfy the size constraint with respect to the second motion vector. In some embodiments, the first motion vector and the second motion vector are represented according to the motion vector signaling accuracy in the bitstream representation. In some embodiments, the first motion vector and the second motion vector are represented according to the storage accuracy for storing the motion vector. In some embodiments, the first motion vector and the second motion vector are represented according to an accuracy different from the motion vector signaling accuracy or the storage accuracy for storing the motion vector.

[0568] FIG. 18 is a flowchart representation of a video processing method 1800 according to the present disclosure. Method 1800 includes, in operation 1802, determining an affine model having six parameters for conversion between a current block of a video and a bitstream representation of the video. The affine model is inherited from the affine coding information of adjacent blocks of the current block. Method 1800 includes, in operation 1804, performing the conversion based on the affine model.

[0569] In some embodiments, the adjacent block is coded using a second affine model having six parameters. The affine model is the same as the second affine model. In some embodiments, the adjacent block is coded using a third affine model having four parameters. In some embodiments, the affine model is determined based on the position of the current block. In some embodiments, the affine model is determined according to the third affine model when the adjacent block is not in the same coding tree unit (CTU) as the current block. In some embodiments, the affine model is determined according to the third affine model when the adjacent block is not in the same CTU line or the same CTU row as the current block.

[0570] In some embodiments, a tile, slice, or picture is divided into a plurality of non-overlapping regions. In some embodiments, the affine model is determined according to a third affine model when an adjacent block is not in the same region as the current block. In some embodiments, the affine model is determined according to a third affine model when an adjacent block is not in the same region line or the same region row as the current block. In some embodiments, each region has a size of 64×64. In some embodiments, the upper left corner of the current block is represented as (x, y), the upper left corner of the adjacent block is represented as (x', y'), and the affine model is determined according to a third affine model when conditions regarding x, y, x', and y' are satisfied. In some embodiments, the condition indicates that x / M ≠ x' / M, where M is a positive integer. In some embodiments, M is 128 or 64. In some embodiments, the condition indicates that y / N ≠ y' / N, where N is a positive integer. In some embodiments, N is 128 or 64. In some embodiments, the condition indicates that x / M ≠ x' / M and y / N ≠ y' / N, where M and N are positive integers. In some embodiments, M = N = 128 or M = N = 64. In some embodiments, the condition indicates that x >> M ≠ x' >> M, where M is a positive integer. In some embodiments, M is 6 or 7. In some embodiments, the condition indicates that y >> N ≠ y' >> N, where N is a positive integer. In some embodiments, N is 6 or 7. In some embodiments, the condition indicates that x >> M ≠ x' >> M and y >> N ≠ y' >> N, where M and N are positive integers. In some embodiments, M = N = 6 or M = N = 7.

[0571] FIG. 19 is a flowchart representation of a video processing method 1900 according to the present disclosure. The method 1900 includes, at operation 1902, determining whether dual prediction coding techniques are applicable to a block of the video for conversion between the block of the video and a bitstream representation of the video, based on the size of the block having a width W and a height H, where W and H are positive integers. The method 1900 includes, at operation 1904, performing the conversion according to the determination.

[0572] In some embodiments, the dual prediction coding techniques are not applicable when W = T1 and H = T2, where T1 and T2 are positive integers. In some embodiments, the dual prediction coding techniques are not applicable when W = T2 and H = T1, where T1 and T2 are positive integers. In some embodiments, the dual prediction coding techniques are not applicable when W = T1 and H ≤ T2, where T1 and T2 are positive integers. In some embodiments, the dual prediction coding techniques are not applicable when W ≤ T2 and H = T1, where T1 and T2 are positive integers. In some embodiments, T1 = 4 and T2 = 16. In some embodiments, the dual prediction coding techniques are not applicable when W ≤ T1 and H ≤ T2, where T1 and T2 are positive integers. In some embodiments, T1 = T2 = 8. In some embodiments, T1 = 8 and T2 = 4. In some embodiments, T1 = T2 = 4. In some embodiments, T1 = 4 and T2 = 8.

[0573] In some embodiments, an indicator indicating information about the dual prediction coding technique is signaled in the bitstream when the dual prediction coding technique is applicable. In some embodiments, an indicator indicating information about the dual prediction coding technique for a block is removed from the bitstream when the dual prediction coding technique is not applicable to that block. In some embodiments, the dual prediction coding technique is not applicable when the block size is one of 4×8 or 8×4. In some embodiments, the dual prediction coding technique is not applicable when the block size is 4×N or N×4, where N is a positive integer and N≦16. In some embodiments, the block size corresponds to the first color component of the block, and whether the dual prediction coding technique is applicable is determined for the first color component and the remaining color components of the block. In some embodiments, the block size corresponds to the first color component of the block, and whether the dual prediction coding technique is applicable is determined only for the first color component. In some embodiments, the first color component includes a luma component.

[0574] In some embodiments, the method further comprises assigning a single motion vector from the first reference list or the second reference list when it is determined that the selected merge candidate is to be coded using the bi-prediction coding technique when the bi-prediction coding technique is not applicable to the current block. In some embodiments, the method further comprises determining that the triangular prediction mode is not applicable to the block when the bi-prediction coding technique is not applicable to the current block. In some embodiments, whether the bi-prediction coding technique is applicable is associated with a prediction direction, the prediction direction is further associated with the uni-prediction coding technique, and the prediction direction is signaled in the bitstream based on the size of the block. In some embodiments, information regarding the uni-prediction coding technique is signaled in the bitstream when (1) W×H<64 or (2) W×H = 64 and W is not equal to H. In some embodiments, information regarding the uni-prediction coding technique or the bi-prediction coding technique is signaled in the bitstream when (1) W×H>64 or (2) W×H = 64 and W is equal to H.

[0575] In some embodiments, the limitation indicates that neither the bi-prediction coding technique nor the uni-prediction coding technique is applicable to the block when the size of the block is 4×4. In some embodiments, the limitation is applicable when the block is coded in affine. In some embodiments, the limitation is applicable when the block is not coded in affine. In embodiments, the limitation is applicable when the block is coded intra. In some embodiments, the limitation is not applicable when the motion vector of the block has integer precision.

[0576] In some embodiments, notifying that a block is generated based on the splitting of a parent block is skipped in the bitstream. The parent block has a size of (1) 8×8 for quadtree splitting, (2) 8×4 or 4×8 for binary tree splitting, or (3) 4×16 or 16×4 for ternary tree splitting. In some embodiments, an indicator indicating that a motion vector has integer precision is set to 1 in the bitstream. In some embodiments, the motion vector of a block is rounded to integer precision.

[0577] In some embodiments, the dual prediction coding technique is applicable to a block. The reference block has a size of (W+N-1-PW)×(H+N-1-PH), and the boundary pixels of the reference block are repeated to generate a second block of size (W+N-1)×(H+N-1) for the interpolation operation, where N represents the interpolation filter taps, and N, PW, and PH are integers. In some embodiments, PH = 0, and at least the left or right boundary pixels are repeated to generate the second block. In some embodiments, PW = 0, and at least the upper or lower boundary pixels are repeated to generate the second block. In some embodiments, PW>0 and PH>0, and the second block is generated by repeating at least the left or right boundary pixels, and then repeating at least the upper or lower boundary pixels. In some embodiments, PW>0 and PH>0, and the second block is generated by repeating at least the upper or lower boundary pixels, and then repeating at least the left or right boundary pixels. In some embodiments, the left boundary pixels are repeated M1 times, and the right boundary pixels are repeated (PW-M1) times. In some embodiments, the upper boundary pixels are repeated M2 times, and the lower boundary pixels are repeated (PH-M2) times. In some embodiments, how the boundary pixels of the reference pixels are repeated is applied to some or all of the reference blocks for transformation. In some embodiments, PW and PH are different for different components of the block.

[0578] In some embodiments, the merge candidate list construction process is performed based on the size of the block. In some embodiments, a merge candidate is considered a uni-prediction candidate that refers to the first reference list in the uni-prediction coding technique when (1) the merge candidate is coded using the bi-prediction coding technique and (2) bi-prediction is not applicable to that block according to the size of the block. In some embodiments, the first reference list has reference list 0 or reference list 1 of the uni-prediction coding technique. In some embodiments, a merge candidate is considered unavailable when (1) the merge candidate is coded using the bi-prediction coding technique and (2) bi-prediction is not applicable to that block according to the size of the block. In some embodiments, unavailable merge candidates are removed from the merge candidate list in the merge candidate list construction process. In some embodiments, the merge candidate list construction process for the triangular prediction mode is invoked when bi-prediction is not applicable to that block according to the size of the block.

[0579] FIG. 20 is a flowchart representation of a video processing method 2000 according to the present disclosure. Method 2000 includes, in operation 2002, determining, according to a coding tree splitting process, whether the coding tree splitting process is applicable to a block for conversion between a block of video and a bitstream representation of the video, based on the size of a sub-block that is a child coding unit of the block according to the coding tree splitting process. The sub-block has a width W and a height H, where W and H are positive integers. Method 2000 also includes, in operation 2004, performing the conversion according to the determination.

[0580] In some embodiments, the coding tree splitting process is not applicable when W = T1 and H = T2, where T1 and T2 are positive integers. In some embodiments, the coding tree splitting process is not applicable when W = T2 and H = T1, where T1 and T2 are positive integers. In some embodiments, the coding tree splitting process is not applicable when W = T1 and H ≤ T2, where T1 and T2 are positive integers. In some embodiments, the coding tree splitting process is not applicable when W ≤ T2 and H = T1, where T1 and T2 are positive integers. In some embodiments, T = 4 and T2 = 16. In some embodiments, the coding tree splitting process is not applicable when W ≤ T1 and H ≤ T2, where T1 and T2 are positive integers. In some embodiments, T1 = T2 = 8. In some embodiments, T1 = 8 and T2 = 4. In some embodiments, T1 = T2 = 4. In some embodiments, T1 = 4. In some embodiments, T2 = 4. In some embodiments, the signaling of the coding tree splitting process is removed from the bitstream if the coding tree splitting process is not applicable to the current block.

[0581] FIG. 21 is a flowchart representation of a video processing method 2100 according to the present disclosure. At operation 2102, it includes determining whether an index of a coding unit level weighted bi-prediction (BCW) coding mode is derived based on rules regarding the position of the current block for conversion between the current block of the video and the bitstream representation of the video. In the BCW coding mode, a set of weights including a plurality of weights is used to generate the bi-predicted value of the current block. Method 2100 also includes, at operation 2104, performing the conversion according to the determination.

[0582] In some embodiments, the dual prediction value of the current block is generated as a non-average weighted sum of two motion vectors when at least one weight in the weight set is applied. In some embodiments, the rule determines that the index is not derived from the adjacent block if the current block and the adjacent block are located in different coding tree units or maximum coding units. In some embodiments, the rule determines that the index is not derived from the adjacent block if the current block and the adjacent block are located in different lines or rows within a coding tree unit. In some embodiments, the rule determines that the index is not derived from the adjacent block if the current block and the adjacent block are located in different non-overlapping regions of a video tile, slice, or picture. In some embodiments, the rule determines that the index is not derived from the adjacent block if the current block and the adjacent block are located in different rows of non-overlapping regions of a video tile, slice, or picture. In some embodiments, each region has a size of 64×64.

[0583] In some embodiments, the upper corner of the current block is represented as (x, y), and the upper corner of the adjacent block is represented as (x’, y’). The rule is defined such that when (x, y) and (x’, y’) satisfy the condition, the index is not derived according to the adjacent block. In some embodiments, the condition indicates that x / M ≠ x’ / M, where M is a positive integer. In some embodiments, M is 128 or 64. In some embodiments, the condition indicates that y / N ≠ y’ / N, where N is a positive integer. In some embodiments, N is 128 or 64. In some embodiments, the condition indicates that (x / M ≠ x’ / M) and (y / N ≠ y’ / N), where M and N are positive integers. In some embodiments, M = N = 128 or M = N = 64. In some embodiments, the condition indicates that x >> M ≠ x’ >> M, where M is a positive integer. In some embodiments, M is 6 or 7. In some embodiments, the condition indicates that y >> N ≠ y’ >> N, where N is a positive integer. In some embodiments, N is 6 or 7. In some embodiments, the condition indicates that (x >> M ≠ x’ >> M) and (y >> N ≠ y’ >> N), where M and N are positive integers. In some embodiments, M = N = 6 or M = N = 7.

[0584] In some embodiments, whether the BCW coding mode is applicable to a picture, slice, tile group, or tile is notified in the picture parameter set, slice header, tile group header, or tile in the bitstream, respectively. In some embodiments, whether the BCW coding mode is applicable to a picture, slice, tile group, or tile is derived based on information associated with the picture, slice, tile group, or tile. In some embodiments, the information has at least a quantization parameter (QP), a temporal layer, or a picture order count (POC) distance.

[0585] FIG. 22 is a flowchart representation of a video processing method 2200 according to the present disclosure. The method 2200 includes, at operation 2202, determining an intra prediction mode of a current block of video coded using a Combined Inter and Intra Prediction (CIIP) coding technique independently from intra prediction modes of adjacent blocks for conversion between the current block of video and a bitstream representation of the video. The CIIP coding technique uses intermediate inter prediction values and intermediate intra prediction values to derive a final prediction value of the current block. The method 2200 also includes, at operation 2204, performing the conversion based on the determination.

[0586] In some embodiments, the intra prediction mode of the current block is determined without reference to the intra prediction mode of any adjacent blocks. In some embodiments, the adjacent blocks are coded using the CIIP coding technique. In some embodiments, the intra prediction of the current block is determined based on the intra prediction mode of a second adjacent block coded using an intra prediction coding technique. In some embodiments, whether to determine the intra prediction mode of the current block according to the second intra prediction mode is based on whether a condition defining a relationship between the current block as a first block and the second adjacent block as a second block is satisfied. In some embodiments, the determination is part of a most probable mode (MPM) construction process of the current block for deriving a list of MPM modes.

[0587] FIG. 23 is a flowchart representation of a video processing method 2300 according to the present disclosure. The method 2300 includes, in operation 2302, determining an intra prediction mode of the current block of the video coded using inter and intra composite prediction (CIIP) coding technology according to a first intra prediction mode of a first adjacent block and a second intra prediction mode of a second adjacent block for conversion between the current block of the video coded using the CIIP coding technology and a bitstream representation of the video. The first adjacent block is coded using intra prediction coding technology, and the second adjacent block is coded using CIIP coding technology. The first intra prediction mode is given a different priority than the second intra prediction mode. The CIIP coding technology uses an intermediate inter prediction value and an intermediate intra prediction value to derive a final prediction value of the current block. The method 2300 also includes, in operation 2304, performing the conversion based on the determination.

[0588] In some embodiments, the determining step is part of the most probable mode (MPM) construction process for the current block to derive a list of MPM modes. In some embodiments, the first intra prediction mode is positioned before the second intra prediction mode in the MPM candidate list. In some embodiments, the first intra prediction mode is positioned after the second intra prediction mode in the MPM candidate list. In some embodiments, the coding of the intra prediction mode bypasses the most probable mode (MPM) construction process of the current block. In some embodiments, the method also includes a step of determining an intra prediction mode of a subsequent block according to the intra prediction mode of the current block. Here, the subsequent block is coded using an intra prediction coding technique, and the current block is coded using a CIIP coding technique. In some embodiments, the determining step is part of the most probable mode (MPM) construction process of the subsequent block. In some embodiments, in the MPM construction process of the subsequent block, the intra prediction mode of the current block is given a lower priority than the intra prediction modes of other adjacent blocks coded using the intra prediction coding technique. In some embodiments, whether to determine the intra prediction mode of the subsequent block according to the intra prediction mode of the current block is based on whether a condition defining the relationship between the subsequent block as the first block and the current block as the second block is satisfied. In some embodiments, the condition has at least one of (1) the first block and the second block are located on the same line of the coding tree unit, (2) the first block and the second block are located in the same CTU, (3) the first block and the second block are located in the same region, or (4) the first block and the second block are on the same line of the region. In some embodiments, the width of the region is the same as the height of the region. In some embodiments, the region has a size of 64×64.

[0589] In some embodiments, only a subset of the list of most probable modes (MPMs) for normal intra coding techniques is used for the current block. In some embodiments, the subset has a single MPM mode within the list of MPM modes for normal intra coding techniques. In some embodiments, the single MPM mode is the first MPM mode in the list. In some embodiments, the index indicating the single MPM mode is omitted in the bitstream. In some embodiments, the subset has the first four MPM modes within the list of MPM modes. In some embodiments, the indices indicating the MPM modes within the subset are signaled in the bitstream. In some embodiments, the coding context for coding an intra-coded block is reused for coding the current block. In some embodiments, the first MPM flag for the intra-coded block and the second MPM flag for the current block share the same coding context in the bitstream. In some embodiments, the intra prediction mode of the current block is selected from the list of MPM modes regardless of the size of the current block. In some embodiments, the MPM configuration process is defaulted to be valid, and the flag indicating the MPM configuration process is omitted in the bitstream. In some embodiments, the MPM list configuration process is not required for the current block.

[0590] In some embodiments, the luma prediction chroma mode is used to process the chroma component of the current block. In some embodiments, the derived mode is used to process the chroma component of the current block. In some embodiments, a plurality of intra prediction modes are used to process the chroma component of the current block. In some embodiments, the plurality of intra prediction modes are used based on the color format of the chroma component. In some embodiments, when the color format is 4:4:4, the plurality of intra prediction modes are the same as the intra prediction mode for the luma component of the current block. In some embodiments, each of the four intra prediction modes is coded using one or more bits, and the four intra prediction modes include a planar mode, a DC mode, a vertical mode, and a horizontal mode. In some embodiments, the four intra prediction modes are coded using 00, 01, 10, and 11. In some embodiments, the four intra prediction modes are 0, They are coded using 10, 110, 111. In some embodiments, the four intra prediction modes are coded using 1, 01, 001, 000. In some embodiments, only a subset of the four intra prediction modes is available for use when the width W and height H of the current block satisfy the conditions. In some embodiments, the subset has a planar mode, a DC mode, and a vertical mode when W > N×N, where N is an integer. In some embodiments, the planar mode, the DC mode, and the vertical mode are coded using 1, 01, and 11. In some embodiments, the planar mode, the DC mode, and the vertical mode are coded using 0, 10, and 00. In some embodiments, the subset has a planar mode, a DC mode, and a horizontal mode when H > N×W, where N is an integer. In some embodiments, the planar mode, the DC mode, and the horizontal mode are coded using 1, 01, and 11. In some embodiments, the planar mode, the DC mode, and the horizontal mode are coded using 0, 10, and 00. In some embodiments, N = 2. In some embodiments, only the DC mode and the planar mode are used for the current block. In some embodiments, an indicator indicating the DC mode or the planar mode is signaled in the bitstream.

[0591] Figure 24 is a flowchart representation of a video processing method 2400 according to the present disclosure. The method 2400 includes, in operation 2402, determining whether an inter and intra combined prediction (CIIP) process is applicable to a color component of the current block for conversion between the current block of the video and the bitstream representation of the video, based on the size of the current block. The CIIP coding technique uses an intermediate inter prediction value and an intermediate intra prediction value to derive a final prediction value of the current block. The method 2400 also includes, in operation 2404, performing the conversion based on the determination.

[0592] In some embodiments, the color component has a chroma component, and the CIIP process is not performed on the chroma component when the width of the current block is less than 4. In some embodiments, the color component has a chroma component, and the CIIP process is not performed on the chroma component when the height of the current block is less than 4. In some embodiments, the intra prediction mode for the chroma component of the current block is different from the intra prediction mode for the luma component of the current block. In some embodiments, the chroma component uses one of a DC mode, a planar mode, or a luma prediction chroma mode. In some embodiments, the intra prediction mode for the chroma component is determined based on the color format of the chroma component. In some embodiments, the color format has 4:2:0 or 4:4:4.

[0593] FIG. 25 is a flowchart representation of a video processing method 2500 according to the present disclosure. The method 2500 includes, in operation 2502, determining whether an inter and intra combined prediction (CIIP) coding technique should be applied to a current block of a video for conversion between the current block of the video and a bitstream representation of the video, based on characteristics of the current block. The CIIP coding technique uses an intermediate inter prediction value and an intermediate intra prediction value to derive a final prediction value of the current block. The method 2500 also includes, in operation 2504, performing the conversion based on the determination.

[0594] In some embodiments, the characteristic has that the size of the current block is width W and height H, where W and H are integers, and the inter-intra prediction coding technique is disabled for the current block when the block size satisfies the condition. In some embodiments, the condition indicates that W is equal to T1, where T1 is an integer. In some embodiments, the condition indicates that H is equal to T1, where T1 is an integer. In some embodiments, T1 = 4. In some embodiments, T1 = 2. In some embodiments, the condition indicates that W is greater than T1 or H is greater than T1, where T1 is an integer. In some embodiments, T1 = 64 or 128. In some embodiments, the condition indicates that W is equal to T1 and H is equal to T2, where T1 and T2 are integers. In some embodiments, the condition indicates that W is equal to T2 and H is equal to T1, where T1 and T2 are integers. In some embodiments, T1 = 4 and T2 = 16.

[0595] In some embodiments, the characteristic has a coding technique applied to the current block, and the CIIP coding technique is disabled for the current block if the coding technique meets the condition. In some embodiments, the condition indicates that the coding technique is a bi-prediction coding technique. In some embodiments, the bi-prediction coded merge candidate is converted to a uni-prediction coded merge candidate to enable the inter-intra prediction coding technique to be applied to the current block. In some embodiments, the converted merge candidate is associated with reference list 0 of the uni-prediction coding technique. In some embodiments, the converted merge candidate is associated with reference list 1 of the uni-prediction coding technique. In some embodiments, only the uni-prediction coded merge candidates of the block are selected for conversion. In some embodiments, the bi-prediction coded merge candidate is discarded to determine the merge index indicating the merge candidate in the bitstream representation. In some embodiments, the inter-intra prediction coding technique is applied to the current block according to the determination. In some embodiments, the merge candidate list construction process for the triangular prediction mode is used to derive the motion candidate list for the current block.

[0596] FIG. 26 is a flowchart representation of a video processing method 2600 according to the present disclosure. Method 2600 includes, at operation 2602, determining whether coding tools should be disabled for the current block for conversion between the current block of video and the bitstream representation of the video, based on whether the current block is coded by inter and intra combined prediction (CIIP) coding technology. The CIIP coding technology uses an intermediate inter prediction value and an intermediate intra prediction value to derive the final prediction value of the current block. The coding tools include at least one of bi-directional optical flow (BDOF), overlapped block motion compensation (OBMC), or decoder-side motion vector refinement process (DMVR). Method 2500 also includes, at operation 2504, performing the conversion based on the determination.

[0597] In some embodiments, the intra prediction process for the current block is different from the intra prediction process for a second block coded using intra prediction coding technology. In some embodiments, in the intra prediction process for the current block, filtering of adjacent samples is skipped. In some embodiments, in the intra prediction process for the current block, position-dependent intra prediction sample filtering is disabled. In some embodiments, in the intra prediction process for the current block, the multi-line intra prediction process is disabled. In some embodiments, in the intra prediction process for the current block, the wide-angle intra prediction process is disabled.

[0598] FIG. 27 is a flowchart representation of a video processing method 2700 according to the present disclosure. Method 2700 includes, in operation 2702, determining a first prediction P1 used for a motion vector for spatial motion prediction and a second prediction P2 used for a motion vector for temporal motion prediction for conversion between a block of video and a bitstream representation of the video. P1 and / or P2 are fractions, and P1 and P2 are different from each other. Method 2700 also includes, in operation 2704, performing the conversion based on the determination.

[0599] In some embodiments, the first accuracy is 1 / 16 luma pixel and the second accuracy is 1 / 4 luma pixel. In some embodiments, the first accuracy is 1 / 16 luma pixel and the second accuracy is 1 / 8 luma pixel. In some embodiments, the first accuracy is 1 / 8 luma pixel and the second accuracy is 1 / 4 luma pixel. In some embodiments, the first accuracy is 1 / 16 luma pixel and the second accuracy is 1 / 4 luma pixel. In some embodiments, the first accuracy is 1 / 16 luma pixel and the second accuracy is 1 / 8 luma pixel. In some embodiments, the first accuracy is 1 / 8 luma pixel and the second accuracy is 1 / 4 luma pixel. In some embodiments, at least one of the first or second accuracy is lower than 1 / 16 luma pixel.

[0600] In some embodiments, at least one of the first or second accuracy is variable. In some embodiments, the first accuracy or the second accuracy is variable according to the profile, level, or tier of the video. In some embodiments, the first accuracy or the second accuracy is variable according to the temporal layer of a picture in the video. In some embodiments, the first accuracy or the second accuracy is variable according to the size of a picture in the video.

[0601] In some embodiments, at least one of the first or second precision is signaled in a video parameter set, sequence parameter set, picture parameter set, slice header, tile group header, tile, coding tree unit, or coding unit in the bitstream representation. In some embodiments, a motion vector is represented as (MVx, MVy), the precision of the motion vector is represented as (Px, Py), Px is associated with MVx, and Py is associated with MVy. In some embodiments, Px and Py are variable according to the profile, level, or tier of the video. In some embodiments, Px and Py are variable according to the temporal layer of a picture in the video. In some embodiments, Px and Py are variable according to the width of a picture in the video. In some embodiments, Px and Py are signaled in a video parameter set, sequence parameter set, picture parameter set, slice header, tile group header, tile, coding tree unit, or coding unit in the bitstream representation. In some embodiments, a decoded motion vector is represented as (MVx, MVy), and the motion vector is adjusted according to a second regime before the motion vector is stored as a temporal motion prediction motion vector. In some embodiments, the temporal motion prediction motion vector is adjusted to be (Shift(Mvx, P1 - P2), Shift(MVy, P1 - P2)), where P1 and P2 are integers, P1 ≥ P2, and Shift represents a right shift operation for an unsigned number. In some embodiments, the temporal motion prediction motion vector is adjusted to be (SignShift(MVx, P1 - P2), SignShift(MVy, P1 - P2)), where P1 and P2 are integers, P ≥ P2, and SignShfit represents a right shift operation for a signed number. In some embodiments, the temporal motion prediction motion vector is adjusted to be (MVx << (P1 - P2), MVy << (P1 - P2)), where P1 and P2 are integers, P1 ≥ P2, and << represents a left shift operation for a signed or unsigned number.

[0602] FIG. 28 is a flowchart representation of a video processing method 2800 according to the present disclosure. The method 2800 includes, in operation 2802, determining motion vectors (Mvx, MVy) by prediction (Px, Py) for conversion between a block of video and a bitstream representation of the video. Px is associated with MVx, and Py is associated with MVy. MVx and MVy are represented using N bits, where MinX ≦ MVx ≦ MaxX and MinY ≦ MVy ≦ MaxY, and MinX, MaxX, MinY, and MaxY are real numbers. The method 2800 also includes, in operation 2804, performing the conversion based on the determination.

[0603] In some embodiments, MinX = MinY. In some embodiments, MinX ≠ MinY. In some embodiments, MaxX = MaxY. In some embodiments, MaxX ≠ MaxY.

[0604] In some embodiments, at least one of MinX or MaxX is based on Px. In some embodiments, the motion vector has an accuracy represented by (Px, Py), and at least one of MinY or MaxY is based on Py. In some embodiments, at least one of MinX, MaxX, MinY, or MaxY is based on N. In some embodiments, at least one of MinX, MaxX, MinY, or MaxY for the spatial motion prediction motion vector is different from the corresponding MinX, MaxX, MinY, or MaxY for the temporal motion prediction motion vector. In some embodiments, at least one of MinX, MaxX, MinY, or MaxY is variable according to the profile, level, or tier of the video. In some embodiments, at least one of MinX, MaxX, MinY, or MaxY is variable according to the temporal layer of the picture in the video. In some embodiments, at least one of MinX, MaxX, MinY, or MaxY is variable according to the size of the picture in the video. In some embodiments, at least one of MinX, MaxX, MinY, or MaxY is signaled in a video parameter set, sequence parameter set, picture parameter set, slice header, tile group header, tile, coding tree unit, or coding unit in the bitstream representation. In some embodiments, MVx is clipped to [MinY, MaxX] before being used for spatial motion prediction or temporal motion prediction. In some embodiments, MVy is clipped to [MinY, MaxY] before being used for spatial motion prediction or temporal motion prediction.

[0605] FIG. 29 is a flowchart representation of a video processing method 2900 according to the present disclosure. The method 2900 includes, in operation 2902, determining whether a shared merge list is applicable to a current block of a video for conversion between the current block of the video and a bitstream representation of the video, according to a coding mode of the current block. The method 2900 also includes, in operation 2904, performing the conversion based on the determination.

[0606] In some embodiments, the shared merge list is not applicable when the current block is coded using a regular merge mode. In some embodiments, the shared merge list is applicable when the current block is coded using an Intra Block Copy (IBC) mode. In some embodiments, the method further includes, prior to performing the conversion, maintaining a motion candidate table based on past conversions of the video and the bitstream representation, and, after performing the conversion, disabling an update of the motion candidate table when the current block is a child of a parent block to which the shared merge list is applicable, where the current block is coded using a regular merge mode.

[0607] FIG. 30 is a flowchart representation of a video processing method 3000 according to the present disclosure. The method 3000 includes, in operation 3002, determining a second block of dimension (W+N-1)×(H+N-1) for motion compensation during conversion for conversion between a current block of a video having a size of W×H and a bitstream representation of the video. The second block is determined based on a reference block of dimension (W+N-1-PW)×(H+N-1-PH). N represents a filter size, and W, H, N, PW, and PH are non-negative integers. Neither PW nor PH is equal to 0. The method 3000 also includes, in operation 3004, performing the conversion based on the determination.

[0608] In some embodiments, the pixels in the second block located outside the reference block are determined by repeating one or more boundaries of the reference block. In some embodiments, PH = 0, and at least the left or right boundary of the reference block is repeated to generate the second block. In some embodiments, PW = 0, and at least the upper or lower boundary of the reference block is repeated to generate the second block. In some embodiments, PW>0 and PH>0, and the second block is generated by repeating at least the left or right boundary of the reference block and then repeating at least the upper or lower boundary of the reference block. In some embodiments, PW>0 and PH>0, and the second block is generated by repeating at least the upper or lower boundary of the reference block and then repeating at least the left or right boundary of the reference block.

[0609] In some embodiments, the left boundary of the reference block is repeated M1 times, the right boundary of the reference block is repeated (PW - M1) times, and M is a positive integer. In some embodiments, the upper boundary of the reference block is repeated M2 times, the lower boundary of the reference block is repeated (PH - M2) times, and M2 is a positive integer. In some embodiments, at least one of PW or PH is different for different color components of the current block, and the color components include at least a luma component or one or more chroma components. In some embodiments, at least one of PW or PH is variable according to the size or shape of the current block. In some embodiments, at least one of PW or PH is variable according to the coding characteristics of the current block, and the coding characteristics include single-prediction coding or dual-prediction coding.

[0610] In some embodiments, pixels in a second block located outside the reference block are set to a single value. In some embodiments, the single value is 1<<(BD-1), where BD is the bit depth of the pixel samples in the reference block. In some embodiments, BD is 8 or 10. In some embodiments, the single value is derived based on the pixel samples of the reference block. In some embodiments, the single value is signaled in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, a tile group header, a tile, a coding tree unit row, a coding tree unit, a coding unit, or a prediction unit. In some embodiments, the padding of the pixels in the second block located outside the reference block is disabled when the current block is affine coded.

[0611] FIG. 31 is a flowchart representation of a video processing method 3100 according to the present disclosure. Method 3100 includes, at operation 3102, determining a second block of dimension (W+N-1)×(H+N-1) for motion compensation during conversion for conversion between a current block of a video having a size of W×H and a bitstream representation of the video. W and H are non-negative integers, and N is a non-negative integer based on a filter size. During the conversion, refined motion vectors are determined based on a multi-point search according to a motion vector refinement operation on the original motion vectors, and the pixel length boundaries of the reference block are determined by repeating one or more non-boundary pixels. Method 3100 also includes, at operation 3104, performing the conversion based on the determination.

[0612] In some embodiments, processing the current block involves applying a filter to the current block in a motion vector refinement operation. In some embodiments, whether a reference block is applicable to the processing of the current block is determined based on the dimensions of the current block. In some embodiments, interpolating the current block involves interpolating a plurality of sub-blocks of the current block based on a second block. Each sub-block has a size of W1×H1, where W1 and H1 are non-negative integers. In some embodiments, W1 = H1 = 4, W = H = 8, and PW = PH = 0. In some embodiments, the second block is determined as a whole based on the integer part of the motion vector of at least one of the plurality of sub-blocks. In some embodiments, when the maximum difference between the integer parts of the motion vectors of all the plurality of sub-blocks is 1 pixel or less, the reference block is determined based on the integer part of the motion vector of the upper left sub-block of the current block, and each of the right and lower boundaries of the reference block is repeated once to obtain the second block. In some embodiments, when the maximum difference between the integer parts of the motion vectors of all the plurality of sub-blocks is 1 pixel or less, the reference block is determined based on the integer of the motion vector of the lower right sub-block of the current block, and each of the left and upper boundaries of the reference block is repeated once to obtain the second block. In some embodiments, the second block is determined as a whole based on a modified motion vector of one of the plurality of sub-blocks.

[0613] In some embodiments, when the maximum difference between the integer parts of the motion vectors of all the plurality of sub-blocks is 2 pixels or less, the motion vector of the upper left sub-block of the current block is modified by adding a distance of 1 integer pixel to each component to obtain a modified motion vector. The reference block is determined based on the modified motion vector, and each of the left, right, upper, and lower boundaries of the reference block is repeated once to obtain the second block.

[0614] In some embodiments, when the maximum difference between the integer parts of the motion vectors of all the plurality of sub-blocks is 2 pixels or less, the motion vector of the lower-right sub-block of the current block is changed by subtracting 1 integer pixel distance from each component so as to obtain a changed motion vector. The reference block is determined based on the changed motion vector, and each of the left boundary, right boundary, upper boundary, and lower boundary of the reference block is iterated once to obtain a second block.

[0615] FIG. 32 is a flowchart representation of a video processing method 3200 according to the present disclosure. The method 3200 includes, at operation 3202, determining a predicted value at a position within a block based on a weighted sum of an inter-predicted value and an intra-predicted value at that position for conversion between a block of video coded using an inter-intra composite prediction (CIIP) coding technique and a bitstream representation of the video. The weighted sum is based on adding an offset to an initial sum obtained based on the inter-predicted value and the intra-predicted value, and the offset is added before a right shift operation performed to determine the weighted sum. The method 3200 also includes, at operation 3204, performing the conversion based on the determination.

[0616] In some embodiments, a position within a block is represented as (x, y), the inter-predicted value at the position (x, y) is represented as Pinter(x, y), the intra-predicted value at the position (x, y) is represented as Pintra(x, y), the inter-prediction weight at the position (x, y) is represented as w_inter(x, y), and the intra-prediction weight at the position (x, y) is represented as w_intra(x, y). The predicted value at the position (x, y) is determined to be (Pintra(x, y) × w_intra(x, y) + Pinter(x, y) × w_inter(x, y) + offset(x, y)) >> N. Here, w_intra(x, y) + w_inter(x, y) = 2 N and offset(x, y) = 2 (N-1) and N is a positive integer. In some embodiments, N = 2.

[0617] In some embodiments, the weighted sum is determined using equal weights for the inter prediction value and the intra prediction value at that position. In some embodiments, a zero weight is used according to the position within the block to determine the weighted sum. In some embodiments, the zero weight is applied to the inter prediction value. In some embodiments, the zero weight is applied to the intra prediction value.

[0618] FIG. 33 is a flowchart representation of a video processing method 3300 according to the present disclosure. Method 3300 includes, at operation 3302, determining, based on whether conditions related to the size of the current block are satisfied, a manner in which coding information of the current block is represented in the bitstream representation for conversion between the current block of the video and the bitstream representation of the video. Method 3300 also includes, at operation 3304, performing the conversion based on the determination.

[0619] In some embodiments, the coding information is signaled in the bitstream representation if the conditions are satisfied. In some embodiments, the coding information is omitted in the bitstream representation if the conditions are satisfied. In some embodiments, if the conditions are satisfied, the Intra Block Copy (IBC) coding tool or the inter prediction coding tool is disabled for the current block. In some embodiments, the coding information is omitted in the bitstream representation if the IBC coding tool and the inter prediction coding tool are disabled. In some embodiments, the coding information is signaled in the bitstream representation if the IBC coding tool is enabled and the inter prediction coding tool is disabled.

[0620] In some embodiments, the coding information has an operation mode related to an IBC coding tool and / or an inter-prediction coding tool, and the operation mode has at least a normal mode or a skip mode. In some embodiments, the coding information has an indicator indicating the use of a prediction mode related to a coding tool. In some embodiments, the coding tool is not used when the coding information is omitted in the bitstream representation. In some embodiments, the coding tool is used when the coding information is omitted in the bitstream representation. In some embodiments, the coding tool has an intra-block copy (IBC) coding tool or an inter-prediction coding tool.

[0621] In some embodiments, when an indicator is omitted in the bitstream, one or more indicators indicating one or more other coding tools are signaled in the bitstream representation. In some embodiments, the one or more other coding tools have at least an intra-coding tool or a palette coding tool. In some embodiments, the indicator discriminates between an inter-mode and an intra-mode, and the one or more indicators include a first indicator indicating an IBC mode and a second indicator indicating a palette mode, which are signaled in the bitstream when the inter-prediction coding tool is disabled and the IBC coding tool is enabled.

[0622] In some embodiments, the coding information has a third indicator for a skip mode that is an inter skip mode or an IBC skip mode. In some embodiments, the third indicator for the skip mode is signaled when the inter prediction coding tool is disabled but the IBC coding tool is enabled. In some embodiments, the prediction mode is determined to be the IBC mode when the inter prediction coding tool is disabled and the IBC coding tool is enabled even though an indicator indicating the use of the IBC mode is omitted in the bitstream representation, and the skip mode is applicable to the current block.

[0623] In some embodiments, the coding information has a triangle mode. In some embodiments, the coding information has an inter prediction direction. In some embodiments, the coding information is omitted in the bitstream representation when the current block is coded using a uni-prediction coding tool. In some embodiments, the coding information includes an indicator indicating the use of the Symmetric Motion Vector Difference (SMVD) method. In some embodiments, when conditions are satisfied, the current block is set to be uni-predicted even though there is an indicator indicating the use of the SMVD method in the bitstream representation. In some embodiments, only list 0 or list 1 associated with the uni-prediction coding tool is used in the motion compensation process when an indicator indicating the use of the SMVD method is removed from the bitstream.

[0624] In some embodiments, the condition indicates that the width of the current block is T1 and the height of the current block is T2, where T1 and T2 are positive integers. In some embodiments, the condition indicates that the width of the current block is T2 and the height of the current block is T1, where T1 and T2 are positive integers. In some embodiments, the condition indicates that the width of the current block is T1 and the height of the current block is less than or equal to T2, where T1 and T2 are positive integers. In some embodiments, the condition indicates that the width of the current block is T2 and the height of the current block is less than or equal to T1, where T1 and T2 are positive integers. In some embodiments, the condition indicates that the width of the current block is less than or equal to T1 and the height of the current block is less than or equal to T2, where T1 and T2 are positive integers. In some embodiments, the condition indicates that the width of the current block is greater than or equal to T1 and the height of the current block is greater than or equal to T2, where T1 and T2 are positive integers. In some embodiments, the condition indicates that the width of the current block is greater than or equal to T1 or the height of the current block is greater than or equal to T2, where T1 and T2 are positive integers. In some embodiments, T1 = 4 and T2 = 16. In some embodiments, T1 = T2 = 8. In some embodiments, T1 = 8 and T2 = 4. In some embodiments, T1 = T2 = 4. In some embodiments, T1 = T2 = 32. In some embodiments, T1 = T2 = 64. In some embodiments, T1 = T2 = 128.

[0625] In some embodiments, the condition indicates that the current block has a size of 4×8, 8×4, 4×16, or 16×4. In some embodiments, the condition indicates that the current block has a size of 4×8 or 8×4. In some embodiments, the condition indicates that the size of the current block is 4×N or N×4, where N is a positive integer and N≦16. In some embodiments, the condition indicates that the size of the current block is 8×N or N×8, where N is a positive integer and N≦16. In some embodiments, the condition indicates that the color components of the current block contain N or fewer samples. In some embodiments, N = 16. In some embodiments, N = 32.

[0626] In some embodiments, the condition indicates that the color components of the current block have a width equal to the height. In some embodiments, the condition indicates that the color components of the current block have a size of 4×8, 8×4, 4×4, 4×16, or 16×4. In some embodiments, the color components have luma components or one or more chroma components. In some embodiments, the inter prediction tool is disabled if the condition indicates that the current block has a width of 4 and a height of 4. In some embodiments, the IBC prediction tool is disabled if the condition indicates that the current block has a width of 4 and a height of 4.

[0627] In some embodiments, list 0 associated with the uni-prediction coding tool is used in the motion compensation process when (1) the current block has a width of 4 and a height of 8, or (2) the current block has a width of 8 and a height of 4. The bi-prediction coding tool is disabled in the motion compensation process.

[0628] FIG. 34 is a flowchart representation of a video processing method 3400 according to the present disclosure. Method 3400 includes, in operation 3402, determining a modified motion vector set for conversion between a current block of video and a bitstream representation of the video. Method 3400 includes, in operation 3404, performing the conversion based on the modified motion vector set. The modified motion vector set is a modified version of the motion vector set associated with the current block by the current block satisfying a condition.

[0629] In some embodiments, the motion vector set is derived using one of a regular merge coring technique, a sub-block Temporal Motion Vector Prediction coding technique (sbTMVP), or a Merge with Motion Vector Difference (MMVD) coding technique. In some embodiments, the motion vector set has block vectors used in an Intra Block Copy (IBC) coding technique. In some embodiments, the condition specifies that the size of the luma component of the current block is the same as a predefined size. In some embodiments, the predefined size has at least one of 8×4 or 4×8.

[0630] In some embodiments, changing the motion vector set comprises converting the motion vector set to a uni-directional motion vector when it is determined that the motion information of the current block is bi-directional and has a first motion vector that references a first reference picture in a first reference picture list for a first prediction direction and a second motion vector that references a second reference picture in a second reference picture list for a second prediction direction. In some embodiments, the motion information of the current block is derived based on the motion information of adjacent blocks. In some embodiments, the method comprises discarding information for one of the first prediction direction or the second prediction direction. In some embodiments, discarding the information comprises changing the motion vector in the discarded prediction direction to (0,0). In some embodiments, discarding the information comprises changing the reference index in the discarded prediction direction to -1. In some embodiments, discarding the information comprises changing the discarded prediction direction to the other prediction direction.

[0631] In some embodiments, the method further comprises deriving a new motion candidate based on information for the first prediction direction and information for the second prediction direction. In some embodiments, deriving the new motion candidate comprises determining a motion vector of the new motion candidate based on an average of the first motion vector in the first prediction direction and the second motion vector in the second prediction direction. In some embodiments, deriving the new motion candidate comprises determining a scaled motion vector by scaling the first motion vector in the first prediction direction according to information for the second prediction direction, and determining a motion vector of the new motion candidate based on an average of the scaled motion vector and the second motion vector in the second prediction direction.

[0632] In some embodiments, the first reference picture list is List 0, and the second reference picture list is List 1. In some embodiments, the first reference picture list is List 1, and the second reference picture list is List 0. In some embodiments, the method includes updating a table of motion candidates determined based on past transformations using the transformed motion vectors. In some embodiments, the table has a History Motion Vector Prediction (HMVP) table.

[0633] In some embodiments, the method includes updating a table of motion candidates determined based on past transformations using information for a first prediction direction and a second prediction direction before changing a set of motion vectors. In some embodiments, the transformed motion vectors are used for a motion compensation operation for a current block. In some embodiments, the transformed motion vectors are used to predict motion vectors of subsequent blocks. In some embodiments, the transformed motion vectors are used for a deblocking process for a current block.

[0634] In some embodiments, the information for the first prediction direction and the second prediction direction before changing the set of motion vectors is used for a motion compensation operation for a current block. In some embodiments, the information before changing the set of motion vectors is used to predict motion vectors of subsequent blocks. In some embodiments, the information for the first prediction direction and the second prediction direction before changing the set of motion vectors is used for a deblocking process for a current block.

[0635] In some embodiments, the converted motion vectors are used for the motion refinement operation for the current block. In some embodiments, the information for the first prediction direction and the second prediction direction before changing the motion vector set is used for the refinement operation for the current block. In some embodiments, the refinement operation has an optical flow coding operation including at least a bi-directional optical flow (BDOF) operation or a prediction refined with optical flow (PROF) operation refined by optical flow.

[0636] In some embodiments, the method includes generating an intermediate prediction block based on the information for the first prediction direction and the information for the second prediction direction, refining the intermediate prediction block, and determining a final prediction block based on one of the intermediate prediction blocks.

[0637] FIG. 35 is a flowchart representation of a video processing method 3500 according to the present disclosure. The method 3500 includes, in operation 3502, determining a unidirectional motion vector from a bidirectional motion vector when the block size condition is satisfied for the conversion between the current block of the video and the bitstream representation of the video. The unidirectional motion vector is then used as a merge candidate for the conversion. The method 3500 also includes, in operation 3504, performing the conversion based on the determination.

[0638] In some embodiments, the transformed one - direction motion vector is used as a basic merge candidate in motion vector difference based merge (MMVD) coding technique. In some embodiments, the transformed one - direction motion vector is then inserted into the merge list. In some embodiments, the one - direction motion vector is transformed based on one prediction direction of the bi - direction motion vector. In some embodiments, the one - direction motion vector is transformed based on only one prediction direction, and the only one prediction direction is related to reference picture list 0. In some embodiments, when the video data unit where the current block is located is bi - predicted, the only one prediction direction for the first candidate in the candidate list is related to reference picture list 0, and the only one prediction direction for the second candidate in the candidate list is related to reference picture list 1. The candidate list has a merge candidate list or an MMVD basic candidate list.

[0639] In some embodiments, the video data unit has the current slice, tile group, or picture of the video. In some embodiments, the one - direction motion vector is determined based on reference picture list 1 when all reference pictures of the video data unit are past pictures in display order. In some embodiments, the one - direction motion vector is determined based on reference picture list 0 when at least the first reference picture among the reference pictures of the video data unit is a past picture and at least the second reference picture among the reference pictures of the video data unit is a future picture in display order. In some embodiments, one - direction motion vectors determined based on different reference picture lists are positioned in an interleaved manner in the merge list.

[0640] In some embodiments, the unidirectional motion vector is determined based on a low-latency check-in indicator. In some embodiments, the unidirectional motion vector is determined prior to the motion compensation process during transformation. In some embodiments, the unidirectional motion vector is determined after the motion candidate list construction process during transformation. In some embodiments, the unidirectional motion vector is determined before adding the motion vector difference in the motion vector difference based merge (MMVD) coding process. In some embodiments, the unidirectional motion vector is determined prior to the sample refinement process. In some embodiments, the unidirectional motion vector is determined prior to the bidirectional optical flow (BDOF) process. In some embodiments, the BDOF process is subsequently disabled based on a determination to use the unidirectional motion vector.

[0641] In some embodiments, the unidirectional motion vector is determined prior to the prediction refinement with optical flow (PROF) process. In some embodiments, the unidirectional motion vector is determined prior to the decoder-side motion vector refinement (DMVR) process. In some embodiments, the DMVR process is subsequently disabled based on the unidirectional motion vector.

[0642] In some embodiments, the block dimension condition indicates that the width of the current block is T1 and the height of the current block is T2, where T1 and T2 are positive integers. In some embodiments, the block dimension condition indicates that the width of the current block is T2 and the height of the current block is T1, where T1 and T2 are positive integers. In some embodiments, the block dimension condition indicates that the width of the current block is T1 and the height of the current block is less than or equal to T2, where T1 and T2 are positive integers. In some embodiments, the block dimension condition indicates that the width of the current block is T2 and the height of the current block is less than or equal to T1, where T1 and T2 are positive integers.

[0643] FIG. 36 is a flowchart representation of a video processing method 3600 according to the present disclosure. The method 3600 includes, at operation 3602, determining, for conversion between a current block of a video and a bitstream representation of the video, that the motion candidates of the current block are restricted to be unidirectional prediction candidates based on the dimensions of the current block. The method 3600 includes, at operation 3604, performing the conversion based on the determination.

[0644] In some embodiments, when the dimensions of the current block satisfy the conditions, all motion candidates of the current block are restricted to be unidirectional prediction candidates. In some embodiments, bidirectional motion candidates are converted to be unidirectional prediction candidates when the dimensions of the current block satisfy the conditions. In some embodiments, the dimensions of the current block indicate that the width of the current block is T1 and the height of the current block is T2, where T1 and T2 are positive integers. In some embodiments, the condition for the dimensions of the block indicates that the width of the current block is T2 and the height of the current block is T1, where T1 and T2 are positive integers. In some embodiments, the dimensions of the current block indicate that the width of the current block is T1 and the height of the current block is less than or equal to T2, where T1 and T2 are positive integers. In some embodiments, the dimensions of the current block indicate that the width of the current block is T2 and the height of the current block is less than or equal to T1, where T1 and T2 are positive integers. In some embodiments, T1 is equal to 4 and T2 is equal to 8. In some embodiments, T1 is equal to 4 and T2 is equal to 4. In some embodiments, T1 is equal to 8 and T2 is equal to 8.

[0645] In some embodiments, the step of performing the conversion by the above method includes generating a bitstream representation based on the current block of the video. In some embodiments, the step of performing the conversion by the above method includes generating the current block of the video from the bitstream representation.

[0646] As used herein, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits that are either at the same position within the bitstream or spread out at different locations, as defined by the syntax, for example. For example, a macroblock may be encoded using the transformed and coded error residue values, and further, bits in headers and other fields within the bitstream. Further, during conversion, the decoder may parse the bitstream based on a determination, knowing the possibility of the presence or absence of some fields, as explained in the above solution. Similarly, the encoder may determine whether a particular syntax field should be included, and accordingly, generate the bitstream representation by including or excluding the syntax field from the bitstream representation.

[0647] The disclosures and other solutions, examples, embodiments, modules, and functional operations described in this specification can be implemented in digital electronic circuits or in computer software, firmware, or hardware that includes the structures disclosed herein and their structural equivalents, or in combinations of one or more of them. The disclosures and other embodiments can be implemented as one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage carrier, a memory device, a composition that provides a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes, by way of example, all apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., an electrical, optical, or electromagnetic signal generated by a machine, that is generated to encode information for transmission to an appropriate receiving device.

[0648] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code), in portions of files that hold other programs or data (e.g., one or more scripts stored in a markup language document). A computer program can be executed on one computer, or located at one location, or deployed to be executed on multiple computers distributed across multiple locations and interconnected by a communication network.

[0649] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data to generate output. The processes and logic flows can also be executed by dedicated logic circuitry, such as a field programmable gate array (FPGA) or application specific integrated circuit (ASIC), and the apparatus can be implemented as such.

[0650] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or be operatively coupled to receive data from or transfer data to one or more such mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and all forms of non-volatile memory, media, and memory devices including CDROM and DVD-ROM disks. The processor and the memory may be enhanced or incorporated in a special purpose logic circuit.

[0651] This specification includes many details, but they should be construed as descriptions of features that may be specific to particular embodiments of a particular technology rather than as limitations on the scope of any subject or of what may be claimed. The particular features described in this specification in connection with separate embodiments may be implemented in combination with a single embodiment. Conversely, the various features described in connection with a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, features may have been described above as acting in a particular combination and even initially claimed as such, but one or more features from a claimed combination may in some cases be excised from that combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0652] Similarly, operations are represented in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in that particular order or in a sequential order to achieve the desired result, or that all of the operations shown be performed. Further, the separation of various system components in the embodiments described in this specification should not be understood as requiring such separation in all embodiments.

[0653] Only a few implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.

[0654] [Cross - reference to related applications] This application is a divisional application of Japanese Patent Application No. 2021-549770, which is a national phase application of International Patent Application No. PCT / CN2020 / 078107, claiming priority and the benefit thereof with respect to International Patent Application No. PCT / CN2019 / 077179 filed on March 6, 2019, International Patent Application No. PCT / CN2019 / 078939 filed on March 20, 2019, and International Patent Application No. PCT / CN2019 / 079397 filed on March 24, 2019. The entire disclosure of the above applications is incorporated herein by reference as part of the disclosure of this application.

Claims

1. 1. A method for processing video data, comprising the steps of: determining a first set of motion vectors for a current block of video; changing the first motion vector set to a second motion vector set in response to the current block satisfying a first condition, the first condition comprising a size of a luma component of the current block being the same as a predefined size; performing a conversion between the current block and the video bitstream based on the second set of motion vectors; having the second set of motion vectors is used to update a motion vector prediction history table; The method further comprises the step of determining whether flags of the current block are present in the bitstream based on whether a same condition associated with a size of the current block is satisfied, and the transformation is performed further based on the determination; the plurality of flags include a skip flag indicating whether a skip mode is used for the current block, and a flag indicating whether an inter prediction mode is used for the current block; The same conditions include that the width and height of the current block are each equal to 4. method.

2. the first set of motion vectors is derived using one of a conventional merge coding technique, a sub-block temporal motion vector predictive coding technique, or a merge coding technique using motion vector differentials; The predefined size comprises at least one of: 8×4 or 4×8. The method of claim 1.

3. Modifying the first set of motion vectors includes: converting the first motion vector set into unidirectional motion vectors when the current block is determined to be bidirectional and the first motion vector set has a first motion vector referencing a first reference picture in a first reference picture list for a first prediction direction and a second motion vector referencing a second reference picture in a second reference picture list for a second prediction direction. The method according to claim 1 or 2.

4. The motion information of one of the first prediction direction or the second prediction direction is discarded; Discarding the motion information comprises: changing the reference index in the discarded prediction direction to -1; The method according to claim 3.

5. the motion information of the second prediction direction is discarded, and the second reference picture list is list 1; The method according to claim 4.

6. the second set of motion vectors is used for a motion compensation operation for the current block, or the second set of motion vectors is used for the deblocking process for the current block, or the second motion vector set is used for a motion refinement operation for the current block, the refinement operation comprising at least an optical flow coding operation including a bidirectional optical flow operation or an optical prediction refinement operation.

6. The method according to any one of claims 1 to 5.

7. In response to the skip flag having a value of one, the bitstream does not include a merge data syntax structure. The method of claim 1.

8. The method of claim 7, wherein the plurality of flags includes the step of: whether the intra block copy flag is present in the bitstream is also based on the same conditions. The method of claim 1.

9. In response to the same condition being satisfied and the skip flag having a value of one, the intra block copy flag is not present in the bitstream. The method according to claim 8.

10. The intra block copy flag is a coding unit level flag, pred_mode_ibc_flag.

10. The method of claim 9.

11. In response to the same condition being satisfied and a value of a video sequence level flag indicating that intra block copy mode is not allowed for the entire video sequence, the skip flag is not presented in the bitstream. The method of claim 1.

12. The video sequence level flag is sps_ibc_enabled_flag. The method of claim 11.

13. The method of claim 12, wherein the plurality of flags includes a flag related to whether a triangular mode is used for the current block. The method of claim 1.

14. The method of claim 13, further comprising: providing a coding unit that is a coding unit that is a coding unit of a coding mode; The method of claim 1.

15. The method of claim 14, wherein if the size of the current block is 8×4 or 4×8, then bi-prediction is not used for the current block. The method of claim 14.

16. If the size of the current block is 8×4 or 4×8, then a context increment used to code the inter-prediction direction flag is set equal to a fixed integer and is independent of the width and height of the current block. The method of claim 15.

17. The method of claim 17, wherein if the size of the current block is 8×4 or 4×8, a weight index is set to 0 for the current block, and the weight index is used to find a weight from a weight set to generate a weighted sum from the predicted samples of list 0 and list 1. The method of claim 15.

18. the transforming includes encoding the current block into the bitstream.

18. A method according to any one of claims 1 to 17.

19. the converting includes decoding the current block from the bitstream.

18. A method according to any one of claims 1 to 17.

20. 1. An apparatus for processing video data, comprising: a processor and a non-transitory memory storing instructions; The instructions, when executed by the processor, cause the processor to: determining a first set of motion vectors for a current block of video; changing the first motion vector set to a second motion vector set in response to the current block satisfying a first condition, the first condition comprising a size of a luma component of the current block being the same as a predefined size; performing a conversion between the current block and the video bitstream based on the second set of motion vectors; Run the command, the second set of motion vectors is used to update a motion vector prediction history table; The instructions further cause the processor to determine whether flags for the current block are present in the bitstream based on whether a same condition associated with a size of the current block is satisfied, the transformation being performed further based on the determination; the plurality of flags include a skip flag indicating whether a skip mode is used for the current block, and a flag indicating whether an inter prediction mode is used for the current block; The same conditions include that the width and height of the current block are each equal to 4. Device.

21. The processor: determining a first set of motion vectors for a current block of video; changing the first motion vector set to a second motion vector set in response to the current block satisfying a first condition, the first condition comprising a size of a luma component of the current block being the same as a predefined size; performing a conversion between the current block and the video bitstream based on the second set of motion vectors; The instruction to execute the the second set of motion vectors is used to update a motion vector prediction history table; The instructions further cause the processor to determine whether flags for the current block are present in the bitstream based on whether a same condition associated with a size of the current block is satisfied, the transformation being performed further based on the determination; the plurality of flags include a skip flag indicating whether a skip mode is used for the current block, and a flag indicating whether an inter prediction mode is used for the current block; The same conditions include that the width and height of the current block are each equal to 4. A non-transitory computer-readable storage medium.

22. 1. A method for storing a video bitstream, comprising: determining a first set of motion vectors for a current block of video; changing the first motion vector set to a second motion vector set in response to the current block satisfying a first condition, the first condition comprising a size of a luma component of the current block being the same as a predefined size; generating the bitstream based on the second set of motion vectors; storing the bitstream on a non-transitory computer-readable recording medium; having the second set of motion vectors is used to update a motion vector prediction history table; The method further comprises the step of determining whether flags of the current block are present in the bitstream based on whether the same condition associated with a size of the current block is satisfied, and generating the bitstream is further based on the determining; the plurality of flags include a skip flag indicating whether a skip mode is used for the current block, and a flag indicating whether an inter prediction mode is used for the current block; The same conditions include that the width and height of the current block are each equal to 4. method.

Citation Information

Patent Citations

  • Moving image decoder, moving image encoder, moving image decoding method, moving image encoding method, moving image decoding program and moving image encoding program

    JP2012191298A

  • Use of converted one-way prediction candidates

    JP2022521554A

  • Restriction of prediction units in b slices to uni-directional inter prediction

    US20130202037A1