Video signal processing method and device using motion compensation

The method optimizes video signal processing by determining merge modes and parsing syntax elements to enhance coding efficiency and motion compensation in video compression.

JP2025172859AActive Publication Date: 2025-11-26WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025141507
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-11
Filing Date
2025-08-27
Publication Date
2025-11-26
Estimated Expiration
2040-01-20

AI Technical Summary

Technical Problem

Existing video signal processing methods lack efficiency in coding, particularly in the use of merge mode signaling during video compression.

Method used

A method and apparatus for video signal processing that includes determining the application of merge modes based on predefined conditions, parsing syntax elements to derive motion information, and generating prediction blocks, optimizing the coding process.

Benefits of technology

Improves coding efficiency by reducing unnecessary parsing and enhancing the accuracy of motion compensation in video signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025172859000001_ABST
    Figure 2025172859000001_ABST
Patent Text Reader

Abstract

To provide a video signal processing method and device that encode or decode a video signal.SOLUTION: The method includes, when a merge mode is applied to a current block, determining whether to parse a second syntax element on the basis of a first condition defined in advance. The second syntax element indicates whether a first mode or a second mode is applied to the current block. The method further includes, when the first mode and the second mode are not applied to the current block, determining whether to parse a third syntax element on the basis of a second condition defined in advance. The third syntax element indicates either one of a third mode or a fourth mode applied to the current block. The method further includes determining a mode applied to the current block on the basis of the second syntax element or the third syntax element, and generating a prediction block of the current block by using motion information of the current block.SELECTED DRAWING: Figure 34
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for processing a video signal, and more particularly to a method and apparatus for processing a video signal that encodes or decodes a video signal using motion compensation. [Background technology]

[0002] Compression coding refers to a series of signal processing techniques for transmitting digitized information over communication lines or storing it in a form suitable for storage media. Compression coding can be applied to audio, video, text, etc., and the technology for compressing video signals is called video compression. Video signals are compressed by removing redundant information by taking into account spatial correlation, temporal correlation, stochastic correlation, etc. However, with the recent development of various media and data transmission media, more efficient video signal processing methods and devices are desired. Summary of the Invention [Problem to be solved by the invention]

[0003] SUMMARY OF THE INVENTION It is an object of the present invention to improve the coding efficiency of video signals, and to provide an efficient merge mode signaling method. [Means for solving the problem]

[0004] In order to solve the above problems, the present invention provides the following video signal processing device and video signal processing method.

[0005] According to an embodiment of the present invention, there is provided a video signal processing method, comprising: a first syntax element indicating whether a merge mode is applied to a current block; determining, based on a first predefined condition, whether to parse a second syntax element if the merge mode is applied to the current block, wherein the second syntax element indicates whether a first mode or a second mode is applied to the current block; determining, based on a second predefined condition, whether to parse a third syntax element if neither the first mode nor the second mode is applied to the current block, wherein the third syntax element indicates a third mode or a fourth mode to be applied to the current block; determining, based on the second syntax element or the third syntax element, a mode to be applied to the current block; deriving motion information of the current block based on the determined mode; and generating a prediction block for the current block using the motion information of the current block, wherein the first condition includes at least one of a condition under which the third mode can be used and a condition under which the fourth mode can be used.

[0006] As an example, the third mode and the fourth mode may be positioned after the first mode in the decoding order in the merge data syntax.

[0007] As an example, the method may include parsing the second syntax element if the first condition is met, and inferring the second syntax element to be 1 if the first condition is not met.

[0008] As an example, if the first condition is not met, the second syntax element may be inferred based on a fourth syntax element indicating whether a sub-block based merge mode is applied to the current block.

[0009] As an example, the second condition may include a condition under which the fourth mode is available.

[0010] As an embodiment, the second condition may include at least one of whether the third mode is available in the current sequence, whether the fourth mode is available in the current sequence, whether the maximum number of candidates for the fourth mode is greater than 1, whether the width of the current block is smaller than a predefined first size, and whether the height of the current block is smaller than a predefined second size.

[0011] As an example, if the second syntax element is 1, the method may include obtaining a fifth syntax element indicating whether the first mode or the second mode is applied to the current block.

[0012] According to an embodiment of the present invention, a video signal processing apparatus includes a processor, the processor being configured to generate a first syntax element indicating whether a merge mode is applied to a current block. and, if the merge mode is applied to the current block, determining whether to parse a second syntax element based on a first predefined condition, wherein the second syntax element indicates whether a first mode or a second mode is applied to the current block. If the first mode and the second mode are not applied to the current block, determining whether to parse a third syntax element based on a second predefined condition, wherein the third syntax element indicates a third mode or a fourth mode to be applied to the current block. The video signal processing device determines a mode to be applied to the current block based on the second syntax element or the third syntax element, derives motion information of the current block based on the determined mode, and generates a predicted block for the current block using the motion information of the current block. The first condition includes at least one of a condition under which the third mode is available and a condition under which the fourth mode is available.

[0013] As an example, the third mode and the fourth mode may be positioned after the first mode in the decoding order in the merge data syntax.

[0014] As an example, the processor may parse the second syntax element if the first condition is met, and may infer that the second syntax element is 1 if the first condition is not met.

[0015] As an example, if the first condition is not met, the second syntax element may be inferred based on a fourth syntax element indicating whether a sub-block based merge mode is applied to the current block.

[0016] As an example, the second condition may include a condition under which the fourth mode is available.

[0017] As an embodiment, the second condition may include at least one of whether the third mode is available in the current sequence, whether the fourth mode is available in the current sequence, whether the maximum number of candidates for the fourth mode is greater than 1, whether the width of the current block is smaller than a predefined first size, and whether the height of the current block is smaller than a predefined second size.

[0018] As an example, the processor may obtain a fifth syntax element indicating whether the first mode or the second mode is applied to the current block when the second syntax element is 1.

[0019] According to an embodiment of the present invention, there is provided a video signal processing method, comprising: a step of encoding a first syntax element indicating whether a merge mode is applied to a current block; a step of determining whether to encode a second syntax element based on a first predefined condition if the merge mode is applied to the current block, wherein the second syntax element indicates whether a first mode or a second mode is applied to the current block; a step of determining whether to encode a third syntax element based on a second predefined condition if neither the first mode nor the second mode is applied to the current block, wherein the third syntax element indicates a third mode or a fourth mode to be applied to the current block; a step of determining a mode to be applied to the current block based on the second syntax element or the third syntax element; a step of deriving motion information of the current block based on the determined mode; and a step of generating a prediction block of the current block using the motion information of the current block, wherein the first condition includes at least one of a condition under which the third mode can be used and a condition under which the fourth mode can be used. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a schematic block diagram of a video signal encoding apparatus according to an embodiment of the present invention; [Figure 2] 1 is a schematic block diagram of a video signal decoding device according to an embodiment of the present invention; [Figure 3] FIG. 1 illustrates an example of how coding tree units are divided into coding units within a picture. [Figure 4] FIG. 1 illustrates an embodiment of a method for signaling the splitting of quadtrees and multi-type trees. [Figure 5] FIG. 1 is a diagram illustrating inter prediction according to an embodiment of the present invention. [Figure 6]1 is a diagram illustrating a motion vector signaling method according to an embodiment of the present invention. [Figure 7] 1 is a diagram illustrating a method for signaling adaptive motion vector resolution information according to one embodiment of the present invention. [Figure 8] FIG. 2 is a diagram illustrating a coding unit syntax according to an embodiment of the present invention. [Figure 9] FIG. 1 illustrates a coding unit syntax according to one embodiment of the present invention. [Figure 10] FIG. 1 illustrates a merge mode signaling method according to one embodiment of the present invention. [Figure 11] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 12] FIG. 10 is a diagram illustrating merge data syntax according to one embodiment of the present invention. [Figure 13] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 14] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 15] FIG. 10 is a diagram illustrating a merge data syntax structure according to one embodiment of the present invention. [Figure 16] FIG. 10 is a diagram illustrating a merge data syntax structure according to one embodiment of the present invention. [Figure 17] FIG. 10 is a diagram illustrating a merge data syntax structure according to one embodiment of the present invention. [Figure 18] FIG. 10 is a diagram illustrating a merge data syntax structure according to one embodiment of the present invention. [Figure 19] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 20] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 21]FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 22] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 23] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 24] FIG. 10 is a diagram illustrating a merge data syntax structure according to one embodiment of the present invention. [Figure 25] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 26] FIG. 10 is a diagram illustrating a merge data syntax structure according to one embodiment of the present invention. [Figure 27] FIG. 10 is a diagram illustrating a merge data syntax structure according to one embodiment of the present invention. [Figure 28] FIG. 10 illustrates a merge data syntax structure according to one embodiment of the present invention. [Figure 29] FIG. 1 illustrates a merge mode signaling method according to one embodiment of the present invention. [Figure 30] FIG. 10 illustrates a merge data syntax according to an embodiment of the present invention. [Figure 31] FIG. 10 illustrates a merge data syntax according to an embodiment of the present invention. [Figure 32] FIG. 1 illustrates a geometric merge mode according to an embodiment of the present invention. [Figure 33] FIG. 10 is a diagram illustrating merge data syntax according to one embodiment of the present invention. [Figure 34] 1 is a diagram illustrating a video signal processing method according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0021] The terms used in this specification have been selected as widely used and general terms as possible, taking into consideration the functions of the present invention, but these may vary depending on the intentions of engineers in the field, customs, or the emergence of new technologies. In addition, in certain cases, the applicant may have arbitrarily selected terms, and in such cases, the meanings of these terms will be described in the relevant section on the mode for carrying out the invention. Therefore, it is made clear that the terms used in this specification should be interpreted not simply as terms, but based on the substantive meanings of the terms and the overall content of this specification.

[0022] In this specification, some terms may be interpreted as follows: "Coding" may be interpreted as "Encoding" or "Decoding" in some cases. In this specification, an apparatus that encodes a video signal to generate a video signal bitstream is referred to as an encoding apparatus or encoder, and an apparatus that decodes a video signal bitstream to restore a video signal is referred to as a decoding apparatus or decoder. In this specification, "video signal processing apparatus" is used as a conceptual term that includes both an encoder and a decoder. "Information" is a term that includes values, parameters, coefficients, elements, etc., and may be interpreted differently in some cases, so the present invention is not limited thereto. "Unit" is used interchangeably to refer to a basic unit of image processing or a specific position in a picture, and refers to an image area including both luma and chroma components. "Block" refers to an image area including specific components of luma and chroma components (i.e., Cb and Cr). However, depending on the embodiment, terms such as "unit," "block," "partition," and "area" may be used interchangeably. In this specification, the term "unit" is used as a concept including a coding unit, a prediction unit, and a transform unit, and the term "picture" refers to a field or a frame, and these terms may be used interchangeably depending on the embodiment.

[0023] 1 is a schematic block diagram of a video signal encoding apparatus according to an embodiment of the present invention. Referring to FIG. 1, the encoding apparatus 100 includes a transform unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transform unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.

[0024] The transform unit 110 transforms a residual signal, which is the difference between the input video signal and the prediction signal generated by the prediction unit 150, to obtain a transform coefficient value. For example, a discrete cosine transform (DCT), a discrete sine transform (DST), or a wavelet transform may be used. The discrete cosine transform and the discrete sine transform divide the input picture signal into blocks and then transform them. During the transformation, coding efficiency may vary depending on the distribution and characteristics within the transformation domain. The quantization unit 115 quantizes the values ​​of the transform coefficients output from the transform unit 110.

[0025] To improve coding efficiency, instead of directly coding the picture signal, the prediction unit 150 predicts a picture using a pre-coded region and adds the residual values ​​between the original picture and the predicted picture to obtain a reconstructed picture. To avoid mismatch between the encoder and decoder, the encoder should use information available to the decoder when making predictions. To achieve this, the encoder performs a process of further reconstructing the coded current block. The inverse quantization unit 120 inversely quantizes the transform coefficient values, and the inverse transform unit 125 reconstructs the residual values ​​using the inversely quantized transform coefficient values. Meanwhile, the filtering unit 130 performs filtering operations to improve the quality of the reconstructed picture and the coding efficiency. Examples of filtering operations include a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter. The filtered picture is stored in a decoded picture buffer (DPB) 156 for output or use as a reference picture.

[0026] To improve coding efficiency, instead of directly coding a picture signal, the prediction unit 150 predicts a picture using an already coded region and adds a residual value between the original picture and the predicted picture to the predicted picture to obtain a reconstructed picture. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 predicts the current picture using a reference picture stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from a reconstructed region within the current picture and transmits the intra coding information to the entropy coding unit 160. The inter prediction unit 154 may further include a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains a motion vector value for the current region by referring to a specific reconstructed region. The motion estimation unit 154a transmits position information of the reference region (e.g., reference frame, motion vector) to the entropy coding unit 160 so that it can be included in the bitstream. Using the motion vector values ​​transmitted from the motion estimation unit 154a, the motion compensation unit 154b performs inter-frame motion compensation.

[0027] The prediction unit 150 includes an intra prediction unit 152 and an inter prediction unit 154. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 performs inter prediction to predict the current picture using a reference buffer stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from reconstructed samples within the current picture and transmits intra coding information to the entropy coding unit 160. The intra coding information includes at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, and an MPM index. The intra coding information may include information about reference samples. The inter prediction unit 154 includes a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains a motion vector value for the current region by referring to a specific region of the reconstructed reference signal picture. The motion estimation unit 154a transmits a motion information set (reference picture index, motion vector information) for the reference region to the entropy coding unit 160. The motion compensation unit 154b performs motion compensation using the motion vector values ​​transmitted from the motion compensation unit 154a. The inter prediction unit 154 transmits inter coding information including the motion information for the reference region to the entropy coding unit 160.

[0028] According to a further embodiment, the prediction unit 150 includes an intra block copy (BC) prediction unit (not shown). The intra BC prediction unit performs intra BC prediction from reconstructed samples in the current picture and transmits intra BC coding information to the entropy coding unit 160. The intra BC prediction unit obtains block vector values ​​indicating a reference region to be used for predicting the current region by referring to a specific region in the current picture. The intra BC prediction unit performs intra BC prediction using the obtained block vector values. The intra BC prediction unit transmits the intra BC coding information to the entropy coding unit 160. The intra BC prediction unit includes the block vector information.

[0029] After the picture prediction is performed, the transform unit 110 converts residual values ​​between the original picture and the predicted picture to obtain transform coefficient values. The transform is performed in units of specific blocks within the picture, and the size of the specific blocks varies within a predetermined range. The quantization unit 115 quantizes the transform coefficient values ​​generated by the transform unit 110 and transmits the quantized values ​​to the entropy coding unit 160.

[0030] The entropy coding unit 160 generates a video signal bitstream by entropy coding information indicating quantized transform coefficients, intra-coding information, and inter-coding information. The entropy coding unit 160 uses a variable length coding (VLC) scheme and an arithmetic coding scheme. The VLC scheme converts input symbols into successive codewords, where the length of the codewords is variable. For example, frequently occurring symbols are represented by short codewords, and infrequently occurring symbols are represented by long codewords. The variable length coding scheme is a context-based adaptive variable length coding (CAVLC) scheme. The arithmetic coding scheme converts successive data symbols into a single prime number, and obtains the optimal prime number bits required to represent each symbol. The arithmetic coding scheme is a context-based adaptive binary arithmetic coding (CABAC) scheme. For example, the entropy coding unit 160 may binarize information indicating quantized transform coefficients, and may arithmetically code the binarized information to generate a bitstream.

[0031] The generated bitstream is encapsulated in Network Abstraction Layer (NAL) units as basic units. An NAL unit includes an integer number of coded coding tree units. In order for a video decoder to decode the bitstream, the bitstream must first be separated into NAL units and then each separated NAL unit must be decoded. Meanwhile, information required for decoding the video signal bitstream is transmitted via Raw Byte Sequence Payload (RBSP) of higher level sets such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS).

[0032] 1 illustrates an encoding device 100 according to one embodiment of the present invention, with separate blocks illustrating logically distinct elements of encoding device 100. Therefore, the elements of encoding device 100 described above may be implemented on a single chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of encoding device 100 described above is performed by a processor (not shown).

[0033] 2 is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. Referring to FIG. 2, the decoding apparatus 200 of the present invention includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.

[0034] The entropy decoding unit 210 entropy decodes the video signal bitstream to extract transform coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit 210 may obtain a binary code for transform coefficient information of a specific region from the video signal bitstream. The entropy decoding unit 210 also de-binarizes the binary code to obtain quantized transform coefficients. The inverse quantization unit 220 de-quantizes the quantized transform coefficients, and the inverse transform unit 225 restores residual values ​​using the de-quantized transform coefficients. The video signal processing device 200 restores original pixel values ​​by combining the residual values ​​obtained from the inverse transform unit 225 with predicted values ​​obtained from the prediction unit 250.

[0035] Meanwhile, the filtering unit 230 performs filtering on the picture to improve image quality. This includes a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is output or stored in the decoded picture buffer (DPB) 256 to be used as a reference picture for the next picture.

[0036] The prediction unit 250 includes an intra prediction unit 252 and an inter prediction unit 254. The prediction unit 250 generates a predicted picture using the coding type, transform coefficients for each region, intra / inter coding information, etc. decoded by the entropy decoding unit 210. To reconstruct the current block to be decoded, the current picture including the current block or a decoded region of another picture can be used. A picture (or tile / slice) that uses only the current picture for reconstruction, i.e., performs intra prediction or intra BC prediction, is called an intra picture or I picture (or tile / slice), and a picture (or tile / slice) that can perform all of intra prediction, inter prediction, and intra BC prediction is called an inter picture (or tile / slice). Among interpictures (or tiles / slices), a picture (or tile / slice) that uses at most one motion vector and reference picture index to predict sample values ​​for each block is called a predictive picture or P picture (or tile / slice), and a picture (or tile / slice) that uses at most two motion vectors and reference picture indexes is called a bi-predictive picture or B picture (or tile / slice). In other words, a P picture (or tile / slice) uses at most one motion information set to predict each block, and a B picture (or tile / slice) uses at most two motion information sets to predict each block. Here, a motion information set includes one or more motion vectors and one reference picture index.

[0037] The intra prediction unit 252 generates a prediction block using intra coding information and reconstructed samples in the current picture. As described above, the intra coding information includes at least one of an intra prediction mode, a Most Probable Mode (MPM) flag, and an MPM index. The intra prediction unit 252 predicts sample values ​​of the current block using reconstructed samples located to the left and / or above the current block as reference samples. In the present disclosure, the reconstructed samples, reference samples, and samples of the current block refer to pixels. Furthermore, sample values ​​refer to pixel values.

[0038] In one embodiment, the reference samples are samples included in neighboring blocks of the current block. For example, the reference samples are samples adjacent to the left boundary and / or the top boundary of the current block. Furthermore, the reference samples are samples located on a line within a predetermined distance from the left boundary of the current block and / or samples located on a line within a predetermined distance from the top boundary of the current block, among samples in neighboring blocks of the current block. In this case, the neighboring blocks of the current block include at least one of the left (L) block, the top (A) block, the below left (BL) block, the above right (AR) block, and the above left (AL) block adjacent to the current block.

[0039] The inter prediction unit 254 generates a prediction block using reference pictures and inter coding information stored in the decoded picture buffer 256. The inter coding information includes a motion information set (e.g., reference picture index, motion vector, etc.) of the current block relative to the reference block. Inter prediction includes L0 prediction, L1 prediction, and bi-prediction. L0 prediction is prediction using one reference picture included in the L0 picture list, and L1 prediction is prediction using one reference picture included in the L1 picture list. This requires one set of motion information (e.g., motion vector and reference picture index). The bi-prediction method uses up to two reference regions, and these two reference regions may exist in the same reference picture or in different pictures. That is, the bi-prediction method uses up to two sets of motion information (e.g., motion vector and reference picture index), and two motion vectors may correspond to the same reference picture index or different reference picture indexes. In this case, the reference picture may be displayed (or output) either temporally before or after the current picture. According to one embodiment, in a bi-predictive scheme, the two reference regions used may be regions selected from the L0 picture list and the L1 picture list, respectively.

[0040] The inter prediction unit 254 obtains a current reference block using a motion vector and a reference picture index. The reference block exists in a reference picture corresponding to the reference picture index. Furthermore, sample values ​​of a block identified by the motion vector or their interpolated values ​​are used as a predictor for the current block. For motion prediction with sub-pel pixel accuracy, for example, an 8-tab interpolation filter is used for the luma signal and a 4-tab interpolation filter is used for the chroma signal. However, the interpolation filters for sub-pel motion prediction are not limited thereto. In this way, the inter prediction unit 254 performs motion compensation, which predicts the texture of the current unit from a previously reconstructed picture. In this case, the inter prediction unit uses a motion information set.

[0041] According to a further embodiment, the predictor 250 may include an intra BC predictor (not shown). The intra BC predictor may reconstruct the current region by referring to a specific region including reconstructed samples in the current picture. The intra BC predictor obtains intra BC coding information for the current region from the entropy decoding unit 210. The intra BC predictor obtains block vector values ​​of the current region indicating the specific region in the current picture. The intra BC predictor may perform intra BC prediction using the obtained block vector values. The intra BC coding information may include block vector information.

[0042] According to a further embodiment, the predictor 250 may include an intra BC predictor (not shown). The intra BC predictor may reconstruct the current region by referring to a specific region including reconstructed samples in the current picture. The intra BC predictor obtains intra BC coding information for the current region from the entropy decoding unit 210. The intra BC predictor obtains block vector values ​​of the current region indicating the specific region in the current picture. The intra BC predictor may perform intra BC prediction using the obtained block vector values. The intra BC coding information may include block vector information.

[0043] 2 illustrates a decoding device 200 according to one embodiment of the present invention, with separate blocks logically separating elements of the decoding device 200. Thus, the elements of the decoding device 200 described above may be implemented on a single chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of the decoding device 200 described above is performed by a processor (not shown).

[0044] FIG. 3 illustrates an example in which a coding tree unit (CTU) is divided into coding units (CUs) within a picture. During video signal coding, a picture is divided into a sequence of coding tree units (CTUs). A coding tree unit consists of an NXN block of luma samples and two blocks of corresponding chroma samples. A coding tree unit is divided into multiple coding units. A coding tree unit may be a leaf node without being divided. In this case, the coding tree unit itself may be a coding unit. A coding unit refers to a basic unit for processing a picture during the above-mentioned video signal processing, i.e., intra / inter prediction, transform, quantization, and / or entropy coding. Within a picture, the size and shape of coding units are not constant. Coding units have a square or rectangular shape. A rectangular coding unit (or rectangular block) includes a vertical coding unit (or vertical block) and a horizontal coding unit (or horizontal block). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. In addition, in this specification, non-square blocks refer to rectangular blocks, but the present invention is not limited to this.

[0045] Referring to Figure 3, a coding tree unit is first divided into a quad tree (QT) structure. That is, in the quad tree structure, one node having a size of 2N x 2N is divided into four nodes having a size of N x N. In this specification, a quad tree is also referred to as a quaternary tree. The quad tree division is performed recursively, and all nodes do not need to be divided to the same depth.

[0046] Meanwhile, the leaf node of the above-mentioned quad tree is further divided into a multi-type tree (MTT) structure. According to an embodiment of the present invention, in the multi-type tree structure, one node is divided into a horizontally or vertically divided binary or ternary tree structure. That is, there are four division structures in the multi-type tree structure: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to an embodiment of the present invention, in each of the tree structures, the width and height of the node are both powers of 2. For example, in a binary tree (BT) structure, a node of size 2N×2N is divided into two N×2N nodes by vertical binary division and into two 2N×N nodes by horizontal binary division. In addition, in a ternary tree (TT) structure, a node of size 2Nx2N is divided into (N / 2)x2N, Nx2N, and (N / 2)x2N nodes by vertical ternary division, and into 2Nx(N / 2), 2NxN, and 2Nx(N / 2) nodes by horizontal ternary division. Such multi-type tree division is performed recursively.

[0047] The leaf nodes of a multi-type tree can be coding units. If no division for a coding unit is specified or the coding unit is not larger than the maximum transform length, the coding unit is used as the unit of prediction and transformation without further division. Meanwhile, in the above-mentioned quad trees and multi-type trees, at least one of the following parameters is predefined or transmitted via the RBSP of a higher-level set such as a PPS, SPS, or VPS: 1) CTU size: the size of the root node of the quad tree; 2) minimum QT size (MinQtSize): the size of the minimum QT leaf node allowed; 3) maximum BT size (MaxBtSize): the size of the maximum BT root node allowed; 4) maximum TT size (MaxTtSize): the size of the maximum TT root node allowed; 5) maximum MTT depth (MaxMttDepth): the maximum allowed depth of MTT division from the QT leaf node; 6) minimum BT size (MinBtSize): the size of the minimum BT leaf node allowed; 7) minimum TT size: the size of the minimum TT leaf node allowed.

[0048] 4 illustrates an embodiment of a method for signaling the splitting of a quadtree and a multi-type tree. Pre-set flags can be used to signal the splitting of the quadtree and the multi-type tree. Referring to FIG. 4, at least one of a flag 'qt_split_flag' indicating whether to split a quadtree node, a flag 'mtt_split_flag' indicating whether to split a multi-type tree node, a flag 'mtt_split_vertical_flag' indicating the split direction of the multi-type tree node, and a flag 'mtt_split_binary_flag' indicating the split type of the multi-type tree node can be used.

[0049] According to an embodiment of the present invention, a coding tree unit is the root node of a quad tree and can be split into a quad tree structure first. In the quad tree structure, a 'qt_split_flag' is signaled for each node 'QT_node'. If the value of 'qt_split_flag' is 1, the corresponding node is split into four regular rectangular nodes, and if the value of 'qt_split_flag' is 0, the corresponding node becomes a leaf node 'QT_leaf_node' of the quad tree.

[0050] Each quadtree leaf node 'QT_leaf_node' can be further split into a multi-type tree structure. In a multi-type tree structure, 'mtt_split_flag' is signaled for each node 'MTT_node'. If 'mtt_split_flag' is set to 1, the node is split into multiple rectangular nodes, and if 'mtt_split_flag' is set to 0, the node becomes a leaf node 'MTT_leaf_node' of the multi-type tree. If a multi-type tree node 'MTT_node' is split into multiple rectangular nodes (i.e., if 'mtt_split_flag' is set to 1), 'mtt_split_vertical_flag' and 'mtt_split_binary_flag' can be additionally signaled for the node 'MTT_node'. If the value of 'mtt_split_vertical_flag' is 1, vertical split of node 'MTT_node' is indicated, and if the value of 'mtt_split_vertical_flag' is 0, horizontal split of node 'MTT_node' is indicated. Also, if the value of 'mtt_split_binary_flag' is 1, node 'MTT_node' is split into two rectangular nodes, and if the value of 'mtt_split_binary_flag' is 0, node 'MTT_node' is split into three rectangular nodes.

[0051] Picture prediction (motion compensation) for coding is performed on coding units that cannot be further divided (i.e., leaf nodes of the coding unit tree). Such a basic unit for prediction is hereinafter referred to as a prediction unit or a prediction block.

[0052] Hereinafter, the term "unit" used in this specification is used as an alternative term to the prediction unit, which is a basic unit for performing prediction, but the present invention is not limited thereto and can be understood as a concept including the coding unit in a broader sense.

[0053] FIG. 5 illustrates inter-prediction according to an embodiment of the present invention. As described above, a decoder predicts a current block by referring to reconstructed samples of other decoded pictures. Referring to FIG. 5, the decoder obtains a reference block 42 in a reference picture based on motion information of a current block 32. The motion information may include a reference picture index and a motion vector 50. The reference picture index indicates a reference picture of the current block in a reference picture list. The motion vector 50 represents an offset between the coordinate values ​​of the current block 32 in the current picture and the coordinate values ​​of the reference block 42 in the reference picture. The decoder obtains a predictor for the current block 32 based on sample values ​​of the reference block 42 and reconstructs the current block 32 using the predictor.

[0054] Meanwhile, according to an embodiment of the present invention, sub-block-based motion compensation may be used. That is, the current block 32 may be divided into a plurality of sub-blocks, and an independent motion vector may be used for each sub-block. Therefore, each sub-block within the current block 32 may be predicted using a different reference block. According to one embodiment, the sub-blocks may have a predetermined size, such as 4x4 or 8x8. The decoder obtains a predictor for each sub-block of the current block 32 using the motion vector of each sub-block. The predictor for each sub-block may be combined to obtain a predictor for the current block 32, and the decoder may reconstruct the current block 32 using the predictor for the current block 32 obtained in this manner.

[0055] According to an embodiment of the present invention, various methods of subblock-based motion compensation may be performed. Subblock-based motion compensation may include affine model-based motion compensation (hereinafter referred to as affine motion compensation or affine motion prediction) and subblock-based temporal motion vector prediction (Subblock-based Temporal Motion Vector Prediction, SbTMVP). Various embodiments of affine motion compensation and SbTMVP will be described below with reference to the drawings.

[0056] 6 is a diagram illustrating a motion vector signaling method according to an embodiment of the present invention. According to an embodiment of the present invention, a motion vector (MV) may be generated based on a motion vector prediction (or predictor) (MVP). As an example, the MV may be determined as the MVP according to the following Equation 1. In other words, the MV may be determined (or set or induced) to be the same value as the MVP.

[0057]

number

[0058] As another example, the MV may be determined based on the MVP and a motion vector difference (MVD), as shown in the following mathematical formula 2. The encoder can signal MVD information to the decoder to represent a more accurate MV, and the decoder can derive the MV by adding the obtained MVD to the MVP.

[0059]

number

[0060] According to one embodiment of the present invention, the encoder transmits the determined motion information to the decoder, and the decoder generates (or induces) a motion vector (MVP) from the received motion information and generates a prediction block based on the motion vector (MVD). For example, the motion information may include MVP information and MVD information. In this case, components of the motion information may differ depending on the inter-prediction mode. For example, in a merge mode, the motion information may include MVP information but not MVD information. For another example, in an advanced motion vector prediction (AMVP) mode, the motion information may include MVP information and MVD information.

[0061] To determine, transmit, and receive information about MVPs, the encoder and decoder can generate MVP candidates (or MVP candidate lists) in the same manner. For example, the encoder and decoder can generate the same MVP candidates in the same order. Then, the encoder transmits an index indicating (or pointing to) the determined (or selected) MVP from the generated MVP candidates to the decoder, and the decoder can derive the determined MVP and / or MV based on the received index.

[0062] According to an embodiment of the present invention, MVP candidates may include spatial candidates, temporal candidates, etc. The MVP candidates may be referred to as merge candidates when a merge mode is applied, and as AMVP candidates when an AMVP mode is applied. The spatial candidate may be an MV (or motion information) for a block at a specific position relative to the current block. For example, the spatial candidate may be an MV for a block at a position adjacent or non-adjacent to the current block. The temporal candidate may be an MV corresponding to a block in a picture different from the current picture. In addition, for example, the MVP candidate may include an affine MV, ATMVP, STMVP, a combination of the above-mentioned MVs (or candidates), an average MV of the above-mentioned MVs (or candidates), a zero MV, etc.

[0063] In one embodiment, the encoder may signal information indicating the reference picture to the decoder. For example, if the reference picture of the MVP candidate is different from the reference picture of the current block (or the currently processed block), the encoder / decoder may perform motion vector scaling (MV scaling) of the MVP candidate. In this case, the MV scaling may be performed based on the picture order count (POC) of the current picture, the POC of the reference picture of the current block, and the POC of the reference picture of the MVP candidate.

[0064] A specific embodiment of the MVD signaling method is described below. Table 1 below illustrates a syntax structure for MVD signaling.

[0065] [Table 1]

[0066] Referring to Table 1, according to one embodiment of the present invention, the sign and absolute value of the MVD may be coded separately. That is, the sign and absolute value of the MVD may each be coded using different syntax (or syntax elements). The absolute value of the MVD may be coded directly or may be coded stepwise based on a flag indicating whether the absolute value is greater than N, as shown in Table 1. If the absolute value is greater than N, the value of (absolute value - N) may also be signaled. Specifically, in the example of Table 1, abs_mvd_greater0_flag may be transmitted to indicate whether the absolute value is greater than 0. If abs_mvd_greater0_flag indicates (or indicates) that the absolute value is not greater than 0, the absolute value of the MVD may be determined to be 0. If abs_mvd_greater0_flag indicates that the absolute value is greater than 0, an additional syntax (or syntax element) may be present.

[0067] For example, abs_mvd_greater1_flag may be transmitted to indicate whether the absolute value is greater than 1. If abs_mvd_greater1_flag indicates (or indicates) that the absolute value is not greater than 1, the absolute value of the MVD may be determined to be 1. If abs_mvd_greater1_flag indicates that the absolute value is greater than 1, additional syntax may be present. For example, abs_mvd_minus2 may be present. abs_mvd_minus2 may be a value of (absolute value - 2). Because the abs_mvd_greater0_flag and abs_mvd_greater1_flag values ​​determine that the absolute value is greater than 1 (i.e., 2 or greater), a value of (absolute value - 2) may be signaled. In this way, by hierarchically syntactically signaling information about the absolute value, fewer bits can be used than when the absolute value is directly binarized and signaled.

[0068] In one embodiment, the absolute value-related syntax described above may be coded using a variable length binarization method such as Exponential-Golomb, truncated unary, truncated Rice, etc. Also, a flag indicating the sign of the MVD may be signaled by mvd_sign_flag.

[0069] Although the coding method for MVD has been described in the above embodiment, information other than MVD can also be signaled by separating the sign and absolute value. The absolute value may be coded as a flag indicating whether the absolute value is greater than a predefined specific value and a value obtained by subtracting the specific value from the absolute value. In Table 1, [0] and [1] may represent component indexes. For example, they may represent the x-component (i.e., horizontal component) and the y-component (i.e., vertical component).

[0070] FIG. 7 is a diagram illustrating a method for signaling adaptive motion vector resolution information according to an embodiment of the present invention. According to an embodiment of the present invention, the resolution for indicating MV or MVD may vary. For example, the resolution may be expressed based on pixels (or pels). For example, MV or MVD may be signaled in units of 1 / 4, 1 / 2, 1 (integer), 2, or 4 pixels. The encoder may then signal the MV or MVD resolution information to the decoder. For example, 16 may be coded as 64 in 1 / 4 units (1 / 4*64=16), 16 in 1 unit (1*16=16), and 4 in 4 units (4*.4=16). That is, the MV or MVD value may be determined by the following Equation 3:

[0071]

number

[0072] In Equation 3, valueDetermined represents an MV or MVD value. Also, valuePerResolution represents a value signaled based on the determined resolution. If the value signaled by MV or MVD is not divisible by the determined resolution, a rounding process may be applied. Using a high resolution may improve accuracy but consume more bits because the coded values ​​are large. Using a low resolution may reduce accuracy but consume fewer bits because the coded values ​​are small. In one embodiment, the resolution may be individually set for each unit, such as a sequence, a picture, a slice, a coding tree unit (CTU), or a coding unit (CU). That is, the encoder / decoder may adaptively determine / apply the resolution according to a predefined unit among the above units.

[0073] According to one embodiment of the present specification, the above-mentioned resolution information may be signaled from the encoder to the decoder. At this time, the resolution information may be binarized and signaled based on the above-mentioned variable length. In this case, if the resolution information is signaled based on the index corresponding to the smallest value (i.e., the earliest value), signaling overhead can be reduced. As one embodiment, the resolution information may be mapped to the signaling index in order from highest to lowest resolution.

[0074] According to one embodiment of the present specification, Figure 7 illustrates a signaling method assuming that three resolutions are used among a variety of resolutions. In this case, three signaling bits may be 0, 10, and 11, and the three signaling indexes may represent a first resolution, a second resolution, and a third resolution, respectively. Since one bit is required to signal the first resolution and two bits are required to signal the remaining resolutions, signaling overhead can be relatively reduced when signaling the first resolution. In the example of Figure 7, the first resolution, the second resolution, and the third resolution may be defined as 1 / 4, 1, and 4 pixel resolutions, respectively. In the following embodiments, MV resolution may refer to the resolution of MVD.

[0075] 8 shows affine motion compensation according to an embodiment of the present invention. Conventional inter prediction methods are optimized for predicting translational motion because they use a single motion vector for L0 prediction and L1 prediction of a current block. However, to efficiently perform motion compensation for zoom-in / out, rotation, and other irregular motions, reference blocks 44 of various shapes and sizes must be used.

[0076] In the following, a motion compensation method based on merge mode with motion vector difference (MMVD) (or merge MVD) will be described.

[0077] 8 is a diagram illustrating a coding unit syntax according to an embodiment of the present invention. According to an embodiment of the present invention, a syntax element indicating whether MMVD is applied may be signaled based on a syntax element indicating whether merge mode is applied. Referring to FIG. 8, in step S801, the MMVD flag (mmvd_flag) may be signaled when the merge flag (merge_flag) is 0 (i.e., when merge mode is not used). In FIG. 8, the MMVD flag represents a syntax element (or flag) indicating whether MMVD is applied. And the merge flag represents a syntax element (or flag) indicating whether merge mode is applied.

[0078] According to one embodiment of the present invention, when a merge mode is applied, an encoder / decoder may determine a motion vector (MV) based on a motion vector predictor (MVP) and a motion vector difference (MVD). In this specification, the MVP may be referred to as a base motion vector (baseMV). That is, the encoder / decoder may derive a motion vector (i.e., a final motion vector) by adding the motion vector difference to the base motion vector. However, the present invention is not limited to these names, and the MVP may also be referred to as a base motion vector, a provisional motion vector, an initial motion vector, an MMVD candidate motion vector, etc. The MVD may be expressed as a value that refines the MVP, and may be referred to as an improved motion vector (refineMV) or a merge motion vector difference.

[0079] According to an embodiment of the present invention, when MMVD is applied, i.e., in the MMVD mode, the motion vector may be determined based on a base motion vector, a distance parameter (or variable), and a direction parameter (or variable). According to another embodiment of the present invention, the base motion vector may be determined from a candidate list. For example, the base motion vector may be determined from a merge candidate list. The encoder / decoder may also determine the base motion vector from a portion of another candidate list. The portion of the candidate list may be a front portion (a smaller index) of the candidate list. For example, the encoder / decoder may determine the base motion vector using the first and second candidates in the merge candidate list. To this end, a candidate index indicating a specific candidate from the two candidates may be signaled from the encoder to the decoder. Referring to FIG. 21, a base candidate index may be defined as an index for signaling the base motion vector. The encoder / decoder may determine a candidate to be applied to the current block from among the candidates in the candidate list based on the base candidate index, and determine the motion vector of the determined candidate as the base motion vector. In the present invention, the base candidate index is not limited to its name, and may be called a base candidate flag, a candidate index, a candidate flag, an MMVD index, an MMVD candidate index, an MMVD candidate flag, etc.

[0080] Furthermore, according to an embodiment of the present invention, an MVD different from the MVD described in FIGS. 6 and 7 may exist. For example, the MVD in the MMVD may be defined differently from the MVD described in FIGS. 6 and 7. In this specification, the MMVD may represent a merge mode (i.e., a motion compensation mode or method) using a motion vector difference, or may represent a motion vector difference when the MMVD is applied. For example, the encoder / decoder may determine whether or not to apply (or use) the MMVD. If the MMVD is applied, the encoder / decoder may determine a merge candidate to be used for inter-prediction of the current block from the merge candidate list, and determine the motion vector of the current block by deriving the MMVD and applying (or adding) it to the motion vector of the merge candidate.

[0081] In one embodiment, the other MVD may refer to a simplified MVD, an MVD with a different (or smaller) resolution, a MVD with a smaller number of available MVDs, an MVD with a different signaling method, etc. For example, the MVD used in the existing AMVP, affine inter mode, etc. described in FIGS. 6 and 7 can represent all regions in the x and y axes (i.e., horizontal and vertical directions) for a specific signaling unit (e.g., x-pel), e.g., a picture-based region (e.g., a picture region or a region including a picture and its surrounding regions) at uniform intervals, whereas the MMVD may have relatively limited units for representing a specific signaling unit. Also, the regions (or units) signaling the MMVD may not have uniform intervals. Also, the MMVD can indicate only a specific direction for a specific signaling unit.

[0082] Furthermore, according to one embodiment of the present invention, an MMVD may be determined based on distance and direction. Referring to FIG. 21, the distance and direction of the MMVD may be pre-set using a distance index indicating the distance of the MMVD and a direction index indicating the direction of the MMVD. In one embodiment, the distance may indicate the MMVD size (e.g., absolute value) in specific pixel units, and the direction may indicate the direction of the MMVD. Furthermore, the encoder / decoder may signal a relatively small distance with a relatively small index. That is, in the case of signaling that does not use fixed-length binarization, the encoder / decoder may signal a relatively small distance with relatively few bits.

[0083] In one embodiment of the present invention, MMVD-related syntax elements may be signaled when a merge flag (i.e., merge_flag) is 0 (i.e., when merge mode is not used). As described above, MMVD may be a method for signaling an MVD for a base candidate. In this respect, the MMVD mode may be similar to modes such as AMVP and affine AMVP (or affine inter) that signal an MVD. As a result, the MMVD mode may be signaled when the merge flag is 0. In operation S802, the decoder may parse MMVD-related syntax elements when an MMVD is applied to the current block, i.e., when the MMVD flag is 1. As an example, the MMVD-related syntax elements may include at least one of mmvd_merge_flag, mmvd_distance_idx, and mmvd_direction_idx. Here, mmvd_merge_flag represents a flag (or syntax element) indicating the base candidate of the MMVD, mmvd_distance_idx represents an index (or syntax element) indicating the distance value of the MVD, and mmvd_direction_idx represents an index (or syntax element) indicating the direction of the MVD.

[0084] Also, referring to FIG. 8, CuPredMode represents a variable (or value) indicating the prediction mode of the current block. Alternatively, the prediction mode of the current block may be a value indicating whether the current block is intra-predicted or inter-predicted. Alternatively, the prediction mode of the current block may be determined based on pred_mode_flag. Here, pred_mode_flag represents a syntax element indicating whether the current block is coded in inter-prediction mode or intra-prediction mode. If pred_mode_flag is 0, the prediction mode of the current block may be set to a value indicating the use of inter-prediction. The prediction mode value indicating the use of inter-prediction may be MODE_INTER. If pred_mode_flag is 1, the prediction mode of the current block may be set to a value indicating the use of intra-prediction. The prediction mode value indicating the use of intra-prediction may be MODE_INTRA. If pred_mode_flag is not present, CuPredMode may be set to a previously set value. Also, as an example, the previously set value may be MODE_INTRA.

[0085] Also, referring to FIG. 8, cu_cbf may be a value indicating whether a syntax associated with a transform exists. The syntax associated with a transform may be a transform tree syntax structure. The syntax associated with a transform may be a syntax signaled by the transform tree (transform_tree) of FIG. 28. Also, if cu_cbf is 0, the syntax associated with a transform may not exist. If cu_cbf is 1, the syntax associated with a transform may exist. Referring to FIG. 28, in step S803, if cu_cbf is 1, the decoder may invoke the transform tree syntax. If cu_cbf does not exist, the value of cu_cbf may be determined based on cu_skip_flag. For example, if cu_skip_flag is 1, cu_cbf may be 0. Also, if cu_skip_flag is 0, cu_cbf may be 1. As described above, cu_skip_flag is a syntax element that indicates whether or not to use skip mode. When skip mode is applied, the residual signal may not be used. That is, skip mode may be a mode in which a prediction signal is restored without adding a residual. Therefore, when cu_skip_flag is 1, it may mean that no syntax related to the transform is present.

[0086] According to an embodiment of the present invention, in step S802, the decoder may parse cu_cbf when intra prediction is not used. Also, the decoder may parse cu_cbf when cu_skip_flag is 0. Also, the decoder may parse cu_cbf when the merge flag is 0. These conditions may be combined and applied. For example, the decoder may parse cu_cbf when the prediction mode of the current block is not an intra prediction mode and the merge flag is 0. Or, the decoder may parse cu_cbf when the prediction mode of the current block is an inter prediction mode and the merge flag is 0. This is because skip mode may or may not be used in the case of inter prediction that is not a merge mode.

[0087] FIG. 9 is a diagram illustrating a coding unit syntax according to an embodiment of the present invention. According to an embodiment of the present invention, the cu_cbf and transform-related syntax of FIG. 8 may be modified as shown in FIG. 9. That is, according to an embodiment of the present invention, whether to use skip mode may be determined when a specific mode is applied. For example, whether to use skip mode may be determined when MMVD is applied. As an example, referring to FIG. 9, in step S901, a decoder may determine whether to parse cu_cbf based on whether MMVD is applied. That is, whether to use skip mode is determined depending on whether MMVD is applied, and whether to parse cu_cbf may be determined accordingly. If it is clear whether skip mode is to be used, the decoder does not need to parse cu_cbf.

[0088] In one embodiment, when MMVD is used, the encoder / decoder may not use skip mode. Since MMVD cannot accurately represent MVD like AMVP and can only be expressed within a limited range as described above, it can be more accurately restored using residuals. Therefore, by determining whether to parse cu_cbf based on whether MMVD is used, prediction accuracy and compression efficiency can be improved. For example, when MMVD is used, the decoder may not need to parse cu_cbf. If MMVD is not used, the decoder may be able to parse cu_cbf. In step S901, the decoder may be able to parse cu_cbf when the MMVD flag is 0, and may not need to parse cu_cbf when the MMVD flag is 1.

[0089] In one embodiment of the present invention, if cu_cbf is not present, the decoder can infer the cu_cbf value. According to the method described in FIG. 8, the decoder can infer the cu_cbf value based on the cu_skip_flag value. According to an embodiment of the present invention, the cu_cbf value can be inferred based on the merge flag. If the merge flag is 0, the decoder can infer cu_cbf to be 1. As an example, not using merge mode can indicate the presence of syntax related to the transformation. Thus, in the embodiments of FIGS. 8 and 9, when MMVD is used, the decoder can infer cu_cbf to be 1. As an example, the decoder can 1) infer cu_cbf to be 0 if the merge flag is 1 and cu_skip_flag is 1, 2) infer cu_cbf to be 1 if the merge flag is 1 and cu_skip_flag is 0, or 3) infer cu_cbf to be 1 if the merge flag is 0. Alternatively, as shown in FIG. 8, in one embodiment, the decoder can: 1) infer cu_cbf as 0 if cu_skip_flag is 1; and 2) infer cu_cbf as 1 if cu_skip_flag is 0.

[0090] According to yet another embodiment of the present invention, the cu_cbf value can be inferred based on the MMVD flag. When the MMVD flag is 1, cu_cbf can be inferred to be 1. Also, when the MMVD flag is 0, cu_cbf can be inferred to be 0 or 1. By combining this with the inference method described in FIG. 28, 1) when the MMVD flag is 1, cu_cbf can be inferred to be 1, 2) when the MMVD flag is 0 and cu_skip_flag is 1, cu_cbf can be inferred to be 0, and 3) when the MMVD flag is 0 and cu_skip_flag is 0, cu_cbf can be inferred to be 1.

[0091] Also, in one embodiment, the MMVD flag may not be present in the merge data (merge_data) syntax in the embodiment of Figures 8 and 9. The merge data syntax may be the merge data syntax shown in Figures 8 and 9.

[0092] 10 is a diagram illustrating a merge mode signaling method according to an embodiment of the present invention. In one embodiment of the present invention, the merge mode may be signaled based on syntax elements as shown in FIG. 10. Referring to FIG. 10, the merge mode may be signaled based on at least one of a regular flag, an MMVD flag, a subblock flag, and / or a combined inter-picture merge and intra-picture prediction (CIIP) flag. In the present invention, CIIP refers to a prediction method that combines inter-prediction (e.g., merge mode inter-prediction) and intra-prediction, and may be referred to as multi-hypothesis prediction.

[0093] Referring to FIG. 10, (a) and (b) of FIG. 10 may indicate cases corresponding to a non-skip merge mode and a skip merge mode, respectively. In the embodiment of FIG. 10, unlike the merge data syntax of FIGS. 8 and 9 described above, a regular flag may be present. For example, a triangle flag may not be present. The regular flag may be a syntax element indicating the use of a conventional merge mode, and in the present invention, the regular flag may be referred to as a regular merge flag. The conventional merge mode may be the same merge mode used in HEVC. Furthermore, the conventional merge mode may be a merge mode that uses a candidate indicated by a merge index and performs motion compensation without using MVD. In one embodiment, the regular flag, MMVD flag, sub-block flag, and CIIP flag may be signaled in a previously set order. The MMVD flag represents a syntax element indicating whether to use MMVD. The sub-block flag represents a syntax element indicating whether to use a sub-block mode in which sub-block-based prediction is performed. The CIIP flag represents a syntax element that indicates whether the CIIP mode is applied.

[0094] In one embodiment of the present invention, the signaling for indicating whether a corresponding mode is used in the normal flag, MMVD flag, sub-block flag, and CIIP flag may be equal to or less than 1. Thus, if a value of 1 occurs among the normal flag, MMVD flag, sub-block flag, and CIIP flag, the encoder / decoder may determine that a flag acquired subsequently in decoding order is 0. Also, if the normal flag, MMVD flag, sub-block flag, and CIIP flag are all 0, a mode not indicated by the normal flag, MMVD flag, sub-block flag, and CIIP flag may be used. The mode not indicated by the normal flag, MMVD flag, sub-block flag, and CIIP flag may be triangle prediction. That is, in one embodiment, if the normal flag, MMVD flag, sub-block flag, and CIIP flag are all 0, it may be determined that triangle prediction mode is applied.

[0095] 11 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. FIG. 11 illustrates a syntax structure in which the syntax elements previously described in FIG. 10 are used. In FIG. 11, a regular merge flag represents a regular merge flag. For example, the regular merge flag may be the regular flag previously described in FIG. 10.

[0096] According to an embodiment of the present invention, the regular merge flag may be located at the beginning of the merge data syntax. In operation S1101, the decoder may parse the regular merge flag first in the merge data syntax. That is, the regular merge flag may be the first syntax element to parse after confirming that the merge flag is 1. In operation S1102, the decoder may parse the MMVD flag if the regular merge flag is 0. In operations S1103, S1106, and S1105, the decoder may parse at least one of the subblock merge flag, the multiple hypothesis flag, and / or the triangle merge flag if the regular merge flag is 0. In FIG. 11, merge_subblock_flag represents a subblock merge flag indicating whether a subblock merge mode is applied, mh_intra_flag represents a multiple hypothesis prediction flag indicating whether a multiple hypothesis prediction mode is applied, and merge_triangle_flag represents a triangle merge flag indicating whether a triangle merge mode is applied.

[0097] 11, the decoder can parse the MMVD flag if the current block satisfies a predefined specific block size condition. In one embodiment, the triangle merge flag may be defined as (!regular_merge_flag&&!MMVD_flag&&!merge_subbock_flag&&!mh_intra_flag). That is, when the regular merge flag, MMVD flag, subblock merge flag, and multiple hypothesis prediction flag are all 0, the triangle merge flag may be 1. When any one of the regular merge flag, MMVD flag, subblock merge flag, and mh_intra_flag is 1, the triangle merge flag may be 0.

[0098] FIG. 12 illustrates a merge data syntax according to an embodiment of the present invention. FIG. 12 illustrates an example syntax structure in which the syntax elements described above in FIG. 10 are used. Referring to FIG. 12, in one embodiment of the present invention, a decoder first parses a regular merge flag, and if the parsed regular merge flag is 1, it can parse a merge index (S1201). Furthermore, the decoder can parse a merge index if MaxNumMergeCand is greater than 1. Here, MaxNumMergeCand is a variable indicating the maximum number of merge candidates. Furthermore, if the regular merge flag is 0, the decoder can parse at least one of the MMVD flag, sub-block merge flag, multiple hypothesis flag, and / or triangle merge flag. In one embodiment, the value of the triangle merge flag may be determined by the method described above in FIGS. 10 and 11. That is, the triangle merge flag may be determined based on a flag value indicating whether another mode is applied. A decoder can parse syntax (or syntax elements) related to triangle prediction if the triangle merge flag is 1. For example, a decoder can parse a triangle merge index (i.e., merge_triangle_idx) if the triangle merge flag is 1.

[0099] In the example of Figure 11, the merge index required for regular merge mode is located at the end of the merge data syntax, and signaling such as the MMVD flag, sub-block merge flag, and multiple hypothesis prediction flag exists between the regular merge flag and the merge index, which can result in inefficient signaling when using regular merge mode. However, in the example of Figure 12, when the regular merge flag is 1, the merge index can be parsed immediately after the regular merge flag, eliminating the need to parse other signaling unrelated to regular merge mode, which can improve compression efficiency.

[0100] Furthermore, according to an embodiment of the present invention, when various prediction modes are used, whether a specific prediction mode is parsed may be determined based on the application condition of the prediction mode, as described below with reference to Table 2.

[0101] [Table 2] If( A1 && A2 && A3 ) mode_A_flag If( mode_A_flag) { / / mode A related syntax elements } else { if( B1 && B2 && B3 ) mode_B_flag if( mode_B_flag ) { / / mode B related syntax elements } else { / / mode C related syntax elements } }

[0102] Referring to Table 2, it is assumed that Mode A, Mode B, and Mode C are present as prediction modes. It is also assumed that only one of Mode A, Mode B, and Mode C is used for prediction. It is also assumed that a condition for using Mode A may be defined, and that the conditions for using Mode A may be A1, A2, and A3. In this embodiment, if all of A1, A2, and A3 are satisfied, the encoder / decoder can apply Mode A. It is also assumed that the conditions for using Mode B may be B1, B2, and B3. In this embodiment, if all of B1, B2, and B3 are satisfied, the encoder / decoder can apply Mode B. It is also assumed that the conditions for using Mode C may be C1, C2, and C3, and that if all of C1, C2, and C3 are satisfied, the encoder / decoder can apply Mode C. Signaling (or a syntax element) indicating whether a given prediction mode X (Mode X) is used may be mode_X_flag.

[0103] Referring to Table 2, the decoder may parse the associated syntax to determine the prediction mode to be applied to the current block in the order of mode A, mode B, and mode C. Alternatively, the encoder may signal mode_A_flag, mode_B_flag, and mode_C_flag in that order, as shown in Table 2. If the decoder meets the conditions for using mode A, it may parse mode_A_flag. If mode_A_flag is 1, it may parse the syntax associated with mode A and not parse flags and associated syntax associated with the remaining modes. If mode_A_flag is 0, it may be possible to use mode B or mode C. Therefore, if the decoder meets the conditions for using mode B, it may parse mode_B_flag. If mode_B_flag is 1, it may parse the syntax associated with mode B and not parse mode_X_flag and associated syntax associated with the remaining modes (i.e., mode C). If mode_B_flag is 0, the decoder can determine that mode C is to be used. In other words, if all mode_X_flag that do not correspond to mode C are 0, the decoder can determine that mode C is to be used. Then, the decoder can parse syntax related to mode C.

[0104] Furthermore, according to an embodiment of the present invention, when various prediction modes are used, whether a specific prediction mode is parsed may be determined based on the application condition of the prediction mode, as described with reference to Table 3 below.

[0105] [Table 3] If( (A1 && A2 && A3) && !((!B1 || !B2 || !B3) && (!C1 || !C2 || !C3)) ) mode_A_flag If( mode_A_flag) { / / mode A related syntax elements } else { if( B1 && B2 && B3 ) mode_B_flag if( mode_B_flag ) { / / mode B related syntax elements } else { / / mode C related syntax elements } }

[0106] Referring to Table 3, Mode A, Mode B, and Mode C may be defined as prediction modes, similar to Table 2 above, and a syntax element (i.e., mode_X_flag) indicating whether a prediction mode is to be used and / or a syntax element indicating related prediction mode information may be defined. Also, X1, X2, X3, etc., which are conditions for using any mode X, may be defined. As with Table 2 above, it is determined whether mode A, mode B, and mode C are applied in that order, and if they are applied, the syntax element related to the prediction mode may be parsed.

[0107] According to an embodiment of the present invention, if none of the prediction modes determined later than a specific prediction mode can be used, the encoder / decoder may determine to use the specific prediction mode. In this case, the decoder may not need to parse a flag indicating whether the specific prediction mode is applicable (i.e., mode_X_flag when the specific prediction mode is mode X). In one embodiment, whether a prediction mode is unavailable may be determined based on whether the above-mentioned conditions for using a prediction mode are satisfied. For example, if neither mode B nor mode C, which are determined to be available relatively later, can be used, the decoder may not need to parse mode_A_flag and may determine (or determine or infer) to use mode A.

[0108] In the above Tables 2 and 3, the description is given assuming that three prediction modes, Mode A, Mode B, and Mode C, are applied, but the present invention is not limited to this number of prediction modes, and the mode can be determined using the proposed method even when more prediction modes exist. For example, assuming that Mode A, Mode B, Mode C, and Mode D are available, if none of Mode B, Mode C, and Mode D are available, the decoder can determine to use Mode A without additional signaling (or parsing). Furthermore, after determining not to use Mode A, if none of Mode C and Mode D are available, the decoder can determine to use Mode B.

[0109] Referring to Table 3, a condition that a given prediction mode X (i.e., mode X) cannot be used may be when any one of X1, X2, and X3 is not satisfied. That is, when !X1||!X2||!X3, mode X may not be usable. Therefore, when neither mode B nor mode C can be used, this represents a case where the ((!B1||!B2||!B3)&&(!C1||!C2||!C3)) condition is satisfied. When this condition is satisfied, the decoder does not need to parse mode_A_flag and can infer its value as 1. That is, the decoder can determine that mode A will be used. When the ((!B1||!B2||!B3)&&(!C1||!C2||!C3)) condition is not satisfied, the decoder can parse mode_A_flag. In this case, the decoder can also consider the condition for using mode A. That is, the decoder can parse mode_A_flag if !((!B1||!B2||!B3)&&(!C1||!C2||!C3)) is satisfied and (A1&& A2&& A3). In other words, the decoder can parse mode_A_flag if at least one of the conditions for using mode B and the conditions for using mode C is satisfied. The decoder can parse mode_A_flag if (B1&&B2&&B3) or (C1&&C2&&C3) is satisfied.

[0110] Also, if mode_A_flag is not present, the decoder can infer the value of mode_A_flag to be 0 when (B1&&B2&&B3) or (C1&&C2&&C3), and 1 otherwise. That is, if neither mode B nor mode C can be used, the decoder can infer the value of mode_A_flag to be 1 (i.e., apply mode A) when mode_A_flag is not present.

[0111] Tables 2 and 3 above are described assuming that prediction modes A, B, and C are selectively applied, and mode A, B, and C may be defined as specific prediction modes among various prediction modes proposed in the present invention. For example, mode A, B, and C may each be defined as one of normal merge mode, CIIP mode, and triangle merge mode. Alternatively, as described above, Tables 2 and 3 above may also be applied when mode A, B, C, and D are defined. For example, mode A, B, C, and D may each be defined as one of normal merge mode, MMVD mode, CIIP mode, and triangle merge mode.

[0112] 13 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. According to an embodiment of the present invention, the method described above in Table 2 and / or Table 3 may be applied to FIG. 13, and related duplicated descriptions will be omitted. Also, FIG. 13 may be an embodiment related to the regular merge flag as described in FIG. 10 and FIG. 11.

[0113] As described above, according to one embodiment of the present invention, if all modes whose use is determined relatively later than a particular mode in the decoding process order cannot be used, the decoder can determine (or decide, infer) that the particular mode will be used without parsing signaling indicating whether to use the particular mode. For example, if none of the modes whose use is determined relatively later than the use of the sub-block merging mode can be used, the signaling (or syntax element) indicating whether to use the sub-block merging mode may not be parsed. In this case, the decoder can determine that the sub-block merging mode will be used without syntax parsing. For example, the modes whose use is determined relatively later may include multiple hypothesis prediction and triangular prediction.

[0114] In one embodiment, if none of the modes for which the use of MMVD is determined later than the mode for which the use of MMVD is determined in step S1301 can be used, the decoder may determine to use MMVD without parsing the signaling indicating whether to use MMVD. For example, the modes for which the use of MMVD is determined later may include sub-block merging mode, multiple hypothesis prediction, and triangle prediction.

[0115] Furthermore, in the above-described embodiment, the conditions under which multiple hypothesis prediction can be used (i.e., mh_intra_conditions in FIG. 13) may include at least one of 1) sps_mh_intra_enabled_flag, 2) cu_skip_flag[x0][y0]==0, and 3) block size condition. As an example, the block size condition may be defined as ((cbWidth*cbHeight)>=64&&cbWidth<128&&cbHeight<128). Here, the sps_mh_intra_enabled_flag represents a syntax element indicating whether multiple hypothesis prediction can be used in the current sequence. For example, this syntax element may be signaled in a sequence parameter set (SPS). Furthermore, cbWidth and cbHeight are variables indicating the width and height of the current block (current coding block), respectively.

[0116] In addition, in the above-described embodiment, the conditions under which triangle prediction can be used (merge_triangle_conditions in FIG. 13) may include at least one of 1) sps_triangle_enabled_flag, 2) tile_group_type (or slice_type) == B, and 3) block size condition. For example, the block size condition may be defined as (cbWidth*cbHeight>=64). Here, the sps_triangle_enabled_flag represents a syntax element indicating whether triangle prediction is available in the current sequence. For example, the syntax element may be signaled by the SPS.

[0117] In addition, in the above-described embodiment, the conditions under which subblock merging can be used (merge_subblock_conditions in FIG. 13) may include at least one of 1) MaxNumSubblockMergeCand>0 and 2) block size conditions. For example, the block size condition may be defined as (cbWidth>=8&&cbHeight>=8). Here, MaxNumSubblockMergeCand is a variable indicating the maximum number of subblock merging candidates.

[0118] Thus, in one embodiment, the decoder may not parse the sub-block merge flag if (!mh_intra_conditions&& !merge_triangle_conditions), and if the sub-block merge flag is not present, the decoder may infer the sub-block merge flag to be 1 if (!mh_intra_conditions&& !merge_triangle_conditions), and 0 otherwise.

[0119] Also, in one embodiment, the decoder may not parse the MMVD flag if (!merge_subblock_conditions&&!mh_intra_conditions&&!merge_triangle_conditions), and if the MMVD flag is not present, the decoder may infer the MMVD flag to be 1 if (!merge_subblock_conditions&&!mh_intra_conditions&&!merge_triangle_conditions), and 0 otherwise.

[0120] Also, in one embodiment, if (!sps_mh_intra_enabled_flag&&!sps_triangle_enabled_flag), the decoder may not parse the sub-block merge flag and may infer its value as 1. Alternatively, if cu_skip_flag is 1 and tile_group_type (slice_type) is not B, the decoder may not parse the sub-block merge flag and may infer its value as 1. Alternatively, if the width and height are 128 and 128, respectively, and tile_group_type is not B, the decoder may not parse the sub-block merge flag and may infer its value as 1.

[0121] FIG. 14 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment described in FIG. 14 may be the same as the embodiments described in FIGS. 10 to 13, and redundant descriptions will be omitted for convenience. According to an embodiment of the present invention, merge modes may include a regular merge mode, an MMVD flag, a sub-block merge mode, a CIIP mode, a triangle merge mode (or a triangle partitioning mode (TPM))), etc. Furthermore, there may be signaling (or syntax elements) indicating whether a mode is used (or applied), such as a regular merge flag, an MMVD flag, a sub-block merge flag, a CIIP flag, and a triangle merge flag. As described above, prediction modes may include MODE_INTRA, MODE_IBC, and MODE_INTER. MODE_INTRA and MODE_IBC may be prediction modes using a current picture including a current block. MODE_INTRA may be the intra prediction described above. MODE_IBC may be a prediction method using a motion vector or a block vector. MODE_INTER may be a prediction method using another picture, or may be the inter prediction described above.

[0122] If the current prediction mode is MODE_IBC and the merge flag is 1, the decoder can use only the regular merge mode (S1401). In this case, the decoder does not need to parse the regular merge flag. The decoder can infer that the regular merge flag is 1.

[0123] According to an embodiment of the present invention, whether to parse a syntax element may be determined based on a block size. For example, whether to parse a syntax element may be determined based on the block size. For example, if syntax elements are signaled in the order of first mode, second mode, third mode, fourth mode, and fifth mode, third, fourth, and fifth mode may be used as block size conditions under which the third, fourth, and fifth modes can be used, respectively. If condition A, which is a condition in which none of conditions three, four, or fifth is satisfied, is satisfied, the decoder can infer that it will not parse or use syntax elements related to the third, fourth, or fifth modes. Furthermore, if condition A is satisfied, the decoder can determine whether to use syntax elements related to the second mode based on the first mode syntax elements. In this case, if it is determined or inferred that the first mode will not be used, the decoder can determine or infer that it will use the second mode. Then, based on this, it is possible to parse the syntax elements necessary to use the second mode.

[0124] According to one embodiment of the present invention, there may be block size conditions under which the sub-block merge mode, CIIP, and triangle merge mode can be used. For example, this may be as described in the embodiment of FIG. 13. Therefore, blocks of 4x4, 8x4, and 4x8 size may not be able to use the sub-block merge mode, CIIP, and triangle merge mode. Therefore, for blocks of 4x4, 8x4, and 4x8 size, when the merge flag is 1, the only modes available may be the regular merge mode and MMVD. Therefore, in this case, the decoder does not need to parse the MMVD flag. Also, the decoder can determine or infer the MMVD flag value based on the regular merge flag.

[0125] In one embodiment, the decoder may not perform inter prediction for 4x4 blocks. Therefore, the following embodiment can be described without including conditions related to 4x4 blocks, but the embodiment of the present invention can also be applied when 4x4 inter prediction is possible.

[0126] 14, when cbWidth and cbHeight are 8 and 4, or 4 and 8, respectively, the decoder does not need to parse the MMVD flag, sub-block merge flag, or multiple hypothesis prediction flag (S1402, S1403, S1404). Also, although not shown in FIG. 14, when cbWidth and cbHeight are 4 and 4, the decoder does not need to parse the MMVD flag, sub-block merge flag, or multiple hypothesis prediction flag. In this case, other syntax elements related to MMVD, sub-block merge mode, CIIP, and triangles also do not need to be parsed.

[0127] Furthermore, in the present invention, when cbWidth and cbHeight are 4 and 8, or 8 and 4, respectively, it can be indicated that cbWidth+cbHeight is 12. That is, when cbWidth+cbHeight is 12 or less than or equal to 12, it is not necessary to parse the MMVD flag, subblock merge flag, and mh_intra_flag. Furthermore, the present invention is applicable when the prediction mode is MODE_INTER.

[0128] According to an embodiment of the present invention, there may be higher-level signaling indicating whether MMVD is used. The higher-level signaling may be signaling for a unit including a current block. For example, the higher level of the current block may be a CTU, a sequence, a picture, a slice, a tile, a tile group, etc. For example, the higher-level signaling (or a syntax element) indicating whether MMVD is used may be SPS-level signaling. For example, the higher-level signaling indicating whether MMVD is used may be sps_mmvd_enabled_flag. The higher-level signaling indicating whether MMVD is used may indicate whether MMVD can be used. If the higher-level signaling indicating whether MMVD is used is 0, a decoder may not parse MMVD-related syntax elements. Also, if the higher-level signaling indicating whether MMVD is used is 0, a decoder may infer that the MMVD flag is 0. When the higher level signaling indicating whether the MMVD is available is 1, the MMVD flag may be 1 or 0 depending on the block.

[0129] Also, in one embodiment, the sub-block merging mode-related syntax elements may include a sub-block merging flag and a sub-block merging index. The sub-block merging mode may include SbTMVP (sub-block-based temporal motion vector) and affine motion compensation mode. Also, the CIIP-related syntax elements may include mh_intra_flag (CIIP flag) and an index indicating a candidate for the inter-prediction part of the CIIP. The index indicating a candidate for the inter-prediction part of the CIIP may be a merge index. As described above, the CIIP may be a method of performing prediction based on a prediction signal generated from a current picture and a prediction signal generated from another reference picture, and may be referred to as multiple hypothesis prediction.

[0130] In one embodiment, syntax elements related to a triangle merge mode may include merge_triangle_split_dir, merge_triangle_idx0, and merge_triangle_idx1. The triangle merge mode may be a prediction method (or prediction mode) that divides a current block into two parts and uses different motion information for each of the two parts. Each of the two parts may have any polygonal shape other than a rectangle. The present invention is not limited to these names, and the triangle merge mode may have various other names. merge_triangle_split_dir may be a syntax element that indicates how the two parts are divided. merge_triangle_idx0 and merge_triangle_idx1 may be syntax elements that indicate what motion information each of the two parts uses.

[0131] According to one embodiment of the present invention, the MMVD flag may not be present. For example, as described in FIG. 14, the MMVD flag may not be present depending on higher level signaling indicating whether the MMVD is available, block size conditions, etc. The following embodiment will show a method for inferring the absence of the MMVD flag. According to one embodiment of the present invention, if certain conditions are met, the decoder can infer the MMVD flag to be 1. If at least one of the certain conditions is not met, the decoder can infer the MMVD flag to be 0.

[0132] In one embodiment, the specific condition may include a case where higher level signaling (or a syntax element) indicating whether MMVD is used is 1. As described above, the higher level signaling may be included in any one of SPS, PPS, slice header, tile group header, and CTU. The specific condition may also include a block size condition. For example, the specific condition may include a case where the block size is 4x8, 8x4, or 4x4. That is, the specific condition may include a case where cbWidth+cbHeight is 12 or less. If 4x4 inter prediction is not allowed, the 4x4 case may be excluded. The specific condition may also include a case where a regular merge flag is 0. The specific condition may also include a case where a merge flag is 1.

[0133] Also, in one embodiment, if the MMVD flag is not present, the encoder / decoder can infer that the MMVD flag is 1 if 1) sps_mmvd_enabled_flag is 1, 2) cbWidth+cbHeight is 12, and 3) the regular merge flag is 0. Also, if at least one of 1), 2), and 3) is not met, the encoder / decoder can infer that the MMVD flag is 0.

[0134] According to one embodiment of the present invention, if the normal merge flag does not exist, the decoder can infer its value according to a predefined condition. In one embodiment, the decoder can infer the normal merge flag based on the prediction mode of the current block. For example, the decoder can infer the normal merge flag based on the CuPredMode value. For example, if the CuPredMode value is MODE_IBC, the decoder can infer the normal merge flag to be 1. Also, if the CuPredMode value is MODE_INTER, the decoder can infer the normal merge flag to be 0.

[0135] According to a further embodiment, a decoder can infer a canonical merge flag value based on the merge flag. For example, if the merge flag is 1 and CuPredMode is MODE_IBC, the decoder can infer the canonical merge flag value to be 1. Also, if the merge flag is 0, the decoder can infer the canonical merge flag value to be 0.

[0136] FIG. 15 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 15 may be yet another embodiment related to the embodiments described with reference to FIGS. 10 to 13. As described above, in an embodiment of the present invention, multiple modes may be defined as merge modes. When signaling whether a certain mode is to be used, whether to parse signaling indicating whether a certain mode is to be used, or a method for inferring signaling indicating whether a certain mode is to be used may be determined based on the signaling order for the multiple modes and the conditions under which the multiple modes can be used.

[0137] According to an embodiment of the present invention, a decoder can determine whether to parse signaling indicating whether a first mode is used based on higher-level signaling indicating whether a second mode is used. Also, the decoder can infer a signaling value indicating whether a first mode is used based on the higher-level signaling indicating whether a second mode is used. In this case, whether the second mode is used may be determined later than whether the first mode is used.

[0138] In a more specific embodiment, the decoder can determine whether it can parse the regular merge flag based on higher-level signaling indicating whether the MMVD is enabled. The decoder can also infer (or determine) the regular merge flag value based on the higher-level signaling indicating whether the MMVD is enabled. Referring to Figure 15, for example, the decoder can parse the regular merge flag if sps_mmvd_enabled_flag is 1 (S1501).

[0139] In one embodiment, whether to parse signaling indicating whether a specific mode is used can be determined based on the size of the current block. Furthermore, a signaling value indicating whether a specific mode is used can be inferred based on the size of the current block. According to one embodiment, even if the signaling indicating whether a specific mode is used is not parsed based on the size of the current block, the specific mode may still be used. That is, the signaling value indicating whether a specific mode is used can be inferred to be 1.

[0140] In a more specific embodiment, the decoder may determine whether to parse the regular merge flag based on the size of the current block. For example, the decoder may determine whether to parse the regular merge flag based on whether the size of the current block is larger than 4x8 or 8x4. A block size larger than 4x8 or 8x4 may have a sum of width and height greater than 12. Referring to FIG. 15, if the sum of cbWidth and cbHeight is greater than 12, the decoder may parse the regular merge flag (S1501). In addition, there may be a mode whose use is restricted for block sizes smaller than 4x8 and 8x4.

[0141] According to one embodiment of the present invention, if all of a plurality of conditions are satisfied, it is not necessary to parse signaling indicating whether a particular mode is to be used. In this case, the signaling indicating whether a particular mode is to be used can be inferred to be 1. A signaling indicating whether a particular mode is to be used that is 1 may indicate that the mode is to be used. As an example, the plurality of conditions may include a condition related to higher-level signaling indicating whether a second mode different from a first mode is to be used. For example, the plurality of conditions may include a condition that higher-level signaling indicating whether a second mode different from the first mode is to be used is 0. In this case, the second mode may be a mode whose availability is determined later than the first mode, or whose associated syntax element exists later.

[0142] In a more specific embodiment, the decoder may use a signal indicating whether a particular mode is used or not, such as a regular merge flag. Furthermore, the plurality of conditions may be a case in which a higher-level signaling value indicating whether an MMVD is used or not is 0. Furthermore, the plurality of conditions may include a condition related to block size. For example, the plurality of conditions may include a condition that the block size is equal to or less than a threshold value. In the condition in which the block size is equal to or less than a threshold value, one or more other modes whose use or non-use is determined later than the particular mode or whose associated syntax elements exist later may be unavailable.

[0143] More specifically, the signaling indicating whether to use a certain mode may be a regular merge flag. The plurality of conditions may include a case where the sum of the width and height of the current block is 12. Alternatively, the plurality of conditions may include a case where the size of the current block is 4x8 or 8x4. Furthermore, if 4x4 inter prediction is possible, the plurality of conditions may include a case where the size of the current block is 4x8, 8x4, or 4x4.

[0144] Therefore, according to one embodiment, if the higher level signaling value indicating whether or not the MMVD is used is 0 and the current block size is 4x8 or 8x4, the regular merge flag does not need to be parsed. In this case, the regular merge flag value can be inferred to be 1. Alternatively, if the higher level signaling value indicating whether or not the MMVD is used is 1 or the current block size is larger than 4x8 or 8x4, the regular merge flag can be parsed.

[0145] In step S1501, if sps_mmvd_enabled_flag is 1 or cbWidth+cbHeight>12, the decoder can parse the regular merge flag. Otherwise, that is, if sps_mmvd_enalbed_flag is 0 and cbWidth+cbHeight<=12, the decoder does not need to parse the regular merge flag.

[0146] In the above-described embodiment, as previously described with reference to FIGS. 10 to 13, the enablement conditions for modes associated with syntax elements present after the regular merge flag may be relevant. For example, when regular merge mode, MMVD, sub-block merge mode, CIIP, and triangle merge mode are signaled or determined in this order, if higher-level signaling indicating whether or not to use MMVD is 0 in the above-described embodiment, the decoder may not use MMVD. Also, if the block size is equal to or less than a threshold value, the decoder may not use sub-block merge mode, CIIP, or triangle merge mode. Therefore, if all of these conditions are met, the decoder can determine to use regular merge mode without further signaling. Furthermore, this embodiment is applicable to the case of MODE_INTER.

[0147] In one embodiment of the present invention, if certain predefined conditions are met, as shown in Figure 15, the normal merge flag does not need to be parsed, and in this case, the decoder can infer its value to be 1. For example, if the higher level signaling value indicating whether MMVD is available is 0 and the block size is 4x8 or 8x4, the decoder can infer the normal merge flag value to be 1. This may also be done when the merge flag is 1. This may also be done when CuPredMode is MODE_INTER. If the higher level signaling value indicating whether MMVD is available is 1 or the block size is larger than 4x8 or 8x4, the decoder can infer the normal merge flag value to be 0.

[0148] As an example, if the normal merge flag does not exist, the decoder can infer the normal merge flag according to the following conditions: Specifically, if sps_mmvd_enabled_flag is 0 and cbWidth+cbHeight==12, the decoder can infer that the normal merge flag is 1. In this case, if 4x4 inter prediction is allowed, the condition cbWidth+cbHeight==12 may change to cbWidth+cbHeight<=12. Otherwise, the decoder can infer that the normal merge flag is 0.

[0149] In one embodiment of the present invention, if the triangle merge flag, affine inter flag, and sub-block merge flag are all 0, the same motion information can be used for the entire current block. In this case, the following motion information deriving process can be performed. In addition, in this case, if one or more conditions are met, the decoder can set dmvrFlag to 1.

[0150] - If sps_dmvr_enabled_flag is 1,

[0151] - If merge_flag[xCb][yCb] is 1,

[0152] - If predFlagL0[0][0] and predFlagL1[0][0] are 1,

[0153] - If mmvd_flag[xCb][yCb] is 1,

[0154] - DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) and DiffPicOrderCnt(RefPicList[1][refIdxL1],currPic) are the same

[0155] - if cbHeight is greater than or equal to 8

[0156] - cbHeight*cbWidth is greater than or equal to 64

[0157] The motion information deriving process may also be performed for blocks of 4x8 or 8x4 size. If bi-prediction is used for blocks of 4x8 or 8x4 size, the decoder may switch to uni-prediction.

[0158] In addition, in an embodiment of the present invention, when the merge flag is 1 and the regular merge flag is 1, the same motion information can be used for the entire current block. Alternatively, when the merge flag is 1 and the MMVD flag is 1, the same motion information can be used for the entire current block. Alternatively, when the merge flag is 1 and the CIIP flag is 1, the same motion information can be used for the entire current block. Alternatively, when the merge flag is 0 and the inter_affine_flag is 0, the same motion information can be used for the entire current block. In this case, a motion information derivation process for such cases may be performed. Furthermore, if one or more predefined conditions are satisfied, the decoder may set dmvrFlag to 1. In this case, the conditions in the above-described embodiment may be applied. Furthermore, the motion information derivation process may also be performed for blocks of 4x8 or 8x4 size. If bi-prediction is used for a 4x8 or 8x4 size block, the decoder may switch to uni-prediction.

[0159] According to an embodiment of the present invention, among merge modes, CIIP may be the last mode determined or signaled. For example, the merge modes may be determined in the order of normal merge mode, MMVD, sub-block merge mode, triangle merge mode, and CIIP. In this case, if the conditions for using CIIP are not met, the decoder may determine the mode without parsing signaling indicating whether a mode determined earlier in the decoding order (or syntax parsing order) is available. For example, the decoder may not need to parse signaling indicating whether a mode immediately before CIIP is available. In this case, it may be determined to use the immediately previous mode. For example, this case may include a case where cu_skip_flag is 1. Alternatively, this case may be a case where cbWidth is 128 or greater or cbHeight is 128 or greater. Alternatively, this case may include a case where higher-level signaling indicating whether CIIP is available, such as sps_ciip_enabled_flag is 0.

[0160] FIG. 16 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiments of FIGS. 16 to 19 may be applied to the embodiments previously described with reference to FIGS. 10 to 13, and the related redundant description will be omitted. As described above, among the merge modes, CIIP may be the mode that is determined or signaled last. Therefore, a decoder can determine whether to use CIIP without parsing the CIIP flag. For example, if a mode signaled before CIIP is not used, it can be determined that CIIP is to be used. Furthermore, the CIIP flag may be a value derived from other signaling (or syntax elements).

[0161] According to one embodiment of the present invention, there may be multiple signalings indicating whether a mode is used. Referring to FIG. 53, the signaling indicating whether a mode is used may include a regular merge flag, an MMVD flag, a sub-block merge flag, and a triangle merge flag. Furthermore, there may be cases where the regular merge flag, the MMVD flag, the sub-block merge flag, and the triangle merge flag are parsed. For example, if the merge flag value is 1, the signaling indicating whether the mode is used can be parsed. Alternatively, if CuPredMode is MODE_INTER, the signaling indicating whether the mode is used can be parsed. Furthermore, if the merge flag value is 1, the decoder can parse the regular merge flag.

[0162] The decoder can also parse the MMVD flag if the normal merge flag value is 0. The decoder can also parse the MMVD flag if the sps_mmvd_enabled_flag value is 1. The decoder can also parse the MMVD flag if the block size condition is met. For example, the decoder can parse the MMVD flag if the block size is not 4x8, not 8x4, or not 4x4.

[0163] Also, the sub-block merge flag can be parsed when the normal merge flag value is 0. Also, the sub-block merge flag can be parsed when the MMVD flag value is 0. Also, the sub-block merge flag can be parsed when the block size condition is met. For example, the sub-block merge flag can be parsed when the block size is 8x8 or larger. Also, the sub-block merge flag can be parsed when the maximum number of sub-block merge candidates is greater than 0. For example, when the maximum number of sub-block merge candidates is greater than 0, it may indicate that at least one of higher level signaling regarding whether or not candidates that may be included in the sub-block merge candidate list can be used. For example, when sps_affine_enabled_flag or sps_sbtmvp_enabled_flag is 1, the maximum number of sub-block merge candidates may be greater than 0.

[0164] The decoder can parse the triangle merge flag if the normal merge flag value is 0. The decoder can parse the triangle merge flag if the MMVD flag value is 0. The decoder can parse the triangle merge flag if the subblock merge flag value is 0. The decoder can parse the triangle merge flag if the block size condition is met. For example, the triangle merge flag can be parsed if the block size meets the condition (width * height >= 64). The decoder can parse the triangle merge flag if the slice type is B. For example, a slice type of B may mean that two or more pieces of motion information can be used when predicting one sample. The decoder can parse the triangle merge flag if the sps_triangle_enabled_flag value is 1. The decoder can parse the triangle merge flag if the condition based on the maximum number of triangle merge candidates (MaxNumTriangleMergeCand) value is met. For example, the decoder can parse the triangle merge flag if the maximum number of triangle merge candidates is 2 or more. The maximum number of triangle merge candidates may be the maximum number (or length) of candidate lists that can be used in the triangle merge mode.

[0165] If the above-mentioned parsing conditions are met, the decoder can parse the signaling. That is, if any of the mentioned parsing conditions are not met, the signaling does not need to be parsed. Also, it is possible to infer when the signaling should not be parsed. For example, if any of the mentioned parsing conditions are not met, it can be inferred that the signaling value is 0. As yet another example, if any of the mentioned parsing conditions are not met, and the signaling indicating whether the first mode is used is 0, it can be inferred that the signaling value indicating whether the second mode is used is 1. As yet another example, if any of the mentioned parsing conditions are not met, and the signaling indicating whether the first mode is used is 1, it can be inferred that the signaling value indicating whether the second mode is used is 0.

[0166] According to another embodiment of the present invention, if the CIIP flag is not present, the decoder can infer its value. For example, the inferred value may be determined based on a signaling value indicating whether one or more modes are used. The signaling indicating whether a mode is used may include signaling indicating whether a mode determined before the CIIP is used may be used. For example, the signaling indicating whether a mode is used may include signaling indicating whether a regular merge mode is used, signaling indicating whether an MMVD is used, signaling indicating whether a sub-block merge mode is used, and signaling indicating whether a triangle merge mode is used. Furthermore, the signaling indicating whether a mode is used may include signaling indicating whether a merge mode is used.

[0167] According to one embodiment, if the signaling values ​​indicating whether one or more modes are used are all 0, the decoder can infer that the CIIP flag value is 1. The signaling indicating whether one or more modes are used may include a regular merge flag, an MMVD flag, a sub-block merge flag, and a triangle merge flag. Therefore, if the regular merge flag == 0 && the MMVD flag == 0 && the sub-block merge flag == 0 && the triangle merge flag == 0, the decoder can infer that the CIIP flag value is 1. Otherwise, the decoder can infer that the CIIP flag value is 0.

[0168] According to one embodiment, if signaling values ​​indicating whether one or more modes are used are all 0 and the merge flag is 1, the decoder can infer that the CIIP flag value is 1. The signaling indicating whether one or more modes are used may include a regular merge flag, an MMVD flag, a sub-block merge flag, and a triangle merge flag. Therefore, if the regular merge flag == 0 && the MMVD flag == 0 && the sub-block merge flag == 0 && the triangle merge flag == 0 && the merge flag == 1, the decoder can infer that the CIIP flag value is 1. Otherwise, the decoder can infer that the CIIP flag value is 0. Furthermore, a signaling value indicating whether a mode is used may indicate that the mode is used, and a signaling value indicating whether a mode is used may indicate that the mode is not used.

[0169] FIG. 17 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 17 is an efficient signaling method based on the embodiment of FIG. 16, and related redundant description will be omitted. As described above, among the merge modes, CIIP may be the mode that is determined or signaled last. According to an embodiment, in this case, the signaling method previously described with reference to FIGS. 10 to 13 can be used. The embodiments of FIGS. 17 to 19 may be specific examples of the methods described with reference to FIGS. 10 to 13.

[0170] According to one embodiment of the present invention, when mode use is determined or signaled in the order of Mode A, Mode B, Mode C, and Mode D, there may be conditions under which Mode D cannot be used. If at least one of the conditions under which Mode D cannot be used is satisfied, the decoder may not need to parse signaling indicating whether Mode C is used. Alternatively, if there is no signaling indicating whether Mode C is used, the decoder can infer its value. In this case, the inferred value may be based on the conditions under which Mode D cannot be used, the signaling indicating whether Mode A is used, and the signaling indicating whether Mode B is used. Alternatively, if none of the conditions under which Mode D cannot be used is satisfied, the decoder may parse signaling (or syntax elements) indicating whether Mode C is used. If there are multiple conditions under which Mode D cannot be used, only some of them can be used in the signaling method of the present invention, and therefore, the above conditions may be only some of the conditions. For example, when determining whether to parse signaling indicating whether Mode C is in use or not, only some of the conditions can be used to reduce the number of condition checks.

[0171] According to an embodiment, Mode D may be CIIP. Mode A, Mode B, and Mode C may be MMVD, sub-block merge mode, and triangle merge mode, respectively. In this case, Mode A, Mode B, and Mode C may be configured in a different order. FIGS. 17 to 19 assume that Mode A, Mode B, and Mode C are MMVD, sub-block merge mode, and triangle merge mode, respectively. According to an embodiment, the condition for disabling Mode D may be based on higher-level signaling indicating whether Mode D is available. The condition for disabling Mode D may be based on block size. The condition for disabling Mode D may be based on cu_skip_flag. The condition for disabling Mode D may be based on tile group (or slice) type. The condition for disabling Mode D may be based on the maximum number of candidates available for Mode D.

[0172] 17, conditions under which CIIP cannot be used include when sps_ciip_enabled_flag is 0, when cu_skip_flag is 1, when cbWidth is 128 or more, and when cbHeight is 128 or more. Therefore, according to an embodiment of the present invention, when sps_ciip_enabled_flag is 0, when cu_skip_flag is 1, when cbWidth is 128 or more, or when cbHeight is 128 or more, it is not necessary to parse the signaling indicating whether Mode C is being used. That is, in the embodiment of FIG. 17, when sps_ciip_enabled_flag is 0, when cu_skip_flag is 1, when cbWidth is 128 or more, or when cbHeight is 128 or more, it is not necessary to parse the triangle merge flag. Also, when sps_ciip_enabled_flag is 1, cu_skip_flag is 0, cbWidth is less than 128, and cbHeight is less than 128, the signaling indicating whether Mode C is in use can be parsed. That is, in the example of Figure 54, when sps_ciip_enabled_flag is 1, cu_skip_flag is 0, cbWidth is less than 128, and cbHeight is less than 128, the triangle merge flag can be parsed.

[0173] Furthermore, when determining whether to parse the signaling indicating whether Mode C is used, the conditions under which Mode C can be used may be further considered. For example, if the conditions under which Mode C can be used are met, the signaling (or syntax element) indicating whether Mode C is used may be parsed. Referring to FIG. 16, the conditions under which the triangle merge mode can be used include the condition that the sps_triangle_enabled_flag value is 1, the condition that the tile_group_type is B, and the condition that cbWidth*cbHeight>=64.

[0174] In one embodiment of the present invention, an example of an inference method related to the embodiment described in Fig. 17 will be described. This embodiment may be a method for inferring signaling indicating whether or not Mode C is being used, described in Fig. 17. Furthermore, inferring signaling indicating whether or not Mode C is being used can be performed when there is no signaling indicating whether or not Mode C is being used.

[0175] In the embodiment of FIG. 17 , when at least one of the conditions under which Mode D cannot be used is not satisfied, it is not necessary to parse the signaling indicating whether Mode C is used. According to one embodiment of the present invention, when multiple conditions are satisfied, the signaling value indicating whether Mode C is used can be inferred to be 1. For example, a value of 1 can indicate that Mode C is used, and a value of 0 can indicate that Mode C is not used. The multiple conditions can include a condition under which at least one of the conditions under which Mode D cannot be used is satisfied. The multiple conditions can also include a condition under which Mode C can be used. The multiple conditions can also include a condition based on the signaling indicating whether Mode A and Mode B are used. For example, the multiple conditions can include a case where the multiple conditions indicate that neither Mode A nor Mode B is used. If at least one of the multiple conditions is not satisfied, the decoder can infer that the signaling value indicating whether Mode C is used is 0.

[0176] In one embodiment of the present invention, a decoder can infer a triangle merge flag value based on predefined conditions. For example, the decoder can infer a triangle merge flag value of 1 if sps_ciip_enabled_flag is 0, cu_skip_flag is 1, cbWidth is 128 or greater, or cbHeight is 128 or greater. For example, the decoder can infer a triangle merge flag value of 1 only if sps_ciip_enabled_flag is 0, cu_skip_flag is 1, cbWidth is 128 or greater, or cbHeight is 128 or greater. In addition, additional conditions may be satisfied to infer a triangle merge flag value of 1. For example, the additional conditions may include a condition that the regular merge flag is 0, a condition that the MMVD flag is 0, or a condition that the sub-block merge flag is 0. The additional conditions may also include a condition that the merge flag is 1. The additional conditions may include the condition that sps_triangle_enabled_flag is 1, the condition that tile_group_type is B, and the condition that cbWidth*cbHeight>=64. When all the additional conditions are met, the triangle merge flag value can be inferred to be 1.

[0177] In one embodiment, the triangle merge flag value may be inferred as 1 if all of the following conditions are met:

[0178] 1) Regular merge flag == 0

[0179] 2) MMVD flag == 0

[0180] 3) Subblock merge flag == 0

[0181] 4) sps_ciip_enabled_flag==0||cu_skip_flag==1||cbWidth>=128||cbHeight>=128

[0182] 5) sps_triangle_enabled_flag==1&&tile_group_type==B&&cbWidth*cbHeight>=64

[0183] Or, in another embodiment, the triangle merge flag value may be inferred to be 1 if all of the following conditions are met:

[0184] 1) Regular merge flag == 0

[0185] 2) MMVD flag == 0

[0186] 3) Subblock merge flag == 0

[0187] 4) sps_ciip_enabled_flag==0||cu_skip_flag==1||cbWidth>=128||cbHeight>=128

[0188] 5) sps_triangle_enabled_flag==1&&tile_group_type==B&&cbWidth*cbHeight>=64

[0189] 6) Merge flag == 1

[0190] In addition, in one embodiment, if any one of the above conditions is not met, the triangle merge flag value can be inferred to be 0. For example, if sps_ciip_enabled_flag is 1, cu_skip_flag is 0, cbWidth<128, and cbHeight<128, the triangle merge flag value can be inferred to be 0. Alternatively, if the regular merge flag is 1, the triangle merge flag value can be inferred to be 0. Alternatively, if the MMVD flag is 1, the triangle merge flag value can be inferred to be 0. Alternatively, if the subblock merge flag is 1, the triangle merge flag value can be inferred to be 0. Alternatively, if sps_triangle_enalbed_flag is 0, or tile_group_type is not B, or cbWidth*cbHeight<64, the triangle merge flag value can be inferred to be 0. Alternatively, if the merge flag is 0, the triangle merge flag value can be inferred to be 0.

[0191] FIG. 18 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 18 may be a specific embodiment of the method described in FIG. 17. In the description of FIG. 17, when determining whether to parse the signaling indicating whether Mode C is used, as described above, only some of the conditions may be used to reduce the number of condition checks. For example, the embodiment of FIG. 18 may be a method in which the sps_ciip_enabled_flag in FIG. 17 is not checked.

[0192] For example, if cu_skip_flag is 1, or cbWidth >= 128, or cbHeight >= 128, the decoder may not parse the triangle merge flag. The decoder may then infer that the triangle merge flag value is 1. Alternatively, the decoder may infer that the triangle merge flag value is 1 only if the above conditions are met. As described above, the decoder may infer that the triangle merge flag value is 1 if additional conditions are met. The decoder may also parse the triangle merge flag if cu_skip_flag is 0, cbWidth < 128, and cbHeight < 128. If cu_skip_flag is 0, cbWidth < 128, and cbHeight < 128, and the triangle merge flag is not present, the decoder may infer that the value is 0.

[0193] This embodiment has the advantage of being able to reduce the number of condition checking operations in the syntax element parsing process compared to the embodiment of Figure 17. As mentioned above, if the mode signaling order is configured differently, the present invention can be applied to other signaling instead of the triangle merge flag.

[0194] FIG. 19 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 19 may be a specific embodiment of the method described in FIG. 17. In the description of FIG. 17, it was mentioned that only some conditions are used to reduce the number of condition checks when determining whether to parse signaling indicating the use of Mode C. This can be illustrated in FIG. 19. For example, the embodiment of FIG. 19 may be a method in which sps_ciip_enabled_flag is not checked in FIG. 17, and whether cbWidth is less than 128 and whether cbHeight is less than 128 are not checked.

[0195] For example, if cu_skip_flag is 1, the decoder may not parse the triangle merge flag. The decoder may then infer that the triangle merge flag value is 1. Alternatively, the decoder may infer that the triangle merge flag value is 1 only if this condition is met. The decoder may also infer that the triangle merge flag value is 1 if a further condition is met, as described above. The triangle merge flag may also be parsed if cu_skip_flag is 0. The decoder may also infer that the triangle merge flag is not present if its value is 0.

[0196] This embodiment has the advantage of reducing the number of condition checking operations during syntax element parsing compared to the embodiment of Figure 17. As mentioned above, if the mode signaling order is configured differently, the present invention can be applied to other signaling instead of the triangle merge flag.

[0197] FIG. 20 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 20 may be a specific embodiment of the method described in FIG. 17. The embodiments of FIGS. 20 to 24 may be specific embodiments of the above-mentioned invention. For example, the embodiments of FIGS. 20 to 24 may be related to the embodiments described in FIGS. 10 to 13, and related overlapping descriptions will be omitted.

[0198] 20, in the invention from FIG. 20 onward, signaling indicating whether or not MMVD is used may be mmvd_merge_flag. That is, the MMVD flag in the above-mentioned embodiment may be referred to as mmvd_merge_flag in the following description. Also, in the invention from FIG. 20 onward, signaling indicating a base candidate for MMVD may be mmvd_cand_flag. That is, what was previously referred to as mmvd_merge_flag may be mmvd_cand_flag in FIG. 20 onward. Also, values ​​for slice type may be applied to title group type (tile group type), and vice versa. Also, values ​​indicating slice type and title group type (tile group type) may be slice_type and tile_group_type, respectively. Also, signaling indicating whether or not the above-mentioned merge mode is used may be general_merge_flag. That is, what was previously described as a merge flag may be applied to the general_merge_flag, and what was previously described as the general_merge_flag can be applied to the merge flag.

[0199] According to an embodiment of the present invention, the merge mode signaled last among the various merge modes may be the sub-block merge mode. As described above, the merge modes may include the regular merge mode, MMVD, CIIP, triangle merge mode, sub-block merge mode, etc. Furthermore, the merge mode signaled immediately before the sub-block merge mode may be the triangle merge mode.

[0200] Referring to FIG. 20, an encoder / decoder may signal / parse the merge data syntax in the order of normal merge mode, MMVD, CIIP, triangle merge mode, and sub-block merge mode. According to an embodiment of the present invention, whether to parse the normal merge flag may be determined based on general_merge_flag. In this specification, general_merge_flag may be referred to as a general merge flag. According to an embodiment of the present invention, if general_merge_flag is 1, it may be possible to parse the normal merge flag. In this case, additional conditions for parsing may be required. Also, if general_merge_flag is 0, it may not be possible to parse the normal merge flag. In this case, if general_merge_flag is 0, it may not be possible to parse the normal merge flag regardless of other conditions. According to an embodiment of the present invention, if general_merge_flag is 1, it may be possible to parse the merge data structure portion of FIG. 20.

[0201] According to one embodiment of the present invention, a decoder can determine whether to parse mmvd_merge_flag based on general_merge_flag. According to one embodiment of the present invention, if general_merge_flag is 1, mmvd_merge_flag can be parsed. At this time, additional conditions for parsing may be required. Also, if general_merge_flag is 0, it may not be possible to parse mmvd_merge_flag. At this time, if general_merge_flag is 0, it may not be possible to parse mmvd_merge_flag regardless of other conditions.

[0202] According to one embodiment of the present invention, whether to parse the CIIP flag can be determined based on general_merge_flag. According to one embodiment of the present invention, when general_merge_flag is 1, it may be possible to parse the CIIP flag. At this time, additional conditions for parsing may be required. Also, when general_merge_flag is 0, it may not be possible to parse the CIIP flag. At this time, when general_merge_flag is 0, it may not be possible to parse the CIIP flag regardless of other conditions.

[0203] According to one embodiment of the present invention, whether to parse the triangle merge flag can be determined based on general_merge_flag. According to one embodiment of the present invention, when general_merge_flag is 1, it may be possible to parse the triangle merge flag. At this time, additional conditions for parsing may be required. Also, when general_merge_flag is 0, it may not be possible to parse the triangle merge flag. At this time, when general_merge_flag is 0, it may not be possible to parse the triangle merge flag regardless of other conditions.

[0204] According to one embodiment of the present invention, a decoder can determine whether to parse mmvd_merge_flag based on the canonical merge flag. According to one embodiment of the present invention, it may be possible to parse mmvd_merge_flag when the canonical merge flag is 0. In this case, additional conditions for parsing may be required. Also, it may not be possible to parse mmvd_merge_flag when the canonical merge flag is 1. In this case, it may not be possible to parse mmvd_merge_flag when the canonical merge flag is 1, regardless of other conditions.

[0205] According to one embodiment of the present invention, a decoder can determine whether to parse a CIIP flag based on mmvd_merge_flag. According to one embodiment of the present invention, when mmvd_merge_flag is 0, the decoder may be able to parse the CIIP flag. At this time, additional conditions for parsing may be required. Also, when mmvd_merge_flag is 1, the decoder may not be able to parse the CIIP flag. At this time, when mmvd_merge_flag is 1, the decoder may not be able to parse the CIIP flag regardless of other conditions.

[0206] According to one embodiment of the present invention, whether to parse the triangle merge flag can be determined based on the CIIP flag. According to one embodiment of the present invention, when the CIIP flag is 0, it may be possible to parse the triangle merge flag. In this case, additional conditions for parsing may be required. Also, when the CIIP flag is 1, it may not be possible to parse the triangle merge flag. In this case, when the CIIP flag is 1, it may not be possible to parse the triangle merge flag regardless of other conditions.

[0207] According to one embodiment of the present invention, it may be determined whether to parse the sub-block merge flag based on the triangle merge flag. According to one embodiment of the present invention, when the triangle merge flag is 0, it may be possible to parse the sub-block merge flag. In this case, additional conditions for parsing may be required. Also, when the triangle merge flag is 1, it may not be possible to parse the sub-block merge flag. In this case, when the triangle merge flag is 1, it may not be possible to parse the sub-block merge flag regardless of other conditions.

[0208] According to another embodiment of the present invention, the value indicating whether to use the last merge mode signaled among various merge modes can be determined without parsing. For example, referring to Figure 20, the sub-block merge flag may be determined without parsing. For example, when all of the following conditions are met, the sub-block merge flag may be determined to be 1:

[0209] 1) general_merge_flag==1

[0210] 2) If you do not use any of the merge modes that are signaled before the sub-block merge mode.

[0211] 3) When the conditions for using the sub-block merge mode are met

[0212] If this is not the case (ie, if at least one of the above conditions is not met), the sub-block merge flag can be determined to be 0.

[0213] For example, among the above conditions, “2) If all modes among the various merge modes that are signaled before the sub-block merge mode are not used,” may be defined as (or may include) the following condition:

[0214] (regular_merge_flag==0&&mmvd_merge_flag==0&&ciip_flag==0&&merge_triangle_flag==0)

[0215] Also, among the above conditions, "3) if the conditions for using the sub-block merge mode are satisfied" may be (or may include) the following conditions.

[0216] (MaxNumSubblockMergeCand>0&&cbWidth>=8&&cbHeight>=8)

[0217] Alternatively, the condition "3)" may be (or may include) the following condition:

[0218] (If at least one of the methods included in the subblock merge mode is enabled &&cbWidth>=8&&cbHeight>=8)

[0219] In addition, methods that can be included in the sub-block merging mode may include affine motion compensation and sub-block-based temporal motion vector prediction. Also, higher level signaling indicating whether affine motion compensation and sub-block-based temporal motion vector predictor are available may be defined as sps_affine_enabled_flag and sps_sbtmvp_enabled_flag, respectively. In this embodiment, the above conditions have been described using specific values ​​for width and height, but the present invention is not limited thereto and may be a condition based on a general block size.

[0220] FIG. 21 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 21 may be obtained by adding a more efficient signaling method to the embodiment of FIG. 20 described above. According to an embodiment of the present invention, if a mode among various modes does not satisfy the usability conditions for all of one or more modes signaled after the mode, the mode may be determined to be usable without explicitly signaling whether to use the mode. For example, it is not necessary to parse signaling indicating whether to use the mode.

[0221] For example, as shown in FIG. 21, if the conditions for using the sub-block merge mode are not met in the signaled syntax structure, it can be determined whether to use the triangle merge mode without explicitly signaling whether to use it. For example, if the conditions for using the sub-block merge mode are not met, the triangle merge flag does not need to be parsed. In one embodiment, the conditions for using the sub-block merge mode may be the same as when the conditions for using the sub-block merge mode in FIG. 20 are met (or may include the following conditions).

[0222] Therefore, referring to FIG. 21 , the decoder may not parse the triangle merge flag if MaxNumSubblockMergeCand is 0, where MaxNumSubblockMergeCand is a variable indicating the maximum number of subblock merge candidates. Alternatively, the decoder may not parse the triangle merge flag based on the block size. Alternatively, the decoder may not parse the triangle merge flag if cbWidth is less than 8. That is, the decoder may not parse the triangle merge flag if cbWidth is 4 (or less than or equal to 4). Alternatively, the decoder may not parse the triangle merge flag if bHeight is less than 8. That is, the decoder may not parse the triangle merge flag if cbHeight is 4 (or less than or equal to 4).

[0223] Therefore, according to an embodiment of the present invention, a decoder does not need to parse the triangle merge flag for a 4-by-X block or an X-by-4 block (for a block whose width or height is 4). In this case, a method for inferring the triangle merge flag will be described below. In one embodiment of the present invention, the minimum value of cbWidth and cbHeight may be 4. For example, for a luminance block, the minimum value of cbWidth and cbHeight may be 4. Also, cbWidth and cbHeight may be in the form of an exponent of 2. Therefore, for example, cbWidth being 8 or greater may be the same as cbWidth not being 4. As a further example, the maximum value of cbWidth and cbHeight may be 128.

[0224] 21, various merge modes are signaled in the order of triangle merge mode and sub-block merge mode, but the invention is not limited to this and can also be applied to cases where CIIP and sub-block merge mode are signaled in this order. That is, in the above-described embodiment, the triangle merge mode and triangle merge flag can be replaced with the CIIP and CIIP flag.

[0225] Figure 21 is a diagram illustrating a method for determining signaling indicating whether a mode is in use according to an embodiment of the present invention. The embodiment of Figure 21 may be applied to the method described in the embodiment of Figure 20, and related redundant description will be omitted. Referring to Figure 21, if the triangle merge flag is not present, the decoder can infer its value.

[0226] According to one embodiment of the present invention, when there is no signaling indicating whether a certain mode among various merge modes is used, its value can be inferred. As one example, a decoder can infer the value to be 1 if 1) none of the various merge modes signaled before the certain mode among the various merge modes are used, 2) the usability conditions for all of the various merge modes signaled after the certain mode are not met, and 3) the usability conditions for the certain mode are met. Otherwise (i.e., if 1), 2), or 3) is not met), the decoder can infer the value to be 0. Here, in condition "2"), not meeting the usability conditions for all modes may mean not meeting at least one of the usability conditions for each of all modes.

[0227] In addition, the condition 4) of using one of various merge modes may be added to the conditions for inferring that the signaling indicating whether or not a certain mode is used is 1. For example, 4) the case where general_merge_mode is 1 may be added.

[0228] For example, based on the example of Figure 20, if 1) none of the various merge modes signaled before the triangle merge mode are used, 2) the conditions for using all of the various merge modes signaled after the triangle merge mode are not met, 3) the conditions for using the triangle merge mode are met, and 4) general_merge_mode is 1, the decoder can infer that the triangle merge flag is 1. Otherwise (i.e., if 1 or 2 or 3 or 4) are not met), the decoder can infer that the triangle merge flag is 0.

[0229] Referring to FIG. 21, "1) among various merge modes, all modes signaled before the triangle merge mode are not used" may be the following condition.

[0230] (regular_merge_flag==0&&mmvd_merge_flag==0&&ciip_flag==0)

[0231] Also, referring to FIG. 21, "2) The conditions for using all merge modes signaled after the triangle merge mode among various merge modes are not met" may be the case where the conditions for using the sub-block merge mode are not met, and may include the following conditions: For example, a condition for block size.

[0232] (MaxNumSubblockMergeCand==0||cbWidth==4||cbHeight==4)

[0233] Also, referring to FIG. 60, "3) Satisfying the conditions for using the triangle merge mode" may be the following conditions.

[0234] (MaxNumTriangleMergeCand>=2&&sps_triangle_enabled_flag&&slice_type==B&&cbWidth*cbHeight>=64)

[0235] As a further example, in the embodiments of FIGS. 20 and 21, some condition checks may not be included to reduce the calculations required for the condition checks. For example, when parsing or inferring a triangle merge flag, the decoder may not use the condition for which the subblock merge mode can be used or some of the conditions listed in the conditions for which the subblock merge mode can be used. In this case, the conditions used in the parsing step and the inference conditions may be the same. For example, when parsing or inferring a triangle merge flag, the decoder may not check the condition for MaxNumSubblockMergeCand. That is, the triangle merge flag can be parsed even if MaxNumSubblockMergeCand is 0. If the triangle merge flag is not present, the decoder may not check whether MaxNumSubblockMergeCand is 0 when inferring its value.

[0236] FIG. 22 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 22 may be a specific embodiment that combines the embodiments of FIGS. 16 to 19 with the embodiments of FIGS. 20 and 21. According to an embodiment of the present invention, signaling overhead can be reduced. According to an embodiment of the present invention, signaling may be performed in the order of Mode A, Mode B, Mode C, Mode D, and Mode E. In this case, if Mode D or Mode E is available, signaling indicating whether Mode C is used or not can be parsed. If neither Mode D nor Mode E is available, signaling indicating whether Mode C is used or not may not be parsed. Furthermore, if neither Mode D nor Mode E is available, or if neither Mode A nor Mode B is used, it may be determined that signaling indicating whether Mode C is used or not is to be used.

[0237] Referring to FIG. 22, various merge modes may be signaled in the order of regular merge, MMVD, triangle merge mode, sub-block merge mode, and CIIP. In this case, according to an embodiment of the present invention, if CIIP is not available, the decoder does not need to parse the signaling indicating whether or not the sub-block merge mode is used. Alternatively, if CIIP is available, the decoder can parse the signaling indicating whether or not the sub-block merge mode is used. Alternatively, if CIIP is not available, regular merge is not used, MMVD is not used, triangle merge mode is not used, and general_merge_flag is 1, the decoder can infer that the signaling indicating whether or not the sub-block merge mode is used is used. Otherwise, the decoder can infer that the signaling indicating whether or not the sub-block merge mode is not used.

[0238] For example, the conditions for using CIIP may include one or more of the following conditions, arranged in an && (and) relationship: 1) a condition based on higher-level signaling indicating whether CIIP is available, 2) a condition based on cu_skip_flag, and 3) a condition based on block size (width or height). Referring to Figure 61, the conditions for using CIIP may include one or more of the following conditions, arranged in an && (and) relationship: 1) sps_ciip_enabled_flag, 2) cu_skip_flag==0, and 3) cbWidth*cbHeight>=64&&cbWidth<128&&cbHeight<128. Referring to Figure 22, the case where CIIP is available may be (sps_ciip_enabled_flag&&cu_skip_flag==0&&cbWidth*cbHeight>=64&&cbWidth<128&&cbHeight<128).

[0239] Furthermore, according to an embodiment of the present invention, if neither the sub-block merge mode nor the CIIP can be used, it is not necessary to parse the signaling indicating whether the triangle merge mode is used. Furthermore, if either the sub-block merge mode or the CIIP can be used, the signaling indicating whether the triangle merge mode is used can be parsed. Furthermore, if neither the sub-block merge mode nor the CIIP can be used, regular merge is not used, MMVD is not used, and general_merge_flag is 1, it can be inferred that the signaling indicating whether the triangle merge mode is used is used. Otherwise, it can be inferred that the triangle merge mode is not used. For example, the conditions under which the CIIP can be used and the conditions under which the CIIP cannot be used can be referenced above. However, here, overlapping conditions between the conditions under which the triangle merge mode can be used and the conditions under which the CIIP can be used (e.g., cbWidth*cbHeight>=64 in FIG. 22) can be omitted from the conditions under which the CIIP can be used. In addition, the conditions for using the subblock mergeonjungko mode may include one or more of the following conditions, which are combined with an && (and) condition: 1) a condition based on MaxNumSubblockMergeonjungkoCand; and 2) a condition based on the block size.

[0240] Referring to Figure 22, the conditions under which the sub-block merge mode can be used may include one or more of the following conditions, expressed as && (and): 1) MaxNumSubblockMergeCand>0, and 2) cbWdith>=8&& cbHeight>=8. Referring to Figure 2, the conditions under which the sub-block merge mode can be used may be (MaxNumSubblockMergeCand>0&& cbWdith>=8&& cbHeight>=8). If the sub-block merge mode cannot be used, the condition may be a NOT condition for the cases under which the sub-block merge mode can be used.

[0241] FIG. 23 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 23 may be yet another embodiment similar to that of FIG. 22. According to an embodiment of the present invention, signaling may be performed in the order of Mode A, Mode B, Mode C, Mode D, and Mode E. In this case, if Mode D or Mode E is available, the signaling indicating whether Mode C is used can be parsed. If neither Mode D nor Mode E is available, the signaling indicating whether Mode C is used can be omitted. Furthermore, if neither Mode D nor Mode E is available, and if neither Mode A nor Mode B is used, it can be determined that the signaling indicating whether Mode C is used can be used.

[0242] Referring to FIG. 23, various merge modes may be signaled in the order of regular merge, MMVD, triangle merge mode, CIIP, and sub-block merge mode. In this case, according to an embodiment of the present invention, if the sub-block merge mode is unavailable, it is not necessary to parse the signaling indicating whether CIIP is used. Also, if the sub-block merge mode is available, the signaling indicating whether CIIP is used can be parsed. Also, if the sub-block merge mode is unavailable, regular merge is not used, MMVD is not used, triangle merge mode is not used, and general_merge_flag is 1, it can be inferred that the signaling indicating whether CIIP is used is used. Otherwise, it can be inferred that CIIP is not used. For this, see the description of FIG. 21. Also, according to an embodiment of the present invention, if neither CIIP nor sub-block merge mode is available, and CIIP or sub-block merge mode is available, see the description of FIG. 22 for the signaling indicating whether triangle merge mode is used.

[0243] FIG. 24 illustrates a merge data syntax structure according to an embodiment of the present invention. The embodiments of FIGS. 24 and 25 illustrate an embodiment in which conditions for using the triangle merge mode are added to the embodiment of FIG. 17. According to an embodiment of the present invention, the conditions for using the triangle merge mode may include a condition related to the maximum number of triangle merge mode candidates. For example, the value indicating the maximum number of triangle merge mode candidates may be MaxNumTriangleMergeCand. For example, to be able to use the triangle merge mode, (MaxNumTriangleMergeCand>=2) must be satisfied.

[0244] 24, if (MaxNumTriangleMergeCand>=2) is satisfied, the triangle merge flag can be parsed, and if (MaxNumTriangleMergeCand>=2) is not satisfied, the triangle merge flag does not need to be parsed. Other explanations that overlap with those in FIG. 17 will be omitted.

[0245] Therefore, the decoder can parse the triangle merge flag when all of the following conditions are met: If at least one of the following conditions is not met, the decoder does not need to parse the triangle merge flag.

[0246] 1) MaxNumTriangleMergeCand>=2

[0247] 2)sps_triangle_enabled_flag

[0248] 3) slice_type==B

[0249] 4)cbWidth*cbHeight>=64

[0250] 5)sps_ciip_enabled_flag

[0251] 6)cu_skip_flag==0

[0252] 7)cbWidth<128

[0253] 8)cbHeight<128

[0254] In yet another embodiment, some of the conditions may be omitted, perhaps to reduce the computation required to check the conditions. For example, the omitted conditions may be at least one of 5), 6), 7), and 8).

[0255] In addition, in one embodiment of the present invention, the CIIP flag may be determined as follows: If all of the following conditions are met, the CIIP flag may be set to 1:

[0256] a)general_merge_flag==1

[0257] b)regular_merge_flag==0

[0258] c) mmvd_merge_flag==0

[0259] d)merge_subblock_flag==0

[0260] e)merge_triangle_flag==0

[0261] f)sps_ciip_enabled_flag==1

[0262] g)cu_skip_flag==0

[0263] h)cbWidth*cbHeight>=64

[0264] i)cbWidth<128

[0265] j)cbHeight<128

[0266] If at least one of the above conditions is not met, the CIIP flag may be set to 0. For example, conditions h), i), and j) may be replaced with other block size conditions.

[0267] In one embodiment of the present invention, if the triangle merge flag is not present, the triangle merge flag may be inferred by the following process: If all of the following conditions are met, the triangle merge flag may be inferred to be 1:

[0268] 1)regular_merge_flag==0

[0269] 2) mmvd_merge_flag==0

[0270] 3)merge_subblock_flag==0

[0271] 4) sps_ciip_enabled_flag==0||cu_skip_flag==1||cbWidth>=128||cbHeight>=128

[0272] 5)MaxNumTriangleMergeCand>=2&&sps_triangle_enabled_flag==1&&tile_group_type==B&&cbWidth*cbHeight>=64

[0273] 6)general_merge_flag==1

[0274] In other cases, the triangle merge flag value can be inferred to be 0. Among the above conditions, those connected with || (i.e., OR) in condition 4) correspond to conditions 5), 6), 7), and 8) described in FIG. 63. However, if any of 5), 6), 7), and 8) is omitted, it may also be omitted in condition 4) of FIG. 64. As previously described in FIG. 24, the conditions for using the triangle merge mode may include a condition regarding the maximum number of triangle merge mode candidates. Related redundant explanations will be omitted.

[0275] Furthermore, according to an embodiment of the present invention, at least one of the various merge modes may be used for signaling indicating whether or not it is used. For example, when the merge mode is used (general_merge_flag is 1), at least one of the various merge modes may indicate that it uses signaling indicating whether or not it is used. According to an embodiment, the at least one mode may be a pre-configured mode. For example, the at least one mode may be a single mode. For example, the at least one mode may be a normal merge mode.

[0276] According to one embodiment, when a merge mode is used, if signaling indicating whether or not various merge modes are used is set to not be used, signaling indicating whether or not a previously set mode is used can be set to be used. According to another embodiment, when a merge mode is used, if signaling indicating whether or not various merge modes are used is set to not be used except for a previously set mode, signaling indicating whether or not a previously set mode is used can be set to be used. This may be to prevent erroneous signaling and operations resulting therefrom.

[0277] In one embodiment of the present invention, the regular merge flag may be set to 1 if all of the following conditions are met:

[0278] 1)regular_merge_flag==0

[0279] 2) mmvd_merge_flag==0

[0280] 3)merge_subblock_flag==0

[0281] 4)ciip_flag==0

[0282] 5)merge_triangle_flag==0

[0283] 6)general_merge_flag==1

[0284] According to still another embodiment, some of the above conditions may be omitted, for example, condition 1) may be omitted.

[0285] According to another embodiment of the present invention, if the enablement conditions for all modes except one specific mode among the various merge modes are not met, the decoder may infer the value of 1 without parsing the signaling indicating whether the specific mode is used. Alternatively, if the enablement conditions for at least one mode except one specific mode among the various merge modes are met, the decoder may parse the signaling indicating whether the specific mode is used. In one embodiment, this may correspond to the case where a merge mode is used. The specific mode may be a regular merge mode.

[0286] More specifically, a decoder can parse a regular merge flag if at least one of the following conditions 1) to 4) is met. In an embodiment, this may be the case when using merge mode.

[0287] 1) sps_mmvd_enabled_flag&&cbWidth*cbHeight!=32

[0288] 2)MaxNumSubblockMergeCand>0&&cbWidth>=8&&cbHeight>=8

[0289] 3) sps_ciip_enabled_flag&&cu_skip_flag==0&&cbWidth*cbHeight>=64&&cbWidth<128&&cbHeight<128

[0290] 4)MaxNumTriangleMergeCand>=2&&sps_triangle_enabled_flag&&slice_type==B&&cbWidth*cbHeight>=64

[0291] Furthermore, if none of the above conditions 1) to 4) are satisfied, the normal merge flag does not need to be parsed and its value can be inferred to be 1. This may be the case when the merge mode is used. In this case, some of the above conditions may be omitted to reduce the amount of calculation.

[0292] According to another embodiment of the present invention, when signaling indicating the use of two or more merge modes indicates their use, signaling indicating the use of all modes except a pre-set mode among the various merge modes may be set to indicate that they are not to be used, and signaling indicating the use of the pre-set mode may be set to indicate that they are to be used. For example, the pre-set mode may be a normal merge mode. As another example, the pre-set mode may be one of the two or more modes indicated to be used by the signaling indicating the use. In this case, a pre-set method for determining one mode may exist. For example, the first mode among the pre-set modes may be determined. For example, this embodiment may correspond to a case where a merge mode is used.

[0293] For example, when regular_merge_flag==1 and merge_subblock_flag==1, merge_subblock_flag can be set to 0. Or, when ciip_flag==1 and merge_subblock_flag==1, ciip_flag and merge_subblock_flag can be set to 0, and regular_merge_flag can be set to 1. As another example, when ciip_flag==1 and merge_subblock_flag==1, among the already set order of regular merge mode, MMVD, subblock merge mode, CIIP, and triangle merge mode, the merge_subblock_flag that comes first can be set to 1, and the CIIP flag can be set to 1.

[0294] 25 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. The embodiment of FIG. 25 may be a specific embodiment of the invention described above. For example, the embodiment of FIG. 25 may be the same as the methods described in the embodiments of FIGS. 10 to 13, and related redundant descriptions will be omitted.

[0295] As described above, whether to use a plurality of merge modes may be signaled or whether to use a plurality of merge modes may be determined in the order in which they have been set. Referring to FIG. 66, the plurality of merge modes may include regular merge mode, MMVD, sub-block merge mode, CIIP, and triangle merge mode. Also, referring to FIG. 25, whether to use a plurality of merge modes may be signaled or whether to use a plurality of merge modes may be determined in the order of regular merge mode, MMVD, sub-block merge mode, CIIP, and triangle 1 merge mode. Also, referring to FIG. 25, signaling indicating whether to use the regular merge mode, MMVD, sub-block merge mode, CIIP, and triangle merge mode may be regular_merge_flag, mmvd_merge_flag, merge_subblock_flag, ciip_flag, and MergeTriangleFlag, respectively. Also, MergeTriangleFlag may have the same meaning as merge_triangle_flag described above.

[0296] Furthermore, there may be conditions that must be met in order for each mode to be implemented. For example, if a condition that must be met in order for a certain mode to be implemented is not met, the certain mode may not be implemented. In this case, a mode other than the certain mode may be implemented. Alternatively, if a condition that must be met in order for a certain mode to be implemented is met, the certain mode may or may not be implemented. In this case, there may be additional signaling to determine whether or not to implement the certain mode.

[0297] For example, the conditions that must be met to enable a certain mode may be based on higher-level signaling indicating whether the certain mode is available. The higher level may include a sequence level, a sequence parameter set (SPS) level, a slice level, a title level, a tile level, a title group level, a brick level, a CTU level, etc. The higher level may also include the aforementioned sps_mode_enabled_flag. In this case, the mode may be substituted for a certain mode.

[0298] Furthermore, the conditions that must be met for a certain mode to be enabled may include conditions related to block size. For example, conditions based on the width or height of the current block may be included. For example, there may be an upper or lower limit for the width. Or there may be an upper or lower limit for the height. Or there may be an upper or lower limit for area(width*height). Furthermore, the current block may be a CU or PU. Furthermore, the width and height of the current block may be cbWidth and cbHeight, respectively. In the present invention, width and height may be used interchangeably with cbWidth and cbHeight, respectively.

[0299] In addition, the conditions that must be met to enable a certain mode may be based on the slice type or the tile group type, which may have the same meaning.

[0300] Furthermore, the conditions that must be met to enable a certain mode may be based on whether another certain mode is used. The other certain mode may include a skip mode. Whether to use the skip mode may be determined based on cu_skip_flag. The other certain mode may include a mode that is signaled or determined before the certain mode. For example, if the other certain mode is not used, the certain mode may be enabled.

[0301] Furthermore, the condition that must be met to enable a certain mode may be based on the maximum number of candidates. For example, the candidates may be candidates associated with the certain mode. For example, the candidates may be candidates used by the certain mode. For example, a certain mode may be enabled when there are a sufficient number of candidates. For example, a certain mode may be enabled when the maximum number of candidates is equal to or greater than a preset value. For example, the maximum number of candidates may be indicated by a parameter called MaxNumModeCand, and the "mode" in MaxNumModeCand may be replaced with the mode indicated. For example, a MaxNumMergeCand value may exist for merge mode. For example, a MaxNumTriangleMergeCand value may exist for triangle merge mode. For example, a MaxNumSubblockMergeCand value may exist for subblock merge mode. Furthermore, the maximum number of candidates may be based on higher-level signaling indicating whether a certain mode is enabled. For example, MaxNumSubblockMergeCand may be based on sps_affine_enabled_flag or sps_sbtmvp_enabled_flag. sps_sbtmvp_enalbed_flag may be higher level signaling indicating whether subblock based temporal motion (vector) prediction is enabled or not.

[0302] According to an embodiment of the present invention, there may be a condition that must be met to enable the regular merge mode. For example, a condition may be that signaling indicating the use of the merge mode must be true in order to enable the regular merge mode. The signaling indicating the use of the merge mode may be merge_flag or general_merge_flag. Furthermore, for other modes described below, the other modes can only be enabled if signaling indicating the use of the merge mode is true.

[0303] Alternatively, unlike other modes, there may be no conditions that must be met to use the regular merge mode. This may be because the regular merge mode is the most basic mode. When using the merge modes described above, no additional conditions may be required to use the regular merge mode.

[0304] Referring to FIG. 25, if a first condition 2501 is satisfied, the decoder can parse the regular merge flag (i.e., regular_merge_flag). The first condition 2501 may be a case where sps_mmvd_enabled_flag is 1 or width*height is not 32. Alternatively, if the first condition 2501 is not satisfied, the decoder may not parse regular_merge_flag. In this case, the decoder can infer that the value is 1. For example, if sps_mmvd_enabled_flag is 0 && width*height == 32, the decoder can infer that regular_merge_flag is 1. Alternatively, in order to infer that the value is 1, a condition that general_merge_flag is 1 may be included. This may be because if the first condition 2501 is not satisfied, all of the conditions that must be satisfied to enable other modes belonging to the merge mode cannot be satisfied. Additionally, if regular_merge_flag is not present and the above condition for inferring 1 cannot be satisfied, the decoder can infer its value as 0. Furthermore, the width or height may each be a power of 2. Furthermore, the width or height may be a positive number. Therefore, a width*height of 32 may mean that the width and height are 4 and 8, or 8 and 4, respectively. Furthermore, a width*height that is not 32 may mean that the width and height are not 4 and 8, or 8 and 4, respectively. Furthermore, a width*height that is not 32 may also represent a width and height that are 8 or greater. This may correspond to inter prediction, for example, because inter prediction may not be allowed for 4x4 blocks.

[0305] According to one embodiment of the present invention, there may be a condition that must be met for MMVD to be enabled. For example, this may be based on higher-level signaling indicating whether MMVD is available. For example, the higher-level signaling indicating whether MMVD is available may be sps_mmvd_enabled_flag. Referring to FIG. 25, if the second condition 2502 is met, mmvd_merge_flag can be parsed. Also, if the second condition 2502 is not met, mmvd_merge_flag does not need to be parsed and its value can be inferred. The second condition 2502 may be (sps_mmvd_enabled_flag && cbWidth * cbHeight != 32). That is, when sps_mmvd_enabled_flag is 1 and the block size condition is met, mmvd_merge_flag can be parsed, and when sps_mmvd_enabled_flag is 0 or the block size condition is not met, mmvd_merge_flag does not need to be parsed. Also, when sps_mmvd_enabled_flag is 1 and the block size is not met, mmvd_merge_flag can be inferred to be 1. For example, when sps_mmvd_enabled_flag is 1, width*height is 32, regular_merge_flag is 0, and general_merge_flag is 1, mmvd_merge_flag can be inferred to be 1. The block size condition may be related to a condition under which a mode signaled after MMVD cannot be used.

[0306] According to an embodiment of the present invention, there may be conditions that must be met to enable the sub-block merging mode. For example, the conditions may be based on higher-level signaling indicating whether the sub-block merging mode is available. Alternatively, the conditions may be based on higher-level signaling indicating whether a mode belonging to the sub-block merging mode is available. For example, the sub-block merging mode may include affine motion prediction, sub-block-based temporal motion vector prediction, etc. Therefore, whether the sub-block merging mode is enabled may be determined based on higher-level signaling indicating whether affine motion prediction is available (e.g., sps_affine_enabled_flag).

[0307] Alternatively, whether the sub-block merging mode is performed may be determined based on higher-level signaling (e.g., sps_sbtmvp_enabled_flag) indicating whether sub-block-based temporal motion vector prediction is available. Alternatively, a condition based on the maximum number of candidates for the sub-block merging mode must be satisfied in order for the sub-block merging mode to be available. For example, the sub-block merging mode may be available when the maximum number of candidates for the sub-block merging mode is greater than 0. The maximum number of candidates for the sub-block merging mode may also be based on higher-level signaling indicating whether a mode belonging to the sub-block merging mode is available. For example, the maximum number of candidates for the sub-block merging mode may be greater than 0 only if at least one of the higher-level signaling indicating whether a mode belonging to a plurality of sub-block merging modes is available is 1. Alternatively, a condition based on the block size must be satisfied in order for the sub-block merging mode to be available. For example, there may be lower limits for width and height. For example, the sub-block merging mode may be available when the width is 8 or more and the height is 8 or more.

[0308] 25, if the third condition 2503 is satisfied, the subblock merge flag can be parsed. If the third condition 2503 is not satisfied, the merge_subblock_flag does not need to be parsed and its value can be inferred to be 0. The third condition 2503 may be (MaxNumSubblockMergeCand>0&&width>=8&&height>=8).

[0309] According to an embodiment of the present invention, conditions that must be met for CIIP to be performed may exist. For example, whether CIIP can be performed may be determined based on higher-level signaling (e.g., spsXBT_ciip_enabled_flag) indicating whether CIIP is available. Whether CIIP can be performed may also be determined based on whether skip mode is used. For example, CIIP may not be possible when skip mode is used. Whether CIIP can be performed may also be determined based on block size. For example, whether CIIP can be performed may be determined based on whether the block size is greater than or equal to a lower limit and less than or equal to an upper limit. For example, CIIP may be possible when width*height is greater than or equal to a lower limit, the width is less than or equal to an upper limit, and the height is less than or equal to an upper limit. For example, CIIP may be possible when width*height is greater than or equal to 64, the width is less than 128, and the height is less than 128.

[0310] 25, when the fourth condition 2504 is met, the CIIP flag can be parsed. When the fourth condition 2504 is not met, the CIIP flag does not need to be parsed and its value can be inferred to be 0. The fourth condition 2504 may be (sps_ciip_enabled_flag&&cu_skip_flag==0&&width*height>=64&&width<128&&height<128).

[0311] According to an embodiment of the present invention, there may be conditions that must be met to enable the triangle merge mode. For example, whether the triangle merge mode can be enabled may be determined based on higher-level signaling (e.g., sps_triangle_enabled_flag) indicating whether the triangle merge mode is available. Whether the triangle merge mode can be enabled may also be determined based on the slice type. For example, if the slice type is B, the triangle merge mode may be enabled. This may be because two or more pieces of motion information or two or more reference pictures are required to enable the triangle merge mode. Whether the triangle merge mode can be enabled may also be determined based on the maximum number of candidates for the triangle merge mode. The maximum number of candidates for the triangle merge mode may be indicated by a MaxNumTriangleMergeCand value. For example, if the maximum number of candidates for the triangle merge mode is two or more, the triangle merge mode may be enabled. This may be because two or more pieces of motion information or two or more reference pictures are required to enable the triangle merge mode. Furthermore, according to an embodiment of the present invention, when higher level signaling indicating whether the triangle merge mode is available indicates that it is available, the maximum number of candidates for the triangle merge mode is always 2 or more, and when higher level signaling indicating whether the triangle merge mode is available indicates that it is not available, the maximum number of candidates for the triangle merge mode may always be less than 2 or 0.

[0312] Therefore, in this case, whether the triangle merge mode can be performed may be determined based on the maximum number of candidates for the triangle merge mode, rather than based on higher-level signaling indicating whether the triangle merge mode is available. This reduces the number of calculations required to check conditions. Whether the triangle merge mode can be performed may also be determined based on the block size. For example, whether the triangle merge mode can be performed may be determined based on whether the block size is greater than or equal to the lower limit and less than or equal to the upper limit. For example, when width*height is greater than or equal to the lower limit, the width is less than or equal to the upper limit, and the height is less than or equal to the upper limit, the triangle merge mode may be performed. For example, when width*height is greater than or equal to 64, the triangle merge mode may be performed. Also, when the width is less than 128 and the height is less than 128, the triangle merge mode may be performed.

[0313] Referring to FIG. 25 , if a fifth condition 2505 is satisfied, the triangle merge mode may be used. Alternatively, if the fifth condition 2505 is not satisfied, the triangle merge mode may be used. The fifth condition 2505 may be a case where MergeTriangleFlag is 1. Furthermore, there may be other conditions that must be satisfied to satisfy the fifth condition 2505. For example, this may include (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2&&width*height>=64). If (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2&&width*height>=64) is true, MergeTriangleFlag may be 1 or 0. In this case, whether it is 1 or 0 may be determined based on additional conditions. The further condition may be when a mode signaled or determined before the triangle merge mode (e.g., normal merge mode, MMVD, subblock merge mode, CIIP) is not used or when a merge mode is used (e.g., general_merge_flag==1). If the further condition is met, MergeTriangleFlag may be 1, and if the further condition is not met, MergeTriangleFlag may be 0. Also, if (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2&&width*height>=64) is false, MergeTriangleFlag may be 0.

[0314] 26 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention, which may be a specific example of the signaling method described in FIGS.

[0315] According to an embodiment of the present invention, if at least one of a signaled mode or a determined mode after a certain mode can be performed, signaling indicating whether or not the certain mode is to be used can be parsed. Also, if neither a signaled mode nor a determined mode after a certain mode can be performed, the signaling indicating whether or not the certain mode is to be used may not be parsed. Also, if neither a signaled mode nor a determined mode after a certain mode can be performed, it may be determined to use the signaling value indicating whether or not the certain mode is to be used.

[0316] Furthermore, whether a mode signaled or determined after a certain mode can be performed may depend on whether the conditions that must be satisfied for the mode to be performed, as described in FIG. 26, are satisfied. Alternatively, whether a mode signaled or determined after a certain mode can be performed may depend on whether some of the conditions that must be satisfied for the mode to be performed, as described in FIG. 26, are satisfied. For example, some of the conditions that must be satisfied for the mode to be performed may be omitted when determining signaling for another mode. This may reduce the number of operations required to check the conditions, and for example, it may be possible to omit a condition that is often true.

[0317] More specifically, to use a certain mode, higher level signaling indicating whether the certain mode is available must be true and the slice type must be a specific value. However, when checking the conditions for using the certain mode to determine whether to parse the signaling indicating whether the certain mode and other modes are used, it can be determined whether to parse the signaling indicating whether the other modes are used based on the higher level signaling and not based on the slice type. This may be because the slice type is often a specific value. Therefore, even if the slice type is not a specific value, if the higher level signaling indicating whether the certain mode is available is true, the signaling indicating whether the certain mode and other modes are used can be parsed.

[0318] 26, mmvd_condition, subblock_merge_condition, ciip_condition, and triangle_merge_condition may exist. For example, mmvd_condition, subblock_merge_condition, ciip_condition, and triangle_merge_condition may be conditions that must be satisfied to enable the MMVD, subblock merge mode, CIIP, and triangle merge mode described in FIG. 25, respectively. Or, mmvd_condition, subblock_merge_condition, ciip_condition, and triangle_merge_condition may be part of the conditions that must be satisfied to enable the MMVD, subblock merge mode, CIIP, and triangle merge mode described in FIG. 25, respectively. For example, mmvd_condition, subblock_merge_condition, and ciip_condition may be the second condition 2502, the third condition 2503, and the fourth condition 2504 described in FIG. 25, respectively, or part thereof.

[0319] In addition, if the conditions that must be met for a mode that is signaled or determined after a certain mode to be able to be performed overlap with the conditions that must be met for the mode to be able to be performed, the mode can only be used if the conditions that must be met for the mode to be able to be performed are met, and such overlapping conditions can be excluded from the mmvd_condition, subblock_merge_condition, ciip_condition, triangle_merge_condition, etc. of Figure 26.

[0320] Referring to Figure 26, if triangle_merge_condition is satisfied, ciip_flag can be parsed. If triangle_merge_condition is not satisfied, ciip_flag does not need to be parsed. If triangle_merge_condition is not satisfied, ciip_flag can be inferred to be 1. In order to infer 1, the conditions for CIIP to be performed must be satisfied, and if no modes signaled or determined before CIIP are used (e.g., regular_merge_flag==0&&mmvd_merge_flag==0&&merge_subblock_flag==0), the condition for using merge mode (general_merge_flag==1) must be satisfied. In other cases, if ciip_flag does not exist, 0 can be inferred.

[0321] Referring to Figure 26, if either the ciip_condition or the triangle_merge_condition is satisfied, merge_subblock_flag can be parsed. If the ciip_condition and triangle_merge_condition are not satisfied, merge_subblock_flag does not need to be parsed. If the ciip_condition and triangle_merge_condition are not satisfied, merge_subblock_flag can be inferred as 1. To infer a value of 1, the conditions for subblock merge mode must be satisfied. If no modes signaled or determined before the subblock merge mode are used (e.g., regular_merge_flag == 0 && mmvd_merge_flag == 0), the condition for using merge mode (general_merge_flag == 1) must be satisfied. Otherwise, if merge_subblock_flag does not exist, it can be inferred as 0.

[0322] Referring to FIG. 26, if the subblock_merge_condition, ciip_condition, or triangle_merge_condition is satisfied, mmvd_merge_flag can be parsed. Furthermore, if the subblock_merge_condition, ciip_condition, and triangle_merge_condition are not satisfied, mmvd_merge_flag does not need to be parsed. Furthermore, if the subblock_merge_condition, ciip_condition, and triangle_merge_condition are not satisfied, mmvd_merge_flag can be inferred as 1. In order to infer a value of 1, the conditions for MMVD must be satisfied, and the conditions for using merge mode (general_merge_flag == 1) and not using any modes signaled or determined before MMVD must be satisfied. Furthermore, there may be cases where mmvd_merge_flag described in FIG. 66 is inferred as 1. Otherwise, when mmvd_merge_flag is not present, it can be inferred to be 0.

[0323] 26, if the mmvd_condition is satisfied, the subblock_merge_condition is satisfied, the ciip_condition is satisfied, or the triangle_merge_condition is satisfied, the regular_merge_flag can be parsed. Also, if the mmvd_condition is not satisfied, the subblock_merge_condition is not satisfied, the ciip_condition is not satisfied, and the triangle_merge_condition is not satisfied, the regular_merge_flag does not need to be parsed. Also, if the mmvd_condition is not satisfied, the subblock_merge_condition is not satisfied, the ciip_condition is not satisfied, and the triangle_merge_condition is not satisfied, the regular_merge_flag can be inferred to be 1. In this case, to infer a value of 1, the conditions for regular merge mode must be met (this condition does not have to exist for regular merge mode), all modes signaled or determined before regular merge mode must not be used (this condition does not have to exist for regular merge mode), and the condition for using merge mode (general_merge_flag == 1) must be met. Also, there may be cases where regular_merge_flag, as described in Figure 25, is inferred to be 1. In other cases, if regular_merge_flag does not exist, it can be inferred to be 0.

[0324] Also, in Figure 26, signaling was described for multiple modes depending on whether a mode that will be signaled or decided later can be performed, but the signaling method can be used for only some of the multiple modes. That is, at least one of the first condition 2501, second condition 2502, third condition 2503, and fourth condition 2504 in Figure 25 can be used, and the method in Figure 26 can be used for the rest. That is, the first condition 2501 in Figure 25, the second condition 2602 in Figure 26, the third condition 2603, and the fourth condition 2604 in Figure 26 can be used.

[0325] FIG. 27 illustrates a merge data syntax structure according to an embodiment of the present invention. FIG. 27 may be a specific example of the signaling method described with reference to FIGS. 12, 13, and 26. Referring to FIG. 27, if (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2), ciip_flag can be parsed. In this case, ciip_flag may be parsed only if the condition for performing CIIP is met. Also, if (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) is not met, ciip_flag does not need to be parsed. Also, if (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) is not met, ciip_flag can be inferred to be 1. In this case, if the merge mode is used, the conditions for performing CIIP are met, and none of the modes that are signaled or determined before CIIP are used, it can be inferred that ciip_flag is 1.

[0326] For example, if (general_merge_flag==1&&sps_ciip_enabled_flag&&cu_skip_flag==0&&width*height>=64&&width<128&&height<128&®ular_merge_flag==0&&mmvd_merge_flag==0&&merge_subblock_flag==0) and (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) is not satisfied, ciip_flag can be inferred to be 1. In this case, only some of the conditions (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) can be used. When only some of the conditions are used, some of the conditions must match when determining whether to parse and when making inferences. For example, when slice_type is not used, it may be possible to parse ciip_flag when (sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>=2) is satisfied, and it is not necessary to parse ciip_flag when (sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>=2) is not satisfied. Also, when (sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>=2) is not satisfied, it is possible to infer that ciip_flag is 1 when using merge mode, satisfying the conditions for CIIP, and not using any mode that is signaled or determined before CIIP.

[0327] Referring to FIG. 27, if ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)), merge_subblock_flag can be parsed. In this case, merge_subblock_flag may be parsed only if the condition for performing the subblock merge mode is met. Furthermore, if ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)) is not met, merge_subblock_flag does not need to be parsed. Also, if ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)) is not satisfied, merge_subblock_flag can be inferred to be 1. In this case, merge mode is used, the conditions for being able to perform subblock merge mode are satisfied, and if all modes that are signaled or determined before the subblock merge mode are not used, merge_subblock_flag can be inferred to be 1.

[0328] For example, if (general_merge_flag==1&&MaxNumSubblockMergeCand>0&&width>=8&&height>=8&®ular_merge_flag==0&&mmvd_merge_flag==0) and ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)) is not satisfied, merge_subblock_flag can be inferred to be 1. In this case, you can use only some of the conditions (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) and some of the conditions (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128). If you use only some of the conditions, they must match when inferring whether to parse or not. Also, the above describes the case where ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)) is not satisfied, but this may be the same as the case where (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) is not satisfied (&&) and (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128) is not satisfied.

[0329] 27, mmvd_merge_flag can be parsed if ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)). In this case, mmvd_merge_flag may be parsed only if the conditions for performing MMVD are met. Also, if the following conditions are not met: ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)), mmvd_merge_flag does not need to be parsed. Also, if ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)) is not satisfied, mmvd_merge_flag can be inferred to be 1.

[0330] In this case, if the merge mode is used, the conditions for performing MMVD are met, and none of the modes that are signaled or determined before MMVD are used, mmvd_merge_flag can be inferred to be 1. For example, if (general_merge_flag==1&&sps_mmvd_enabled_flag&®ular_merge_flag==0) and ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)) are not met, mmvd_merge_flag can be inferred to be 1. In this case, you can use only some of the conditions (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2), only some of the conditions (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128), and only some of the conditions (MaxNumSubblockMergeCand>0&&width>=8&&height>=8). If you use only some of the conditions, they must match when inferring whether to parse.Also, the above mentions a case where ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)) is not satisfied, but this may be the same as a case where (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) is not satisfied (&&), (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128) is not satisfied (&&), and (MaxNumSubblockMergeCand>0&&width>=8&&height>=8) is not satisfied.

[0331] 25, there may be cases where mmvd_merge_flag is inferred to be 1. For example, if (sps_mmvd_enabled_flag==1&&general_merge_flag==1&&width*height==32&general_merge_flag==0), mmvd_merge_flag can be inferred to be 1. Also, as described in Figure 26, if (general_merge_flag==1&&sps_mmvd_enabled_flag&®ular_merge_flag==0) and ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)) is not satisfied, mmvd_merge_flag can be inferred to be 1. Otherwise, mmvd_merge_flag can be inferred to be 0.

[0332] 27, regular_merge_flag can be parsed if ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag). In this case, regular_merge_flag may be parsed only if regular merge mode is the only possible merge mode. Also, if the following conditions are not met: ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag), regular_merge_flag does not need to be parsed.

[0333] Also, if ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag) is not satisfied, regular_merge_flag can be inferred to be 1. In this case, if merge mode is used, regular_merge_flag can be inferred to be 1. For example, if (general_merge_flag==1) and ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag) is not satisfied, regular_merge_flag can be inferred to be 1. In this case, you can use only some of the conditions (sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2), only some of the conditions (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128), and only some of the conditions (MaxNumSubblockMergeCand>0&&width>=8&&height>=8). If you use only some of the conditions, they must match when inferring whether to parse.

[0334] More specifically, a low-complexity encoder may not use various merge tools, and for such an encoder, if (sps_triangle_enabled_flag||sps_affine_enabled_flag||sps_sbtmvp_enabled_flag||sps_ciip_enabled_flag||sps_mmvd_enabled_flag) is not satisfied, regular_merge_flag can be inferred to have a value of 1 without parsing it. Also, above we mentioned the case where ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag) is not satisfied, but this is because (sps_trian It may be the same as the case where (gle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2) is not satisfied (&&), (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128) is not satisfied (&&), (MaxNumSubblockMergeCand>0&&width>=8&&height>=8) is not satisfied (&&), and sps_mmvd_enabled_flag is not satisfied.

[0335] As described in FIG. 25, there may be cases where regular_merge_flag is inferred to be 1. For example, if (sps_mmvd_enabled_flag==0&&general_merge_flag==1&&width*height==32), regular_merge_flag can be inferred to be 1. Also, as described in FIG. 68, if (general_merge_flag==1) and ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag) is not satisfied, regular_merge_flag can be inferred to be 1. Otherwise, regular_merge_flag can be inferred as 0.

[0336] 27 includes the conditions (1) width*height>=64 and (2) width>=8, height>=8, so in (1) or (2), the width and height must not be 4 and 8, respectively, or the width and height must not be 8 and 4, respectively. Therefore, in the first condition 2701 and the second condition 2702, it is not necessary to check whether width*height is 32. Therefore, the second condition 2702 can only use ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag). In such a case, regular_merge_flag can be inferred to be 1 if general_merge_flag is 1 and ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)||sps_mmvd_enabled_flag) is not satisfied. In other cases, regular_merge_flag can be inferred to be 0.

[0337] Also, the second condition 2702 can only use sps_mmvd_enabled_flag&&((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)). In this case, if (general_merge_flag==1&&sps_mmvd_enabled_flag&®ular_merge_flag==0) and ((sps_triangle_enabled_flag&&slice_type==B&&MaxNumTriangleMergeCand>=2)||(sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128)||(MaxNumSubblockMergeCand>0&&width>=8&&height>=8)) is not satisfied, mmvd_merge_flag can be inferred to be 1. In other cases, mmvd_merge_flag can be inferred to be 0.

[0338] Also, as explained in Fig. 26, the signaling method can be used for only some of the multiple modes in Fig. 27. That is, some of the first condition 2501, second condition 2502, third condition 2503, and fourth condition 2504 in Fig. 25 can be used, and the method in Fig. 27 can be used for the rest. That is, the first condition 2501 in Fig. 25 and the second condition 2702, third condition 2703, and fourth condition 2704 in Fig. 27 can be used.

[0339] Figure 28 is a diagram illustrating a merge data syntax structure according to an embodiment of the present invention. Figure 28 may be a specific example of the signaling method described in Figures 12, 13, and 26. In addition, the example of Figure 28 may be an example in which overlapping conditions in the example described in Figure 27 are removed.

[0340] 28, compared to FIG. 27, the condition width*height>=64 may not be present in the second condition 2802. This may be because if the condition (width>=8&&height>=8) is not satisfied, the condition width*height>=64 will not be satisfied and the condition width*height!=32 already exists. In other words, this may be because width*height>=64 is always satisfied due to other conditions.

[0341] As explained in FIG. 26, the signaling method can be used for only some of the multiple modes in FIG. 28 as well. That is, some of the first condition 2501, second condition 2502, third condition 2503, and fourth condition 2504 in FIG. 25 can be used, and the method in FIG. 28 can be used for the rest. That is, the first condition 2501 in FIG. 25 and the second condition 2802, third condition 2803, and fourth condition 2804 in FIG. 28 can be used. Alternatively, the first to fourth conditions in FIGS. 25 to 28 may be used in combination. That is, the first condition 2701 in FIG. 27, the second condition 2802 in FIG. 28, the third condition 2801 in FIG. 28, and the fourth condition 2801 in FIG. 28 can be used.

[0342] Figure 29 is a diagram illustrating a merge mode signaling method according to one embodiment of the present invention. A sequential signaling method has been described in the merge mode signaling described above. For example, a sequential signaling method such as that shown in Figure 25 can be used. Figure 29(a) shows such a sequential signaling method. In Figure 29, bold text indicates a mode to be determined, and italics may indicate signaling. This signaling may be a flag and may have a value of 0 or 1.

[0343] Furthermore, this signaling may be explicit or implicit, depending on the case. For example, in (a) of FIG. 29, regular_merge_flag may be signaled, and it is possible to determine whether the regular merge mode is in effect based on the value of regular_merge_flag. If regular_merge_flag indicates that the regular merge mode is not in effect, mmvd_merge_flag may be signaled, and it is possible to determine whether the MMVD mode is in effect based on mmvd_merge_flag. If the MMVD mode is not in effect, merge_subblock_flag may be signaled, and it is possible to determine whether the subblock merge mode is in effect based on merge_subblock_flag. If the subblock merge mode is not in effect, ciip_flag may be signaled, and it is possible to determine whether the CIIP mode is in effect based on ciip_flag. It is also possible to determine whether the TPM (triangular merge mode, triangular partitioning mode) mode is in effect based on ciip_flag. 29(a) shows an example in which the signals are sent in the order of normal merge mode, MMVD, sub-block merge mode, CIIP, and triangle merge mode, but the present invention is not limited to this and other orders may be used. The figures above also show examples in which the signals are sent in other orders.

[0344] Another merge mode signaling method is the grouping method. Figure 29(b) shows an example of the grouping method. For example, group_1_flag can be signaled, and based on group_1_flag, it can be determined whether the selected mode belongs to group 1. If group_1_flag indicates that the selected mode does not belong to group 1, group_2_flag can be signaled. Furthermore, based on group_2_flag, it can be determined whether the selected mode belongs to group 2. This operation can also be performed when there are multiple groups. Furthermore, signaling may exist to indicate which mode is indicated within a group. The grouping method allows for a reduction in signaling depth compared to sequential signaling. Furthermore, it allows for a reduction in the maximum signaling length (maximum codeword length).

[0345] According to one embodiment of the present invention, there may be three groups. Each group may have one mode. For example, there may be one mode in group 1. There may be two modes in group 2 and group 3. Referring to (b) of FIG. 29, the subblock merge mode may belong to group 1, the regular merge mode and MMVD may belong to group 2, and the CIIP and triangle merge mode may belong to group 3. Furthermore, group_1_flag may be merge_subblock_flag, and group_2_flag may be regular_merge_flag. Furthermore, ciip_flag and mmvd_merge_flag may be present as signaling indicating which mode is indicated within a group. For example, merge_subblock_flag is signaled, and it can be determined whether or not the subblock merge mode is selected based on merge_subblock_flag. If the subblock merge mode is not selected, regular_merge_flag may be signaled. Based on regular_merge_flag, it is possible to determine whether it is group 2 (regular merge mode or MMVD) or group 3 (CIIP or triangle merge mode). Also, if it indicates group 2, it is possible to determine whether it is regular merge mode or MMVD based on mmvd_merge_flag. Also, if it indicates group 3, it is possible to determine whether it is CIIP or triangle merge mode based on ciip_flag. In other words, merge_subblock_flag, regular_merge_flag, mmvd_merge_flag, and ciip_flag in (a) and (b) of Figure 29 may have somewhat different meanings.

[0346] Figure 30 is a diagram illustrating merge data syntax according to an embodiment of the present invention. The embodiment of Figure 30 may use the grouping method described in Figure 29(b). In this embodiment, descriptions that overlap with the above content may be omitted.

[0347] According to one embodiment of the present invention, when the merge mode is used, merge_subblock_flag can be signaled. When the merge mode is used, it may be the same as described above, or it may be when general_merge_flag is 1. The present invention may also apply when CuPredMode is not MODE_IBC or when CuPredMode is MODE_INTER. Whether to parse merge_subblock_flag can be determined based on MaxNumSubblockMergeCand and the block size, which may be based on the conditions for using the subblock merge mode described above. If merge_subblock_flag is 1, it can be determined that the subblock merge mode is used, and further, a candidate index can be determined based on merge_subblock_idx.

[0348] Also, if merge_subblock_flag is 0, regular_merge_flag can be parsed. At this time, there may be conditions for parsing regular_merge_flag. For example, a condition based on block size may be included. A condition based on higher level signaling indicating whether a mode is available may also be included. The higher level signaling indicating whether a mode is available may include sps_ciip_enabled_flag and sps_triangle_enabled_flag. A condition based on slice type may also be included. A condition based on cu_skip_flag may also be included. Referring to FIG. 71, regular_merge_flag can be parsed only if (width*height>=64&&width<128&&height<128) is satisfied. Also, if (width*height>=64&&width<128&&height<128) is not satisfied, regular_merge_flag does not need to be parsed.

[0349] Furthermore, the conditions under which CIIP can be used may include (sps_ciip_enabled_flag&&cu_skip_flag==0). Furthermore, the block size conditions under which CIIP can be used may include (width*height>=64&&width<128&&height<128). Furthermore, the conditions under which triangle merge mode can be used may include (sps_triangle_enabled_flag&&slice_type==B). Furthermore, the block size conditions under which triangle merge mode can be used may include (width*height>=64&&width<128&&height<128). If the conditions under which CIIP can be used or the conditions under which triangle merge mode can be used are met, regular_merge_flag can be parsed. Furthermore, if neither the conditions under which CIIP can be used nor the conditions under which triangle merge mode can be used are met, regular_merge_flag does not need to be parsed.

[0350] According to an embodiment of the present invention, if regular_merge_flag is not present, its value can be inferred to be 1. For example, it can always be inferred to be 1. In this embodiment, if regular_merge_flag is 1, regular merge mode or MMVD may be used. Therefore, if the block size conditions for using the CIIP and the block size conditions for using the triangle merge mode are not all met, the usable modes may be regular merge mode and MMVD, and regular_merge_flag can be determined to be 1 without parsing. In the embodiment shown in Figure 71, the block size conditions for using the CIIP and the triangle merge mode are the same. That is, it may not be possible to use both the CIIP and the triangle merge mode for a block with a width or height of 128.

[0351] Also, even if the conditions for using the CIIP and the conditions for using the triangle merge mode are not all met, as mentioned above, the possible modes may be regular merge mode or MMVD, so regular_merge_flag can be inferred to be 1 without parsing it.

[0352] Referring to FIG. 30, when regular_merge_flag is 1, syntax elements can be parsed based on the sps_mmvd_enabled_flag value. As described above, sps_mmvd_enabled_flag may be higher-level signaling indicating whether MMVD is available. When sps_mmvd_enabled_flag is 0, MMVD may not be available. Referring to FIG. 71, if sps_mmvd_enabled_flag is 0, mmvd_merge_flag, mmvd_cand_flag, mmvd_distance_idx, mmvd_direction_idx, and merge_idx do not need to be parsed. Also, if mmvd_merge_flag does not exist, its value can be inferred to be 0. Also, if merge_idx does not exist, its value can be inferred using a previously configured method. For example, when merge_idx does not exist, if mmvd_merge_flag is 1, it can be inferred as mmvd_cand_flag, and if mmvd_merge_flag is 0, it can be inferred as 0. Therefore, in the example of FIG. 30, if sps_mmvd_enabled_flag is 0 and regular_merge_flag is 1, the merge_idx value can always be 0, and regular merge mode prediction can be performed using the candidate with index 0 in the merge candidate list. Therefore, there is no freedom in selecting candidates, which may result in reduced coding efficiency. Also, if sps_mmvd_enabled_flag is 1, mmvd_merge_flag can be parsed, and if mmvd_merge_flag is 0, merge_idx can be parsed based on MaxNumMergeCand.

[0353] Also, referring to FIG. 30, when regular_merge_flag is 0, if the conditions for using the CIIP and the conditions for using the triangle merge mode are all met, ciip_flag can be parsed. If ciip_flag is 1, CIIP can be used, and if ciip_flag is 0, triangle merge mode can be used. If the conditions for using the CIIP or the conditions for using the triangle merge mode are not met, ciip_flag does not need to be parsed. If ciip_flag does not exist and regular_merge_flag is 1, ciip_flag can be inferred to be 0. If ciip_flag does not exist and regular_merge_flag is 0, ciip_flag can be inferred to be (sps_ciip_enabled_flag && cu_skip_flag == 0). For B slices, MergeTriangleFlag can be set to !ciip_flag. Also, MergeTriangleFlag can be set to 0 for P slices.

[0354] Figure 31 illustrates a merge data syntax according to an embodiment of the present invention. In this embodiment, overlapping content with that described above can be omitted. As described in Figure 30, when regular_merge_flag is 1 and higher-level signaling indicating whether an MMVD can be used indicates that an MMVD cannot be used, the degree of freedom in candidate selection is reduced. However, the embodiment of Figure 31 can solve this problem.

[0355] Referring to Figure 31, whether to parse merge_idx may be independent of sps_mmvd_enabled_flag. That is, whether to parse merge_idx can be determined regardless of the value of sps_mmvd_enabled_flag. According to an embodiment of the present invention, when regular_merge_flag is 1, mmvd_merge_flag is 0, and MaxNumMergeCand>1, merge_idx can be parsed. Also, when regular_merge_flag is 1 and mmvd_merge_flag is 1, merge_idx does not need to be parsed. Also, when regular_merge_flag is 1 and MaxNumMergeCand is 1, merge_idx does not need to be parsed. For example, if sps_mmvd_enabled_flag is 1, regular_merge_flag is 1, mmvd_merge_flag is 0, and MaxNumMergeCand>1, merge_idx can be parsed. Similarly, if sps_mmvd_enabled_flag is 0, regular_merge_flag is 1, mmvd_merge_flag is 0, and MaxNumMergeCand>1, merge_idx can be parsed.

[0356] In addition, in the embodiment of FIG. 30 , the block size condition for using the triangle merge mode is (width*height>=64&&width<128&&height<128), but it may be possible to use the triangle merge mode when the width or height is 128. For example, this is because prediction of the triangle merge mode can be useful for improving coding efficiency when the width or height is 128. In FIG. 72 , the triangle merge mode may be usable even when the width or height is 128.

[0357] In FIG. 31, the block size condition for using CIIP may be (width*height>=64&&width<128&&height<128). Also, the block size condition for using triangle merge mode may be (width*height>=64). Therefore, referring to FIG. 72, if (width*height>=64) is not satisfied, regular_merge_flag does not need to be parsed. Also, if the width is 128 (or the height is 128 or greater), or if the height is 128 (or the height is 128 or greater), regular_merge_flag may be parsed. For example, if (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128), regular_merge_flag can be parsed if (width*height>=64) is satisfied. Also, if (sps_triangle_enabled_flag&&slice_type==B), regular_merge_flag can be parsed if (width*height>=64) is satisfied. Also, if (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128) is not satisfied and (sps_triangle_enabled_flag&&slice_type==B) is not satisfied, regular_merge_flag does not need to be parsed.

[0358] Also, referring to FIG. 31, a condition based on the block size may be necessary when determining whether to parse ciip_flag. For example, if width<128 and height<128, the decoder can parse ciip_flag. If the width is 128 (or is 128 or greater) or the height is 128 (or is 128 or greater), ciip_flag does not need to be parsed. This may be because, when the width or height is 128 (or is 128 or greater), one of CIIP and triangle merge mode cannot be used, but the other can be used. This may be because, when the width or height is 128 (or is 128 or greater), CIIP cannot be used, but triangle merge mode can be used. In the example of Figure 30, if the width or height is 128 (or greater than or equal to 128), neither CIIP nor triangle merge mode can be used, which means that regular_merge_flag is inferred to be 1 without being parsed, which is different from Figure 31.

[0359] In one embodiment of the present invention, the value of regular_merge_flag can be inferred based on merge_subblock_flag. In this embodiment, regular_merge_flag, merge_subblock_flag, and ciip_flag may be regular_merge_flag, merge_subblock_flag, and ciip_flag described in Figure 29(b), Figure 30, and Figure 31.

[0360] In the description of FIG. 30 , when regular_merge_flag is not present, its value is always inferred to be 1. However, there may be cases where merge_subblock_flag, which is signaled before regular_merge_flag, is 1. In this case, prediction is made based on merge_subblock_flag or regular_merge_flag, but since both values ​​are 1, ambiguity may occur as to which prediction to make. For this reason, in this embodiment, regular_merge_flag can be inferred based on merge_subblock_flag. For example, when merge_subblock_flag is 1, regular_merge_flag can be inferred to be 0. Furthermore, when merge_subblock_flag is 0, regular_merge_flag can be inferred to be 1. Alternatively, a general_merge_flag condition may be added here. For example, when merge_subblock_flag is 0 and general_merge_flag is 1, regular_merge_flag can be inferred to be 1.

[0361] Also, Figure 30 shows a method for inferring the value of ciip_flag when it does not exist. However, if the usable block size conditions for CIIP and triangle merge mode are different, using the ciip_flag inference method described in Figure 30 signals that a certain mode is to be used for a block size where the certain mode cannot be used. That is, for example, if the width or height is 128, CIIP cannot be used, but ciip_flag may be set to 1. This embodiment can solve this problem.

[0362] According to one embodiment of the present invention, if ciip_flag is not present, its value can be inferred based on the block size. Also, if ciip_flag is not present, its value can be inferred based on regular_merge_flag. For example, if regular_merge_flag is 1, ciip_flag can be inferred to be 0. Also, if regular_merge_flag is 0, ciip_flag can be inferred based on the block size. For example, if regular_merge_flag is 0, ciip_flag can be inferred based on the block size, sps_ciip_enabled_flag, and cu_skip_flag. If regular_merge_flag is 0, ciip_flag can be inferred as (sps_ciip_enabled_flag && cu_skip_flag == 0 && width < 128 && height < 128). Therefore, if regular_merge_flag is 0 and either the width or height is 128, ciip_flag can be inferred to be 0. Also, to infer ciip_flag to be 1, the condition that general_merge_flag is 1 may be included. If general_merge_flag is 0, ciip_flag can be inferred to be 0. That is, if regular_merge_flag is 0, ciip_flag can be inferred to be (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128&&general_merge_flag==1). Otherwise, it can be inferred to be 0.

[0363] Alternatively, if regular_merge_flag is 0 and general_merge_flag is 1, ciip_flag can be inferred as (sps_ciip_enabled_flag&&cu_skip_flag==0&&width<128&&height<128). If general_merge_flag is 0, ciip_flag can be inferred as 0.

[0364] Also, in the embodiment of FIG. 30, the method for setting MergeTriangleFlag is described, but it is set regardless of the value of regular_merge_flag. Therefore, there may be cases where both regular_merge_flag and MergeTriangleFlag are 1, which causes ambiguity in the prediction method. Therefore, in this embodiment, MergeTriangleFlag can be set based on regular_merge_flag. For example, when regular_merge_flag is 1, MergeTriangleFlag can be set to 0. Also, when regular_merge_flag is 0, MergeTriangleFlag can be set to !ciip_flag. Furthermore, when regular_merge_flag is 0, MergeTriangleFlag can be set taking into account the conditions under which the triangle merge mode can be used. For example, if regular_merge_flag is 0, you can set MergeTriangleFlag to (!ciip_flag&&sps_triangle_enabled_flag&&slice_type==B).

[0365] Therefore, it is possible to prevent a situation where MergeTriangleFlag is set to 1 when sps_triangle_enabled_flag is 0 or slice_type is not B. Furthermore, in order to determine MergeTriangleFlag to 1, a condition that general_merge_flag is 1 may be included. If general_merge_flag is 0, MergeTriangleFlag can be set to 0. That is, if regular_merge_flag is 0, MergeTriangleFlag can be set to (!ciip_flag&&sps_triangle_enabled_flag&&slice_type==B&&general_merge_flag==1). Otherwise, MergeTriangleFlag can be set to 0.

[0366] Or, if regular_merge_flag is 0 and general_merge_flag is 1, MergeTriangleFlag can be set to (!ciip_flag&&sps_triangle_enabled_flag&&slice_type==B). If general_merge_flag is 0, MergeTriangleFlag can be set to 0.

[0367] FIG. 32 illustrates a geometric merge mode according to an embodiment of the present invention. According to an embodiment of the present invention, the geometric merge mode may be variously referred to as a geometric partitioning mode, a GEO mode, a GEO merge mode, or GEO partitioning. According to an embodiment of the present invention, the geometric merge mode may be a method of partitioning a coding unit (CU) or a coding block (CB). For example, the geometric merge mode may be a method of partitioning a square or rectangular CU or CB into non-square or non-rectangular partitions. Referring to FIG. 32, an example of geometric partitioning is shown. As shown in FIG. 32, a rectangular CU may be partitioned into triangular and trapezoidal partitions (or polygons) through geometric partitioning. Furthermore, signaling for a method of performing the geometric merge mode may be signaled to the CU. Furthermore, in the geometric merge mode, motion compensation and prediction may be performed based on two pieces of motion information. Furthermore, the two pieces of motion information may be obtained from merge candidates. According to an embodiment of the present invention, signaling for indicating the two pieces of motion information used in the geometric merge mode may be present. For example, two indexes may be signaled to indicate two pieces of motion information to be used in the geometric merge mode. More specifically, for example, two merge candidate indexes may be signaled to indicate two pieces of motion information to be used in the geometric merge mode. Furthermore, two predictors may be blended in the geometric merge mode. For example, two predictors may be blended near an inner boundary within a CU in the geometric merge mode. Blending two predictors may mean that the two predictors are weighted summed.

[0368] As an example, the syntax elements for indicating two pieces of motion information used in the geometric merge mode may be merge_triangle_idx0 and merge_triangle_idx1. In this case, two indexes m and n can be derived from the syntax elements. For example, they may be derived as follows:

[0369] m=merge_triangle_idx0

[0370] n=merge_triangle_idx1+((merge_triangle_idx1>=m)?1:0)

[0371] That is, index m may be the same as merge_triangle_idx0, and index n may be merge_triangle_idx1+1 if merge_triangle_idx1 is greater than or equal to merge_triangle_idx0, or may be merge_triangle_idx1 if merge_triangle_idx1 is less than merge_triangle_idx0.

[0372] Also, referring to FIG. 32, the split boundary of the geometric merge mode may be represented by an angle φ (phi) and a distance offset ρ (rho). The angle φ may represent a quantized angle, and the distance offset ρ may represent a quantized offset. The angle and the distance offset may be signaled by merge_geo_idx. For example, they may be defined by a look-up table. The geometric merge mode may generate a prediction signal from two predictors based on two pieces of motion information. In this case, when two predictors are weighted together, the weight may be based on the angle and the distance offset. Alternatively, when two predictors are weighted together, the weight may be based on a position (coordinate) within a block. Alternatively, when two predictors are weighted together, the weight may be based on the block width and height.

[0373] In geometric merge mode, the number of possible split types may be greater than in TPM. For example, the number of possible split types in geometric merge mode may be greater than two. For example, there may be 80 possible split types. The geometric merge mode may be a type of merge mode. That is, when in geometric merge mode, the general_merge_flag value may be 1.

[0374] FIG. 33 is a diagram illustrating merge data syntax according to one embodiment of the present invention.

[0375] In the embodiment of FIG. 33, the details explained in FIGS. 29 to 32 or previously explained may be omitted.

[0376] As described above, there may be a method for signaling multiple merge modes. The multiple merge modes may include sub-block merge mode, regular merge mode, MMVD, CIIP, geometric merge mode, etc. The multiple merge modes may not include triangle partitioning mode. Alternatively, triangle partitioning mode may be included in geometric merge mode. When signaling merge modes using the signaling method of this embodiment, codewords of different lengths may be used, and coding efficiency may be improved by using codewords of shorter lengths for specific modes. The signaling method of this embodiment may eliminate redundant signaling, thereby improving coding efficiency. Furthermore, parsing complexity may be reduced by omitting redundant conditional checks in the signaling of this embodiment.

[0377] According to an embodiment of the present invention, there may be conditions under which the CIIP can be used. The conditions under which the CIIP can be used may be referred to as CIIP_conditions. The CIIP_conditions may be true if all of the following conditions are met:

[0378] Condition 1.sps_ciip_enabled_flag

[0379] Condition 2.cu_skip_flag==0

[0380] Condition 3.cbWidth*cbHeight>=64

[0381] Condition 4.cbWidth<128

[0382] Condition 5.cbHeight<128

[0383] Also, CIIP_conditions may be false if at least one of the conditions is not satisfied. The conditions have been described in the previous embodiment and therefore may be omitted.

[0384] According to an embodiment of the present invention, there may be conditions under which the geometric merge mode can be used. The conditions under which the geometric merge mode can be used may be referred to as GEO_conditions. GEO_conditions may be true when all of the following conditions are met:

[0385] Condition 1.sps_triangle_enabled_flag

[0386] Condition 2.MaxNumTriangleMergeCand>1

[0387] Condition 3.slice_type==B

[0388] Condition 4.cbWidth>=8

[0389] Condition 5. cbHeight>=8

[0390] Additionally, GEO_conditions may be false when at least one of the conditions is not met.

[0391] In still other embodiments, the slice_type condition may not be necessary. This may be because the slice_type-based condition is satisfied if another condition, for example, a condition based on MaxNumTriangleMergeCand, is satisfied. According to embodiments of the present invention, there may be conditions under which the geometric merge mode can be used. The conditions under which the geometric merge mode can be used may be referred to as GEO_conditions. GEO_conditions may be true when all of the following conditions are satisfied:

[0392] Condition 1.sps_triangle_enabled_flag

[0393] Condition 2.MaxNumTriangleMergeCand>1

[0394] Condition 3.cbWidth>=8

[0395] Condition 4. cbHeight>=8

[0396] Additionally, GEO_conditions may be false when at least one of the conditions is not met.

[0397] The above conditions have been explained in the previous embodiment and can be omitted. However, although sps_triangle_enabled_flag and MaxNumTriangleMergeCand were explained in the previous embodiment as values ​​related to the TPM, in this embodiment they may be values ​​related to the geometric merge mode. That is, sps_triangle_enabled_flag may be higher-level signaling indicating whether the geometric merge mode can be used. Furthermore, MaxNumTriangleMergeCand may be the maximum number of candidate lists to be used in the geometric merge mode.

[0398] According to an embodiment of the present invention, regular_merge_flag can be parsed if either CIIP_conditions or GEO_conditions are satisfied. Furthermore, if neither CIIP_conditions nor GEO_conditions are satisfied, regular_merge_flag does not need to be parsed. Referring to Figure 75, condition 2 represents (CIIP_conditions||GEO_conditions). That is, regular_merge_flag can be parsed if at least one of the following conditions is satisfied:

[0399] Condition 1 (CIIP_conditions).sps_ciip_enabled_flag&&cu_skip_flag==0&&cbWidth*cbHeight>=64&&cbWidth<128&&cbHeight<128

[0400] Condition 2 (GEO_conditions).sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>1&&cbWidth>=8&&cbHeight>=8

[0401] Also, if none of the above conditions are met, there is no need to parse regular_merge_flag. Also, if regular_merge_flag does not exist, its value can be inferred as general_merge_flag&&!merge_subblock_flag.

[0402] As another example, the condition 2 (GEO_conditions) may be as follows, including the slice_type condition, as described above.

[0403] Condition 2 (GEO_conditions).sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>1&&slice_type==B&&cbWidth>=8&&cbHeight>=8

[0404] However, if the slice_type condition is always satisfied when other conditions are satisfied, then the slice_type condition may not be checked further to reduce the complexity of parsing condition checking.

[0405] According to an embodiment of the present invention, ciip_flag can be parsed if all of the CIIP_conditions and GEO_conditions are satisfied. Also, if the CIIP_conditions or GEO_conditions are not satisfied, ciip_flag does not need to be parsed. That is, ciip_flag can be parsed if all of the following conditions are satisfied, and ciip_flag does not need to be parsed if at least one of the following conditions is not satisfied.

[0406] Condition 1 (CIIP_conditions).sps_ciip_enabled_flag&&cu_skip_flag==0&&cbWidth*cbHeight>=64&&cbWidth<128&&cbHeight<128

[0407] Condition 2 (GEO_conditions).sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>1&&cbWidth>=8&&cbHeight>=8

[0408] As mentioned above, it is also possible to include a condition based on slice_type in condition 2 (GEO_conditions). The condition can be as follows:

[0409] Condition 2 (GEO_conditions).sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>1&&slice_type==B&&cbWidth>=8&&cbHeight>=8

[0410] According to an embodiment of the present invention, the conditions for parsing ciip_flag may be changed to reduce parsing complexity. For example, some of the block size conditions may be omitted. In the present invention, when the block size conditions for using the geometric merge mode are satisfied, some of the block size conditions for using the CIIP may be satisfied. Therefore, referring to condition 3 in FIG. 75, according to an embodiment of the present invention, ciip_flag can be parsed if all of the following conditions are satisfied, and ciip_flag does not need to be parsed if at least one of the following conditions is not satisfied.

[0411] Condition 1 (CIIP_conditions).sps_ciip_enabled_flag&&cu_skip_flag==0&&cbWidth<128&&cbHeight<128

[0412] Condition 2 (GEO_conditions).sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>1&&cbWidth>=8&&cbHeight>=8

[0413] Also, if ciip_flag does not exist, it can be inferred as 1 if all of the following conditions are met, and as 0 if at least one of the following conditions is not met.

[0414] Condition 1.sps_ciip_enabled_flag==1

[0415] Condition 2.general_merge_flag==1

[0416] Condition 3.merge_subblock_flag==0

[0417] Condition 4.regular_merge_flag==0

[0418] Condition 5.cu_skip_flag==0

[0419] Condition 6.cbWidth<128

[0420] Condition 7.cbHeight<128

[0421] Condition 8.cbWidth*cbHeight>=64

[0422] That is, the conditions for parsing and the conditions for inference may be different. Also, conditions omitted during parsing (e.g., conditions based on block size) may need to be included in the conditions for inference.

[0423] Furthermore, merge_geo_flag, which is a value indicating whether or not to use the geometric merge mode, can be determined to be 1 if all of the following conditions are met, and can be determined to be 0 if at least one of the following conditions is not met.

[0424] Condition 1.sps_triangle_enabled_flag==1

[0425] Condition 2.general_merge_flag==1

[0426] Condition 3.merge_subblock_flag==0

[0427] Condition 4.regular_merge_flag==0

[0428] Condition 5.ciip_flag==0

[0429] Condition 6.MaxNumTriangleMergeCand>=2

[0430] Condition 7.cbWidth>=8

[0431] Condition 8. cbHeight >= 8

[0432] As a further example, condition 9 may be added: slice_type==B.

[0433] Therefore, referring to FIG. 33, the following signaling structure can be obtained. If the first condition 3301 is satisfied, merge_subblock_flag can be parsed. If merge_subblock_flag is 1, the subblock merge mode may be used, merge_subblock_idx can be parsed, and regular_merge_flag, mmvd_merge_flag, and ciip_flag do not need to be parsed. If merge_subblock_flag is 0, if the second condition 3302 is satisfied, regular_merge_flag can be parsed. If regular_merge_flag is 1, the regular merge mode or MMVD may be used, and mmvd_merge_flag may be parsed. The contents described with reference to FIGS. 29 to 32 can be applied to this. If regular_merge_flag is 0, if the second condition 3302 is satisfied, ciip_flag can be parsed. If ciip_flag is 1, CIIP may be used. If CIIP is used, merge_idx can be parsed if MaxNumMergeCand is greater than 1. If ciip_flag is 0, merge_geo_flag may be determined to be 1. Also, if ciip_flag is 0, geometric merge mode may be used. If geometric merge mode is used, merge_geo_idx, merge_triangle_idx0, and merge_triangle_idx1 can be parsed. Alternatively, if geometric merge mode is used, merge_geo_idx, merge_triangle_idx0, and merge_triangle_idx1 can be parsed if MaxNumTriangleMergeCand is greater than 1.

[0434] Therefore, according to an embodiment of the present invention, when a block with a width or height of 4, i.e., a block of size 4xN or Nx4, uses CIIP, it can be signaled as follows: merge_subblock_flag may be 0, which satisfies the second condition 3302, so regular_merge_flag can be parsed, or its value may be 0, which does not satisfy the third condition 3303, so ciip_flag does not need to be parsed and its value can be inferred to be 1 according to the above. Also, when using geometric merge mode, signaling may be as follows: merge_subblock_flag may be 0, regular_merge_flag may be 0, and ciip_flag may be 0.

[0435] In the embodiment described in Fig. 33, there may be no syntax element other than ciip_flag that indicates whether or not to use the geometric merge mode. Also, in the embodiment described in Fig. 33, the geometric merge mode cannot be used for blocks whose width or height is smaller than 8, but the present invention is not limited to this, and the embodiment can also be applied to cases where the geometric merge mode cannot be used for other block sizes (for example, block sizes smaller than a threshold).

[0436] 34 is a diagram illustrating a video signal processing method according to an embodiment to which the present invention is applied. For convenience of explanation, the description will be focused on a decoder, but the present invention is not limited thereto, and the multiple hypothesis prediction-based video signal processing method according to this embodiment can be applied to an encoder in substantially the same manner.

[0437] The decoder parses a first syntax element that indicates whether a merge mode is applied to the current block (S3401).

[0438] If the merge mode is applied to the current block, the decoder determines whether to parse a second syntax element based on a predefined first condition (S3402). For example, the second syntax element may indicate whether a first mode or a second mode is applied to the current block.

[0439] If the first mode and the second mode are not applied to the current block, the decoder determines whether to parse a third syntax element based on a predefined second condition (S3403). As an example, the third syntax element may indicate a mode to be applied to the current block, either a third mode or a fourth mode.

[0440] The decoder determines the mode to be applied to the current block based on the second syntax element or the third syntax element (S3404).

[0441] The decoder derives motion information of the current block based on the determined mode (S3405).

[0442] The decoder generates a prediction block of the current block using the motion information of the current block (S3406).

[0443] The first condition includes at least one of a condition under which the third mode can be used and a condition under which the fourth mode can be used.

[0444] As described above, in an embodiment, the third mode and the fourth mode may be positioned after the first mode in the decoding order in the merge data syntax.

[0445] As mentioned above, an example method may include parsing the second syntax element if the first condition is met, and inferring the second syntax element to be 1 if the first condition is not met.

[0446] As mentioned above, as an embodiment, if the first condition is not met, the second syntax element can be inferred based on the fourth syntax element indicating whether a sub-block based merge mode is applied to the current block.

[0447] As described above, as an example, the second condition may include a condition under which the fourth mode is available.

[0448] As described above, in an embodiment, the second condition may include at least one of whether the third mode is available in the current sequence, whether the fourth mode is available in the current sequence, whether the maximum number of candidates for the fourth mode is greater than 1, whether the width of the current block is smaller than a predefined first size, and whether the height of the current block is smaller than a predefined second size.

[0449] As described above, an example may include a step of obtaining a fifth syntax element indicating whether the first mode or the second mode is applied to the current block when the second syntax element is 1.

[0450] The above-described embodiments of the present invention may be implemented by various means, such as hardware, firmware, software, or a combination thereof.

[0451] In the case of a hardware implementation, the method according to an embodiment of the present invention may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), processors, controllers, microcontrollers, microprocessors, etc.

[0452] In the case of implementation by firmware or software, the methods according to the embodiments of the present invention may be implemented in the form of modules, procedures, or functions that perform the functions or operations described above. The software code may be stored in a memory and driven by a processor. The memory may be located inside or outside the processor, and data may be exchanged with the processor by various means known in the art.

[0453] Some embodiments may be embodied in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available medium accessible by a computer, including volatile and nonvolatile media, detachable and non-detachable media. Computer-readable media may also include any computer storage media and communication media. Computer storage media includes any volatile and non-volatile, detachable and non-detachable media embodied in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Communication media typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as a program module, or other transmission mechanism, and includes any information delivery media.

[0454] The above description of the present invention is for illustrative purposes only, and those skilled in the art will understand that the present invention can be easily modified into other specific forms without changing the technical spirit or essential features of the present invention. Therefore, the above-described embodiments should be construed as illustrative in all respects and not restrictive. For example, each component described as a single component may be implemented in a distributed form, and similarly, each component described as a distributed component may be implemented in a combined form.

[0455] The scope of the present invention is indicated by the claims below rather than the above detailed description, and any modifications or variations derived from the meaning and scope of the claims and their equivalents should be construed as being included within the scope of the present invention. [Industrial Applicability]

[0456] The above-described preferred embodiments of the present invention have been disclosed for illustrative purposes, and those skilled in the art will be able to improve, modify, substitute or add various other embodiments within the technical spirit and scope of the present invention as disclosed in the appended claims.

Claims

1. 1. A method for processing a video signal, comprising: Parsing a first syntax element indicating whether a merge mode is applied to the current block; determining whether to parse a second syntax element based on a first predefined condition if the merge mode is applied to the current block, wherein the second syntax element indicates whether a first mode or a second mode is applied to the current block; determining whether to parse a third syntax element based on a second predefined condition when the first mode and the second mode are not applied to the current block, wherein the third syntax element indicates a third mode or a fourth mode to be applied to the current block; determining a mode to be applied to the current block based on the second syntax element or the third syntax element; deriving motion information of the current block based on the determined mode; and generating a predicted block for the current block using motion information of the current block; The video signal processing method, wherein the first condition includes at least one of a condition under which the third mode can be used and a condition under which the fourth mode can be used.

2. 2. The video signal processing method of claim 1, wherein the third mode and the fourth mode are located after the first mode in a decoding order in a merge data syntax.

3. parsing the second syntax element if the first condition is met; 2. The method of claim 1, wherein the second syntax element is inferred to be 1 if the first condition is not met.

4. 4. The video signal processing method of claim 3, wherein if the first condition is not satisfied, the second syntax element is inferred based on a fourth syntax element indicating whether a sub-block based merge mode is applied to the current block.

5. 2. The video signal processing method according to claim 1, wherein the second condition includes a condition under which the fourth mode is available.

6. 2. The video signal processing method of claim 1, wherein the second condition includes at least one of whether the third mode is available in the current sequence, whether the fourth mode is available in the current sequence, whether a maximum number of candidates for the fourth mode is greater than 1, whether a width of the current block is smaller than a predefined first size, and whether a height of the current block is smaller than a predefined second size.

7. 2. The video signal processing method of claim 1, further comprising: if the second syntax element is 1, obtaining a fifth syntax element indicating whether the first mode or the second mode is applied to the current block.

8. 1. A video signal processing device, comprising: a processor; The processor: Parse a first syntax element that indicates whether a merge mode is applied to the current block; If the merge mode is applied to the current block, determining whether to parse a second syntax element based on a predefined first condition, where the second syntax element indicates whether a first mode or a second mode is applied to the current block; If the first mode and the second mode are not applied to the current block, determining whether to parse a third syntax element based on a predefined second condition, wherein the third syntax element indicates a mode applied to the current block, among a third mode and a fourth mode; determining a mode to be applied to the current block based on the second syntax element or the third syntax element; deriving motion information of the current block based on the determined mode, and generating a predicted block of the current block using the motion information of the current block; The video signal processing device according to claim 1, wherein the first condition includes at least one of a condition under which the third mode can be used and a condition under which the fourth mode can be used.

9. 9. The video signal processing apparatus of claim 8, wherein the third mode and the fourth mode are positioned after the first mode in a decoding order in a merge data syntax.

10. 9. The video signal processing device of claim 8, wherein the processor parses the second syntax element if the first condition is satisfied, and the second syntax element is inferred to be 1 if the first condition is not satisfied.

11. 11. The video signal processing apparatus of claim 10, wherein, if the first condition is not satisfied, the second syntax element is inferred based on a fourth syntax element indicating whether a sub-block based merge mode is applied to the current block.

12. 9. The video signal processing device according to claim 8, wherein the second condition includes a condition under which the fourth mode is available.

13. 9. The video signal processing device of claim 8, wherein the second condition includes at least one of whether the third mode is available in the current sequence, whether the fourth mode is available in the current sequence, whether a maximum number of candidates for the fourth mode is greater than 1, whether a width of the current block is smaller than a predefined first size, and whether a height of the current block is smaller than a predefined second size.

14. The video signal processing device of claim 8 , wherein the processor obtains a fifth syntax element indicating whether the first mode or the second mode is applied to the current block when the second syntax element is 1.

15. 1. A method for processing a video signal, comprising: encoding a first syntax element indicating whether a merge mode is applied to a current block; determining whether to encode a second syntax element based on a first predefined condition when the merge mode is applied to the current block, wherein the second syntax element indicates whether a first mode or a second mode is applied to the current block; determining whether to encode a third syntax element based on a second predefined condition when the first mode and the second mode are not applicable to the current block, wherein the third syntax element indicates a mode to be applied to the current block, either a third mode or a fourth mode; determining a mode to be applied to the current block based on the second syntax element or the third syntax element; deriving motion information of the current block based on the determined mode; and generating a predicted block for the current block using motion information of the current block; The video signal processing method, wherein the first condition includes at least one of a condition under which the third mode can be used and a condition under which the fourth mode can be used.

Citation Information

Patent Citations

  • Method for signaling video information and method for decoding video information using the video information signaling method.

    JP2014501090A

  • Method and apparatus for video coding with automatic motion information refinement

    WO2018002024A1