Video signal processing method and apparatus using adaptive motion vector resolution

KR103024582B1Active Publication Date: 2026-09-29WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020217034935
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-17
Filing Date
2020-05-04
Publication Date
2026-09-29
Estimated Expiration
2040-05-04

Smart Images

  • Figure 112021123213009-PCT00001_ABST
    Figure 112021123213009-PCT00001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for processing a video signal, and more specifically, is characterized by comprising the steps of: parsing an Adaptive Motion Vector Resolution (AMVR) enable flag (sps_amvr_enabled_flag) from a bitstream indicating whether an Adaptive Motion Vector Difference Resolution is used; parsing an Affine Enabled flag (sps_affine_enabled_flag) from a bitstream indicating whether an Affine Motion Compensation can be used; determining whether an Affine Motion Compensation can be used based on the Affine Enabled flag; determining whether an Adaptive Motion Vector Difference Resolution is used based on the AMVR enabled flag when an Affine Motion Compensation can be used; and, when an Adaptive Motion Vector Difference Resolution is used, parsing an Affine AMVR Enabled flag (sps_affine_amvr_enabled_flag) from a bitstream indicating whether an Adaptive Motion Vector Difference Resolution can be used for an Affine Motion Compensation.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to a method and apparatus for processing a video signal, and more specifically, to a video signal processing method and apparatus for encoding or decoding a video signal. Background Technology

[0002] Compression encoding refers to a series of signal processing techniques used to transmit digitized information over communication lines or store it in a form suitable for storage media. Targets of compression encoding include voice, video, and text; specifically, the technology of performing compression encoding on video is called video image compression. Compression encoding of video signals is achieved by removing redundant information by considering spatial correlation, temporal correlation, and probabilistic correlation. However, due to recent advancements in various media and data transmission media, there is a demand for more efficient video signal processing methods and devices. The problem to be solved

[0003] The purpose of the present disclosure is to increase the coding efficiency of video signals. means of solving the problem

[0004] A method for processing a video signal according to one embodiment of the present disclosure comprises: a step of parsing an Adaptive Motion Vector Resolution (AMVR) enable flag (sps_amvr_enabled_flag) from a bitstream indicating whether Adaptive Motion Vector Difference Resolution is used; a step of parsing an Affine Enabled Flag (sps_affine_enabled_flag) from a bitstream indicating whether Affine Motion Compensation can be used; a step of determining whether Affine Motion Compensation can be used based on the Affine Enabled Flag (sps_affine_enabled_flag); a step of determining whether Adaptive Motion Vector Difference Resolution is used based on the AMVR enabled Flag (sps_amvr_enabled_flag) when Affine Motion Compensation can be used; and a step of parsing an Affine AMVR Enabled Flag (sps_affine_amvr_enabled_flag) from a bitstream indicating whether Adaptive Motion Vector Difference Resolution can be used for Affine Motion Compensation when Adaptive Motion Vector Difference Resolution is used.

[0005] A method for processing a video signal according to one embodiment of the present disclosure is characterized in that one of an AMVR-enabled flag (sps_amvr_enabled_flag), an affine-enabled flag (sps_affine_amvr_enabled_flag), or an affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) is signaled to one of a Coding Tree Unit, a slice, a tile, a tile group, a picture, or a sequence unit.

[0006] A method for processing a video signal according to one embodiment of the present disclosure, wherein affine motion compensation may be used and, when adaptive motion vector difference resolution is not used, the affine AMVR enable flag (sps_affine_amvr_enabled_flag) implies that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0007] In a method for processing a video signal according to one embodiment of the present disclosure, when affine motion compensation cannot be used, the affine AMVR enable flag (sps_affine_amvr_enabled_flag) is characterized by implying that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0008] A method for processing a video signal according to one embodiment of the present disclosure comprises the steps of: parsing information about the resolution of motion vector differences from a bitstream when an AMVR-enabled flag (sps_amvr_enabled_flag) indicates the use of adaptive motion vector difference resolution, an inter-affine flag (inter_affine_flag) obtained from a bitstream indicates that affine motion compensation is not used for the current block, and at least one of a plurality of motion vector differences for the current block is not zero; and modifying a plurality of motion vector differences for the current block based on information about the resolution of motion vector differences.

[0009] A method for processing a video signal according to one embodiment of the present disclosure is characterized by comprising the steps of: parsing information about the resolution of the motion vector difference from a bitstream when an affine AMVR capable flag indicates that an adaptive motion vector difference resolution can be used for affine motion compensation, an inter_affine_flag obtained from a bitstream indicates the use of affine motion compensation for the current block, and at least one of a plurality of control point motion vector differences for the current block is not zero; and modifying a plurality of control point motion vector differences for the current block based on information about the resolution of the motion vector difference.

[0010] A method for processing a video signal according to one embodiment of the present disclosure comprises: a step of obtaining information (inter_pred_idc) about a reference picture list for a current block; a step of parsing a motion vector predictor index (mvp_l1_flag) of a first reference picture list (list 1) from a bitstream if the information (inter_pred_idc) about the reference picture list indicates that only a zero reference picture list (list 0) is used; a step of generating motion vector predictor candidates; a step of obtaining a motion vector predictor from the motion vector predictor candidates based on the motion vector predictor index; and a step of predicting the current block based on the motion vector predictor.

[0011] A method for processing a video signal according to one embodiment of the present disclosure further comprises the step of obtaining a motion vector difference zero flag (mvd_l1_zero_flag) from a bitstream indicating whether the motion vector difference and a plurality of control point motion vector differences are set to zero for a first reference picture list, and the step of parsing a motion vector predictor index (mvp_l1_flag) is characterized by including the step of parsing the motion vector predictor index (mvp_l1_flag) regardless of whether the motion vector difference zero flag (mvd_l1_zero_flag) is 1 and information about the reference picture list (inter_pred_idc) indicates the use of both the zero reference picture list and the first reference picture list.

[0012] A method for processing a video signal according to one embodiment of the present disclosure comprises: a step of parsing first information (six_minus_max_num_merge_cand) related to the maximum number of candidates for merge motion vector prediction from a bitstream in sequence units; a step of obtaining the maximum number of merge candidates based on the first information; a step of parsing second information from a bitstream indicating whether a block can be partitioned for inter prediction; and a step of parsing third information from a bitstream related to the maximum number of merge mode candidates for a partitioned block when the second information indicates 1 and the maximum number of merge candidates is greater than 2.

[0013] A method for processing a video signal according to one embodiment of the present disclosure further comprises the steps of: obtaining a maximum number of merge mode candidates for a partitioned block by subtracting a third information from a maximum number of merge candidates when the second information indicates 1 and the maximum number of merge candidates is greater than or equal to 3; setting a maximum number of merge mode candidates for a partitioned block to 2 when the second information indicates 1 and the maximum number of merge candidates is 2; and setting a maximum number of merge mode candidates for a partitioned block to 0 when the second information is 0 or the maximum number of merge candidates is 1.

[0014] An apparatus for processing a video signal according to one embodiment of the present disclosure comprises a processor and a memory, wherein the processor, based on instructions stored in the memory, parses an Adaptive Motion Vector Resolution (AMVR) enable flag (sps_amvr_enabled_flag) indicating whether Adaptive Motion Vector Difference Resolution is used from a bitstream, parses an Affine Enabled flag (sps_affine_enabled_flag) indicating whether Affine Motion Compensation can be used from a bitstream, determines whether Affine Motion Compensation can be used based on the Affine Enabled flag (sps_affine_enabled_flag), and if Affine Motion Compensation can be used, determines whether Adaptive Motion Vector Difference Resolution is used based on the AMVR enabled flag (sps_amvr_enabled_flag), and if Adaptive Motion Vector Difference Resolution is used, determines whether Adaptive Motion Vector Difference Resolution can be used for Affine Motion Compensation from a bitstream, an Affine AMVR Enabled flag indicating whether Adaptive Motion Vector Difference Resolution can be used It is characterized by parsing the flag (sps_affine_amvr_enabled_flag).

[0015] An apparatus for processing a video signal according to one embodiment of the present disclosure is characterized in that one of an AMVR-enabled flag (sps_amvr_enabled_flag), an affine-enabled flag (sps_affine_amvr_enabled_flag), or an affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) is signaled as one of a Coding Tree Unit, a slice, a tile, a tile group, a picture, or a sequence unit.

[0016] An apparatus for processing a video signal according to one embodiment of the present disclosure, wherein affine motion compensation may be used and, when adaptive motion vector difference resolution is not used, the affine AMVR enable flag (sps_affine_amvr_enabled_flag) implies that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0017] In an apparatus for processing a video signal according to one embodiment of the present disclosure, when affine motion compensation cannot be used, the affine AMVR enable flag (sps_affine_amvr_enabled_flag) is characterized by implying that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0018] An apparatus for processing a video signal according to one embodiment of the present disclosure, wherein the processor, based on instructions stored in memory, parses information regarding the resolution of motion vector differences from the bitstream when the AMVR-enabled flag (sps_amvr_enabled_flag) indicates the use of adaptive motion vector difference resolution, the inter-affine flag (inter_affine_flag) obtained from the bitstream indicates that affine motion compensation is not used for the current block, and at least one of a plurality of motion vector differences for the current block is not zero; and modifies a plurality of motion vector differences for the current block based on the information regarding the resolution of motion vector differences.

[0019] In an apparatus for processing a video signal according to one embodiment of the present disclosure, the processor is characterized by, based on instructions stored in memory, parsing information about the resolution of the motion vector difference from the bitstream and modifying the multiple control point motion vector difference for the current block based on the information about the resolution of the motion vector difference, wherein an affine AMVR capable flag indicates that an adaptive motion vector difference resolution can be used for affine motion compensation, an inter_affine_flag obtained from the bitstream indicates the use of affine motion compensation for the current block, and at least one of a plurality of control point motion vector difference differences for the current block is not zero.

[0020] In an apparatus for processing a video signal according to one embodiment of the present disclosure, the processor obtains information (inter_pred_idc) about a reference picture list for a current block based on instructions stored in memory, and if the information (inter_pred_idc) about the reference picture list indicates that only the 0th reference picture list (list 0) is used, parses the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1) from the bitstream, generates motion vector predictor candidates, obtains a motion vector predictor from the motion vector predictor candidates based on the motion vector predictor index, and predicts the current block based on the motion vector predictor.

[0021] In an apparatus for processing a video signal according to one embodiment of the present disclosure, the processor obtains a motion vector difference zero flag (mvd_l1_zero_flag) from a bitstream based on instructions stored in memory, which indicates whether the motion vector difference and a plurality of control point motion vector differences are set to zero for a first reference picture list; and parses a motion vector predictor index (mvp_l1_flag) regardless of whether the motion vector difference zero flag (mvd_l1_zero_flag) is 1 and information about the reference picture list (inter_pred_idc) indicates the use of both the zero reference picture list and the first reference picture list.

[0022] An apparatus for processing a video signal according to one embodiment of the present disclosure comprises a processor and a memory, wherein the processor parses first information (six_minus_max_num_merge_cand) related to the maximum number of candidates for merge motion vector prediction from a bitstream in sequence units based on instructions stored in memory, obtains the maximum number of merge candidates based on the first information, parses second information from the bitstream indicating whether a block can be partitioned for inter prediction, and if the second information indicates 1 and the maximum number of merge candidates is greater than 2, parses third information from the bitstream related to the maximum number of merge mode candidates for a partitioned block.

[0023] In an apparatus for processing a video signal according to one embodiment of the present disclosure, the processor, based on an instruction stored in memory, obtains a maximum number of merge mode candidates for a partitioned block by subtracting a third information from the maximum number of merge candidates when the second information indicates 1 and the maximum number of merge candidates is greater than or equal to 3, sets the maximum number of merge mode candidates for a partitioned block to 2 when the second information indicates 1 and the maximum number of merge candidates is 2, and sets the maximum number of merge mode candidates for a partitioned block to 0 when the second information is 0 or the maximum number of merge candidates is 1.

[0024] A method for processing a video signal according to one embodiment of the present disclosure comprises: generating an Adaptive Motion Vector Resolution (AMVR) enable flag (sps_amvr_enabled_flag) indicating whether Adaptive Motion Vector Difference Resolution is used; generating an Affine Enabled Flag (sps_affine_enabled_flag) indicating whether Affine Motion Compensation can be used; determining whether Affine Motion Compensation can be used based on the Affine Enabled Flag (sps_affine_enabled_flag); if Affine Motion Compensation can be used, determining whether Adaptive Motion Vector Difference Resolution is used based on the AMVR enabled Flag (sps_amvr_enabled_flag); if Adaptive Motion Vector Difference Resolution is used, generating an Affine AMVR Enabled Flag (sps_affine_amvr_enabled_flag) indicating whether Adaptive Motion Vector Difference Resolution can be used for Affine Motion Compensation; and the AMVR enabled Flag (sps_amvr_enabled_flag). It is characterized by including the step of generating a bitstream by entropy coding the affine-enabled flag (sps_affine_enabled_flag) and the AMVR-enabled flag (sps_amvr_enabled_flag).

[0025] A method for processing a video signal according to one embodiment of the present disclosure comprises: generating first information (six_minus_max_num_merge_cand) related to the maximum number of candidates for merge motion vector prediction based on the maximum number of merge candidates; generating second information indicating whether a block can be partitioned for inter prediction; generating third information related to the maximum number of merge mode candidates for a partitioned block when the second information indicates 1 and the maximum number of merge candidates is greater than 2; and generating a bitstream in sequence units by entropy coding the first information (six_minus_max_num_merge_cand), the second information, and the third information. Effects of the invention

[0026] According to an embodiment of the present disclosure, the coding efficiency of a video signal can be increased. Brief explanation of the drawing

[0027] FIG. 1 is a schematic block diagram of a video signal encoder device according to an embodiment of the present disclosure. FIG. 2 is a schematic block diagram of a video signal decoder device according to an embodiment of the present disclosure. FIG. 3 is a drawing showing one embodiment of the present disclosure of dividing a coding unit. FIG. 4 is a drawing illustrating an example of a method for hierarchically representing the partitioned structure of FIG. 3. FIG. 5 is a drawing showing an additional embodiment of the present disclosure dividing a coding unit. FIG. 6 is a diagram illustrating a method for obtaining reference pixels for in-frame prediction. FIG. 7 is a diagram illustrating an example of prediction modes used for in-frame prediction. FIG. 8 is a diagram showing an inter prediction according to one embodiment of the present disclosure. FIG. 9 is a diagram illustrating a motion vector signaling method according to one embodiment of the present disclosure. FIG. 10 is a diagram showing motion vector difference syntax according to one embodiment of the present disclosure. FIG. 11 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure. FIG. 12 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure. FIG. 13 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure. FIG. 14 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure. FIG. 15 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure. FIG. 16 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure. FIG. 17 is a diagram showing affine motion prediction according to one embodiment of the present disclosure. FIG. 18 is a diagram showing affine motion prediction according to one embodiment of the present disclosure. FIG. 19 is an equation representing a motion vector field according to one embodiment of the present disclosure. FIG. 20 is a drawing showing affine motion prediction according to one embodiment of the present disclosure. FIG. 21 is an equation representing a motion vector field according to one embodiment of the present disclosure. FIG. 22 is a drawing showing affine motion prediction according to one embodiment of the present disclosure. FIG. 23 is a diagram showing a mode of affine motion prediction according to one embodiment of the present disclosure. FIG. 24 is a diagram showing a mode of affine motion prediction according to one embodiment of the present disclosure. FIG. 25 is a drawing showing an affine motion predictor derivation according to one embodiment of the present disclosure. FIG. 26 is a drawing showing an affine motion predictor derivation according to one embodiment of the present disclosure. FIG. 27 is a drawing showing an affine motion predictor derivation according to one embodiment of the present disclosure. FIG. 28 is a drawing showing an affine motion predictor derivation according to one embodiment of the present disclosure. FIG. 29 is a diagram illustrating a method for generating a control point motion vector according to one embodiment of the present disclosure. FIG. 30 is a diagram showing a method for determining the motion vector difference through the method described in FIG. 29. FIG. 31 is a diagram illustrating a method for generating a control point motion vector according to one embodiment of the present disclosure. FIG. 32 is a diagram illustrating a method for determining motion vector difference through the method described in FIG. 31. FIG. 33 is a diagram showing motion vector difference syntax according to one embodiment of the present disclosure. FIG. 34 is a drawing showing an upper-level signaling structure according to one embodiment of the present disclosure. FIG. 35 is a diagram showing a coding unit syntax structure according to one embodiment of the present disclosure. FIG. 36 is a drawing showing an upper-level signaling structure according to one embodiment of the present disclosure. FIG. 37 is a drawing showing an upper-level signaling structure according to one embodiment of the present disclosure. FIG. 38 is a diagram showing a coding unit syntax structure according to one embodiment of the present disclosure. FIG. 39 is a diagram showing MVD default value settings according to one embodiment of the present disclosure. FIG. 40 is a diagram showing MVD default value settings according to one embodiment of the present disclosure. FIG. 41 is a diagram showing an AMVR-related syntax structure according to one embodiment of the present disclosure. FIG. 42 is a diagram showing an inter prediction related syntax structure according to one embodiment of the present disclosure. FIG. 43 is a diagram showing an inter prediction related syntax structure according to one embodiment of the present disclosure. FIG. 44 is a diagram showing an inter prediction-related syntax structure according to one embodiment of the present disclosure. FIG. 45 is a diagram showing inter prediction-related syntax according to an embodiment of the present invention. FIG. 46 is a diagram showing a triangle partitioning mode according to an embodiment of the present invention. FIG. 47 is a diagram showing merge data syntax according to an embodiment of the present invention. FIG. 48 is a diagram showing upper-level signaling according to an embodiment of the present invention. FIG. 49 is a diagram showing the maximum number of candidates used in a TPM according to one embodiment of the present invention. FIG. 50 is a diagram showing upper-level signaling regarding a TPM according to an embodiment of the present invention. FIG. 51 is a diagram showing the maximum number of candidates used in a TPM according to one embodiment of the present invention. FIG. 52 is a diagram showing TPM-related syntax elements according to an embodiment of the present invention. FIG. 53 is a diagram showing TPM candidate index signaling according to an embodiment of the present invention. FIG. 54 is a diagram showing TPM candidate index signaling according to an embodiment of the present invention. FIG. 55 is a diagram showing TPM candidate index signaling according to an embodiment of the present invention. Specific details for implementing the invention

[0028] The terms used in this specification have been selected to be as widely used as possible, taking into account their functions in the present invention; however, these may vary depending on the intent, convention, or emergence of new technologies of those skilled in the art. In addition, in certain cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in the relevant description of the invention. Therefore, it should be noted that the terms used in this specification should be interpreted based on their actual meanings and the overall content of this specification, rather than merely their names.

[0029] In this disclosure, the following terms may be interpreted according to the following criteria, and even terms not explicitly stated may be interpreted according to the intent below. "Coding" may be interpreted as encoding or decoding depending on the case, and "information" is a term that includes values, parameters, coefficients, elements, etc., and since its meaning may be interpreted differently depending on the case, this disclosure is not limited thereto. "Unit" is used to refer to a basic unit of image (picture) processing or a specific location within a picture, and may be used interchangeably with terms such as "block," "partition," or "region" depending on the case. Furthermore, in this specification, "unit" may be used as a concept that includes coding units, prediction units, and transformation units.

[0030] FIG. 1 is a schematic block diagram of a video signal encoding device according to one embodiment of the present disclosure. Referring to FIG. 1, the encoding device (100) of the present disclosure mainly comprises a conversion unit (110), a quantization unit (115), an inverse quantization unit (120), an inverse conversion unit (125), a filtering unit (130), a prediction unit (150), and an entropy coding unit (160).

[0031] The conversion unit (110) converts pixel values ​​of the input video signal to obtain conversion coefficient values. For example, the Discrete Cosine Transform (DCT) or Wavelet Transform may be used. In particular, the Discrete Cosine Transform divides the input picture signal into blocks of a certain size to perform the conversion. In the conversion, the coding efficiency may vary depending on the distribution and characteristics of the values ​​within the conversion area.

[0032] The quantization unit (115) quantizes the conversion coefficient value output from the conversion unit (110). The inverse quantization unit (120) inversely quantizes the conversion coefficient value, and the inverse conversion unit (125) restores the original pixel value using the inversely quantized conversion coefficient value.

[0033] The filtering unit (130) performs filtering operations to improve the quality of the restored picture. For example, it may include a deblocking filter and an adaptive loop filter. The filtered picture is stored in a Decoded Picture Buffer (156) to be output or used as a reference picture.

[0034] To increase coding efficiency, instead of coding the picture signal as is, a method is used in which the picture is predicted using an area already coded through the prediction unit (150), and the residual value between the original picture and the predicted picture is added to the predicted picture to obtain the restored picture. The intra prediction unit (152) performs intra-frame prediction within the current picture, and the inter prediction unit (154) predicts the current picture using a reference picture stored in the decoded picture buffer (156). The intra prediction unit (152) performs intra-frame prediction from the restored areas within the current picture and transmits intra-frame encoding information to the entropy coding unit (160). The inter prediction unit (154) may again be configured to include a motion estimation unit (154a) and a motion compensation unit (154b). The motion estimation unit (154a) obtains the motion vector value of the current area by referencing a specific restored area. The motion estimation unit (154a) transmits location information of the reference area (reference frame, motion vector, etc.) to the entropy coding unit (160) so that it can be included in the bitstream. Using the motion vector value transmitted from the motion estimation unit (154a), the motion compensation unit (154b) performs inter-frame motion compensation.

[0035] The entropy coding unit (160) generates a video signal bitstream by entropy coding the quantized conversion coefficients, inter-frame encoding information, intra-frame encoding information, and reference area information input from the inter-prediction unit (154). Here, the entropy coding unit (160) may use a variable length coding (VLC) method and arithmetic coding. The variable length coding (VLC) method converts input symbols into a continuous codeword, and the length of the codeword may be variable. For example, frequently occurring symbols are represented as short codewords, and infrequently occurring symbols are represented as long codewords. As a variable length coding method, a context-based adaptive variable length coding (CAVLC) method may be used. Arithmetic coding converts continuous data symbols into a single prime number, and arithmetic coding can obtain the optimal prime number bit required to represent each symbol. Context-based Adaptive Binary Arithmetic Code (CABAC) can be used as arithmetic coding.

[0036] The generated bitstream is encapsulated with Network Abstraction Layer (NAL) units as the basic unit. The NAL unit contains encoded slice segments, which consist of an integer number of Coding Tree Units. To decode the bitstream in a video decoder, the bitstream must first be separated into NAL units, and then each separated NAL unit must be decoded.

[0037] FIG. 2 is a schematic block diagram of a video signal decoding device (200) according to one embodiment of the present disclosure. Referring to FIG. 2, the decoding device (200) of the present disclosure mainly includes an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (225), a filtering unit (230), and a prediction unit (250).

[0038] The entropy decoding unit (210) entropies decodes the video signal bitstream to extract transformation coefficients, motion information, etc. for each region. The inverse quantization unit (220) inversely quantizes the entropy-decoded transformation coefficients, and the inverse transformation unit (225) restores the original pixel values ​​using the inversely quantized transformation coefficients.

[0039] Meanwhile, the filtering unit (230) performs filtering on the picture to improve image quality. This may include a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is stored in a Decoded Picture Buffer (256) to be output or used as a reference picture for the next frame.

[0040] Additionally, the prediction unit (250) of the present disclosure includes an intra prediction unit (252) and an inter prediction unit (254), and restores a predicted picture by utilizing the encoding type decoded through the aforementioned entropy decoding unit (210), a conversion coefficient for each region, motion information, etc.

[0041] In this regard, the intra prediction unit (252) performs in-frame prediction from the decoded samples within the current picture. The inter prediction unit (254) generates a predicted picture using the reference picture and motion information stored in the decoded picture buffer (256). The inter prediction unit (254) may again be configured to include a motion estimation unit (254a) and a motion compensation unit (254b). The motion estimation unit (254a) obtains a motion vector representing the positional relationship between the current block and the reference block of the reference picture used for coding, and transmits it to the motion compensation unit (254b).

[0042] A restored video frame is generated by adding the predicted value output from the intra prediction unit (252) or the inter prediction unit (254) and the pixel value output from the inverse transformation unit (225).

[0043] Hereinafter, regarding the operation of the encoding device (100) and the decoding device (200), a method of dividing the coding unit and the prediction unit, etc., will be explained with reference to FIGS. 3 to 5.

[0044] A coding unit refers to a basic unit for processing a picture during the video signal processing described above, such as intra / inter-frame prediction, transform, quantization, and / or entropy coding. The size of the coding unit used to code a single picture may not be constant. A coding unit may have a rectangular shape, and a single coding unit can be divided into multiple coding units.

[0045] FIG. 3 illustrates an embodiment of the present disclosure for dividing a coding unit. For example, a single coding unit having a size of 2N x 2N may be divided into four coding units having a size of NXN. This division of the coding unit may be performed recursively, and not all coding units need to be divided into the same form. However, for convenience in the coding and processing process, there may be limitations on the size of the maximum coding unit and / or the minimum coding unit.

[0046] For a single coding unit, information indicating whether the coding unit is divided can be stored. FIG. 4 illustrates an example of a method for hierarchically representing the division structure of a coding unit illustrated in FIG. 3 using a flag value. Information indicating whether a coding unit is divided can be assigned a value of '1' if the unit is divided and '0' if it is not divided. As illustrated in FIG. 4, if the flag value indicating division is 1, the coding unit corresponding to the node is divided again into four coding units, and if it is 0, it is not divided further, and a processing process for the coding unit can be performed.

[0047] The structure of the coding unit described above can be represented using a recursive tree structure. That is, with a single picture or maximum-size coding unit as the root, a coding unit that is divided into other coding units will have as many child nodes as the number of divided coding units. Therefore, a coding unit that is no longer divided becomes a leaf node. Assuming that only square division is possible for a single coding unit, since a single coding unit can be divided into up to four other coding units, the tree representing the coding unit can take the form of a quad tree.

[0048] In an encoder, the optimal coding unit size is selected based on the characteristics of the video picture (e.g., resolution) or by considering coding efficiency, and information regarding this or information that can derive it may be included in the bitstream. For example, the maximum coding unit size and the maximum depth of the tree may be defined. When performing square partitioning, the height and width of the coding unit become half the height and width of the parent node's coding unit; therefore, the minimum coding unit size can be calculated using such information. Conversely, the minimum coding unit size and the maximum depth of the tree can be defined in advance and used to derive the maximum coding unit size. Since the unit size changes in the form of multiples of 2 in square partitioning, the actual coding unit size can be represented as a logarithm with base 2 to improve transmission efficiency.

[0049] The decoder can obtain information indicating whether the current coding unit has been split. Efficiency can be improved by obtaining (transmitting) this information only under specific conditions. For example, the condition under which the current coding unit can be split is that the sum of the current coding unit sizes at the current position is smaller than the picture size, and the current unit size is larger than a preset minimum coding unit size; therefore, information indicating whether the current coding unit has been split can be obtained only in such cases.

[0050] If the above information indicates that the coding unit has been divided, the size of the coding unit to be divided becomes half of the current coding unit, and it is divided into four square coding units based on the current processing position. The above processing can be repeated for each divided coding unit.

[0051] FIG. 5 illustrates an additional embodiment of the present disclosure for dividing a coding unit. According to an additional embodiment of the present disclosure, the aforementioned quad tree-shaped coding unit may be further divided into a binary tree structure of horizontal or vertical division. That is, a square quad tree division may first be applied to the root coding unit, and a rectangular binary tree division may be additionally applied to the leaf nodes of the quad tree. According to one embodiment, the binary tree division may be a symmetric horizontal division or a symmetric vertical division, but the present disclosure is not limited thereto.

[0052] At each split node of the binary tree, a flag indicating the split type (i.e., horizontal split or vertical split) may be additionally signaled. According to one embodiment, if the value of the flag is '0', a horizontal split is indicated, and if the value of the flag is '1', a vertical split is indicated.

[0053] However, in the embodiments of the present disclosure, the method of dividing the coding unit is not limited to the methods described above, and asymmetric horizontal / vertical division, a triple tree divided into three rectangular coding units, etc. may be applied.

[0054] Picture prediction (motion compensation) for coding is performed on indivisible coding units (i.e., leaf nodes of the coding unit tree). The basic unit that performs this prediction is referred to hereinafter as a prediction unit or prediction block.

[0055] Hereinafter, the term "unit" as used in this specification may be used as a substitute for the prediction unit, which is the basic unit for performing predictions. However, this disclosure is not limited thereto, and can be understood more broadly as a concept including the coding unit.

[0056] To restore the current unit for which decoding is performed, the decoded portions of the current picture containing the current unit or other pictures may be used. A picture (slice) that uses only the current picture for restoration, that is, performs only intra-frame prediction, is called an intra-picture or I picture (slice), and a picture (slice) that can perform both intra-frame prediction and inter-frame prediction is called an inter-picture (slice). Among the inter-pictures (slices), a picture (slice) that uses at most one motion vector and reference index to predict each unit is called a predictive picture or P picture (slice), and a picture (slice) that uses at most two motion vectors and reference indices is called a bi-predictive picture or B picture (slice).

[0057] The intra prediction unit performs intra-prediction, which predicts the pixel value of a target unit from restored regions within the current picture. For example, the pixel value of the current unit can be predicted from the restored pixels of units located to the left and / or above the current unit. In this case, the units located to the left of the current unit may include the left unit adjacent to the current unit, the top-left unit, and the bottom-left unit. Additionally, the units located above the current unit may include the top unit adjacent to the current unit, the top-left unit, and the top-right unit.

[0058] Meanwhile, the inter-prediction unit performs inter-frame prediction, which predicts the pixel values ​​of the target unit using information from other restored pictures rather than the current picture. In this case, the picture used for prediction is called the reference picture. During the inter-frame prediction process, which reference region is used to predict the current unit can be indicated using information such as the index representing the reference picture containing that reference region and motion vector information.

[0059] Inter-frame prediction may include L0 prediction, L1 prediction, and bi-prediction. L0 prediction refers to prediction using one reference picture included in L0 (the 0th reference picture list), and L1 prediction refers to prediction using one reference picture included in L1 (the 1st reference picture list). For this, one set of motion information (e.g., motion vectors and reference picture indices) may be required. In the bi-prediction method, up to two reference regions can be used, and these two reference regions may exist in the same reference picture or in different pictures. That is, in the bi-prediction method, up to two sets of motion information (e.g., motion vectors and reference picture indices) can be used, and two motion vectors may correspond to the same reference picture index or to different reference picture indices. In this case, the reference pictures may be displayed (or output) both before and after the current picture in time.

[0060] The reference unit of the current unit can be obtained using the motion vector and the reference picture index. The reference unit exists within the reference picture having the reference picture index. Additionally, the pixel value or interpolated value of the unit specified by the motion vector can be used as the predictor of the current unit. For motion prediction with sub-pel level pixel accuracy, for example, an 8-tap interpolation filter may be used for the luminance signal and a 4-tap interpolation filter may be used for the chrominance signal. However, the interpolation filters for sub-pel level motion prediction are not limited thereto. In this way, motion compensation is performed to predict the texture of the current unit from the previously decoded picture using motion information.

[0061] Hereinafter, an intra-frame prediction method according to an embodiment of the present disclosure will be described in more detail with reference to FIGS. 6 and FIGS. 7. As described above, the intra-frame prediction unit predicts the pixel value of the current unit by using adjacent pixels located to the left and / or top of the current unit as reference pixels.

[0062] As illustrated in FIG. 6, when the size of the current unit is NXN, reference pixels can be established using up to 4N+1 adjacent pixels located to the left and / or top of the current unit. If at least some of the adjacent pixels to be used as reference pixels have not yet been restored, the intra prediction unit can obtain reference pixels by performing a reference sample padding process according to a pre-established rule. Additionally, the intra prediction unit can perform a reference sample filtering process to reduce the error in intra-frame prediction. That is, reference pixels can be obtained by performing filtering on adjacent pixels and / or pixels obtained by the reference sample padding process. The intra prediction unit predicts the pixels of the current unit using the reference pixels obtained in this way.

[0063] FIG. 7 illustrates an example of prediction modes used for intra-frame prediction. For intra-frame prediction, intra-frame prediction mode information indicating the direction of intra-frame prediction may be signaled. When the current unit is an intra-frame prediction unit, the video signal decoding device extracts intra-frame prediction mode information of the current unit from the bitstream. The intra-prediction unit of the video signal decoding device performs intra-frame prediction for the current unit based on the extracted intra-frame prediction mode information.

[0064] According to one embodiment of the present disclosure, the intra-frame prediction mode may include a total of 67 modes. Each intra-frame prediction mode may be indicated by a pre-set index (i.e., an intra-mode index). For example, as illustrated in FIG. 7, an intra-mode index 0 indicates a planar mode, an intra-mode index 1 indicates a DC mode, and intra-mode indices 2 through 66 may each indicate different directional modes (i.e., angle modes). The intra-frame prediction unit determines reference pixels and / or interpolated reference pixels to be used for intra-frame prediction of the current unit based on the intra-frame prediction mode information of the current unit. When the intra-mode index indicates a specific directional mode, a reference pixel or interpolated reference pixel corresponding to the specific direction from the current pixel of the current unit is used for the prediction of the current pixel. Accordingly, different sets of reference pixels and / or interpolated reference pixels may be used for intra-frame prediction depending on the intra-frame prediction mode.

[0065] After the intra-frame prediction of the current unit is performed using reference pixels and intra-frame prediction mode information, the video signal decoding device restores the pixel values ​​of the current unit by adding the residual signal of the current unit obtained from the inverse transform unit to the intra-frame prediction value of the current unit.

[0066] FIG. 8 is a diagram showing an inter prediction according to one embodiment of the present disclosure.

[0067] As explained earlier, when encoding or decoding the current picture or block, it is possible to make predictions based on other pictures or blocks. In other words, it is possible to encode or decode based on similarity with other pictures or blocks. Parts similar to other pictures or blocks can be encoded or decoded using signaling that is omitted in the current picture or block, and this will be explained further below. It is possible to make predictions at the block level.

[0068] Referring to FIG. 8, there is a Reference picture on the left and a Current picture on the right. The Current picture or a part of the Current picture can be predicted using the similarity to the Reference picture or a part of the Reference picture. If the rectangle indicated by solid lines within the Current picture in FIG. 8 represents the block currently being encoded or decoded, the Current block can be predicted from the rectangle indicated by dotted lines in the Reference picture. In this case, there may be information indicating the block that the Current block must reference (Reference block); this information may be directly signaled or generated by some convention to reduce signaling overhead. The information indicating the block that the Current block must reference may include a motion vector. This may be a vector indicating the relative position within the picture between the Current block and the Reference block. Referring to FIG. 8, there is a part indicated by dotted lines in the Reference picture, and the motion vector may be a vector indicating how the Current block moves to reach the block that must be referenced in the Reference picture. That is, the block that appears when the current block is moved according to the motion vector may be the part indicated by the dotted line in the current picture of Fig. 8, and the position of this part indicated by the dotted line within the current picture may be the same as the position of the reference block in the reference picture.

[0069] Additionally, the information indicating the block that the current block must reference may include information representing a reference picture. The information representing the reference picture may include a reference picture list and a reference picture index. The reference picture list is a list containing reference pictures, and it is possible to use a reference block from a reference picture included in the reference picture list. That is, it is possible to predict the current block from a reference picture included in the reference picture list. Furthermore, the reference picture index may be an index for indicating the reference picture to be used.

[0070] FIG. 9 is a diagram illustrating a motion vector signaling method according to one embodiment of the present disclosure.

[0071] According to one embodiment of the present disclosure, a motion vector (MV) can be generated based on a motion vector predictor (MVP). For example, the motion vector predictor can be a motion vector as follows.

[0072] MV = MVP

[0073] As another example, the motion vector can be based on the motion vector difference (MVD) as shown below. The motion vector difference (MVD) can be added to the motion vector predictor to represent the accurate motion vector.

[0074] MV = MVP + MVD

[0075] Additionally, in video coding, motion vector information determined by the encoder is transmitted to the decoder, and the decoder can generate motion vectors from the received motion vector information and determine prediction blocks. For example, the motion vector information may include information regarding motion vector predictors and motion vector differences. In this case, the components of the motion vector information may vary depending on the mode. For example, in merge mode, the motion vector information may include information regarding motion vector predictors but may not include motion vector differences. As another example, in AMVP (advanced motion vector prediction) mode, the motion vector information may include information regarding motion vector predictors and include motion vector differences.

[0076] In order to determine, transmit, and receive information regarding a motion vector predictor, the encoder and the decoder can generate MVP candidates in the same way. For example, the encoder and the decoder can generate the same MVP candidates in the same order. Then, the encoder transmits an index (mvp_lx_flag) representing the MVP determined from among the generated MVP candidates to the decoder, and the decoder can determine the determined MVP and MV based on this index (mvp_lx_flag). The index (mvp_lx_flag) may include the motion vector predictor index (mvp_l0_flag) of the 0th reference picture list (list 0) and the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1). A method for receiving the index (mvp_lx_flag) is described in FIGS. 42 through 45.

[0077] MVP candidates and methods for generating MVP candidates may include spatial candidates, temporal candidates, etc. Spatial candidates may be motion vectors for blocks located at a certain position relative to the current block. For example, they may be motion vectors corresponding to blocks or positions that are adjacent to or not adjacent to the current block. Temporal candidates may be motion vectors corresponding to blocks within a picture different from the current picture. Alternatively, MVP candidates may include affine motion vectors, ATMVP, STMVP, combinations of the previously described motion vectors, average vectors of the previously described motion vectors, zero motion vectors, etc.

[0078] In addition, information representing the reference picture described earlier can also be transmitted from the encoder to the decoder. Furthermore, motion vector scaling can be performed when the reference picture corresponding to the MVP candidate does not correspond to the information representing the reference picture. Motion vector scaling can be a calculation based on the picture order count (POC) of the current picture, the POC of the reference picture of the current block, the POC of the reference picture of the MVP candidate, or the MVP candidate.

[0079] FIG. 10 is a diagram showing motion vector difference syntax according to one embodiment of the present disclosure.

[0080] Motion vector difference can be coded by separating the sign and absolute value of the motion vector difference. That is, the sign and absolute value of the motion vector difference can have different syntax. Additionally, while the absolute value of the motion vector difference can be coded directly, it can also be coded to include a flag indicating whether the absolute value is greater than N, as shown in Fig. 10. If the absolute value is greater than N, the value of (absolute value - N) can be signaled along with it. In the example of Fig. 10, abs_mvd_greater0_flag can be transmitted, and this flag can be a flag indicating whether the absolute value is greater than 0. If abs_mvd_greater0_flag indicates that the absolute value is not greater than 0, it can be determined that the absolute value is 0. Furthermore, if abs_mvd_greater0_flag indicates that the absolute value is greater than 0, additional syntax may exist. For example, abs_mvd_greater1_flag may exist, which may indicate whether the absolute value is greater than 1. If abs_mvd_greater1_flag indicates that the absolute value is not greater than 1, it can be determined that the absolute value is 1. If abs_mvd_greater1_flag indicates that the absolute value is greater than 1, additional syntax may exist. For example, abs_mvd_minus2 may exist, which may be the value of (absolute value - 2).Since it was determined through the previously explained abs_mvd_greater0_flag and abs_mvd_greater1_flag that the absolute value is greater than 1 (i.e., 2 or greater), it represents (absolute value - 2). This is intended to signal with fewer bits when abs_mvd_minus2 is binarized to variable length. For example, variable-length binarization methods such as Exp-Golomb, truncated unary, and truncated Rice exist. Additionally, mvd_sign_flag can be a flag indicating the sign of the motion vector difference.

[0081] In this embodiment, the coding method was explained through the motion vector difference, but information other than the motion vector difference can also be divided into sign and absolute value. The absolute value can be coded as a flag indicating whether the absolute value is greater than a certain value and as the absolute value minus the said certain value. Additionally, [0] and [1] in FIG. 10 can represent component indices. For example, they can represent x-component and y-component.

[0082] FIG. 11 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure.

[0083] According to one embodiment of the present disclosure, the resolution representing a motion vector or a motion vector difference may vary. In other words, the resolution at which the motion vector or motion vector difference is coded may vary. For example, the resolution may be represented based on pixels. For example, the motion vector or motion vector difference may be signaled in units such as 1 / 4 (quarter), 1 / 2 (half), 1 (integer), 2, or 4 pixels. For example, when you want to represent 16, if you use a 1 / 4 unit, it is coded as 64 (1 / 4 * 64 = 16); if you use a 1 unit, it is coded as 16 (1 * 16 = 16); and if you use a 4 unit, it is coded as 4 (4 * 4 = 16). That is, the value can be determined as follows.

[0084] valueDetermined = resolution*valuePerResolution

[0085] Here, valueDetermined is a value to be transmitted, and in this embodiment, it may be a motion vector or a motion vector difference. Additionally, valuePerResolution may be a value representing valueDetermined in units of [ / resolution].

[0086] In this case, if the signal value represented by the motion vector or motion vector difference is not divisible by the resolution, an inaccurate value may be sent instead of the motion vector or motion vector difference that offers the best prediction performance, due to rounding or other methods. Using high resolution can reduce inaccuracy but allows for the use of more bits because the encoded value is large, whereas using low resolution may increase inaccuracy but allows for the use of fewer bits because the encoded value is small.

[0087] In addition, it is possible to set the above resolution differently in units such as blocks, CUs, and slices. Therefore, the resolution can be applied adaptively to suit the unit.

[0088] The above resolution can be signaled from the encoder to the decoder. In this case, the signaling for the resolution may be a signal binarized by the variable length explained earlier. In such a case, signaling with the index corresponding to the smallest value (the first value) reduces the signaling overhead.

[0089] In one embodiment, signaling indices can be matched in order from high resolution (detailed signaling) to low resolution.

[0090] Figure 11 illustrates signaling for three resolutions. In this case, the three signals can be 0, 10, and 11, and each of the three signals can correspond to resolution 1, resolution 2, and resolution 3. Since 1 bit is required to signal resolution 1 and 2 bits are required to signal the other resolutions, the signaling overhead is low when signaling resolution 1. In the example of Figure 11, resolutions 1, 2, and 3 are 1 / 4, 1, and 4 pel, respectively.

[0091] In the following disclosures, motion vector resolution may mean the resolution of motion vector difference.

[0092] FIG. 12 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure.

[0093] As explained in Fig. 11, since the number of bits required to signal a resolution may vary depending on the resolution, it is possible to change the signaling method according to the situation. For example, the signaling value signaling a certain resolution may vary depending on the situation. For example, the signaling index and the resolution can be matched in a different order depending on the situation. For example, the resolutions corresponding to signaling 0, 10, 110, ... may be resolution 1, resolution 2, resolution 3, ... in some situations, respectively, but in other situations, they may not be resolution 1, resolution 2, resolution 3, ... but in a different order. It is also possible to define two or more situations.

[0094] Referring to Fig. 12, the resolutions corresponding to 0, 10, and 11 may be resolution 1, resolution 2, and resolution 3, respectively, in Case 1, and resolution 2, resolution 1, and resolution 3, respectively, in Case 2. In this case, there may be two or more cases.

[0095] FIG. 13 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure.

[0096] As explained in FIG. 12, motion vector resolution can be signaled differently depending on the situation. For example, assuming there are resolutions of 1 / 4, 1, and 4 pel, in some situations, the signaling described in FIG. 11 may be used, while in other situations, the signaling shown in FIG. 13(a) or FIG. 13(b) may be used. It is possible for only two of FIG. 11, FIG. 13(a), and FIG. 13(b) to exist, or for all three to exist. Therefore, in some situations, it is possible to signal a resolution other than the highest resolution with fewer bits.

[0097] FIG. 14 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure.

[0098] According to one embodiment of the present disclosure, the available resolutions in the adaptive motion vector resolution described in FIG. 11 may change depending on the situation. For example, the resolution values ​​may change depending on the situation. In one embodiment, it is possible to use resolution 1, resolution 2, resolution 3, resolution 4, ... in some situations and resolution A, resolution B, resolution C, resolution D, ... in other situations. Furthermore, there may be an intersection that is not an empty set between {resolution 1, resolution 2, resolution 3, resolution 4, ...} and {resolution B, resolution C, resolution D, ...}. That is, a certain resolution value may be used in two or more situations, and the set of available resolution values ​​may differ in the two or more situations. In addition, the number of available resolution values ​​may differ depending on the situation.

[0099] Referring to Fig. 14, in Case 1, Resolution 1, Resolution 2, and Resolution 3 can be used, and in Case 2, Resolution A, Resolution B, and Resolution C can be used. For example, Resolution 1, Resolution 2, and Resolution 3 can be 1 / 4, 1, and 4 pel. Also, for example, Resolution A, Resolution B, and Resolution C can be 1 / 4, 1 / 2, and 1.

[0100] FIG. 15 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure.

[0101] Referring to FIG. 15, resolution signaling can be varied depending on the motion vector candidate or motion vector predictor candidate. For example, the case described in FIG. 12 to 14 may be selected, or it may be about which candidate the signaled motion vector or motion vector predictor is. The method of varying the signaling may follow the method of FIG. 12 to 14.

[0102] For example, the situation can be defined differently depending on the position a candidate holds among the candidates. Alternatively, the situation can be defined differently depending on how the candidate was created.

[0103] An encoder or decoder may generate a candidate list containing at least one MV candidate (motion vector candidate) or at least one MVP candidate (motion vector predictor candidate). In the candidate list for the MV candidate or MVP candidate, the MVP (motion vector predictor) at the beginning may have high accuracy, while the MVP (motion vector predictor) at the end of the candidate list may have low accuracy. This may be because the one at the beginning of the candidate list signals with fewer bits, and the MVP (motion vector predictor) at the beginning of the candidate list is designed to have higher accuracy. In the embodiments of the present disclosure, if the accuracy of the MVP (motion vector predictor) is high, the MVD (motion vector difference) value for representing a motion vector with good prediction performance may be small, and if the accuracy of the MVP (motion vector predictor) is low, the MVD (motion vector difference) value for representing a motion vector with good prediction performance may be large. Therefore, when the accuracy of the MVP (motion vector predictor) is low, it is possible to signal at low resolution to reduce the number of bits required to represent the motion vector difference value (e.g., a value representing the difference value based on resolution).

[0104] Based on this principle, low resolution can be used when the accuracy of the MVP (motion vector predictor) is low. Therefore, according to one embodiment of the present disclosure, it is possible to promise to signal with the lowest resolution bits, rather than the highest resolution, depending on the MVP candidate (motion vector predictor candidate). For example, when 1 / 4, 1, or 4 pel are possible resolutions, 1 or 4 can be signaled with the fewest bits (1-bit). Referring to FIG. 15, for Candidate 1 and Candidate 2, 1 / 4 pel, which is high resolution, is signaled with the fewest bits, and for Candidate N, which is after Candidate 1 and Candidate 2, a resolution other than 1 / 4 pel is signaled with the fewest bits.

[0105] FIG. 16 is a diagram showing adaptive motion vector resolution signaling according to one embodiment of the present disclosure.

[0106] As explained in FIG. 15, the motion vector resolution signaling can be varied depending on which candidate the determined motion vector or motion vector predictor is. The method of varying the signaling can follow the method of FIG. 12 to FIG. 14.

[0107] Referring to Fig. 16, for some candidates, high resolution is signaled with the fewest bits, and for others, a resolution other than high resolution is signaled with the fewest bits. For example, a candidate that signals a resolution other than high resolution with the fewest bits may be an inaccurate candidate. For example, a candidate that signals a resolution other than high resolution with the fewest bits may be a temporal candidate, a zero motion vector, a non-adjacent spatial candidate, a candidate depending on whether or not there is a refinement process.

[0108] A temporal candidate can be a motion vector from another picture. A zero motion vector can be a motion vector in which all vector components are zero. A non-adjacent spatial candidate can be a motion vector referenced from a location that is not adjacent to the current block. The refinement process can be the process of refining motion vector predictors, and can be refined, for example, through template matching, bilateral matching, etc.

[0109] According to one embodiment of the present disclosure, it is possible to add the motion vector difference after refining the motion vector predictor. This may be intended to reduce the value of the motion vector difference by making the motion vector predictor accurate. In this case, the motion vector resolution signaling can be different for a candidate that has not undergone a refinement process. For example, a resolution other than the highest resolution can be signaled with the fewest bits.

[0110] According to another embodiment, a refinement process can be performed after adding the motion vector difference to the motion vector predictor. In this case, the motion vector resolution signaling can be done differently for the candidate undergoing the refinement process. For example, a resolution other than the highest resolution can be signaled with the fewest bits. This may be because, since the refinement process is performed after adding the motion vector difference, it is possible to reduce the prediction error through the refinement process even if the MV difference is not signaled as accurately as possible (to minimize the prediction error).

[0111] As another example, the motion vector resolution signaling can be changed when the selected candidate differs from other candidates by more than a certain amount.

[0112] In another embodiment, the subsequent motion vector refinement process may be varied depending on which candidate the determined motion vector or motion vector predictor is. The motion vector refinement process may be a process aimed at finding a more accurate motion vector. For example, it may be a process of finding a block that matches the current block according to a defined convention from a reference point (e.g., template matching or bilateral matching). The reference point may be a location corresponding to the determined motion vector or motion vector predictor. In this case, the degree of movement from the reference point may vary according to the defined convention; thus, varying the motion vector refinement process may mean varying the degree of movement from the reference point. For example, for an accurate candidate, a detailed refinement process may be started, while for an inaccurate candidate, a less detailed refinement process may be started. The accuracy and inaccuracy of a candidate can be determined by its position in the candidate list or the method in which the candidate was generated. The method in which the candidate was generated may be the location from which the spatial candidate was taken. For this, refer to the explanation in Fig. 16. Additionally, the detailed and less detailed refinement may be whether the matching block is searched by moving it slightly or by moving it significantly from the reference point.Additionally, in cases where searching involves significant movement, you can add a process of searching by moving slightly further from the most matching block found through extensive movement.

[0113] In another embodiment, the motion vector resolution signaling can be changed based on the POC of the current picture and the POC of the reference picture of the motion vector or motion vector predictor candidate. The method of changing the signaling can follow the method of FIGS. 12 to 14.

[0114] For example, if there is a large difference between the Picture Order Count (POC) of the current picture and the POC of the reference picture of the motion vector or motion vector predictor candidate, the motion vector or motion vector predictor may be inaccurate and may signal a resolution other than high resolution with the fewest bits.

[0115] As another example, motion vector resolution signaling can be modified based on whether motion vector scaling is required. For instance, if the selected motion vector or motion vector predictor is a candidate for motion vector scaling, a non-high resolution can be signaled with the fewest bits. Motion vector scaling may occur when the reference picture for the current block differs from the reference picture of the referenced candidate.

[0116] FIG. 17 is a diagram showing affine motion prediction according to one embodiment of the present disclosure.

[0117] In conventional prediction methods such as those described in Fig. 8, the current block can be predicted from a block at a position where it has been moved without rotation or scaling. It can be predicted from a reference block of the same size, shape, and angle as the current block. Fig. 8 is merely a translation motion model. However, the content contained in actual video can have more complex movements, and prediction performance can be improved if predicted from various shapes.

[0118] Referring to Fig. 17, the block currently being predicted is indicated by a solid line in the current picture. The current block can be predicted by referencing a block with a different shape, size, and angle from the block currently being predicted, and the reference block is indicated by a dotted line in the reference picture. The same position as the reference block within the picture is indicated by a dotted line in the current picture. In this case, the reference block may be a block represented by an affine transformation of the current block. Through this, it is possible to represent stretching (scaling), rotation, shearing, reflection, orthogonal projection, etc.

[0119] The number of parameters representing affine motion and affine transformation can vary. Using more parameters allows for a wider variety of motions than using fewer parameters, but it may result in overhead in signaling or calculation.

[0120] For example, an affine transformation can be represented by 6 parameters. Or, an affine transformation can be represented by 3 control point motion vectors.

[0121] FIG. 18 is a diagram showing affine motion prediction according to one embodiment of the present disclosure.

[0122] As shown in Fig. 17, it is possible to represent complex motion using affine transformation, but to reduce signaling overhead and computation for this, simpler affine motion prediction or affine transformation can be used. Simpler affine motion prediction can be achieved by restricting motion. Restricting motion may limit the shape in which the current block transforms into the reference block.

[0123] Referring to Fig. 18, affine motion prediction can be performed using the control point motion vectors v0 and v1. Using two vectors, v0 and v1, is equivalent to using four parameters. The two vectors v0 and v1 or the four parameters can indicate what shape the current block is predicted from a reference block. By using such a simple affine transformation, the rotation and scaling (zoom in / out) movements of the block can be represented. Referring to Fig. 18, the current block, indicated by the solid line, can be predicted from the position indicated by the dotted line in Fig. 18 within the reference picture. Through the affine transformation, each point (pixel) of the current block can be mapped to a different point.

[0124] FIG. 19 is an equation representing a motion vector field according to one embodiment of the present disclosure. The control point motion vector v0 in FIG. 18 is (v_0x, v_0y) and may be the motion vector of a top-left corner control point. Additionally, the control point motion vector v1 is (v_1x, v_1y) and may be the motion vector of a top-right corner control point. In this case, the motion vector (v_x, v_y) at the (x, y) position may be as in FIG. 19. Therefore, the motion vector for each pixel position or any position can be estimated according to the equation of FIG. 19, and this is based on v0 and v1.

[0125] In addition, in the formula of Fig. 19, (x, y) may be relative coordinates within the block. For example, (x, y) may be the position when the left top position of the block is (0, 0).

[0126] If v0 is a control point motion vector for position (x0, y0) on the picture and v1 is a control point motion vector for position (x1, y1) on the picture, to represent the position (x, y) within the block using the same coordinates as the positions of v0 and v1, x and y in the formula of FIG. 19 can be changed to (x-x0) and (y-y0), respectively. Also, w (width of the block) can be (x1-x0).

[0127] FIG. 20 is a diagram showing affine motion prediction according to one embodiment of the present disclosure.

[0128] According to one embodiment of the present disclosure, affine motion can be represented using a plurality of control point motion vectors or a plurality of parameters.

[0129] Referring to FIG. 20, affine motion prediction can be performed using control point motion vectors v0, v1, and v2. Using the three vectors v0, v1, and v2 may be equivalent to using six parameters. The three vectors v0, v1, and v2 or the six parameters can indicate what shape of reference block the current block is predicted from. Referring to FIG. 20, the current block represented by the solid line can be predicted from the position represented by the dotted line in FIG. 20 in the reference picture. Through affine transformation, each point (pixel) of the current block can be mapped to another point.

[0130] FIG. 21 is an equation representing a motion vector field according to one embodiment of the present disclosure. In FIG. 20, the control point motion vector v0 may be (mv_0^x, mv_0^y) and may be the motion vector of a top-left corner control point, the control point motion vector v1 may be (mv_1^x, mv_1^y) and may be the motion vector of a top-right corner control point, and the control point motion vector v2 may be (mv_2^x, mv_2^y) and may be the motion vector of a bottom-left corner control point. In that case, the motion vector (mv_x, mv_y) at the (x, y) position may be as in FIG. 21. Thus, the motion vector for each pixel position or any position can be estimated according to the equation of FIG. 21, which is based on v0, v1, and v2.

[0131] In addition, in the formula of FIG. 21, (x, y) may be relative coordinates within the block. For example, (x, y) may be the position when the left top position of the block is (0, 0). Therefore, if v0 is the control point motion vector for position (x0, y0), v1 is the control point motion vector for position (x1, y1), and v2 is the control point motion vector for position (x2, y2), and (x, y) is to be represented using the same coordinates as the positions of v0, v1, and v2, then x and y in the formula of FIG. 21 can be changed to (x-x0) and (y-y0), respectively. Also, w (width of the block) may be (x1-x0) and h (height of the block) may be (y2 - y0).

[0132] FIG. 22 is a diagram showing affine motion prediction according to one embodiment of the present disclosure.

[0133] As explained earlier, a motion vector field exists and a motion vector can be calculated for each pixel, but to make it simpler, an affine transform can be performed based on subblocks as shown in FIG. 22. For example, the small square in FIG. 22 (a) is a subblock, and a representative motion vector can be created for the subblock, and the said representative motion vector can be used for the pixels of that subblock. Also, as complex movements were expressed in FIG. 17, FIG. 18, FIG. 20, etc., the subblock can express such movements and correspond to the reference block, or it can be made simpler by applying only translation motion to the subblock. In FIG. 22 (a), v0, v1, and v2 can be control point motion vectors.

[0134] At this time, the subblock size may be M*N, and M and N may be as shown in FIG. 22 (b). Also, MvPre may be motion vector fraction accuracy. In addition, (v_0x, v_0y), (v_1x, v_1y), and (v_2x, v_2y) may be motion vectors of the top-left, top-right, and bottom-left control points, respectively. (v_2x, v_2y) may be the motion vector of the bottom-left control point of the current block, and, for example, in the case of 4-parameters, it may be the MV (motion vector) for the bottom-left calculated by the equation in FIG. 19.

[0135] In addition, when generating a representative motion vector for a subblock, it is possible to calculate the representative motion vector using the center sample position of the subblock. Furthermore, when generating the motion vector for a subblock, a motion vector with higher accuracy than the normal motion vector can be used, and to achieve this, motion compensation interpolation filters can be applied.

[0136] In another embodiment, the size of the subblock may not be variable and may be fixed to a specific size. For example, the size of the subblock may be fixed to a 4x4 size.

[0137] FIG. 23 is a diagram showing a mode of affine motion prediction according to one embodiment of the present disclosure.

[0138] According to one embodiment of the present disclosure, an affine inter mode may be an example of affine motion prediction. There may be a flag indicating that it is an affine inter mode. Referring to FIG. 23, there may be blocks at positions A, B, C, D, and E near v0 and v1, and the motion vectors corresponding to each block may be vA, vB, vC, vD, and vE. Using this, a candidate list for the following motion vectors or motion vector predictors can be created.

[0139] {(v0, v1)|v0 = {vA, vB, vC}, v1 = {vD, vE}}

[0140] In other words, a (v0, v1) pair can be formed using v0 selected from vA, vB, and vC, and v1 selected from vD and vE. In this case, the motion vector can be scaled according to the Picture Order Count (POC) of the neighbor block's reference, the POC of the reference to the current CU (current coding unit; current block), and the POC of the current CU. Once a candidate list is created using motion vector pairs as described above, it is possible to signal which candidate from the list has been selected or whether it has been selected. Furthermore, if the candidate list is not sufficiently filled, it is possible to fill it with candidates from other inter-predictions. For example, it can be filled using AMVP (Advanced Motion Vector Prediction) candidates. In addition, instead of directly using v0 and v1 selected from the candidate list as control point motion vectors for affine motion prediction, it is possible to signal the difference to be corrected, thereby creating better control point motion vectors. That is, in the decoder, it is possible to use v0' and v1', created by adding the difference to v0 and v1 selected from the candidate list, as control point motion vectors for affine motion prediction.

[0141] In one embodiment, it is also possible to use affine inter mode for CUs (coding units) of a specific size or larger.

[0142] FIG. 24 is a diagram showing a mode of affine motion prediction according to one embodiment of the present disclosure.

[0143] According to one embodiment of the present disclosure, an affine merge mode may be an example of affine motion prediction. There may be a flag indicating that it is an affine merge mode. In an affine merge mode, if affine motion prediction was used in the vicinity of the current block, the control point motion vector of the current block can be calculated from the motion vectors of the surrounding blocks. For example, when checking whether the surrounding blocks used affine motion prediction, the surrounding blocks that serve as candidates may be as shown in FIG. 24(a). In addition, it is possible to check whether affine motion prediction was used in the order of A, B, C, D, and E, and if a block using affine motion prediction is found, the control point motion vector of the current block can be calculated using the motion vector of that block or the motion vector around that block. A, B, C, D, and E can be left, above, above right, left bottom, and above left, respectively, as shown in FIG. 24(a).

[0144] In one embodiment, as shown in FIG. 24(b), if a block at position A uses affine motion prediction, v0 and v1 can be calculated using the motion vectors of that block or around that block. The motion vectors of that block or around that block may be v2, v3, and v4.

[0145] In the previous embodiment, the order of the surrounding blocks to be referenced is fixed. However, a control point motion vector derived from a specific location is not always more effective. Therefore, in another embodiment, it is possible to signal which block to reference to derive the control point motion vector. For example, the candidate locations for the control point motion vector derivation are set in the order of A, B, C, D, and E in FIG. 24(a), and it is possible to signal which of them to reference.

[0146] In another embodiment, when deriving control point motion vectors, accuracy can be improved by taking from the nearest block for each control point motion vector. For example, referring to FIG. 24, the left block can be referenced when deriving v0, and the above block can be referenced when deriving v1. Alternatively, A, D, or E can be referenced when deriving v0, and B or C can be referenced when deriving v1.

[0147] FIG. 25 is a diagram showing an affine motion predictor derivation according to one embodiment of the present disclosure.

[0148] For affine motion prediction, control point motion vectors may be required, and based on these control point motion vectors, a motion vector field—that is, a motion vector for a subblock or a specific location—can be calculated. The control point motion vector may also be called a seed vector.

[0149] In this case, the control point MV (control point motion vector) can be based on the predictor. For example, the predictor can become the control point MV (control point motion vector). As another example, the control point MV (control point motion vector) can be calculated based on the predictor and the difference. Specifically, the control point MV (control point motion vector) can be calculated by adding or subtracting the difference from the predictor.

[0150] At this time, in the process of creating a predictor for a control point MV, it is possible to derive it from the control point MVs (control point motion vectors) or MVs (motion vectors) of a block that has undergone affine motion prediction (affine motion compensation (MC)) in the surrounding area. For example, if a block corresponding to a preset position has undergone affine motion prediction, a predictor for affine motion compensation of the current block can be derived from the control point MVs or MVs of that block. Referring to FIG. 25, the preset positions may be A0, A1, B0, B1, and B2. Alternatively, the preset positions may include positions adjacent to the current block and positions that are not adjacent. Furthermore, it is possible to reference the control point MVs (control point motion vectors) or MVs (motion vectors) of the preset positions (spatial), and it is possible to reference the temporal control point MVs or MVs of the preset positions.

[0151] Candidates for affine MC (affine motion compensation) can be created in the same way as in the embodiment of FIG. 25, and such candidates may be called inherited candidates. Alternatively, such candidates may be called merge candidates. Also, when referencing preset positions in the method of FIG. 25, references can be made following a preset order.

[0152] FIG. 26 is a diagram showing an affine motion predictor derivation according to one embodiment of the present disclosure.

[0153] Control point motion vectors may be required for affine motion prediction, and based on these control point motion vectors, a motion vector field—that is, a motion vector for a subblock or a specific location—can be calculated. Control point motion vectors may also be referred to as seed vectors.

[0154] In this case, the control point MV (control point motion vector) can be based on the predictor. For example, the predictor can become the control point MV (control point motion vector). As another example, the control point MV (control point motion vector) can be calculated based on the predictor and the difference. Specifically, the control point MV (control point motion vector) can be calculated by adding or subtracting the difference from the predictor.

[0155] In this process of creating the predictor for the control point MV, it is possible to derive it from surrounding MVs. In this case, surrounding MVs may include MVs (motion vectors) that are not affine motion-compensated MVs (affine motion-compensated MVs). For example, when deriving each control point MV (control point motion vector) of the current block, the MV at a pre-configured location for each control point MV can be used as the predictor for the control point MV. For example, the pre-configured location may be a part included in a block adjacent to that part.

[0156] Referring to FIG. 26, control point MV (control point motion vector) mv0, mv1, and mv2 can be determined. In this case, according to one embodiment of the present disclosure, the MV (motion vector) corresponding to preset positions A, B, and C can be used as a predictor for mv0. Additionally, the MV (motion vector) corresponding to preset positions D and E can be used as a predictor for mv1. The MV (motion vector) corresponding to preset positions F and G can be used as a predictor for mv2.

[0157] In addition, when determining each predictor of control point MV (control point motion vector) mv0, mv1, and mv2 according to the embodiment of FIG. 26, the order of reference to preset positions for each control point position may be determined. In addition, there may be multiple preset positions referenced as predictors of control point MV for each control point, and possible combinations of preset positions may be determined.

[0158] Candidates for affine MC (affine motion compensation) can be created in the same manner as in the embodiment of FIG. 26, and such candidates may be called constructed candidates. Alternatively, such candidates may be called inter candidates or virtual candidates. In addition, when referencing preset positions in the method of FIG. 41, references can be made following a preset order.

[0159] According to one embodiment of the present disclosure, a candidate list of affine MC (affine motion compensation) or a control point MV candidate list of affine MC (affine motion compensation) can be generated by the embodiments described in FIGS. 23 to 26 or a combination thereof.

[0160] FIG. 27 is a drawing showing an affine motion predictor derivation according to one embodiment of the present disclosure.

[0161] As described in FIGS. 24 and 25, a control point MV (control point motion vector) for affine motion prediction of the current block can be derived from a surrounding affine motion predicted block. In this case, a method such as that shown in FIG. 27 can be used. In the equation of FIG. 27, the top-left, top-right, and bottom-left MV (motion vector) or control point MV (control point motion vector) of the surrounding affine motion predicted block can be (v_E0x, v_E0y), (v_E1x, v_E1y), and (v_E2x, v_E2y), respectively. Additionally, the top-left, top-right, and bottom-left coordinates of the surrounding affine motion predicted block can be (x_E0, y_E0), (x_E1, y_E1), and (x_E2, y_E2), respectively. At this time, according to FIG. 27, the predictor of the control point MV (control point motion vector) of the current block or the control point MV (control point motion vector), (v_0x, v_0y) and (v_1x, v_1y) can be calculated.

[0162] FIG. 28 is a diagram showing an affine motion predictor derivation according to one embodiment of the present disclosure.

[0163] As explained earlier, multiple control motion MVs or multiple control point MV predictors may be required for affine motion compensation. In this case, other control motion MVs or control point MV predictors can be derived from any control motion MV or control point MV predictor.

[0164] For example, when two control point MVs (control point motion vectors) or two control point MV predictors (control point motion vector predictors) are created using the method described in the previous drawings, another control point MV (control point motion vector) or another control point MV predictor (control point motion vector predictor) can be created based on this.

[0165] Referring to FIG. 28, a method for generating mv0, mv1, and mv2, which are control point MV predictors (control point motion vector predictors) or control point MVs (control point motion vectors) for top-left, top-right, and bottom-left is shown. x and y in the figure represent the x-component and y-component, respectively, and the current block size may be w*h.

[0166] FIG. 29 is a diagram illustrating a method for generating a control point motion vector according to one embodiment of the present disclosure.

[0167] According to one embodiment of the present disclosure, in order to affine MC (affine motion compensation) the current block, it is possible to determine the control point MV (control point motion vector) by creating a predictor for the control point MV and adding a difference thereto. According to one embodiment, it is possible to create a predictor for the control point MV by the method described in FIGS. 23 to 26. The difference can be signaled from the encoder to the decoder.

[0168] Referring to Fig. 29, there may be differences for each control point MV (control point motion vector). Additionally, differences for each control point MV (control point motion vector) may be signaled. Fig. 29 (a) shows a method for determining mv0 and mv1, which are control point MVs of a 4-parameter model, and Fig. 29 (b) shows a method for determining mv0, mv1, and mv2, which are control point MVs (control point motion vectors) of a 6-parameter model. The control point MV (control point motion vector) is determined by adding mvd0, mvd1, and mvd2, which are differences for each control point MV (control point motion vector), to the predictor.

[0169] The top bar in Fig. 29 may be a predictor of the control point MV (control point motion vector).

[0170] FIG. 30 is a diagram showing a method for determining motion vector difference through the method described in FIG. 29.

[0171] In one embodiment, a motion vector difference can be signaled by the method described in FIG. 10. The motion vector difference determined by the signaled method may be lMvd of FIG. 30. Also, the signaled mvd (motion vector difference) shown in FIG. 29, i.e., values ​​such as mvd0, mvd1, and mvd2, may be lMvd of FIG. 30. As described in FIG. 29, the signaled mvd (motion vector difference) can be determined as the difference with the predictor of the control point MV (control point motion vector), and the determined difference may be MvdL0 and MvdL1 of FIG. 30. L0 may represent reference list 0 (the zeroth reference picture list), and L1 may represent reference list 1 (the first reference picture list). compIdx is a component index that can represent x, y components, etc.

[0172] FIG. 31 is a diagram illustrating a method for generating a control point motion vector according to one embodiment of the present disclosure.

[0173] According to one embodiment of the present disclosure, in order to affine MC (affine motion compensation) the current block, it is possible to determine the control point MV (control point motion vector) by creating a predictor for the control point MV (control point motion vector) and adding a difference thereto. According to one embodiment, it is possible to create a predictor for the control point MV (control point motion vector) by the method described in FIGS. 23 to 26. The difference can be signaled from the encoder to the decoder.

[0174] Referring to Fig. 31, there may be a predictor for the difference for each control point MV (control point motion vector). For example, the difference of another control point MV (control point motion vector) can be determined based on the difference of a certain control point MV (control point motion vector). This may be based on the similarity between the differences for the control point MV (control point motion vector). If the predictor is determined because of the similarity, the difference with the predictor may be reduced. In this case, the difference predictor for the control point MV (control point motion vector) is signaled, and the difference with the difference predictor for the control point MV (control point motion vector) can be signaled.

[0175] Figure 31 (a) shows a method for determining mv0 and mv1, which are control point MVs (control point motion vectors) of a 4-parameter model, and Figure 31 (b) shows a method for determining mv0, mv1, and mv2, which are control point MVs (control point motion vectors) of a 6-parameter model.

[0176] Referring to Fig. 31, the difference for each control point MV and the control point MV are determined based on the difference (mvd0) of mv0, which is control point MV 0. mvd0, mvd1, and mvd2 shown in Fig. 31 can be signaled from the encoder to the decoder. Compared to the method described in Fig. 29, the method in Fig. 31 may have different values ​​for mvd1 and mvd2 being signaled even if the same predictors are used for mv0, mv1, and mv2 as in Fig. 29. If the difference between the control point MV mv0, mv1, and mv2 and the predictors is similar, there is a possibility that the absolute values ​​of mvd1 and mvd2 will be smaller when using the method in Fig. 31 than when using the method in Fig. 29, and accordingly, the signaling overhead of mvd1 and mvd2 can be reduced. Referring to Figure 31, the difference with the predictor of mv1 can be determined as (mvd1+mvd0), and the difference with the predictor of mv2 as (mvd2+mvd0).

[0177] The top bar in Fig. 31 may be a predictor of the control point MV.

[0178] Figure 32 is a diagram showing a method for determining the motion vector difference through the method described in Figure 31.

[0179] In one embodiment, the motion vector difference may be signaled in the manner described in FIG. 10 or FIG. 33. The motion vector difference determined based on the signaled parameters may be the lMvd of FIG. 32. Additionally, values ​​such as the signaled mvd shown in FIG. 31, namely mvd0, mvd1, and mvd2, may be the lMvd of FIG. 32.

[0180] MvdLX in Fig. 32 may be the difference between each control point MV (control point motion vector) and the predictor. That is, it may be (mv - mvp). In this case, as explained in Fig. 31, for control point MV 0 (mv_0), the signaled motion vector difference can be used directly as the difference (MvdLX) for the control point MV, and for other control point MVs (mv_1, mv_2), MvdLX, which is the difference of the control point MV, can be determined and used based on the signaled motion vector difference (mvd1, mvd2 in Fig. 31) and the signaled motion vector difference for control point MV 0 (mv_0) (mvd0 in Fig. 31).

[0181] In FIG. 32, LX may represent reference list X (reference picture list X). compIdx may represent component index x, y component, etc. cpIdx may represent control point index. cpIdx may mean 0, 1 or 0, 1, 2 as shown in FIG. 31.

[0182] The resolution of the motion vector difference can be considered for the values ​​shown in Figures 10, 12, 13, etc. For example, when the resolution is R, the value of lMvd*R can be used for lMvd in the drawing.

[0183] FIG. 33 is a diagram showing motion vector difference syntax according to one embodiment of the present disclosure.

[0184] Referring to Fig. 33, the motion vector difference can be coded in a manner similar to that described in Fig. 10. In this case, the coding can be done separately depending on cpIdx and the control point index.

[0185] FIG. 34 is a diagram showing an upper-level signaling structure according to one embodiment of the present disclosure.

[0186] According to one embodiment of the present disclosure, one or more upper-level signalings may exist. Upper-level signalings may refer to signaling at an upper level. An upper level may be a unit that includes a certain unit. For example, an upper level of a current block or a current coding unit may include a CTU, slice, tile, tile group, picture, sequence, etc. An upper-level signaling may affect lower levels of that upper level. For example, if the upper level is a sequence, it may affect CTU, slice, tile, tile group, and picture units that are lower to the sequence. Here, "affecting" means that the upper-level signaling affects encoding or decoding for lower levels.

[0187] Additionally, the upper-level signaling may include a signaling indicating whether a mode is available. Referring to FIG. 34, the upper-level signaling may include sps_modeX_enabled_flag. According to one embodiment, it may be determined whether the mode modeX is available based on sps_modeX_enabled_flag. For example, when sps_modeX_enabled_flag is a certain value, the mode modeX may not be available. Also, when sps_modeX_enabled_flag is a different value, it may be possible to use the mode modeX. Additionally, when sps_modeX_enabled_flag is a different value, it may be determined whether the mode modeX is available based on additional signaling. For example, the said certain value may be 0 and the said other certain value may be 1. However, it is not limited thereto, and the said certain value may be 1 and the said other value may be 0.

[0188] According to one embodiment of the present disclosure, there may be a signaling indicating whether affine motion compensation can be used. For example, this signaling may be a high-level signaling. Referring to FIG. 34, this signaling may be sps_affine_enabled_flag. Referring to FIG. 2 and FIG. 7, the signaling may refer to a signal transmitted from an encoder to a decoder through a bitstream. The decoder may parse the affine-enabled flag from the bitstream.

[0189] For example, if sps_affine_enabled_flag is 0, the syntax may be restricted so that affine motion compensation is not used. Also, if sps_affine_enabled_flag is 0, inter_affine_flag and cu_affine_type_flag may not exist.

[0190] For example, the inter_affine_flag may be a signaling indicating whether affine MC (affine motion compensation) is used in the block. Referring to FIGS. 2 and 7, the signaling may refer to a signal transmitted from the encoder to the decoder through the bitstream. The decoder may parse the inter_affine_flag from the bitstream.

[0191] Additionally, cu_affine_type_flag (coding unit affine type flag) can be a signal indicating which type of affine MC (affine motion compensation) is used in the block. Also, here, type can indicate whether it is a 4-parameter affine model or a 6-parameter affine model. Additionally, if sps_affine_enabled_flag (affine enable flag) is 1, affine model compensation (affine motion compensation) can be used.

[0192] Affine motion compensation may mean affine model based motion compensation or affine model based motion compensation for inter prediction.

[0193] Additionally, according to one embodiment of the present disclosure, there may be a signaling indicating whether a specific type of a certain mode can be used. For example, there may be a signaling indicating whether a specific type of affine motion compensation can be used. For example, this signaling may be a high-level signaling. Referring to FIG. 34, this signaling may be sps_affine_type_flag. Also, the specific type may mean a 6-parameter affine model. For example, if sps_affine_type_flag is 0, the syntax may be restricted so that a 6-parameter affine model is not used. Also, if sps_affine_type_flag is 0, cu_affine_type_flag may not exist. Also, if sps_affine_type_flag is 1, a 6-parameter affine model may be used. If sps_affine_type_flag does not exist, its value may be inferred to 0.

[0194] Additionally, according to one embodiment of the present disclosure, a signaling indicating whether a specific type of a certain mode is available may exist when there is a signaling indicating whether the certain mode is available. For example, if the signaling value indicating whether the certain mode is available is 1, it is possible to parse the signaling indicating whether a specific type of the certain mode is available. Additionally, if the signaling value indicating whether the certain mode is available is 0, it is possible not to parse the signaling indicating whether a specific type of the certain mode is available. For example, the signaling indicating whether the certain mode is available may include sps_affine_enabled_flag. Additionally, the signaling indicating whether a specific type of the certain mode is available may include sps_affine_type_flag. Referring to Fig. 34, when sps_affine_enabled_flag is 1, it is possible to parse sps_affine_type_flag. Also, when sps_affine_enabled_flag is 0, it is possible not to parse sps_affine_type_flag and to infer its value to 0.

[0195] In addition, as explained earlier, it is possible to use adaptive motion vector resolution (AMVR). It is possible to use different AMVR resolution sets depending on the situation. For example, it is possible to use different AMVR resolution sets depending on the prediction mode. For instance, the AMVR resolution set may differ when using regular inter prediction such as AMVP compared to when using affine MC. Furthermore, the AMVR applied to regular inter prediction such as AMVP may be applied to the motion vector difference. Alternatively, the AMVR applied to regular inter prediction such as AMVP (Advanced Motion Vector Prediction) may be applied to the motion vector predictor. Additionally, the AMVR applied to affine MC may be applied to the control point motion vector or the control point motion vector difference.

[0196] Additionally, according to one embodiment of the present disclosure, there may be a signaling indicating whether AMVR is available. This signaling may be a high-level signaling. Referring to FIG. 34, sps_amvr_enabled_flag (AMVR enabled flag) may exist. Referring to FIG. 2 and FIG. 7, the signaling may refer to a signal transmitted from an encoder to a decoder through a bitstream. The decoder may parse the AMVR enabled flag from the bitstream.

[0197] According to one embodiment of the present disclosure, sps_amvr_enabled_flag (AMVR enable flag) may indicate whether adaptive motion vector difference resolution is used. Additionally, according to one embodiment of the present disclosure, sps_amvr_enabled_flag (AMVR enable flag) may indicate whether adaptive motion vector difference resolution can be used. For example, according to one embodiment of the present disclosure, if sps_amvr_enabled_flag (AMVR enable flag) is 1, AMVR may be used for motion vector coding. Additionally, according to one embodiment of the present disclosure, if sps_amvr_enabled_flag (AMVR enable flag) is 1, AMVR may be used for motion vector coding. Additionally, if sps_amvr_enabled_flag (AMVR enable flag) is 1, there may be additional signaling to indicate which resolution is used. Additionally, if sps_amvr_enabled_flag (AMVR enable flag) is 0, AMVR may not be used for motion vector coding. Also, if sps_amvr_enabled_flag (AMVR enable flag) is 0, AMVR may not be available for motion vector coding. According to one embodiment of the present disclosure, the AMVR corresponding to sps_amvr_enabled_flag (AMVR enable flag) may mean that it is used for regular inter prediction. For example, the AMVR corresponding to sps_amvr_enabled_flag (AMVR enable flag) may not mean that it is used for affine MC (affine motion compensation). Additionally, whether affine MC (affine motion compensation) is used may be indicated by inter_affine_flag (inter affine flag).That is, the AMVR corresponding to sps_amvr_enabled_flag (AMVR enable flag) means that it is used when inter_affine_flag (inter affine flag) is 0, or it may not mean that it is used when inter_affine_flag (inter affine flag) is 1.

[0198] Additionally, according to one embodiment of the present disclosure, there may be a signaling indicating whether AMVR can be used in affine MC (affine motion compensation). This signaling may be a high-level signaling. Referring to FIG. 34, there may be sps_affine_amvr_enabled_flag (affine AMVR enable flag), which is a signaling indicating whether AMVR can be used in affine MC (affine motion compensation). Referring to FIG. 2 and FIG. 7, the signaling may refer to a signal transmitted from an encoder to a decoder through a bitstream. The decoder may parse the affine AMVR enable flag from the bitstream.

[0199] sps_affine_amvr_enabled_flag (Affine AMVR Enabled Flag) may indicate whether adaptive motion vector difference resolution is used for affine motion compensation. Additionally, sps_affine_amvr_enabled_flag (Affine AMVR Enabled Flag) may indicate whether adaptive motion vector difference resolution can be used for affine motion compensation. According to one embodiment, if sps_affine_amvr_enabled_flag (Affine AMVR Enabled Flag) is 1, AMVR may be used for affine inter-mode motion vector coding. Additionally, if sps_affine_amvr_enabled_flag (Affine AMVR Enabled Flag) is 1, AMVR may be used for affine inter-mode motion vector coding. Also, if sps_affine_amvr_enabled_flag is 0, AMVR may not be available for affine inter-mode motion vector coding. If sps_affine_amvr_enabled_flag is 0, AMVR may not be used for affine inter-mode motion vector coding.

[0200] For example, if sps_affine_amvr_enabled_flag is 1, the AMVR corresponding to the case where inter_affine_flag is 1 can be used. Additionally, if sps_affine_amvr_enabled_flag is 1, there may be additional signaling to indicate which resolution is used. Also, if sps_affine_amvr_enabled_flag is 0, the AMVR corresponding to the case where inter_affine_flag is 1 may not be used.

[0201] FIG. 35 is a diagram showing a coding unit syntax structure according to one embodiment of the present disclosure.

[0202] As described in FIG. 34, there may be additional signaling to indicate resolution based on upper-level signaling using AMVR. Referring to FIG. 34, the additional signaling to indicate resolution may include amvr_flag or amvr_precision_flag. amvr_flag or amvr_precision_flag may be information about the resolution of the motion vector difference.

[0203] According to one embodiment, there may be a signaling that indicates when amvr_flag is 0. Also, when amvr_flag is 1, amvr_precision_flag may exist. Also, when amvr_flag is 1, the resolution may be determined based on amvr_precision_flag. For example, when amvr_flag is 0, the resolution may be 1 / 4. Also, if amvr_flag does not exist, the amvr_flag value may be inferred based on CuPredMode. For example, if CuPredMode is MODE_IBC, the amvr_flag value may be inferred as 1, and if CuPredMode is not MODE_IBC or CuPredMode (Coding unit prediction mode) is MODE_INTER, the amvr_flag value may be inferred as 0.

[0204] Additionally, 1-pel resolution can be used if inter_affine_flag is 0 and amvr_precision_flag is 0. Additionally, 1 / 16-pel resolution can be used if inter_affine_flag is 1 and amvr_precision_flag is 0. Additionally, 4-pel resolution can be used if inter_affine_flag is 0 and amvr_precision_flag is 1. Additionally, 1-pel resolution can be used if inter_affine_flag is 1 and amvr_precision_flag is 1.

[0205] If amvr_precision_flag is 0, the value can be inferred as 0.

[0206] According to one embodiment, the resolution may be applied by the MvShift value. Additionally, MvShift may be determined by amvr_flag and amvr_precision_flag, which are information about the resolution of the motion vector difference. For example, when inter_affine_flag is 0, the MvShift value may be determined as follows.

[0207] MvShift = (amvr_flag + amvr_precision_flag) << 1

[0208] Additionally, the Mvd (motion vector difference) value can be shifted based on the MvShift value. For example, the Mvd is shifted as shown below, and the AMVR resolution may be applied accordingly.

[0209] MvdLX = MvdLX << (MvShift + 2)

[0210] As another example, when inter_affine_flag is 1, the MvShift value can be determined as follows.

[0211] MvShift = amvr_precision_flag ? (amvr_precision_flag << 1) : ( -(amvr_flag << 1) )

[0212] Additionally, the MvdCp (control point motion vector difference) value can be shifted based on the MvShift value. MvdCp can be the control point motion vector difference or the control point motion vector. For example, MvdCp (control point motion vector difference) is shifted as shown below, and the AMVR resolution may be applied accordingly.

[0213] MvdCpLX = MvdCpLX << (MvShift + 2)

[0214] Also, Mvd or MvdCp can be a value signaled by mvd_coding.

[0215] Referring to Fig. 35, when CuPredMode is MODE_IBC, amvr_flag, which is information about the resolution of the motion vector difference, may not exist. Also, when CuPredMode is MODE_INTER, amvr_flag, which is information about the resolution of the motion vector difference, may exist, and in this case, if certain conditions are satisfied, it may be possible to parse amvr_flag.

[0216] According to one embodiment of the present disclosure, it is possible to determine whether to parse AMVR-related syntax elements based on a high-level signaling value indicating whether AMVR can be used. For example, if the high-level signaling value indicating whether AMVR can be used is 1, it is possible to parse AMVR-related syntax elements. Additionally, if the high-level signaling value indicating whether AMVR can be used is 0, AMVR-related syntax elements may not be parsed. Referring to FIG. 35, if CuPredMode is MODE_IBC and sps_amvr_enabled_flag (AMVR enable flag) is 0, amvr_precision_flag, which is information about the resolution of the motion vector difference, may not be parsed. Additionally, if CuPredMode is MODE_IBC and sps_amvr_enabled_flag (AMVR enable flag) is 1, amvr_precision_flag, which is information about the resolution of the motion vector difference, may be parsed. At this time, additional conditions for parsing may be considered.

[0217] For example, it is possible to parse amvr_precision_flag if there is at least one non-zero value among MvdLX (multiple motion vector differences). MvdLX may be an Mvd value for reference list LX. Also, Mvd (motion vector difference) may be signaled through mvd_coding. LX may include L0 (the zeroth reference picture list) and L1 (the first reference picture list). Additionally, MvdLX may have components corresponding to the x-axis and y-axis, respectively. For example, the x-axis may correspond to the horizontal axis of the picture, and the y-axis may correspond to the vertical axis of the picture. Referring to FIG. 35, [0] and [1] in MvdLX[x0][y0][0] and MvdLX[x0][y0][1] may indicate that they correspond to the x-axis and y-axis components, respectively. In addition, when CuPredMode is MODE_IBC, it is possible to use only L0. Referring to FIG. 35, amvr_precision_flag can be parsed if 1) sps_amvr_enabled_flag is 1 and 2) MvdL0[x0][y0][0] or MvdL0[x0][y0][1] is not 0. In addition, amvr_precision_flag can be not parsed if 1) sps_amvr_enabled_flag (AMVR enable flag) is 0 or 2) both MvdL0[x0][y0][0] and MvdL0[x0][y0][1] are 0.

[0218] Additionally, there may be cases where CuPredMode is not MODE_IBC. In this case, referring to FIG. 35, if sps_amvr_enabled_flag (AMVR enable flag) is 1, inter_affine_flag (inter affine flag) is 0, and there is at least one non-zero value among MvdLX (multiple motion vector difference flags), amvr_flag can be parsed. Here, amvr_flag may be information about the resolution of the motion vector difference. Also, as previously explained, sps_amvr_enabled_flag (AMVR enable flag) being 1 may indicate the use of adaptive motion vector difference resolution. Additionally, inter_affine_flag being 0 may indicate that affine motion compensation is not used for the current block. In addition, as described in FIG. 34, the plurality of motion vector differences for the current block can be modified based on information regarding the resolution of the motion vector differences, such as amvr_flag. This condition can be referred to as conditionA.

[0219] Additionally, amvr_flag can be parsed if sps_affine_amvr_enabled_flag (Affine AMVR Enabled Flag) is 1, inter_affine_flag (Inter Affine Flag) is 1, and there is at least one non-zero value among MvdCpLX (Multiple Control Point Motion Vector Differences). Here, amvr_flag may be information about the resolution of the motion vector difference. Also, as previously explained, sps_affine_amvr_enabled_flag (Affine AMVR Enabled Flag) being 1 may indicate that an adaptive motion vector difference resolution can be used for affine motion compensation. Also, inter_affine_flag being 1 may indicate the use of affine motion compensation for the current block. In addition, as described in FIG. 34, the motion vector difference of the plurality of control points for the current block can be modified based on information regarding the resolution of the motion vector difference, such as amvr_flag. This condition can be referred to as conditionB.

[0220] In addition, amvr_flag may be parsed if condition A or condition B is satisfied. In addition, amvr_flag, which is information about the resolution of motion vector differences, may not be parsed if neither condition A nor condition B is satisfied. That is, amvr_flag may not be parsed if 1) sps_amvr_enabled_flag (AMVR enable flag) is 0, inter_affine_flag (inter affine flag) is 1, or MvdLX (multiple motion vector differences) are all 0, and 2) sps_affine_amvr_enabled_flag is 0, inter_affine_flag (inter affine flag) is 0, or MvdCpLX (multiple control point motion vector differences) are all 0.

[0221] In addition, you can determine whether to parse amvr_precision_flag based on the amvr_flag value. For example, if the amvr_flag value is 1, amvr_precision_flag can be parsed. Also, if the amvr_flag value is 0, amvr_precision_flag can not be parsed.

[0222] Additionally, MvdCpLX(multiple control point motion vector differences) may represent differences for control point motion vectors. Furthermore, MvdCpLX(multiple control point motion vector differences) may be signaled via mvd_coding. LX may include L0 (the 0th reference picture list) and L1 (the 1st reference picture list). Additionally, MvdCpLX may contain components corresponding to control point motion vectors 0, 1, 2, etc. For example, control point motion vectors 0, 1, 2, etc. may be control point motion vectors corresponding to pre-set positions relative to the current block.

[0223] Referring to FIG. 35, it can be indicated that [0], [1], and [2] in MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][1][], and MvdCpLX[x0][y0][2][] correspond to control point motion vectors 0, 1, and 2, respectively. Additionally, the value of MvdCpLX (multiple control point motion vector differences) corresponding to control point motion vector 0 can be used for other control point motion vectors. For example, control point motion vector 0 can be used as in FIG. 47. Furthermore, MvdCpLX (multiple control point motion vector differences) may have components corresponding to the x-axis and y-axis, respectively. For example, the x-axis may correspond to the horizontal axis of the picture, and the y-axis may correspond to the vertical axis of the picture. Referring to Fig. 35, it can be indicated that [0] and [1] in MvdCpLX[x0][y0][][0] and MvdCpLX[x0][y0][][1] correspond to the x-axis and y-axis components, respectively.

[0224] FIG. 36 is a diagram showing an upper-level signaling structure according to one embodiment of the present disclosure.

[0225] Higher-level signalings such as those described in FIGS. 34 and 35 may exist. For example, sps_affine_enabled_flag (affine enable flag), sps_affine_amvr_enabled_flag, sps_amvr_enabled_flag (AMVR enable flag), sps_affine_type_flag, etc. may exist.

[0226] According to one embodiment of the present disclosure, the upper-level signals may have a parsing dependency. For example, it may be determined whether to parse another upper-level signal based on a certain upper-level signal value.

[0227] According to one embodiment of the present disclosure, it is possible to determine whether affine AMVR can be used based on whether affine MC (affine motion compensation) can be used. For example, it is possible to determine whether affine AMVR can be used based on upper-level signaling indicating whether affine MC can be used. More specifically, it is possible to determine whether to parse upper-level signaling indicating whether affine AMVR can be used based on upper-level signaling indicating whether affine MC can be used.

[0228] In one embodiment, it is possible to use an affine AMVR when an affine MC can be used. Also, if an affine MC cannot be used, an affine AMVR may not be used.

[0229] More specifically, if the upper-level signaling indicating whether affine MC is available is 1, affine AMVR may be available. In this case, additional signaling may exist. Also, if the upper-level signaling indicating whether affine MC is available is 0, affine AMVR may not be available. For example, if the upper-level signaling indicating whether affine MC is available is 1, the upper-level signaling indicating whether affine AMVR is available can be parsed. Also, if the upper-level signaling indicating whether affine MC is available is 0, the upper-level signaling indicating whether affine AMVR is available may not be parsed. Furthermore, if the upper-level signaling indicating whether affine AMVR is available does not exist, its value can be inferred. For example, it can be inferred as 0. As another example, inferring can be based on the upper-level signaling indicating whether affine MC is available. As another example, it can be inferred based on high-level signaling indicating whether AMVR is available.

[0230] According to one embodiment, the affine AMVR may be an AMVR used for the affine MC described in FIGS. 34 and 35. For example, a high-level signaling indicating whether the affine MC is available may be sps_affine_enabled_flag (affine-enabled flag). Additionally, a high-level signaling indicating whether the affine AMVR is available may be sps_affine_amvr_enabled_flag.

[0231] Referring to Fig. 36, sps_affine_amvr_enabled_flag can be parsed when sps_affine_enabled_flag is 1. Also, sps_affine_amvr_enabled_flag can be not parsed when sps_affine_enabled_flag is 0. Additionally, if sps_affine_amvr_enabled_flag does not exist, its value can be inferred as 0.

[0232] The embodiments of the present disclosure may be meaningful because affine AMVR can be used when affine MC is used.

[0233] FIG. 37 is a diagram showing an upper-level signaling structure according to one embodiment of the present disclosure.

[0234] Higher-level signalings such as those described in FIGS. 34 and 35 may exist. For example, sps_affine_enabled_flag (affine enable flag), sps_affine_amvr_enabled_flag (affine AMVR enable flag), sps_amvr_enabled_flag (AMVR enable flag), sps_affine_type_flag, etc. may exist.

[0235] According to one embodiment of the present disclosure, the upper-level signals may have a parsing dependency. For example, it may be determined whether to parse another upper-level signal based on a certain upper-level signal value.

[0236] According to one embodiment of the present disclosure, it is possible to determine whether an affine AMVR can be used based on whether an AMVR can be used. For example, it is possible to determine whether an affine AMVR can be used based on upper-level signaling indicating whether an AMVR can be used. More specifically, it is possible to determine whether to parse upper-level signaling indicating whether an affine AMVR can be used based on upper-level signaling indicating whether an AMVR can be used.

[0237] In one embodiment, it is possible to use affine AMVR when AMVR can be used. Also, if AMVR cannot be used, affine AMVR may not be used.

[0238] More specifically, if the upper-level signaling indicating whether AMVR is available is 1, affine AMVR may be available. In this case, additional signaling may exist. Also, if the upper-level signaling indicating whether AMVR is available is 0, affine AMVR may not be available. For example, if the upper-level signaling indicating whether AMVR is available is 1, the upper-level signaling indicating whether affine AMVR is available can be parsed. Also, if the upper-level signaling indicating whether AMVR is available is 0, the upper-level signaling indicating whether affine AMVR is available may not be parsed. Furthermore, if the upper-level signaling indicating whether affine AMVR is available does not exist, its value can be inferred. For example, it can be inferred as 0. As another example, it can be inferred based on the upper-level signaling indicating whether affine MC is available. As another example, it can be inferred based on high-level signaling indicating whether AMVR is available.

[0239] According to one embodiment, the affine AMVR may be an AMVR used for the affine MC (affine motion compensation) described in FIGS. 34 and 35. For example, a high-level signaling indicating whether the AMVR is available may be sps_amvr_enabled_flag (AMVR enable flag). Additionally, a high-level signaling indicating whether the affine AMVR is available may be sps_affine_amvr_enabled_flag (affine AMVR enable flag).

[0240] Referring to FIG. 37 (a), if sps_amvr_enabled_flag (AMVR enable flag) is 1, sps_affine_amvr_enabled_flag (affine AMVR enable flag) can be parsed. Also, if sps_amvr_enabled_flag (AMVR enable flag) is 0, sps_affine_amvr_enabled_flag (affine AMVR enable flag) can be not parsed. Also, if sps_affine_amvr_enabled_flag (affine AMVR enable flag) does not exist, its value can be inferred as 0.

[0241] The embodiments of the present disclosure may be such that the efficiency of adaptive resolution may vary depending on the sequence.

[0242] In addition, whether affine AMVR is available can be determined by considering both whether affine MC is available and whether AMVR is available. For example, it is possible to determine whether to parse the upper-level signaling indicating whether affine AMVR is available based on the upper-level signaling indicating whether affine MC is available and the upper-level signaling indicating whether AMVR is available. According to one embodiment, if both the upper-level signaling indicating whether affine MC is available and the upper-level signaling indicating whether AMVR is available are 1, the upper-level signaling indicating whether affine AMVR is available can be parsed. In addition, if the upper-level signaling indicating whether affine MC is available or the upper-level signaling indicating whether AMVR is available is 0, the upper-level signaling indicating whether affine AMVR is available may not be parsed. In addition, if there is no upper-level signaling indicating whether affine AMVR is available, that value can be inferred.

[0243] Referring to FIG. 37 (b), sps_affine_enabled_flag and sps_amvr_enabled_flag can be parsed when both are 1. Additionally, sps_affine_amvr_enabled_flag can be parsed when at least one of sps_affine_enabled_flag and sps_amvr_enabled_flag is 0. Additionally, sps_affine_amvr_enabled_flag can be not parsed when sps_affine_amvr_enabled_flag does not exist. Furthermore, if sps_affine_amvr_enabled_flag does not exist, its value can be inferred as 0.

[0244] More specifically, to explain FIG. 37(b), whether affine motion compensation can be used can be determined based on sps_affine_enabled_flag (affine enable flag) at line (3701). As already explained in FIG. 34 and 35, if sps_affine_enabled_flag (affine enable flag) is 1, it may mean that affine motion compensation can be used. Also, if sps_affine_enabled_flag (affine enable flag) is 0, it may mean that affine motion compensation cannot be used.

[0245] If it is determined that affine motion compensation is used at line (3701) of FIG. 37(b), it may be determined whether adaptive motion vector difference resolution is used based on sps_amvr_enabled_flag (AMVR enable flag) at line (3702). As already explained in FIG. 34 and 35, if sps_amvr_enabled_flag (AMVR enable flag) is 1, it may mean that adaptive motion vector difference resolution is used. If sps_amvr_enabled_flag (AMVR enable flag) is 0, it may mean that adaptive motion vector difference resolution is not used.

[0246] If it is determined that affine motion compensation is not used in line (3701), it may not be determined whether adaptive motion vector difference resolution is used based on sps_amvr_enabled_flag (AMVR-enabled flag). That is, line (3702) of FIG. 37(b) may not be performed. Specifically, if affine motion compensation is not used, sps_affine_amvr_enabled_flag (affine AMVR-enabled flag) may not be transmitted from the encoder to the decoder. That is, the decoder may not receive sps_affine_amvr_enabled_flag (affine AMVR-enabled flag), and sps_affine_amvr_enabled_flag (affine AMVR-enabled flag) may not be parsed by the decoder. In this case, sps_affine_amvr_enabled_flag (affine AMVR enable flag) does not exist and can be inferred as 0. As previously explained, if sps_affine_amvr_enabled_flag (affine AMVR enable flag) is 0, it may indicate that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0247] If it is determined that adaptive motion vector difference resolution is used at line (3702) of FIG. 37(b), sps_affine_amvr_enabled_flag (affine AMVR enabled flag) indicating whether adaptive motion vector difference resolution can be used for affine motion compensation from the bitstream at line (3703) can be parsed.

[0248] If it is determined that adaptive motion vector difference resolution is not used in line (3702), sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be parsed from the bitstream. That is, line (3703) of FIG. 37(b) may not be performed. More specifically, if affine motion compensation is used and adaptive motion vector difference resolution is not used, sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be transmitted from the encoder to the decoder. That is, the decoder may not receive sps_affine_amvr_enabled_flag (affine AMVR enable flag). The decoder may not parse sps_affine_amvr_enabled_flag (affine AMVR enable flag) from the bitstream. In this case, sps_affine_amvr_enabled_flag (affine AMVR enable flag) does not exist and can be inferred as 0. As previously explained, if sps_affine_amvr_enabled_flag (affine AMVR enable flag) is 0, it may indicate that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0249] As shown in Fig. 37(b), by checking sps_affine_enabled_flag first and then sps_amvr_enabled_flag next, unnecessary processing steps can be reduced, thereby increasing efficiency. For example, if sps_amvr_enabled_flag is checked first and sps_affine_enabled_flag is checked next, sps_affine_enabled_flag may need to be checked again to derive sps_affine_type_flag on line 7. However, by checking sps_affine_enabled_flag first and then sps_amvr_enabled_flag next, this unnecessary process can be reduced.

[0250] FIG. 38 is a diagram showing a coding unit syntax structure according to one embodiment of the present disclosure.

[0251] As described in FIG. 35, whether to parse AMVR-related syntax can be determined based on whether there is at least one non-zero value among MvdLX or MvdCpLX. However, the MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) used may differ depending on which reference list (reference picture list) is used, how many parameters the affine model uses, etc. If the initial value of MvdLX or MvdCpLX is not zero, unnecessary AMVR-related syntax elements are signaled because the MvdLX or MvdCpLX not used in the current block is non-zero, which can cause a mismatch between the encoder and the decoder. FIG. 38 and 39 illustrate a method to prevent a mismatch between the encoder and the decoder.

[0252] According to one embodiment, inter_pred_idc (information about the reference picture list) may indicate which reference list is used or what the prediction direction is. For example, inter_pred_idc (information about the reference picture list) may be a value of PRED_L0, PRED_L1, or PRED_BI. If inter_pred_idc (information about the reference picture list) is PRED_L0, only reference list 0 (the 0th reference picture list) may be used. Also, if inter_pred_idc (information about the reference picture list) is PRED_L1, only reference list 1 (the 1st reference picture list) may be used. Also, if inter_pred_idc (information about the reference picture list) is PRED_BI, both reference list 0 (the 0th reference picture list) and reference list 1 (the 1st reference picture list) may be used. If inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_L1, it may be uni-prediction. Also, if inter_pred_idc (information about the reference picture list) is PRED_BI, it may be bi-prediction.

[0253] In addition, it is possible to determine which affine model is used based on the MotionModelIdc value. In addition, it is possible to determine whether an affine MC is used based on the MotionModelIdc value. For example, MotionModelIdc may represent translational motion, 4-parameter affine motion, or 6-parameter affine motion. For example, when the MotionModelIdc value is 0, 1, or 2, it may indicate translational motion, 4-parameter affine motion, and 6-parameter affine motion, respectively. In addition, according to one embodiment, MotionModelIdc may be determined based on inter_affine_flag and cu_affine_type_flag. For example, when merge_flag is 0 (when not in merge mode), MotionModelIdc may be determined based on inter_affine_flag and cu_affine_type_flag. For example, MotionModelIdx can be (inter_affine_flag + cu_affine_type_flag). According to another embodiment, MotionModelIdc can be determined by merge_subblock_flag. For example, when merge_flag is 1 (in merge mode), MotionModelIdc can be determined by merge_subblock_flag. For example, the MotionModelIdc value can be set to the merge_subblock_flag value.

[0254] For example, if inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_BI, it is possible to use the value corresponding to L0 in MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference). Therefore, when parsing AMVR-related syntax, MvdL0 or MvdCpL0 can be considered only when inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_BI. That is, if inter_pred_idc (information about the reference picture list) is PRED_L1, it is possible not to consider MvdL0 (motion vector difference for the zeroth reference picture list) or MvdCpL0 (control point motion vector difference for the zeroth reference picture list).

[0255] Additionally, if inter_pred_idc (information about the reference picture list) is PRED_L1 or PRED_BI, it is possible to use the value corresponding to L1 in MvdLX or MvdCpLX. Therefore, when parsing AMVR-related syntax, MvdL1 (motion vector difference for the first reference picture list) or MvdCpL1 (control point motion vector difference for the first reference picture list) can be considered only when inter_pred_idc (information about the reference picture list) is PRED_L1 or PRED_BI. That is, if inter_pred_idc (information about the reference picture list) is PRED_L0, it is possible not to consider MvdL1 or MvdCpL1.

[0256] Referring to Fig. 38, MvdL0 and MvdCpL0 can determine whether to parse AMVR-related syntax only when inter_pred_idc is not PRED_L1 and its value is not zero. That is, when inter_pred_idc is PRED_L1, AMVR-related syntax may not be parsed even if there is a non-zero value in MvdL0 or a non-zero value in MvdCpL0.

[0257] In addition, when MotionModelIdc is 1, it is possible to consider only MvdCpLX[x0][y0][0][] and MvdCpLX[x0][y0][1][] among MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][2][]. That is, when MotionModelIdc is 1, it is possible not to consider MvdCpLX[x0][y0][2][]. For example, when MotionModelIdc is 1, whether or not there is a non-zero value in MvdCpLX[x0][y0][2][] may not affect AMVR-related syntax parsing.

[0258] Also, when MotionModelIdc is 2, it is possible to consider all of MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][1][], and MvdCpLX[x0][y0][2][]. That is, when MotionModelIdc is 2, it is possible to consider MvdCpLX[x0][y0][2][].

[0259] In addition, MotionModelIdc 1 and 2 in the above embodiment may also be represented as cu_affine_type_flag being 0 and 1. This may be because it is possible to determine whether or not to use affine MC. For example, whether or not to use affine MC can be determined through inter_affine_flag (inter affine flag).

[0260] Referring to FIG. 38, whether there is a non-zero value among MvdCpLX[x0][y0][2][] can be considered only when MotionModelIdc is 2. If MvdCpLX[x0][y0][0][] and MvdCpLX[x0][y0][1][] are both 0, and there is at least one non-zero value among MvdCpLX[x0][y0][2][] (according to the preceding embodiment, L0 and L1 may be separated here to consider only the value corresponding to either L0 or L1), then if MotionModelIdc is not 2, AMVR-related syntax may not be parsed.

[0261] FIG. 39 is a diagram showing the MVD default value setting according to one embodiment of the present disclosure.

[0262] As previously explained, MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) can be signaled via mvd_coding. Additionally, the lMvd value can be signaled via mvd_coding, and MvdLX or MvdCpLX can be set to the lMvd value. Referring to Fig. 39, MvdLX can be set via the lMvd value when MotionModelIdc is 0. Also, MvdCpLX can be set via the lMvd value when MotionModelIdc is not 0. Furthermore, depending on the refList value, it can be determined which of the LX actions to perform.

[0263] In addition, there may be an mvd_coding as described in FIG. 10 or FIG. 33, FIG. 34, FIG. 35, etc. Also, mvd_coding may include a step of parsing or determining abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, mvd_sign_flag, etc. Also, lMvd can be determined through abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, mvd_sign_flag, etc.

[0264] Referring to Fig. 39, lMvd can be set as follows.

[0265] lMvd = abs_mvd_greater0_flag * ( abs_mvd_minus2 + 2 ) * ( 1 - 2*mvd_sign_flag)

[0266] It is possible to set the default value of MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) to a preset value. According to one embodiment of the present disclosure, it is possible to set the default value of MvdLX or MvdCpLX to 0. Alternatively, it is possible to set the default value of lMvd to a preset value. Alternatively, it is possible to set the default value of a related syntax element so that the value of lMvd, MvdLX, or MvdCpLX becomes a preset value. The default value of a syntax element may mean a value inferred when the syntax element does not exist. The preset value may be 0.

[0267] According to one embodiment of the present disclosure, abs_mvd_greater0_flag may indicate whether the absolute value of the MVD is greater than 0. Also, according to one embodiment of the present disclosure, if abs_mvd_greater0_flag does not exist, its value may be inferred to 0. In this case, the lMvd value may be set to 0. Also, in this case, the MvdLX or MvdCpLX value may be set to 0.

[0268] Alternatively, according to one embodiment of the present disclosure, if the lMvd or MvdLX or MvdCpLX value is not set, the value can be set to a preset value. For example, it can be set to 0.

[0269] In addition, according to one embodiment of the present disclosure, if abs_mvd_greater0_flag does not exist, the corresponding lMvd or MvdLX or MvdCpLX value can be set to 0.

[0270] Additionally, abs_mvd_greater1_flag can indicate whether the absolute value of the MVD is greater than 1. Also, if abs_mvd_greater1_flag does not exist, its value can be inferred as 0.

[0271] Also, abs_mvd_minus2 + 2 can indicate the absolute value of the MVD. Additionally, if abs_mvd_minus2 does not exist, it can be inferred as -1.

[0272] Additionally, mvd_sign_flag can indicate the sign of the MVD. If mvd_sign_flag is 0 or 1, it indicates that the corresponding MVD has a positive or negative value, respectively. If mvd_sign_flag does not exist, its value can be inferred as 0.

[0273] FIG. 40 is a diagram showing MVD default value settings according to one embodiment of the present disclosure.

[0274] In order to prevent a mismatch between the encoder and the decoder by setting the initial value of MvdLX or MvdCpLX to 0, the values ​​of MvdLX (multiple motion vector differences) or MvdCpLX (multiple control point motion vector differences) may be initialized as described in FIG. 39. Additionally, the initialized value may be 0. The embodiment of FIG. 40 explains this in more detail.

[0275] According to one embodiment of the present disclosure, the MvdLX (multiple motion vector differences) or MvdCpLX (multiple control point motion vector differences) values ​​may be initialized to preset values. Additionally, the initialization location may be prior to the location where AMVR-related syntax elements are parsed. The AMVR-related syntax elements may include amvr_flag, amvr_precision_flag, etc. of FIG. 40. amvr_flag or amvr_precision_flag may be information regarding the resolution of the motion vector differences.

[0276] In addition, the resolution of the MVD (motion vector difference) or MV (motion vector), or the signaling resolution of the MVD or MV, may be determined by the AMVR-related syntax elements mentioned above. Also, the preset value initialized in the embodiments of the present disclosure may be 0.

[0277] By initializing, the MvdLX and MvdCpLX values ​​between the encoder and decoder can be the same when parsing AMVR-related syntax elements, so a mismatch between the encoder and decoder can be prevented. Additionally, by initializing to a value of 0, AMVR-related syntax elements can be avoided unnecessarily in the bitstream.

[0278] According to one embodiment of the present disclosure, MvdLX may be defined for a reference list (L0, L1, etc.), an x- or y-component, etc. Additionally, MvdCpLX may be defined for a reference list (L0, L1, etc.), an x- or y-component, control points 0, 1, 2, etc.

[0279] According to one embodiment of the present disclosure, it is possible to initialize both MvdLX and MvdCpLX values. Additionally, the initialization location may be prior to performing mvd_coding for the corresponding MvdLX or MvdCpLX. For example, it is possible to initialize the MvdLX and MvdCpLX corresponding to L0 even in a prediction block that uses only L0. Condition checks may be required to initialize only those that are essential for initialization, but by initializing without such distinction, the burden of condition checks can be reduced.

[0280] According to another embodiment of the present disclosure, it is possible to initialize a value corresponding to an unused value among MvdLX or MvdCpLX. Here, "unused" may mean that it is not used in the current block. For example, a value corresponding to a reference list that is not currently in use among MvdLX or MvdCpLX may be initialized. For example, if L0 is not used, the values ​​corresponding to MvdL0 and MvdCpL0 may be initialized. L0 not being used may be when inter_pred_idc is PRED_L1. Additionally, if L1 is not used, the values ​​corresponding to MvdL1 and MvdCpL1 may be initialized. L1 not being used may be when inter_pred_idc is PRED_L0. L0 being used may be when inter_pred_idc is PRED_L0 or PRED_BI, and L1 being used may be when inter_pred_idc is PRED_L1 or PRED_BI. Referring to FIG. 40, when inter_pred_idc is PRED_L1, MvdL0[x0][y0][0], MvdL0[x0][y0][1], MvdCpL0[x0][y0][0][0], MvdCpL0[x0][y0][0][1], MvdCpL0[x0][y0][1][0], MvdCpL0[x0][y0][1][1], MvdCpL0[x0][y0][2][0], and MvdCpL0[x0][y0][2][1] can be initialized. Additionally, the initialized value may be 0. Additionally, if inter_pred_idc is PRED_L0, MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1] can be initialized.Also, the initialization value can be 0.

[0281] MvdLX[x][y][compIdx] may be the (x,y) position relative to reference list LX and the motion vector difference for component index compIdx. MvdCpLX[x][y][cpIdx][compIdx] may be the motion vector difference relative to reference list LX. Additionally, MvdCpLX[x][y][cpIdx][compIdx] may be the position (x,y), the control point motion vector index cpIdx, and the motion vector difference for component index compIdx. Here, component may represent an x ​​or y component.

[0282] In addition, according to one embodiment of the present disclosure, the non-use of MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) may depend on whether or not affine motion compensation is used. For example, MvdLX may be initialized when affine motion compensation is used. Also, MvdCpLX may be initialized when affine motion compensation is not used. For example, there may be a signaling indicating whether affine motion compensation is used. Referring to FIG. 40, inter_affine_flag may be a signaling indicating whether affine motion compensation is used. For example, if inter_affine_flag is 1, affine motion compensation may be used. Or MotionModelIdc may be a signaling indicating whether affine motion compensation is used. For example, if MotionModelIdc is not 0, affine motion compensation may be used.

[0283] In addition, according to one embodiment of the present disclosure, the unused part of MvdLX or MvdCpLX may relate to which affine motion model is used. For example, the unused MvdLX or MvdCpLX may differ depending on whether a 4-parameter affine model or a 6-parameter affine model is used. For example, the unused part of cpIdx in MvdCpLX[x][y][cpIdx][compIdx] may differ depending on which affine motion model is used. For example, if a 4-parameter affine model is used, only some parts of MvdCpLX[x][y][cpIdx][compIdx] may be used. Or, if a 6-parameter affine model is not used, only some parts of MvdCpLX[x][y][cpIdx][compIdx] may be used. Therefore, unused MvdCpLX can be initialized to a preset value. In this case, the unused MvdCpLX may correspond to a cpIdx that is used in a 6-parameter affine model but not in a 4-parameter affine model. For example, when using a 4-parameter affine model, the value where cpIdx in MvdCpLX[x][y][cpIdx][compIdx] is 2 may not be used, and this can be initialized to a preset value. Additionally, as explained earlier, there may be signals or parameters indicating whether a 4-parameter or 6-parameter affine model is being used. For example, whether a 4-parameter or 6-parameter affine model is being used can be determined by MotionModelIdc or cu_affine_type_flag.A MotionModelIdc value of 1 or 2 may indicate the use of a 4-parameter affine model and a 6-parameter affine model, respectively. Referring to FIG. 40, when MotionModelIdc is 1, it is possible to initialize MvdCpL0[x0][y0][2][0], MvdCpL0[x0][y0][2][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1] to preset values. Additionally, it is possible to use the condition where MotionModelIdc is not 2 instead of the case where MotionModelIdc is 1. The preset value may be 0.

[0284] In addition, according to one embodiment of the present disclosure, whether to use MvdLX or MvdCpLX may be based on the value of mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list). For example, if mvd_l1_zero_flag is 1, it is possible to initialize MvdL1 and MvdCpL1 to preset values. Additionally, additional conditions may be considered. For example, it may be determined whether to use MvdLX or MvdCpLX based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list). For example, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, it is possible to initialize MvdL1 and MvdCpL1 to preset values. For example, mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) may be a high-level signaling that indicates that the MVD values ​​(e.g., MvdLX or MvdCpLX) for reference list L1 are zero. Signaling may refer to a signal transmitted from the encoder to the decoder through the bitstream. The decoder may parse mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) from the bitstream.

[0285] According to another embodiment of the present disclosure, it is possible to initialize all Mvd and MvdCp to preset values ​​before performing mvd_coding for a block. In this case, mvd_coding may refer to all mvd_coding for a CU. Thus, it is possible to prevent the initialization of Mvd or MvdCp values ​​determined by parsing the mvd_coding syntax, and the problem described can be solved by initializing all Mvd and MvdCp values.

[0286] FIG. 41 is a diagram showing an AMVR-related syntax structure according to one embodiment of the present disclosure.

[0287] The embodiment of FIG. 41 may be based on the embodiment of FIG. 38.

[0288] As described in Fig. 38, it is possible to check whether there is at least one non-zero value among MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference). The MvdLX or MvdCpLX checked at this time can be determined based on mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list). Therefore, it is possible to determine whether AMVR-related syntax parsing is performed based on mvd_l1_zero_flag. As previously explained, if mvd_l1_zero_flag is 1, it may indicate that the MVDs (motion vector differences) for reference list L1 are 0, so in this case, MvdL1 (motion vector difference for the first reference picture list) or MvdCpL1 (control point motion vector difference for the first reference picture list) may not be considered. For example, when mvd_l1_zero_flag is 1, AMVR-related syntax can be parsed based on whether there is at least one zero value among MvdL0 (motion vector difference for the zero reference picture list) or MvdCpL0 (control point motion vector difference for the zero reference picture list), regardless of whether MvdL1 or MvdCpL1 is 0 or not. According to an additional embodiment, mvd_l1_zero_flag indicating that the MVDs for reference list L1 are zero may be for a bi-prediction block only. Thus, AMVR-related syntax can be parsed based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list). For example, MvdLX or MvdCpLX can be determined based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list) to determine whether at least one non-zero value exists.For example, MvdL1 or MvdCpL1 may not be considered based on mvd_l1_zero_flag and inter_pred_idc. For example, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, MvdL1 or MvdCpL1 may not be considered. For example, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, AMVR-related syntax may be parsed based on whether there is at least one zero value among MvdL0 or MvdCpL0, regardless of whether MvdL1 or MvdCpL1 is zero or not.

[0289] Referring to Fig. 41, when mvd_l1_zero_flag is 1 and inter_pred_idc[x0][y0] == PRED_BI, operation can be performed regardless of whether there is a non-zero value among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. For example, if mvd_l1_zero_flag is 1 and inter_pred_idc[x0][y0] == PRED_BI, if there is no non-zero value among MvdL0 or MvdCpL0, AMVR-related syntax will not be parsed even if a non-zero value exists among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. It is possible.

[0290] In addition, if mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0] != PRED_BI, you can consider whether there exists a non-zero value among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. For example, if mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0] != PRED_BI, AMVR-related syntax can be parsed if there is a non-zero value among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], MvdCpL1[x0][y0][2][1]. For example, if mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0] != PRED_BI, and a non-zero value exists among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1], then AMVR-related syntax can be parsed even if both MvdL0 and MvdCpL0 are 0. there is.

[0291] FIG. 42 is a diagram showing an inter prediction related syntax structure according to one embodiment of the present disclosure.

[0292] According to one embodiment of the present disclosure, mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) may be a signaling indicating that the Mvd (motion vector difference) values ​​for reference list L1 (the first reference picture list) are zero. Additionally, this signaling may be signaled at a level higher than the current block. Thus, based on the value of mvd_l1_zero_flag, the Mvd values ​​for reference list L1 in a plurality of blocks may be zero. For example, if the value of mvd_l1_zero_flag is 1, the Mvd values ​​for reference list L1 may be zero. Or, based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list), the Mvd values ​​for reference list L1 may be zero. For example, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, the Mvd values ​​for reference list L1 may be 0. In this case, the Mvd values ​​may be MvdL1[x][y][compIdx]. Also, in this case, the Mvd values ​​may not represent the control point motion vector difference. That is, in this case, the Mvd values ​​may not represent the MvdCp values.

[0293] According to another embodiment of the present disclosure, mvd_l1_zero_flag may be a signaling indicating that the Mvd and MvdCp values ​​for reference list L1 are 0. Additionally, this signaling may be signaled at a level higher than the current block. Thus, based on the value of mvd_l1_zero_flag, the Mvd and MvdCp values ​​for reference list L1 in a plurality of blocks may be 0. For example, if the value of mvd_l1_zero_flag is 1, the Mvd and MvdCp values ​​for reference list L1 may be 0. Or, based on mvd_l1_zero_flag and inter_pred_idc, the Mvd and MvdCp values ​​for reference list L1 may be 0. For example, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, the Mvd and MvdCp values ​​for reference list L1 may be 0. In this case, the Mvd values ​​may be MvdL1[x][y][compIdx]. Also, the MvdCp values ​​may be MvdCpL1[x][y][cpIdx][compIdx].

[0294] Alternatively, Mvd or MvdCp values ​​being 0 may mean that the corresponding mvd_coding syntax structure is not parsed. That is, for example, if the mvd_l1_zero_flag value is 1, the mvd_coding syntax structure corresponding to MvdL1 or MvdCpL1 may not be parsed. Also, if the mvd_l1_zero_flag value is 0, the mvd_coding syntax structure corresponding to MvdL1 or MvdCpL1 may be parsed.

[0295] According to one embodiment of the present disclosure, if the Mvd or MvdCp values ​​are 0 based on mvd_l1_zero_flag, the signaling indicating MVP may not be parsed. The signaling indicating MVP may include mvp_l1_flag. Additionally, according to the description of mvd_l1_zero_flag described above, the mvd_l1_zero_flag signaling may mean that both Mvd and MvdCp are 0. For example, if the condition that the Mvd or MvdCp values ​​for reference list L1 are 0 is satisfied, the signaling indicating MVP may not be parsed. In such a case, the signaling indicating MVP may be inferred to a preset value. For example, if the signaling indicating MVP does not exist, its value may be inferred to 0. Additionally, based on mvd_l1_zero_flag, if the condition indicating that Mvd or MvdCp values ​​are 0 is not satisfied, a signaling indicating an MVP can be parsed. However, in this embodiment, the degree of freedom to select an MVP when Mvd or MvdCp values ​​are 0 may be lost, and consequently, coding efficiency may decrease.

[0296] More specifically, if the condition is satisfied that the Mvd or MvdCp values ​​for reference list L1 are 0 and affine MC is used, the signal representing the MVP may not be parsed. In this case, the signal representing the MVP can be inferred to a preset value.

[0297] Referring to FIG. 42, if mvd_l1_zero_flag is 1 and inter_pred_idc value is PRED_BI, mvp_l1_flag may not be parsed. In this case, the mvp_l1_flag value may be inferred to 0. Alternatively, if mvd_l1_zero_flag is 0 or inter_pred_idc value is not PRED_BI, mvp_l1_flag may be parsed.

[0298] In this embodiment, determining whether to parse the signaling representing the MVP based on mvd_l1_zero_flag may occur when specific conditions are satisfied. For example, the specific condition may include a condition where general_merge_flag is 0. For example, general_merge_flag may have the same meaning as the merge_flag described earlier. Additionally, the specific condition may include a condition based on CuPredMode. More specifically, the specific condition may include a condition where CuPredMode is not MODE_IBC. Or, the specific condition may include a condition where CuPredMode is MODE_INTER. If CuPredMode is MODE_IBC, it is possible to use a prediction that references the current picture. Also, if CuPredMode is MODE_IBC, a block vector or motion vector corresponding to that block may exist. If CuPredMode is MODE_INTER, it is possible to use a prediction that references a picture other than the current picture. If CuPredMode is MODE_INTER, a motion vector corresponding to that block may exist.

[0299] Accordingly, according to an embodiment of the present disclosure, if general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag is 1, and inter_pred_idc is PRED_BI, mvp_l1_flag may not be parsed. In addition, if mvp_l1_flag does not exist, its value may be inferred as 0.

[0300] More specifically, if general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag is 1, inter_pred_idc is PRED_BI, and affine MC is used, mvp_l1_flag may not be parsed. Also, if mvp_l1_flag does not exist, its value may be inferred as 0.

[0301] Referring to FIG. 42, sym_mvd_flag may be a signaling representing a symmetric MVD. In the case of a symmetric MVD, another MVD can be determined based on a certain MVD. In the case of a symmetric MVD, another MVD can be determined based on an explicitly signaled MVD. For example, in the case of a symmetric MVD, an MVD for another reference list can be determined based on an MVD for one reference list. For example, in the case of a symmetric MVD, an MVD for reference list L1 can be determined based on an MVD for reference list L0. When determining another MVD based on a certain MVD, it is possible to determine the other MVD by reversing the sign of the said certain MVD.

[0302] FIG. 43 is a diagram showing an inter prediction related syntax structure according to one embodiment of the present disclosure.

[0303] The embodiment of FIG. 43 may be an embodiment for solving the problem described in FIG. 42.

[0304] According to one embodiment of the present disclosure, it is possible to parse a signaling indicating an MVP (motion vector predictor) when the Mvd (motion vector difference) or MvdCp (control point motion vector difference) values ​​are 0 based on mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list). The signaling indicating an MVP may include mvp_l1_flag (motion vector predictor index for the first reference picture list). Additionally, according to the description of mvd_l1_zero_flag described above, the mvd_l1_zero_flag signaling may mean that both Mvd and MvdCp are 0. For example, a signaling indicating an MVP can be parsed when the condition is satisfied that the Mvd or MvdCp values ​​for reference list L1 (the first reference picture list) are 0. Therefore, it is possible not to infer the signaling indicating the MVP. Accordingly, there is a degree of freedom to select the MVP even when the Mvd or MvdCp value is 0 based on mvd_l1_zero_flag. This can improve coding efficiency. Additionally, the signaling indicating the MVP can be parsed even when the condition indicating that the Mvd or MvdCp values ​​are 0 based on mvd_l1_zero_flag is not satisfied.

[0305] More specifically, we can parse signals representing an MVP when using affine MC, satisfying the condition that the Mvd or MvdCp values ​​for reference list L1 are 0.

[0306] Referring to line (4301) of FIG. 43, information (inter_pred_idc) about the reference picture list for the current block can be obtained. Referring to line (4302), if the information (inter_pred_idc) about the reference picture list indicates that only the 0th reference picture list (list 0) is used, the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1) can be parsed from the bitstream at line (4303).

[0307] mvd_l1_zero_flag (motion vector difference zero flag) can be obtained from the bitstream. mvd_l1_zero_flag (motion vector difference zero flag) can indicate whether MvdLX (motion vector difference) and MvdCpLX (multiple control point motion vector differences) are set to 0 for the first reference picture list. Signaling may refer to a signal transmitted from the encoder to the decoder through the bitstream. The decoder can parse mvd_l1_zero_flag (motion vector difference zero flag) from the bitstream.

[0308] mvp_l1_flag (motion vector difference zero flag) can be parsed if mvd_l1_zero_flag (motion vector difference zero flag) is 1 and inter_pred_idc (information about reference picture list) value is PRED_BI, where PRED_BI may indicate the use of both List 0 (zero reference picture list) and List 1 (first reference picture list). Alternatively, mvp_l1_flag (motion vector difference zero flag) can be parsed if mvd_l1_zero_flag (motion vector difference zero flag) is 0 or inter_pred_idc (information about reference picture list) value is not PRED_BI. That is, mvp_l1_flag (motion vector predictor index) can be parsed regardless of whether mvd_l1_zero_flag (motion vector difference zero flag) is 1 and inter_pred_idc (information about reference picture list) indicates using both the 0th reference picture list and the 1st reference picture list.

[0309] In this embodiment, determining Mvd and MvdCp based on mvd_l1_zero_flag (the motion vector difference zero flag of the first reference picture list) and parsing the signaling representing the MVP can occur when specific conditions are satisfied. For example, the specific condition may include a condition where general_merge_flag is 0. For example, general_merge_flag may have the same meaning as the merge_flag described earlier. Additionally, the specific condition may include a condition based on CuPredMode. More specifically, the specific condition may include a condition where CuPredMode is not MODE_IBC. Or, the specific condition may include a condition where CuPredMode is MODE_INTER. If CuPredMode is MODE_IBC, it is possible to use a prediction that references the current picture. Also, if CuPredMode is MODE_IBC, a block vector or motion vector corresponding to that block may exist. If CuPredMode is MODE_INTER, it is possible to use a prediction that references a picture other than the current picture. If CuPredMode is MODE_INTER, a motion vector corresponding to that block may exist.

[0310] Accordingly, according to an embodiment of the present disclosure, when general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag is 1, and inter_pred_idc is PRED_BI, mvp_l1_flag (motion vector predictor index for the first reference picture list) can be parsed. Thus, mvp_l1_flag (motion vector predictor index for the first reference picture list) exists, and its value may not be inferred.

[0311] More specifically, mvp_l1_flag can be parsed when general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) is 1, inter_pred_idc (information about the reference picture list) is PRED_BI, and affine MC is used. Also, mvp_l1_flag (motion vector predictor index for the first reference picture list) exists, and its value may not be inferred.

[0312] In addition, it is possible to carry out the embodiments of FIG. 43 and FIG. 40 together. For example, it is possible to parse mvp_l1_flag after initializing Mvd or MvdCp. In this case, the initialization of Mvd or MvdCp may be the initialization described in FIG. 40. Also, the parsing of mvp_l1_flag may follow the description of FIG. 43. For example, based on mvd_l1_zero_flag, if the Mvd and MvdCp values ​​for reference list L1 are not 0, and the MotionModelIdc value is 1, the MvdCpL1 values ​​of control point index 2 may be initialized and mvp_l1_flag may be parsed.

[0313] FIG. 44 is a diagram showing an inter prediction related syntax structure according to one embodiment of the present disclosure.

[0314] The embodiment of FIG. 44 may be an embodiment for increasing coding efficiency by not eliminating the degree of freedom to select the MVP. Additionally, the embodiment of FIG. 44 may be a different representation of the embodiment described in FIG. 43. Therefore, descriptions that overlap with the embodiment of FIG. 43 may be omitted.

[0315] In the embodiment of FIG. 44, a signaling indicating an MVP can be parsed based on mvd_l1_zero_flag, indicating that the Mvd or MvdCp value is 0. The signaling indicating an MVP may include mvp_l1_flag.

[0316] FIG. 45 is a diagram showing an inter prediction related syntax according to one embodiment of the present disclosure.

[0317] According to one embodiment of the present disclosure, the inter prediction method may include skip mode, merge mode, inter mode, etc. According to one embodiment, in skip mode, the residual signal may not be transmitted. In addition, in skip mode, an MV determination method such as that of merge mode may be used. Whether to use skip mode may be determined by a skip flag. Referring to FIG. 33, whether to use skip mode may be determined by the value of cu_skip_flag.

[0318] According to one embodiment, the motion vector difference may not be used in merge mode. The motion vector can be determined based on the motion candidate index. Whether to use merge mode may be determined by the merge flag. Referring to FIG. 33, whether to use merge mode may be determined by the merge_flag value. In addition, it is possible to use merge mode when skip mode is not used.

[0319] In Skip mode or merge mode, it is possible to selectively use one or more types of candidate lists. For example, it is possible to use merge candidates or subblock merge candidates. Additionally, merge candidates may include spatial neighboring candidates, temporal candidates, etc. Furthermore, merge candidates may include candidates that use motion vectors for the entire current block (CU; Coding Unit). That is, they may include candidates where the motion vectors of each subblock belonging to the current block are the same. Additionally, subblock merge candidates may include subblock-based temporal MVs, affine merge candidates, etc. Furthermore, subblock merge candidates may include candidates that allow different motion vectors to be used for each subblock of the current block (CU). An affine merge candidate may be a method created by determining the control point motion vector of affine motion prediction without using motion vector differences. Additionally, subblock merge candidates may include methods that determine motion vectors at the subblock level within the current block. For example, in addition to the aforementioned subblock-based temporal MV and affine merge candidates, subblock merge candidates may include planar MVs, regression-based MVs, STMVPs, etc.

[0320] According to one embodiment, the motion vector difference can be used in inter mode. A motion vector predictor can be determined based on the motion candidate index, and a motion vector can be determined based on the motion vector predictor and the motion vector difference. Whether to use inter mode can be determined based on whether other modes are used. In another embodiment, whether to use inter mode can be determined by a flag. FIG. 45 illustrates an example of using inter mode when other modes, such as skip mode and merge mode, are not used.

[0321] Inter mode may include AMVP mode, affine inter mode, etc. Inter mode may be a mode that determines a motion vector based on the difference between a motion vector predictor and a motion vector. Affine inter mode may be a method that uses the motion vector difference when determining the control point motion vector of an affine motion prediction.

[0322] Referring to FIG. 45, after determining whether to use a subblock merge candidate or a merge candidate, it can be decided whether to use a subblock merge candidate or a merge candidate. For example, a merge_subblock_flag indicating whether to use a subblock merge candidate when a specific condition is satisfied can be parsed. In addition, the specific condition may be a condition related to block size. For example, it may be a condition regarding width, height, area, etc., and these may be used in combination. Referring to FIG. 45, for example, it may be a condition when the width and height of the current block (CU) are greater than or equal to a specific value. When parsing the merge_subblock_flag, its value can be inferred as 0. If the merge_subblock_flag is 1, a subblock merge candidate may be used, and if it is 0, a merge candidate may be used. When using a subblock merge candidate, the candidate index merge_subblock_idx can be parsed, and when using a merge candidate, the candidate index merge_idx can be parsed. In this case, if the maximum number of candidate lists is 1, parsing may not be performed. If merge_subblock_idx or merge_idx is not parsed, it can be inferred as 0.

[0323] Figure 45 shows the coding_unit function, but the intra prediction-related content may be omitted, and Figure 45 may represent the case where it is determined by inter prediction.

[0324] FIG. 46 is a diagram showing a triangle partitioning mode according to one embodiment of the present disclosure.

[0325] The triangle partitioning mode (TPM) mentioned in the present disclosure may be referred to by various names, such as triangle partition mode, triangle prediction, triangle based prediction, triangle motion compensation, triangular prediction, triangle inter prediction, triangular merge mode, triangle merge mode, etc. Additionally, the TPM may be included in the geometric partitioning mode (GPM).

[0326] As shown in Fig. 46, TPM can be a method of dividing a rectangular block into two triangles. However, GPM can divide a block into two blocks in various ways. For example, GPM can divide a single rectangular block into two triangle blocks as shown in Fig. 46. GPM can also divide a single rectangular block into one pentagonal block and one triangle block. Additionally, GPM can divide a single rectangular block into two square blocks. Here, a rectangle may include a square. For the sake of convenience of explanation, the following description is based on TPM, a simplified version of GPM, but it should be interpreted as including GPM.

[0327] According to one embodiment of the present disclosure, a uni-prediction may exist as a prediction method. A uni-prediction may be a prediction method using a single reference list. The reference list may exist in multiple places, and according to one embodiment, it is possible for two to exist as L0 and L1. When using a uni-prediction, it is possible to use a single reference list in a single block. Additionally, when using a uni-prediction, it is possible to use a single motion information to predict a single pixel. In the present disclosure, a block may refer to a CU (coding unit) or a PU (prediction unit). Additionally, in the present disclosure, a block may refer to a TU (transform unit).

[0328] According to another embodiment of the present disclosure, a bi-prediction method may exist. Bi-prediction may be a prediction method using multiple reference lists. In one embodiment, bi-prediction may be a prediction method using two reference lists. For example, bi-prediction may use L0 and L1 reference lists. When using bi-prediction, it is possible to use multiple reference lists in a single block. For example, when using bi-prediction, it is possible to use two reference lists in a single block. Additionally, when using bi-prediction, it is possible to use multiple motion information to predict a single pixel.

[0329] The above motion information may include a motion vector, a reference index, and a prediction list utilization flag.

[0330] The above reference list may be a reference picture list.

[0331] In the present disclosure, motion information corresponding to uni-prediction or bi-prediction can be defined as a single motion information set.

[0332] According to one embodiment of the present disclosure, it is possible to use multiple motion information sets when using a TPM. For example, it is possible to use two motion information sets when using a TPM. For example, it is possible to use at most two motion information sets when using a TPM. Furthermore, the method of applying two motion information sets within a block using a TPM may be based on location. For example, it is possible to use one motion information set for a pre-set location within a block using a TPM, and another motion information set for another pre-set location. Additionally, it is possible to use two motion information sets together for yet another pre-set location. For example, for the other pre-set location, Prediction 1 based on one motion information set and Prediction 3 based on another motion information set may be used for prediction. For example, Prediction 3 may be a weighted sum of Prediction 1 and Prediction 2.

[0333] Referring to FIG. 46, Partition 1 and Partition 2 may schematically represent the aforementioned preset location and the other preset location. When using TPM, one of two split methods can be used as shown in FIG. 46. The two split methods may include diagonal split and anti-diagonal split. Additionally, it may be divided into two triangle-shaped partitions by splitting. As previously explained, TPM can be included in GPM. Since GPM has already been explained, a redundant explanation is omitted.

[0334] According to one embodiment of the present disclosure, when using a TPM, it is possible to use only uni-prediction for each partition. That is, it is possible to use one motion information for each partition. This may be intended to reduce complexity such as memory access and computational complexity. Therefore, it is possible to use only two motion informations for each CU.

[0335] Additionally, it is possible to determine each motion information from the candidate list. According to one embodiment, the candidate list used in the TPM may be based on a merge candidate list. In another embodiment, the candidate list used in the TPM may be based on an AMVP candidate list. Thus, it is possible to signal the candidate index to use the TPM. Furthermore, for a block using the TPM, it is possible to encode, decode, and parse candidate indices equal to the number of partitions in the TPM, or at most the number of partitions.

[0336] In addition, even if a block is predicted by TPM based on multiple motion information, it is possible to perform transform and quantization on the entire block.

[0337] FIG. 47 is a diagram showing merge data syntax according to one embodiment of the present disclosure.

[0338] According to one embodiment of the present disclosure, the merge data syntax may include signaling regarding various modes. The various modes may include regular merge mode, MMVD (merge with MVD), subblock merge mode, CIIP (combined intra- and inter-prediction), TPM, etc. Regular merge mode may be a mode such as the merge mode in HEVC. Additionally, there may be signaling indicating whether the various modes are used in a block. Furthermore, these signalings may be parsed as syntax elements or implicitly signaled. Referring to FIG. 47, the signaling indicating whether regular merge mode, MMVD, subblock merge mode, CIIP, and TPM are used may be regular_merge_flag, mmvd_merge_flag (or mmvd_flag), merge_subblock_flag, ciip_flag (or mh_intra_flag), and MergeTriangleFlag (or merge_triangle_flag), respectively.

[0339] According to one embodiment of the present disclosure, when using a merge mode, if a signal indicates that all modes excluding a certain mode among the various modes are not used, it may be determined that the certain mode is used. Additionally, when using a merge mode, if a signal indicates that at least one of the modes excluding a certain mode among the various modes is used, it may be determined that the certain mode is not used. Furthermore, there may be a higher-level signaling indicating whether a mode is available. The higher level may be a unit including a block. The higher level may be a sequence, picture, slice, tile group, tile, CTU, etc. If the higher-level signaling indicating whether a mode is available indicates that it is available, there may be an additional signaling indicating whether that mode is used, and the mode may be used or not used. If the higher-level signaling indicating whether a mode is available indicates that it is not available, the mode may not be used. For example, when using a merge mode, if a signal indicates that regular merge mode, MMVD, subblock merge mode, and CIIP are all not used, it may be determined that TPM is used. Additionally, when using merge mode, if it is signaled that at least one of regular merge mode, MMVD, subblock merge mode, or CIIP is being used, it can be determined that the TPM is not being used. Furthermore, there may be a signal indicating whether merge mode is being used. For example, the signal indicating whether merge mode is being used may be general_merge_flag or merge_flag.If merge mode is used, merge data syntax such as Fig. 47 can be parsed.

[0340] In addition, the block size for which TPM can be used may be limited. For example, it is possible to use TPM when both width and height are 8 or greater.

[0341] If a TPM is used, TPM-related syntax elements can be parsed. TPM-related syntax elements may include signalings indicating a split method and signalings indicating a candidate index. The split method may indicate the split direction. For a block using a TPM, there may be multiple signalings (e.g., two) indicating a candidate index. Referring to Fig. 47, the signaling indicating a split method may be merge_triangle_split_dir. Additionally, the signalings indicating a candidate index may be merge_triangle_idx0 and merge_triangle_idx1.

[0342] In the present disclosure, candidate indices for TPM may be denoted as m and n. For example, candidate indices for Partition 1 and Partition 2 of FIG. 46 may be m and n, respectively. According to one embodiment, m and n may be determined based on signaling representing candidate indices as described in FIG. 47. According to one embodiment of the present disclosure, one of m and n may be determined based on one of merge_triangle_idx0 and merge_triangle_idx1, and the other of m and n may be determined based on both merge_triangle_idx0 and merge_triangle_idx1.

[0343] Alternatively, it is possible for one of m and n to be determined based on one of merge_triangle_idx0 and merge_triangle_idx1, and the other of m and n to be determined based on the other of merge_triangle_idx0 and merge_triangle_idx1.

[0344] More specifically, it is possible for m to be determined based on merge_triangle_idx0, and for n to be determined based on merge_triangle_idx0 (or m) and merge_triangle_idx1. For example, m and n can be determined as follows.

[0345] m = merge_triangle_idx0

[0346] n = merge_triangle_idx1 + (merge_triangle_idx1 >= m) ? 1:0

[0347] According to one embodiment of the present disclosure, m and n may not be the same. This is because in a TPM, if two candidate indices are the same—that is, if two motion informations are the same—the effect of partitioning may not be obtained. Therefore, the above signaling method may be intended to reduce the number of signaling bits when n > m when signaling n. Since m among all candidates will not be n, it can be excluded from signaling.

[0348] If the candidate list used in the TPM is called mergeCandList, then mergeCandList[m] and mergeCandList[n] can be used as motion information in the TPM.

[0349] FIG. 48 is a diagram showing upper-level signaling according to one embodiment of the present disclosure.

[0350] According to one embodiment of the present disclosure, there may be a plurality of upper-level signalings. An upper-level signaling may be a signaling transmitted at an upper-level unit. An upper-level unit may include one or more lower-level units. An upper-level signaling may be a signaling applied to one or more lower-level units. For example, a slice or sequence may be an upper-level unit for a CU, PU, ​​TU, etc. Conversely, a CU, PU, ​​or TU may be a lower-level unit for a slice or sequence.

[0351] According to one embodiment of the present disclosure, the upper-level signaling may include a signaling indicating the maximum number of candidates. For example, the upper-level signaling may include a signaling indicating the maximum number of merge candidates. For example, the upper-level signaling may include a signaling indicating the maximum number of candidates used in the TPM. The signaling indicating the maximum number of merge candidates or the signaling indicating the maximum number of candidates used in the TPM may be signaled and parsed when inter prediction is allowed. Whether inter prediction is allowed may be determined by the slice type. The slice type may be I, P, B, etc. For example, if the slice type is I, inter prediction may not be allowed. For example, if the slice type is I, only intra prediction or intra block copy (IBC) may be used. Also, if the slice type is P or B, inter prediction may be allowed. Also, if the slice type is P or B, intra prediction, IBC, etc. may be allowed. In addition, when the slice type is P, it is possible to use at most one reference list to predict the pixel. Also, when the slice type is B, it is possible to use multiple reference lists to predict the pixel. For example, when the slice type is B, it is possible to use at most two reference lists to predict the pixel.

[0352] According to one embodiment of the present disclosure, when signaling a maximum number, it is possible to signal based on a reference value. For example, it is possible to signal (reference value - maximum number). Thus, it is possible to derive the maximum number based on the value parsed by the decoder and the reference value. For example, (reference value - parsed value) can be determined as the maximum number.

[0353] According to one embodiment, the reference value in the signaling representing the maximum number of merge candidates may be 6.

[0354] According to one embodiment, the reference value in the signaling representing the maximum number of candidates used in the TPM may be the maximum number of merge candidates.

[0355] Referring to FIG. 48, the signaling representing the maximum number of merge candidates may be six_minus_max_num_merge_cand. Here, a merge candidate may refer to a candidate for merge motion vector prediction. For convenience of explanation, six_minus_max_num_merge_cand will also be referred to as the first information below. Referring to FIG. 2 and FIG. 7, the signaling may refer to a signal transmitted from an encoder to a decoder via a bitstream. six_minus_max_num_merge_cand (the first information) may be signaled in sequence units. The decoder can parse six_minus_max_num_merge_cand (the first information) from the bitstream.

[0356] Additionally, the signaling representing the maximum number of candidates used in the TPM may be max_num_merge_cand_minus_max_num_triangle_cand. Referring to FIGS. 2 and 7, the signaling may refer to a signal transmitted from the encoder to the decoder via a bitstream. The decoder may parse max_num_merge_cand_minus_max_num_triangle_cand (third information) from the bitstream. max_num_merge_cand_minus_max_num_triangle_cand (third information) may be information related to the maximum number of merge mode candidates for a partitioned block.

[0357] Additionally, the maximum number of merge candidates can be MaxNumMergeCand (maximum number of merge candidates), and this value can be based on six_minus_max_num_merge_cand (first information). Also, the maximum number of candidates used in TPM can be MaxNumTriangleMergeCand, and this value can be based on max_num_merge_cand_minus_max_num_triangle_cand. MaxNumMergeCand (maximum number of merge candidates) can be used in merge mode and is information that can be used when a block is partitioned for motion compensation or when it is not partitioned. Although the above explanation was based on TPM, it can be explained in the same way for GPM.

[0358] According to one embodiment of the present disclosure, there may be a high-level signaling indicating whether the TPM mode is available. Referring to FIG. 48, the high-level signaling indicating whether the TPM mode is available may be sps_triangle_enabled_flag (second information). The information indicating whether the TPM mode is available may be the same as the information indicating whether the block can be partitioned for inter prediction as in FIG. 46. Since GPM includes TPM, the information indicating whether the block can be partitioned may be the same as the information indicating whether the GPM mode is available. Being available for inter prediction may indicate that motion compensation is performed. That is, sps_triangle_enabled_flag (second information) may be information indicating whether the block can be partitioned for inter prediction. If the second information indicating whether the block can be partitioned is 1, it may indicate that the TPM or GPM is available. Additionally, if the second information is 0, it may indicate that TPM or GPM cannot be used. However, it is not limited to this, and if the second information is 0, it may indicate that TPM or GPM can be used. Additionally, if the second information is 1, it may indicate that TPM or GPM cannot be used.

[0359] Referring to FIGS. 2 and FIGS. 7, signaling may refer to a signal transmitted from an encoder to a decoder through a bitstream. The decoder can parse sps_triangle_enabled_flag (second information) from the bitstream.

[0360] According to one embodiment of the present disclosure, it is possible to use a TPM only when there are candidates used in the TPM that are greater than or equal to the number of partitions of the TPM. For example, when a TPM is partitioned into two, it is possible to use a TPM only when there are two or more candidates used in the TPM. According to one embodiment, the candidates used in the TPM may be based on merge candidates. Therefore, according to one embodiment of the present disclosure, it may be possible to use a TPM when the maximum number of merge candidates is two or more. Therefore, it is possible to parse signals related to the TPM when the maximum number of merge candidates is two or more. The signals related to the TPM may be signals indicating the maximum number of candidates used in the TPM.

[0361] Referring to Fig. 48, it is possible to parse max_num_merge_cand_minus_max_num_triangle_cand when sps_triangle_enabled_flag is 1 and MaxNumMergeCand is 2 or greater. It is also possible not to parse max_num_merge_cand_minus_max_num_triangle_cand when sps_triangle_enabled_flag is 0 or MaxNumMergeCand is less than 2.

[0362] FIG. 49 is a diagram showing the maximum number of candidates used in a TPM according to one embodiment of the present disclosure.

[0363] Referring to Fig. 49, the maximum number of candidates used in the TPM may be MaxNumTriangleMergeCand. Additionally, the signaling representing the maximum number of candidates used in the TPM may be max_num_merge_cand_minus_max_num_triangle_cand. Furthermore, the details described in Fig. 48 may have been omitted.

[0364] According to one embodiment of the present disclosure, the maximum number of candidates used in the TPM may exist in the range (inclusive) from "the number of partitions in the TPM" to "the reference value in the signaling representing the maximum number of candidates used in the TPM". Thus, when the number of partitions in the TPM is 2 and the reference value is the maximum number of merge candidates, MaxNumTriangleMergeCand may exist in the range (inclusive) from 2 to MaxNumMergeCand, as shown in FIG. 49.

[0365] According to one embodiment of the present disclosure, if there is no signaling indicating the maximum number of candidates used in the TPM, it is possible to infer a signaling indicating the maximum number of candidates used in the TPM or to infer the maximum number of candidates used in the TPM. For example, if there is no signaling indicating the maximum number of candidates used in the TPM, the maximum number of candidates used in the TPM can be inferred to 0. Alternatively, if there is no signaling indicating the maximum number of candidates used in the TPM, the signaling indicating the maximum number of candidates used in the TPM can be inferred as a reference value.

[0366] In addition, if there is no signaling indicating the maximum number of candidates used in the TPM, it is possible not to use the TPM. Alternatively, if the maximum number of candidates used in the TPM is less than the number of partitions in the TPM, it is possible not to use the TPM. Alternatively, if the maximum number of candidates used in the TPM is 0, it is possible not to use the TPM.

[0367] However, according to the embodiments of FIGS. 48 and 49, when the number of partitions of the TPM and the "reference value in the signaling representing the maximum number of candidates used in the TPM" are the same, there may only be one possible value for the maximum number of candidates used in the TPM. However, according to the embodiments of FIGS. 48 and 49, even in such a case, the signaling representing the maximum number of candidates used in the TPM can be parsed, and this may be unnecessary. If MaxNumMergeCand is 2, referring to FIG. 49, the only possible value for MaxNumTriangleMergeCand may be 2. However, referring to FIG. 48, even in such a case, max_num_merge_cand_minus_max_num_triangle_cand can be parsed.

[0368] Referring to Fig. 49, MaxNumTriangleMergeCand can be determined as (MaxNumMergeCand - max_num_merge_cand_minus_max_num_triangle_cand).

[0369] FIG. 50 is a diagram showing upper-level signaling regarding a TPM according to one embodiment of the present disclosure.

[0370] According to one embodiment of the present disclosure, if the number of partitions in the TPM and the "reference value in signaling representing the maximum number of candidates used in the TPM" are the same, the signaling representing the maximum number of candidates used in the TPM may not be parsed. Additionally, according to the embodiment described above, the number of partitions in the TPM may be 2. Also, the "reference value in signaling representing the maximum number of candidates used in the TPM" may be the maximum number of merge candidates. Therefore, if the maximum number of merge candidates is 2, the signaling representing the maximum number of candidates used in the TPM may not be parsed.

[0371] Alternatively, if the "criteria value in the signaling representing the maximum number of candidates used in the TPM" is less than or equal to the number of partitions in the TPM, the signaling representing the maximum number of candidates used in the TPM may not be parsed. Therefore, if the maximum number of merge candidates is 2 or less, the signaling representing the maximum number of candidates used in the TPM may not be parsed.

[0372] Referring to line (5001) of FIG. 50, if MaxNumMergeCand (maximum number of merge candidates) is 2 or if MaxNumMergeCand (maximum number of merge candidates) is 2 or less, max_num_merge_cand_minus_max_num_triangle_cand (third information) may not be parsed. Also, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) may be parsed. Also, if sps_triangle_enabled_flag (second information) is 0, max_num_merge_cand_minus_max_num_triangle_cand (third information) may not be parsed. Therefore, if sps_triangle_enabled_flag (second information) is 0 or MaxNumMergeCand (maximum number of merge candidates) is 2 or less, max_num_merge_cand_minus_max_num_triangle_cand (third information) may not be parsed.

[0373] FIG. 51 is a diagram showing the maximum number of candidates used in a TPM according to one embodiment of the present disclosure.

[0374] The embodiment of FIG. 51 can be carried out together with the embodiment of FIG. 50. Also, the previously described items may be omitted from this drawing.

[0375] According to one embodiment of the present disclosure, if the "reference value in the signaling indicating the maximum number of candidates used in the TPM" is the number of partitions of the TPM, the maximum number of candidates used in the TPM can be inferred and set to the number of partitions of the TPM. Additionally, the inferred and set may be in the case where there is no signaling indicating the maximum number of candidates used in the TPM. According to the embodiment of FIG. 50, if the "reference value in the signaling indicating the maximum number of candidates used in the TPM" is the number of partitions of the TPM, the signaling indicating the maximum number of candidates used in the TPM may not be parsed, and if there is no signaling indicating the maximum number of candidates used in the TPM, the maximum number of candidates used in the TPM can be inferred to the number of partitions of the TPM. Additionally, it is possible to perform this when an additional condition is satisfied. The additional condition may be a condition where the upper-level signaling indicating whether the TPM mode can be used is 1.

[0376] In addition, while this embodiment describes inferring and setting the maximum number of candidates used in the TPM, it is also possible to infer and set a signal representing the maximum number of candidates used in the TPM so that the described maximum number of candidates used in the TPM is derived, instead of inferring and setting the maximum number of candidates used in the TPM.

[0377] Referring to FIG. 50, when sps_triangle_enabled_flag (second information) indicates 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) can be received. At this time, referring to line (5101) of FIG. 51, MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) can be obtained using the clearly signaled max_num_merge_cand_minus_max_num_triangle_cand (third information). To summarize, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than or equal to 3, MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) can be obtained by subtracting the third information (max_num_merge_cand_minus_max_num_triangle_cand) from MaxNumMergeCand (maximum number of merge candidates).

[0378] Referring to line (5102) of Fig. 51, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is 2, it is possible to set MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) to 2. To explain in more detail, as previously explained, when sps_triangle_enabled_flag (second information) indicates 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) can be received; therefore, when sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is 2 as in line (5102), max_num_merge_cand_minus_max_num_triangle_cand (third information) may not be received. In this case, umTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) can be determined without max_num_merge_cand_minus_max_num_triangle_cand (third information).

[0379] Additionally, referring to line (5103) of FIG. 51, if sps_triangle_enabled_flag (second information) is 0 or MaxNumMergeCand (maximum number of merge candidates) is not 2, it is possible to infer and set MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) to 0. At this time, as already explained in line (5101), if sps_triangle_enabled_flag (second information) indicates 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than or equal to 3, max_num_merge_cand_minus_max_num_triangle_cand (third information) will be signaled, so the case in which MaxNumTriangleMergeCand is inferred and set to 0 may be when the second information is 0 or the maximum number of merge candidates is 1. To summarize, if sps_triangle_enabled_flag (secondary information) is 0 or MaxNumMergeCand (maximum number of merge candidates) is 1, MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) can be set to 0.

[0380] MaxNumMergeCand(maximum number of merge candidates) and MaxNumTriangleMergeCand(maximum number of merge mode candidates for partitioned blocks) can be used for different purposes. For example, MaxNumMergeCand(maximum number of merge candidates) can be used when a block is partitioned for motion compensation or not partitioned. However, MaxNumTriangleMergeCand(maximum number of merge mode candidates for partitioned blocks) is information that can be used when a block is partitioned. The number of candidates for a partitioned block in merge mode cannot exceed MaxNumTriangleMergeCand(maximum number of merge mode candidates for partitioned blocks).

[0381] It is also possible to use another embodiment. In the embodiment of FIG. 51, cases where MaxNumMergeCand is not 2 include cases where it is greater than 2. In such cases, the meaning of inferring and setting MaxNumTriangleMergeCand to 0 may be unclear, but since there is a signal indicating the maximum number of candidates used in the TPM, inferring is not performed, so there is no problem with the operation. However, in this embodiment, inferring can be performed while preserving the meaning.

[0382] If sps_triangle_enabled_flag (secondary information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is 2 or greater, it is possible to infer and set MaxNumTriangleMergeCand to 2 (or MaxNumMergeCand). Otherwise (i.e., sps_triangle_enabled_flag is 0 or MaxNumMergeCand is less than 2), it is possible to infer and set MaxNumTriangleMergeCand to 0.

[0383] Alternatively, if sps_triangle_enabled_flag is 1 and MaxNumMergeCand is 2, it is possible to infer and set MaxNumTriangleMergeCand to 2. Otherwise, if sps_triangle_enabled_flag is 0, it is possible to infer and set MaxNumTriangleMergeCand to 0.

[0384] Accordingly, in the above embodiments, according to the embodiment of FIG. 50, when MaxNumMergeCand is 0 or 1 or 2, there may not be a signal indicating the maximum number of candidates used in the TPM, and when MaxNumMergeCand is 0 or 1, the maximum number of candidates used in the TPM can be inferred and set to 0. When MaxNumMergeCand is 2, the maximum number of candidates used in the TPM can be inferred and set to 2.

[0385] FIG. 52 is a diagram showing TPM-related syntax elements according to one embodiment of the present disclosure.

[0386] As explained earlier, there may be a maximum number of candidates used in the TPM, and the number of partitions in the TPM may be pre-configured. Additionally, the candidate indexes used in the TPM may differ from one another.

[0387] According to one embodiment of the present disclosure, if the maximum number of candidates used in the TPM is equal to the number of partitions of the TPM, a different signaling can be performed compared to cases where it is not. For example, if the maximum number of candidates used in the TPM is equal to the number of partitions of the TPM, a different candidate index signaling can be performed compared to cases where it is not. Accordingly, it is possible to signal with fewer bits. Alternatively, if the maximum number of candidates used in the TPM is less than or equal to the number of partitions of the TPM, a different signaling can be performed compared to cases where it is not (among these, if it is less than the number of partitions of the TPM, it may be a case where the TPM cannot be used).

[0388] If the partition of the TPM is 2, two candidate indices can be signaled. If the maximum number of candidates used in the TPM is 2, there may only be two possible combinations of candidate indices. The two combinations could be m and n being 0 and 1, and 1 and 0, respectively. Therefore, two candidate indices can be signaled using only 1-bit signaling.

[0389] Referring to Fig. 52, when MaxNumTriangleMergeCand is 2, a different candidate index signaling can be performed compared to the case where it is not (otherwise, when using TPM, when MaxNumTriangleMergeCand is greater than 2). Alternatively, when MaxNumTriangleMergeCand is 2 or less, a different candidate index signaling can be performed compared to the case where it is not (otherwise, when using TPM, when MaxNumTriangleMergeCand is greater than 2). Referring to Fig. 52, the different candidate index signaling may be merge_triangle_idx_indicator parsing. The different candidate index signaling may be a signaling method that does not parse merge_triangle_idx0 or merge_triangle_idx1. This is further explained in Fig. 53.

[0390] FIG. 53 is a diagram showing TPM candidate index signaling according to one embodiment of the present disclosure.

[0391] According to one embodiment of the present disclosure, when using other index signaling described in FIG. 52, the candidate index can be determined based on the merge_triangle_idx_indicator. Also, when using other index signaling, merge_triangle_idx0 or merge_triangle_idx1 may not exist.

[0392] According to one embodiment of the present disclosure, if the maximum number of candidates used in the TPM is equal to the number of partitions in the TPM, the TPM candidate index can be determined based on the merge_triangle_idx_indicator. Additionally, this may be the case for a block using the TPM.

[0393] More specifically, when MaxNumTriangleMergeCand is 2 (or when MaxNumTriangleMergeCand is 2 and MergeTriangleFlag is 1), TPM candidate indices can be determined based on merge_triangle_idx_indicator. In this case, if merge_triangle_idx_indicator is 0, the TPM candidate indices m and n can be set to 0 and 1, respectively, and if merge_triangle_idx_indicator is 1, the TPM candidate indices m and n can be set to 1 and 0, respectively. Alternatively, merge_triangle_idx0 or merge_triangle_idx1, which are values ​​(syntax elements) that can be parsed to be the same as described, can be inferred and set.

[0394] Referring to the method of setting m and n based on merge_triangle_idx0 and merge_triangle_idx1 described in FIG. 47, and referring to FIG. 53, if merge_triangle_idx0 does not exist, if MaxNumTriangleMergeCand is 2 (or MaxNumTriangleMergeCand is 2 and MergeTriangleFlag is 1), if merge_triangle_idx_indicator is 1, the value of merge_triangle_idx0 can be inferred to 1. Otherwise, the value of merge_triangle_idx0 can be inferred to 0. Also, if merge_triangle_idx1 does not exist, merge_triangle_idx1 can be inferred to 0. Therefore, if MaxNumTriangleMergeCand is 2, when merge_triangle_idx_indicator is 0, merge_triangle_idx0 and merge_triangle_idx1 are 0 and 0, respectively, and accordingly, m and n can be 0 and 1, respectively. Also, if MaxNumTriangleMergeCand is 2, when merge_triangle_idx_indicator is 1, merge_triangle_idx0 and merge_triangle_idx1 are 1 and 0, respectively, and accordingly, m and n can be 1 and 0, respectively.

[0395] FIG. 54 is a diagram showing TPM candidate index signaling according to one embodiment of the present disclosure.

[0396] Fig. 47 describes a method for determining TPM candidate indices, but the embodiment in Fig. 54 describes a different determination method and a signaling method. Descriptions that overlap with those previously explained may be omitted. Additionally, m and n can represent candidate indices as described in Fig. 47.

[0397] According to one embodiment of the present disclosure, the smaller value between m and n can be signaled to a preset syntax element among merge_triangle_idx0 and merge_triangle_idx1. Additionally, a value based on the difference between m and n can be signaled to the other one among merge_triangle_idx0 and merge_triangle_idx1. Furthermore, a value indicating the greater or lesser relationship between m and n can be signaled.

[0398] For example, merge_triangle_idx0 can be the smaller of m and n. Also, merge_triangle_idx1 can be a value based on |mn|. merge_triangle_idx1 can be (|mn| - 1). This is because m and n may not be the same. Also, a value representing the relationship between m and n can be merge_triangle_bigger in Fig. 54.

[0399] Using this relationship, m and n can be determined based on merge_triangle_idx0, merge_triangle_idx1, and merge_triangle_bigger. Referring to Fig. 54, different actions can be performed based on the merge_triangle_bigger value. For example, if merge_triangle_bigger is 0, n may be greater than m. In this case, m may be merge_triangle_idx0. Also, n may be (merge_triangle_idx1 + m + 1). Also, if merge_triangle_bigger is 1, m may be greater than n. In this case, n may be merge_triangle_idx0. Also, m may be (merge_triangle_idx1 + n + 1).

[0400] The method in Fig. 54 has the advantage of reducing signaling overhead compared to the method in Fig. 47 when the smaller value between m and n is not 0 (or is larger). For example, in the case where m and n are 3 and 4, respectively, the method in Fig. 47 may require signaling merge_triangle_idx0 and merge_triangle_idx1 as 3 and 3, respectively. However, in the method in Fig. 54, when m and n are 3 and 4, respectively, merge_triangle_idx0 and merge_triangle_idx1 may require signaling as 3 and 0, respectively (additional signaling indicating the greater / lesser relationship may be required). Therefore, when using variable length signaling, it is possible to use fewer bits because the size of the values ​​being encoded and decoded is reduced.

[0401] FIG. 55 is a diagram showing TPM candidate index signaling according to one embodiment of the present disclosure.

[0402] Fig. 47 describes a method for determining TPM candidate indices, but the embodiment in Fig. 55 describes a different determination method and a signaling method. Descriptions that overlap with those previously explained may be omitted. Also, m and n can represent candidate indices as described in Fig. 47.

[0403] According to one embodiment of the present disclosure, a value based on the larger of m and n can be signaled to a preset syntax element among merge_triangle_idx0 and merge_triangle_idx1. Additionally, a value based on the smaller of m and n can be signaled to the other one among merge_triangle_idx0 and merge_triangle_idx1. Furthermore, a value indicating the greater than or less than relationship between m and n can be signaled.

[0404] For example, merge_triangle_idx0 can be based on the larger value between m and n. According to one embodiment, since m and n are not equal, the larger value between m and n will be 1 or greater. Therefore, considering that the larger value between m and n is 0, it can be signaled with fewer bits. For example, merge_triangle_idx0 can be ((larger value between m and n) - 1). In this case, the maximum value of merge_triangle_idx0 can be (MaxNumTriangleMergeCand - 1 - 1) (-1 because it is a value starting from 0, and -1 because the larger value being 0 can be excluded). The maximum value can be used in binarization, and if the maximum value is reduced, fewer bits may be used. Also, merge_triangle_idx1 can be the smaller value between m and n. Also, the maximum value of merge_triangle_idx1 can be merge_triangle_idx0. Therefore, there may be cases where fewer bits are used than setting the maximum value to MaxNumTriangleMergeCand. Also, when merge_triangle_idx0 is 0, that is, when the larger value between m and n is 1, the smaller value between m and n is 0, so there may be no additional signaling. For example, when merge_triangle_idx0 is 0, that is, when the larger value between m and n is 1, the smaller value between m and n can be determined as 0. Alternatively, when merge_triangle_idx0 is 0, that is, when the larger value between m and n is 1, merge_triangle_idx1 can be inferred and determined as 0. Referring to Fig. 22, whether to parse merge_triangle_idx1 can be determined based on merge_triangle_idx0.For example, if merge_triangle_idx0 is greater than 0, merge_triangle_idx1 can be parsed, and if merge_triangle_idx0 is 0, merge_triangle_idx1 can not be parsed.

[0405] In addition, the value representing the relationship between m and n may be merge_triangle_bigger in Fig. 55.

[0406] Using this relationship, m and n can be determined based on merge_triangle_idx0, merge_triangle_idx1, and merge_triangle_bigger. Referring to Fig. 55, other actions can be performed based on the merge_triangle_bigger value. For example, if merge_triangle_bigger is 0, m may be greater than n. In this case, m may be (merge_triangle_idx0 + 1). Also, n may be merge_triangle_idx1. Also, if merge_triangle_bigger is 1, n may be greater than m. In this case, n may be (merge_triangle_idx0 + 1). Also, m may be merge_triangle_idx1. Additionally, if merge_triangle_idx1 does not exist, its value can be inferred as 0.

[0407] The method in Fig. 55 has the advantage of reducing signaling overhead depending on the values ​​of m and n compared to the method in Fig. 47. For example, when m and n are 1 and 0, respectively, in the method in Fig. 47, merge_triangle_idx0 and merge_triangle_idx1 may need to be signaled as 1 and 0, respectively. However, in the method in Fig. 55, when m and n are 1 and 0, respectively, merge_triangle_idx0 and merge_triangle_idx1 may need to be signaled as 0 and 0, respectively, and merge_triangle_idx1 can be inferred without encoding or parsing (in addition, signaling indicating the magnitude relationship may be required). Therefore, when using variable length signaling, it is possible to use fewer bits because the size of the values ​​being encoded and decoded is reduced. Alternatively, when m and n are 2 and 1, respectively, in the method of FIG. 47, merge_triangle_idx0 and merge_triangle_idx1 may be signaled as 2 and 1, respectively, and in the method of FIG. 55, merge_triangle_idx0 and merge_triangle_idx1 may be signaled as 1 and 1, respectively. However, in the method of FIG. 55, since the maximum value of merge_triangle_idx1 is (3-1-1)=1, 1 can be signaled with fewer bits than when the maximum value is large. For example, the method of FIG. 55 may have advantages when the difference between m and n is small, for example, when the difference is 1.

[0408] In addition, while it was explained that merge_triangle_idx0 is referenced to determine whether to parse merge_triangle_idx1 in the syntax structure of Fig. 55, it is also possible to make a decision based on the larger value between m and n. That is, it can be distinguished into cases where the larger value between m and n is 1 or greater and cases where it is not. However, in this case, merge_triangle_bigger parsing may need to occur before the determination of whether to parse merge_triangle_idx1.

[0409] Although the configuration has been described above through specific embodiments, those skilled in the art may make modifications and changes without departing from the spirit and scope of the present disclosure. Accordingly, anything that can be easily inferred by a person skilled in the art from the detailed description and embodiments of the present disclosure is interpreted as falling within the scope of the rights of the present disclosure.

Claims

Claim 1 A video signal encoding device comprises a processor, wherein the processor acquires a bitstream decoded by a decoder using a decoding method, and the decoding method comprises: a step of parsing a first syntax element indicating whether Adaptive Motion Vector Resolution (AMVR) is enabled; a step of parsing a second syntax element indicating whether affine motion compensation is enabled; a step of parsing a third syntax element indicating whether AMVR is enabled for affine motion compensation when the first syntax element indicates that the AMVR is enabled and the second syntax element indicates that the affine motion compensation is enabled; a step of parsing information related to a reference picture list for a current block; and a step of parsing a fourth syntax element indicating whether the motion vector difference and a plurality of control point motion vector differences for reference picture list 1 are set to zero. A video signal encoding device comprising the step of parsing the motion vector predictor index of the reference picture list 1 when the information indicates that a reference picture list other than reference picture list 0 is available, wherein the motion vector predictor index is parsed independently of the fourth syntax element. Claim 2 A video signal encoding device according to claim 1, wherein at least one of the first syntax element, the second syntax element, and the third syntax element is signaled on an SPS (sequence parameter set) RBSP (raw byte sequence payload) syntax. Claim 3 A video signal encoding device according to claim 1, wherein when the affine motion compensation is activated and the AMVR is not activated, the third syntax element is not parsed and the value of the third syntax element is inferred as a value indicating that the AMVR is not activated for the affine motion compensation. Claim 4 A video signal encoding device according to claim 1, wherein when the affine motion compensation is not activated, the third syntax element is not parsed, and the value of the third syntax element is inferred as a value indicating that the AMVR is not activated for the affine motion compensation. Claim 5 A video signal encoding device according to claim 1, wherein the decoding method further comprises: a step of parsing information related to the resolution of motion vector differences when the first syntax element indicates that the AMVR is activated and the fifth syntax element indicates that affine motion compensation is not used in the current block and at least one of a plurality of motion vector differences for the current block is not zero; and a step of modifying a plurality of motion vector differences for the current block based on the information related to the resolution of motion vector differences. Claim 6 A video signal encoding device according to claim 1, wherein the decoding method further comprises: a step of parsing information related to the resolution of motion vector differences when the third syntax element indicates that the AMVR is activated for the affine motion compensation and the fifth syntax element indicates that the affine motion compensation is used for the current block and at least one of the plurality of control point motion vector differences for the current block is not zero; and a step of modifying the plurality of control point motion vector differences for the current block based on the information related to the resolution of the motion vector differences. Claim 7 A video signal decoding device comprising a processor, wherein the processor parses a first syntax element indicating whether Adaptive Motion Vector Resolution (AMVR) is enabled and parses a second syntax element indicating whether affine motion compensation is enabled; if the first syntax element indicates that the AMVR is enabled and the second syntax element indicates that the affine motion compensation is enabled, the processor parses a third syntax element indicating whether the AMVR is enabled for the affine motion compensation; parses information related to a reference picture list for a current block; parses a fourth syntax element indicating whether the motion vector difference and a plurality of control point motion vector differences for reference picture list 1 are set to 0; and if the information indicates that a reference picture list other than reference picture list 0 is available, the processor parses a motion vector predictor index of the reference picture list 1, wherein the motion vector predictor index is parsed independently of the fourth syntax element. Claim 8 A video signal decoding device according to claim 7, wherein at least one of the first syntax element, the second syntax element, and the third syntax element is signaled on the SPS (sequence parameter set) RBSP (raw byte sequence payload) syntax. Claim 9 A video signal decoding device according to claim 7, wherein when the affine motion compensation is activated and the AMVR is not activated, the third syntax element is not parsed and the value of the third syntax element is inferred as a value indicating that the AMVR is not activated for the affine motion compensation. Claim 10 A video signal decoding device according to claim 7, wherein when the affine motion compensation is not activated, the third syntax element is not parsed, and the value of the third syntax element is inferred as a value indicating that the AMVR is not activated for the affine motion compensation. Claim 11 A video signal decoding device according to claim 7, wherein the processor parses information related to the resolution of motion vector differences when the first syntax element indicates that the AMVR is activated and the fifth syntax element indicates that affine motion compensation is not used in the current block and at least one of the plurality of motion vector differences for the current block is not zero, and modifies the plurality of motion vector differences for the current block based on the information related to the resolution of the motion vector differences. Claim 12 A video signal decoding device according to claim 7, wherein the processor parses information related to the resolution of the motion vector difference when the third syntax element indicates that the AMVR is activated for the affine motion compensation and the fifth syntax element indicates that the affine motion compensation is used for the current block and at least one of the plurality of control point motion vector differences for the current block is not zero, and modifies the plurality of control point motion vector differences for the current block based on the information related to the resolution of the motion vector difference. Claim 13 In a computer-readable non-transient storage medium storing a bitstream, the bitstream is decoded by a decoding method, and the decoding method comprises: a step of parsing a first syntax element indicating whether Adaptive Motion Vector Resolution (AMVR) is enabled; a step of parsing a second syntax element indicating whether affine motion compensation is enabled; a step of parsing a third syntax element indicating whether AMVR is enabled for affine motion compensation when the first syntax element indicates that the AMVR is enabled and the second syntax element indicates that the affine motion compensation is enabled; a step of parsing information related to a reference picture list for a current block; and a step of parsing a fourth syntax element indicating whether the motion vector difference and a plurality of control point motion vector differences for reference picture list 1 are set to zero. A storage medium comprising the step of parsing the motion vector predictor index of the reference picture list 1 when the information indicates that a reference picture list other than reference picture list 0 is available, wherein the motion vector predictor index is parsed independently of the fourth syntax element. Claim 14 In claim 13, a storage medium wherein at least one of the first syntax element, the second syntax element, and the third syntax element is signaled on the SPS (sequence parameter set) RBSP (raw byte sequence payload) syntax. Claim 15 A storage medium according to claim 13, wherein when the affine motion compensation is enabled and the AMVR is not enabled, the third syntax element is not parsed and the value of the third syntax element is inferred as a value indicating that the AMVR is not enabled for the affine motion compensation. Claim 16 A storage medium according to claim 13, wherein when the affine motion compensation is not activated, the third syntax element is not parsed, and the value of the third syntax element is inferred as a value indicating that the AMVR is not activated for the affine motion compensation. Claim 17 In claim 13, the decoding method further comprises: a step of parsing information related to the resolution of motion vector differences when the first syntax element indicates that the AMVR is activated and the fifth syntax element indicates that affine motion compensation is not used in the current block and at least one of a plurality of motion vector differences for the current block is not zero; and a step of modifying a plurality of motion vector differences for the current block based on the information related to the resolution of motion vector differences, a storage medium. Claim 18 A decoding method for a video signal, wherein the decoding method comprises: a step of parsing a first syntax element indicating whether Adaptive Motion Vector Resolution (AMVR) is activated; a step of parsing a second syntax element indicating whether affine motion compensation is activated; a step of parsing a third syntax element indicating whether AMVR is activated for affine motion compensation when the first syntax element indicates that the AMVR is activated and the second syntax element indicates that the affine motion compensation is activated; a step of parsing information related to a reference picture list for a current block; a step of parsing a fourth syntax element indicating whether a motion vector difference and a plurality of control point motion vector differences for reference picture list 1 are set to 0; and a step of parsing a motion vector predictor index of reference picture list 1 when the information indicates that a reference picture list other than reference picture list 0 is available, wherein the motion vector predictor index is parsed independently of the fourth syntax element. Claim 19 delete Claim 20 delete Claim 21 delete Claim 22 delete