Video signal processing method and apparatus utilizing adaptive motion vector resolution

By managing adaptive motion vector resolution and affine motion compensation through flag purging and correction, the method improves coding efficiency in video signal processing, addressing inefficiencies in existing methods.

JP2026083334APending Publication Date: 2026-05-19WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
Filing Date
2026-03-11
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video signal processing methods lack efficiency in coding, particularly in handling adaptive motion vector resolution and affine motion compensation, leading to suboptimal compression and decoding performance.

Method used

The method involves purging certain flags from the bitstream to manage adaptive motion vector resolution and affine motion compensation, determining their availability, and correcting motion vector differences based on these flags, while signaling relevant information in units such as Coding Tree Units, slices, tiles, or sequences.

Benefits of technology

This approach enhances coding efficiency by optimizing the use of adaptive motion vector resolution and affine motion compensation, resulting in improved video signal processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083334000001_ABST
    Figure 2026083334000001_ABST
Patent Text Reader

Abstract

The present invention provides a processing method and apparatus for improving the coding efficiency of video signals. [Solution] The processing method involves purging an AMVR (Adaptive Motion Vector Resolution) enabled flag from the bitstream to indicate whether adaptive motion vector difference resolution is used or not, purging an affine enabled flag from the bitstream to indicate whether affine motion compensation is available or not, determining whether affine motion compensation is available or not based on the affine enabled flag, determining whether adaptive motion vector difference resolution is used or not based on the AMVR enabled flag if affine motion compensation is available, and purging an affine AMVR enabled flag from the bitstream to indicate whether adaptive motion vector difference resolution is available for affine motion compensation or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0003] , , ,

[0004]

[0001] The present invention relates to a method and apparatus for processing video signals, and more particularly, to a method and apparatus for encoding or decoding video signals based on intra prediction.

Background Art

[0002] Compression encoding means a series of signal processing techniques for transmitting digitized information via a communication line or storing it in a form suitable for a storage medium. The targets of compression encoding include audio, video, characters, etc., and in particular, the technique of performing compression encoding on video is called video compression. Compression encoding of video signals is performed by removing redundant information in consideration of spatial correlation, temporal correlation, probabilistic correlation, etc. However, due to the recent development of various media and data transmission media, there is a need for more efficient video signal processing methods and apparatuses.

Summary of the Invention

Problems to be Solved by the Invention

[0003] An object of the present disclosure is to increase the coding efficiency of video signals.

Means for Solving the Problems

[0004] A method for processing a video signal according to one embodiment of the present disclosure is characterized by comprising the steps of: purging an AMVR (Adaptive Motion Vector Resolution) enabled flag (sps_amvr_enabled_flag) from a bitstream indicating whether or not adaptive motion vector difference resolution is used; purging an affine enabled flag (sps_affine_enabled_flag) from a bitstream indicating whether or not affine motion compensation is available; determining whether or not affine motion compensation is available based on the affine enabled flag (sps_affine_enabled_flag); if affine motion compensation is available, determining whether or not adaptive motion vector difference resolution is used based on the AMVR enabled flag (sps_amvr_enabled_flag); and if adaptive motion vector difference resolution is used, purging an affine AMVR enabled flag (sps_affine_amvr_enabled_flag) from the bitstream indicating whether or not adaptive motion vector difference resolution is available for affine motion compensation.

[0005] A method for processing a video signal according to one embodiment of the present disclosure is characterized in that one of the AMVR-enabled flag (sps_amvr_enabled_flag), the affine-enabled flag (sps_affine_amvr_enabled_flag), or the affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) is signaled in one of the Coding Tree Unit, slice, tile, tile group, picture, or sequence units.

[0006] In a method for processing a video signal according to one embodiment of the present disclosure, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) implies that adaptive motion vector difference resolution is unavailable for affine motion compensation when affine motion compensation is available and adaptive motion vector difference resolution is not used.

[0007] In a method for processing a video signal according to one embodiment of the present disclosure, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) is characterized in that, when affine motion compensation is unavailable, adaptive motion vector difference resolution is inferred to be unavailable for affine motion compensation.

[0008] A method for processing a video signal according to one embodiment of the present disclosure is characterized by comprising the steps of: purging information about the resolution of motion vector differences from the bitstream if the AMVR-enabled flag (sps_amvr_enabled_flag) indicates the use of adaptive motion vector difference resolution, the inter-affine flag (inter_affine_flag) obtained from the bitstream indicates that affine motion compensation is not used for the current block, and at least one of a plurality of motion vector differences for the current block is not zero; and correcting a plurality of motion vector differences for the current block based on the information about the resolution of motion vector differences.

[0009] A method for processing a video signal according to one embodiment of the present disclosure is characterized by comprising the steps of: purging information about the resolution of the motion vector difference from the bitstream, unless an affine AMVR-enabled flag indicates that adaptive motion vector difference resolution for affine motion compensation is available; an inter-affine flag obtained from the bitstream indicates the use of affine motion compensation for the current block; and at least one of a plurality of control point motion vector differences for the current block is not zero; and correcting the plurality of control point motion vector differences for the current block based on the information about the resolution of the motion vector difference.

[0010] A method for processing a video signal according to one embodiment of the present disclosure is characterized by comprising the steps of: obtaining information about a reference picture list (inter_pred_idc) for the current block; purging a motion vector predictor index (mvp_l1_flag) of a first reference picture list (list 1) from the bitstream if the information about a reference picture list (inter_pred_idc) indicates that it does not use only the 0th reference picture list (list 0); generating candidate motion vector predictors; obtaining a motion vector predictor from the candidate motion vector predictors based on the motion vector predictor index; and predicting the current block based on the motion vector predictor.

[0011] A method for processing a video signal according to one embodiment of the present disclosure further includes the step of obtaining a motion vector difference zero flag (mvd_l1_zero_flag) from a bitstream indicating whether the motion vector difference and the motion vector differences of multiple control points are set to zero for a first reference picture list, and the step of purging a motion vector predictor index (mvp_l1_flag) is characterized in that the motion vector difference zero flag (mvd_l1_zero_flag) is 1 and the information about the reference picture list (inter_pred_idc) indicates that both the 0th reference picture list and the 1st reference picture list are used.

[0012] A method for processing a video signal according to one embodiment of the present disclosure is characterized by comprising the steps of: purging from the bitstream in sequence units first information (six_minus_max_num_merge_cand) relating to the maximum number of candidates for merge motion vector prediction; obtaining the maximum number of merge candidates based on the first information; purging from the bitstream second information indicating whether or not a block can be partitioned for inter prediction; and, if the second information indicates 1 and the maximum number of merge candidates is greater than 2, purging from the bitstream third information relating to the maximum number of merge mode candidates for a partitioned block.

[0013] A method for processing a video signal according to one embodiment of the present disclosure further comprises the steps of: obtaining the maximum number of merge mode candidates for a partitioned block by subtracting the third information from the maximum number of merge candidates if the second information indicates 1 and the maximum number of merge candidates is greater than or equal to 3; setting the maximum number of merge mode candidates for a partitioned block to 2 if the second information indicates 1 and the maximum number of merge candidates is 2; and setting the maximum number of merge mode candidates for a partitioned block to 0 if the second information is 0 or the maximum number of merge candidates is 1.

[0014] An apparatus for processing a video signal according to one embodiment of the present disclosure includes a processor and memory, wherein the processor, based on instructions stored in memory, purges an AMVR (Adaptive Motion Vector Resolution) enabled flag (sps_amvr_enabled_flag) indicating whether or not adaptive motion vector difference resolution is used from the bitstream, purges an affine enabled flag (sps_affine_enabled_flag) indicating whether or not affine motion compensation is available from the bitstream, determines whether or not affine motion compensation is available based on the affine enabled flag (sps_affine_enabled_flag), if affine motion compensation is available, determines whether or not adaptive motion vector difference resolution is used based on the AMVR enabled flag (sps_amvr_enabled_flag), and if adaptive motion vector difference resolution is used, purges an affine AMVR enabled flag (sps_affine_amvr_enabled_flag) indicating whether or not adaptive motion vector difference resolution is available for affine motion compensation from the bitstream.

[0015] In a device for processing video signals according to one embodiment of the present disclosure, one of the AMVR-enabled flag (sps_amvr_enabled_flag), the affine-enabled flag (sps_affine_amvr_enabled_flag), or the affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) is signaled in one of the following units: Coding Tree Unit, slice, tile, tile group, picture, or sequence.

[0016] In a device for processing video signals according to one embodiment of the present disclosure, if affine motion compensation is available and adaptive motion vector difference resolution is not used, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) implies that adaptive motion vector difference resolution is unavailable for affine motion compensation.

[0017] In a device for processing video signals according to one embodiment of the present disclosure, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) is characterized in that, when affine motion compensation is unavailable, adaptive motion vector difference resolution is inferred to be unavailable for affine motion compensation.

[0018] In a device for processing a video signal according to one embodiment of the present disclosure, the processor, based on instructions stored in memory, if the AMVR-enabled flag (sps_amvr_enabled_flag) indicates the use of adaptive motion vector difference resolution, if the inter-affine flag (inter_affine_flag) obtained from the bitstream indicates that affine motion compensation is not used for the current block, and if at least one of the multiple motion vector differences for the current block is not zero, the processor purges information regarding the resolution of the motion vector differences from the bitstream and corrects the multiple motion vector differences for the current block based on the information regarding the resolution of the motion vector differences.

[0019] In a device for processing a video signal according to one embodiment of the present disclosure, the processor, based on instructions stored in memory, is characterized in that if an affine AMVR-enabled flag indicates that adaptive motion vector difference resolution is available for affine motion compensation, if an inter-affine flag obtained from the bitstream indicates the use of affine motion compensation for the current block, and at least one of the multiple control point motion vector differences for the current block is not zero, it purges information regarding the resolution of the motion vector differences from the bitstream and modifies the multiple control point motion vector differences for the current block based on the information regarding the resolution of the motion vector differences.

[0020] In a device for processing a video signal according to one embodiment of the present disclosure, the processor obtains information about the reference picture list (inter_pred_idc) for the current block based on an instruction word stored in memory, and if the information about the reference picture list (inter_pred_idc) indicates that only the 0th reference picture list (list 0) is used, the processor purges the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1) from the bitstream, generates candidate motion vector predictors, obtains a motion vector predictor from the candidate motion vector predictors based on the motion vector predictor index, and predicts the current block based on the motion vector predictor.

[0021] In a device for processing a video signal according to one embodiment of the present disclosure, the processor obtains a motion vector difference zero flag (mvd_l1_zero_flag) from the bitstream based on an instruction word stored in memory, indicating whether or not the motion vector difference and the motion vector differences of multiple control points are set to zero for the first reference picture list, and purges the motion vector predictor index (mvp_l1_flag) regardless of whether the motion vector difference zero flag (mvd_l1_zero_flag) is 1 and the information regarding the reference picture list (inter_pred_idc) indicates that both the 0th reference picture list and the 1st reference picture list are used.

[0022] An apparatus for processing a video signal according to one embodiment of the present disclosure includes a processor and a memory, wherein the processor, based on instructions stored in the memory, purges from the bitstream in sequence units for first information (six_minus_max_num_merge_cand) relating to the maximum number of candidates for merge motion vector prediction, obtains the maximum number of merge candidates based on the first information, purges from the bitstream for second information indicating whether or not the block can be partitioned for inter prediction, and if the second information indicates 1 and the maximum number of merge candidates is greater than 2, purges from the bitstream for third information relating to the maximum number of merge mode candidates for the partitioned block.

[0023] In an apparatus for processing a video signal according to one embodiment of the present disclosure, the processor, based on an instruction word stored in memory, if the second information indicates 1 and the maximum number of merge candidates is greater than or equal to 3, subtracts the third information from the maximum number of merge candidates to obtain the maximum number of merge mode candidates for the partitioned block; if the second information indicates 1 and the maximum number of merge candidates is 2, sets the maximum number of merge mode candidates for the partitioned block to 2; and if the second information is 0 or the maximum number of merge candidates is 1, sets the maximum number of merge mode candidates for the partitioned block to 0.

[0024] A method for processing a video signal according to one embodiment of the present disclosure includes the steps of: generating an AMVR (Adaptive Motion Vector Resolution) enabled flag (sps_amvr_enabled_flag) indicating whether adaptive motion vector difference resolution is used; generating an affine enabled flag (sps_affine_enabled_flag) indicating whether affine motion compensation is available; determining whether affine motion compensation is available based on the affine enabled flag (sps_affine_enabled_flag); and, if affine motion compensation is available, determining whether adaptive motion vector difference resolution is used based on the AMVR enabled flag (sps_amvr_enabled_flag). The method is characterized by including the steps of: generating an affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) indicating whether adaptive motion vector difference resolution is available for affine motion compensation if adaptive motion vector difference resolution is used; and generating a bitstream by entropy coding the AMVR-enabled flag (sps_amvr_enabled_flag), the affine-enabled flag (sps_affine_enabled_flag), and the AMVR-enabled flag (sps_amvr_enabled_flag).

[0025] A method for processing a video signal according to an embodiment of the present disclosure includes generating first information (six_minus_max_num_merge_cand) regarding the maximum number of candidates for merge motion vector prediction based on the maximum number of merge candidates, generating second information indicating whether a block can be partitioned for inter prediction, and if the second information indicates 1 and the maximum number of merge candidates is greater than 2, generating third information regarding the maximum number of merge mode candidates for the partitioned block, and entropy coding the first information (six_minus_max_num_merge_cand), the second information, and the third information to generate a bitstream in a sequence unit.

Effect of the Invention

[0026] According to an embodiment of the present disclosure, the coding efficiency of a video signal is improved.

Brief Description of the Drawings

[0027] [Figure 1] It is a schematic block diagram of a video signal encoder device according to an embodiment of the present disclosure. [Figure 2] It is a schematic block diagram of a video signal decoder device according to an embodiment of the present disclosure. [Figure 3] It is a diagram showing an embodiment of the present disclosure for dividing a coding unit. [Figure 4] It is a diagram showing an embodiment of a method for hierarchically showing the division structure of FIG. 3. [Figure 5] It is a diagram showing a further embodiment of the present disclosure for dividing a coding unit. [Figure 6] It is a diagram showing a method for obtaining reference pixels for intra prediction. [Figure 7] It is a diagram showing an embodiment of a prediction mode used for intra prediction within a screen. [Figure 8] It is a diagram showing inter prediction according to an embodiment of the present disclosure. [Figure 9]This figure shows a motion vector signaling method according to one embodiment of the present disclosure. [Figure 10] This figure shows a motion vector difference sequence according to one embodiment of the present disclosure. [Figure 11] This figure shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure. [Figure 12] This figure shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure. [Figure 13] This figure shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure. [Figure 14] This figure shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure. [Figure 15] This figure shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure. [Figure 16] This figure shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure. [Figure 17] This figure shows an affine motion prediction according to one embodiment of the present disclosure. [Figure 18] This figure shows an affine motion prediction according to one embodiment of the present disclosure. [Figure 19] This is an equation showing a motion vector field according to one embodiment of the present disclosure. [Figure 20] This figure shows an affine motion prediction according to one embodiment of the present disclosure. [Figure 21] This is an equation showing a motion vector field according to one embodiment of the present disclosure. [Figure 22]This figure shows an affine motion prediction according to one embodiment of the present disclosure. [Figure 23] This figure shows a mode of affine motion prediction according to one embodiment of the present disclosure. [Figure 24] This figure shows a mode of affine motion prediction according to one embodiment of the present disclosure. [Figure 25] This figure shows an affine motion predictor derivation according to one embodiment of the present disclosure. [Figure 26] This figure shows an affine motion predictor derivation according to one embodiment of the present disclosure. [Figure 27] This figure shows an affine motion predictor derivation according to one embodiment of the present disclosure. [Figure 28] This figure shows an affine motion predictor derivation according to one embodiment of the present disclosure. [Figure 29] This figure shows a method for generating a control point motion vector according to one embodiment of the present disclosure. [Figure 30] This figure shows the method for determining the motion vector difference as described in Figure 29. [Figure 31] This figure shows a method for generating a control point motion vector according to one embodiment of the present disclosure. [Figure 32] This figure shows the method for determining the motion vector difference as described in Figure 31. [Figure 33] This figure shows a motion vector difference sequence according to one embodiment of the present disclosure. [Figure 34] This figure shows the structure of upper-level signaling according to one embodiment of the present disclosure. [Figure 35]This figure shows the structure of a coding unit sequence according to one embodiment of the present disclosure. [Figure 36] This figure shows the structure of upper-level signaling according to one embodiment of the present disclosure. [Figure 37] This figure shows the structure of upper-level signaling according to one embodiment of the present disclosure. [Figure 38] This figure shows the structure of a coding unit sequence according to one embodiment of the present disclosure. [Figure 39] This figure shows the setting of MVD base values ​​according to one embodiment of the present disclosure. [Figure 40] This figure shows the setting of MVD base values ​​according to one embodiment of the present disclosure. [Figure 41] This figure shows the syntax of AMVR related to one embodiment of the present disclosure. [Figure 42] This figure shows the syntax related to interpretation according to one embodiment of the present disclosure. [Figure 43] This figure shows the syntax related to interpretation according to one embodiment of the present disclosure. [Figure 44] This figure shows the syntax related to interpretation according to one embodiment of the present disclosure. [Figure 45] This figure shows the syntax related to inter prediction according to one embodiment of the present disclosure. [Figure 46] This figure shows a triangle partitioning mode according to one embodiment of the present disclosure. [Figure 47] This figure shows a merged data column according to one embodiment of the present disclosure. [Figure 48] This figure shows upper-level signaling according to one embodiment of the present disclosure. [Figure 49] This figure shows the maximum candidate number used in a TPM according to one embodiment of the present disclosure. [Figure 50]This figure shows upper-level signaling related to TPM according to one embodiment of the present disclosure. [Figure 51] This figure shows the maximum candidate number used in a TPM according to one embodiment of the present disclosure. [Figure 52] This figure shows a syntax element related to TPM according to one embodiment of the present disclosure. [Figure 53] This figure shows the signaling of the TPM candidate index according to one embodiment of the present disclosure. [Figure 54] This figure shows the signaling of the TPM candidate index according to one embodiment of the present disclosure. [Figure 55] This figure shows the signaling of the TPM candidate index according to one embodiment of the present disclosure. [Modes for carrying out the invention]

[0028] The terminology used herein has been selected as widely used and general terms as possible, taking into account the function of the present invention; however, this may vary depending on the intent of the articulators, conventions, or the emergence of new technologies. In addition, in certain cases, the applicant has arbitrarily selected some terms, in which case their meaning will be described in the section describing the form of implementation of the invention. Therefore, it is important to clarify that the terminology used herein is not merely a set of names, but should be interpreted based on the substantive meaning of the term and the overall content of this specification.

[0029] In this disclosure, the following terms are interpreted according to the following criteria, but any terms not included are also interpreted in the same manner. Coding may be interpreted as encoding or decoding in some cases, and information is a term that includes values, parameters, coefficients, elements, etc., and may be interpreted differently in some cases, and is not limited to this disclosure. Unit is used to mean a basic unit of picture processing or a specific location of a picture, and may be used interchangeably with terms such as block, partition, or region in some cases. In this specification, unit is used as a concept that includes coding units, prediction units, and transformation units.

[0030] Figure 1 is a schematic block diagram of a video signal encoding device according to one embodiment of the present disclosure. Referring to Figure 1, the encoding device 100 of the present disclosure broadly includes a conversion unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse conversion unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.

[0031] The conversion unit 110 converts the pixel values ​​of the input video signal to obtain conversion coefficient values. For example, the discrete cosine transform (DCT) or wavelet transform may be used. In particular, the discrete cosine transform converts the input picture signal by dividing it into blocks of a fixed size. In the conversion, the coding efficiency may differ depending on the distribution and characteristics of the values ​​within the conversion domain.

[0032] The quantization unit 115 quantizes the conversion coefficient values ​​output in the conversion unit 110. The inverse quantization unit 120 inversely quantizes the conversion coefficient values, and the inverse conversion unit 125 uses the inversely quantized conversion coefficient values ​​to restore the original pixel values.

[0033] The filtering unit 130 performs calculations to improve the quality of the restored picture. These include, for example, a deblocking filter and an adaptive loop filter. The filtered picture is saved in the Decoded Picture Buffer 156 for output or use as a reference picture.

[0034] To improve coding efficiency, instead of directly coding the picture signal, a method is used in which the picture is predicted using an already coded region in the prediction unit 150, and the restored picture is obtained by adding the residual value between the original picture and the predicted picture to the predicted picture. The intra-prediction unit 152 performs in-screen prediction within the current picture, and the inter-prediction unit 154 predicts the current picture using a reference picture stored in the decoded picture buffer 156. The intra-prediction unit 152 performs in-screen prediction from the restored region within the current picture and transmits the in-screen coding information to the entropy coding unit 160. The inter-prediction unit 154 may further include a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains the motion vector value of the current region by referring to a specific restored region. The motion estimation unit 154a transmits position information of the reference region (reference frame, motion vector, etc.) to the entropy coding unit 160 so that it can be included in the bitstream. The motion compensation unit 154b performs inter-screen motion compensation using the motion vector values ​​transmitted from the motion estimation unit 154a.

[0035] The entropy coding unit 160 generates a video signal bitstream by entropy coding the quantized conversion coefficients, inter-screen coding information, intra-screen coding information, and reference region information input from the inter-prediction unit 154. Here, the entropy coding unit 160 uses variable length coding (VLC) and arithmetic coding. Variable length coding (VLC) converts input symbols into a sequence of codewords, but the length of the codewords is variable. For example, frequently occurring symbols are represented by short codewords, and less frequently occurring symbols are represented by long codewords. As a variable length coding method, context-based adaptive variable length coding (CAVLC) is used. Arithmetic coding converts a sequence of data symbols into a single prime number, but arithmetic coding obtains the optimal number of prime bits necessary to represent each symbol. As an arithmetic coding method, context-based adaptive binary arithmetic coding (CABAC) is used.

[0036] The generated bitstream is encapsulated in Network Abstraction Layer (NAL) units as its basic units. Each NAL unit contains an encoded slice segment, which consists of an integer number of coding tree units. To decode the bitstream with a video decoder, the bitstream should first be separated into NAL units, and then each separated NAL unit should be decoded.

[0037] Figure 2 is a schematic block diagram of a video signal decoding device 200 according to an embodiment of the present disclosure. Referring to Figure 2, the decoding device 200 of the present disclosure broadly includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 225, a filtering unit 230, and a prediction unit 250.

[0038] The entropy decoding unit 210 entropy-decodes the video signal bitstream and extracts conversion coefficients, motion information, etc., for each region. The inverse quantization unit 220 inversely quantizes the entropy-decoded conversion coefficients, and the inverse conversion unit 225 uses the inversely quantized conversion coefficients to reconstruct the original pixel values.

[0039] Meanwhile, the filtering unit 230 improves image quality by filtering the picture. This includes a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is either output or stored in the Decoded Picture Buffer 256 for use as a reference picture for the next frame.

[0040] Furthermore, the prediction unit 250 of this disclosure includes an intra-prediction unit 252 and an inter-prediction unit 254, and reconstructs the predicted picture by utilizing the encoding type decoded via the entropy decoding unit 210 described above, the conversion coefficients for each region, motion information, etc.

[0041] In connection with this, the intra prediction unit 252 performs in-screen prediction from the currently decoded samples in the picture. The inter prediction unit 254 generates a predicted picture using the reference picture and motion information stored in the composite picture buffer 256. The inter prediction unit 254 further includes a motion estimation unit 254a and a motion compensation unit 254b. The motion estimation unit 254a acquires a motion vector indicating the positional relationship between the current block and the reference block of the reference picture used for coding, and transmits it to the motion compensation unit 254b.

[0042] A video frame is generated by adding the predicted values ​​output from the intra-prediction unit 252 or the inter-prediction unit 254 and the pixel values ​​output from the inverse conversion unit 225.

[0043] In the following, a method for dividing the coding unit, prediction unit, etc., in the operation of the encoding device 100 and the decoding device 200 will be described with reference to Figures 3 to 5.

[0044] A coding unit is a fundamental unit for processing a picture in the video signal processing steps described above, such as intra / inter-frame prediction, transform, quantization, and / or entropy coding. The size of the coding unit used to code a single picture is not fixed. A coding unit has a rectangular shape, and one coding unit can be further divided into several coding units.

[0045] Figure 3 shows an embodiment of the present disclosure in which a coding unit is divided. For example, one coding unit having a size of 2N × 2N is further divided into four coding units having a size of N × N. Such division of coding units is performed recursively, but not all coding units need to be divided into the same form. However, for convenience in the coding and processing process, there may be limitations on the size of the maximum coding unit and / or the size of the minimum coding unit.

[0046] For each coding unit, information is stored indicating whether or not the coding unit can be split. Figure 4 shows one example of a method for hierarchically representing the coding unit splitting structure shown in Figure 3 using a flag value. The information indicating whether a coding unit can be split is assigned a value of "1" if the unit has been split, and "0" if it has not been split. As shown in Figure 4, if the value of the flag indicating whether or not it can be split is 1, the coding unit corresponding to the node is further divided into four coding units, and if it is 0, the processing process for the coding unit is performed without further division.

[0047] The coding unit structure described above is represented using a recursive tree structure. That is, with one picture or the largest coding unit as the root, each coding unit that is divided into other coding units will have as many child nodes as there are coding units it is divided into. Thus, coding units that are not divided further become leaf nodes. Assuming that only square divisions are possible for a single coding unit, one coding unit can be divided into a maximum of four other coding units, so the tree representing coding units becomes a quad tree.

[0048] In an encoder, the optimal coding unit size is selected based on the characteristics of the video picture (e.g., resolution) or considering coding efficiency, and information related to this or information that guides this selection is included in the bitstream. For example, the maximum coding unit size and the maximum tree depth are defined. When performing a square partition, the height and width of a coding unit are half the height and width of the parent node's coding unit, so the minimum coding unit size can be determined using the information mentioned above. Alternatively, the minimum coding unit size and the maximum tree depth can be defined in advance and used to derive the maximum coding unit size. In a square partition, the unit size changes to a form that is a multiple of 2, so the actual coding unit size can be expressed as a base-2 logarithmic value to improve transmission efficiency.

[0049] The decoder obtains information indicating whether or not the coding unit is currently divided. Efficiency can be improved by ensuring that such information is obtained (transmitted) only under specific conditions. For example, the conditions under which the coding unit can be divided are that the sum of the sizes of the current coding units at the current position is less than the size of the picture, and the size of the current units is greater than the size of a preset minimum coding unit. Therefore, information indicating whether the coding unit is currently divided can only be obtained in these cases.

[0050] If the information indicates that the coding unit is divided, the size of the divided coding unit will be half the size of the current coding unit, and it will be divided into four square coding units based on the current processing position. The above process is repeated for each processed coding unit.

[0051] Figure 5 shows a further embodiment of the disclosure for dividing a coding unit. According to a further embodiment of the disclosure, the quad-tree coding unit described above is further divided into a binary tree structure of horizontal or vertical division. That is, a square quad-tree division is first applied to the root coding unit, and then a rectangular binary tree division is further applied to the leaf nodes of the quad-tree. According to one embodiment, the binary tree division is a symmetrical horizontal division or a symmetrical vertical division, but the disclosure is not limited thereto.

[0052] At each partition node of the binary tree, a flag is further signaled to indicate the type of partition (i.e., horizontal or vertical partition). In one embodiment, a flag value of "0" indicates a horizontal partition, and a flag value of "1" indicates a vertical partition.

[0053] However, in the embodiments of this disclosure, the method of dividing the coding unit is not limited to the method described above, and asymmetrical horizontal / vertical division, a triple tree divided into three rectangular coding units, etc., may also be applied.

[0054] Picture prediction (motion compensation) for coding is performed on coding units that cannot be further divided (i.e., leaf nodes in the coding unit tree). The basic unit for performing such predictions is referred to below as a prediction unit or prediction block.

[0055] Hereinafter, the term "unit" as used herein is used as a substitute for the prediction unit, which is the basic unit for making predictions. However, this disclosure is not limited thereto, and in a broader sense, it should be understood as a concept that includes the coding unit.

[0056] To reconstruct the current unit being decoded, the current picture containing the current unit or the decoded portion of another picture is used. A picture (slice) that uses only the current picture for reconstruction, i.e., performs only in-screen prediction, is called an intra-picture or I-picture (slice), while a picture (slice) that can perform both in-screen and inter-screen prediction is called an inter-picture (slice). Among inter-pictures (slice), a picture (slice) that uses up to one motion vector and reference index to predict each unit is called a predictive picture or P-picture (slice), and a picture (slice) that uses up to two motion vectors and reference indexes is called a bi-predictive picture or B-picture (slice).

[0057] The intra-prediction unit performs intra-prediction, which predicts the pixel value of the target unit from the restored area within the current picture. For example, it predicts the pixel value of the current unit from the restored pixels of units located to the left and / or top of the current unit. In this case, the units located to the left of the current unit include the adjacent units to the left of the current unit, the upper left unit, and the lower left unit. Also, the units located at the top of the current unit include the adjacent upper unit, the upper left unit, and the upper right unit.

[0058] Meanwhile, the inter-prediction unit performs inter-frame prediction, which predicts the pixel values ​​of the target unit using information from other restored pictures that are not the current picture. In this case, the picture used for prediction is called the reference picture. During the inter-frame prediction process, which reference region to use to predict the current unit is indicated using an index that shows the reference picture containing the relevant reference region, as well as motion vector information.

[0059] Inter-screen prediction includes L0 prediction, L1 prediction, and bi-prediction. L0 prediction uses one reference picture contained in L0 (the 0th reference picture list), and L1 prediction uses one reference picture contained in L1 (the 1st reference picture list). For this, one set of motion information (e.g., motion vector and reference picture index) is required. In the bi-prediction method, up to two reference regions are used, but these two reference regions may reside in the same reference picture or in different pictures. In other words, in the bi-prediction method, up to two sets of motion information (e.g., motion vector and reference picture index) are used, but the two motion vectors may correspond to the same reference picture index or to different reference picture indices. In this case, the reference picture is displayed (or output) both before and after the current picture in terms of time.

[0060] The reference unit of the current unit is obtained using the motion vector and the reference picture index. The reference unit resides within the reference picture that has the reference picture index. The pixel value or interpolated value of the unit identified by the motion vector is used as the predictor of the current unit. For motion prediction with sub-pel pixel accuracy, for example, an 8-tab interpolation filter is used for the luminance signal and a 4-tab interpolation filter is used for the chrominance signal. However, the interpolation filter for sub-pel motion prediction is not limited to these. In this way, motion compensation is performed to predict the texture of the current unit from a previously decoded picture using motion information.

[0061] The in-screen prediction method according to the embodiment of this disclosure will be described in more detail below with reference to Figures 6 and 7. As described above, the intra-prediction unit uses adjacent pixels located to the left and / or top of the current unit as reference pixels to predict the pixel value of the current unit.

[0062] As shown in Figure 6, if the current unit size is N×N, the reference pixels are set using up to 4N+1 adjacent pixels located on the left and / or top edge of the current unit. If at least some of the adjacent pixels to be used as reference pixels have not yet been restored, the intra-prediction unit performs a reference sample padding process according to a pre-set rule to acquire reference pixels. The intra-prediction unit also performs a reference sample filtering process to reduce the error of the in-screen prediction. That is, it filters the adjacent pixels and / or the pixels acquired through the reference sample padding process to acquire reference pixels. The intra-prediction unit uses the reference pixels thus acquired to predict the pixels of the current unit.

[0063] Figure 7 shows an example of a prediction mode used for in-screen prediction. For in-screen prediction, in-screen prediction mode information indicating the in-screen prediction direction is signaled. If the current unit is an in-screen prediction unit, the video signal decoding device extracts the current unit's in-screen prediction mode information from the bitstream. The intra-prediction unit of the video signal decoding device performs in-screen prediction for the current unit based on the extracted in-screen prediction mode information.

[0064] According to one embodiment of this disclosure, the in-screen prediction mode includes a total of 67 modes. Each in-screen prediction mode is indicated by a preset index (i.e., an intra-mode index). For example, as shown in Figure 7, intra-mode index 0 indicates the planar mode, intra-mode index 1 indicates the DC mode, and intra-mode indices 2 to 66 each indicate different directional modes (i.e., angular modes). The in-screen prediction unit determines the reference pixels and / or interpolated reference pixels to be used for in-screen prediction of the current unit based on the in-screen prediction mode information of the current unit. If the intra-mode index indicates a specific directional mode, the reference pixels or interpolated reference pixels corresponding to that specific direction from the current pixel of the current unit are used for prediction of the current pixel. Thus, different sets of reference pixels and / or interpolated reference pixels are used for in-screen prediction depending on the in-screen prediction mode.

[0065] After the current block's in-screen prediction is performed using the reference pixel and in-screen prediction mode information, the video signal decoding device adds the residual signal of the current unit obtained from the inverse transformer to the in-screen prediction value of the current unit to restore the pixel value of the current unit.

[0066] Figure 8 shows an inter prediction according to one embodiment of the present disclosure. As mentioned above, when encoding and decoding a picture or block, predictions are made based on other pictures or blocks. In other words, encoding and decoding can be done based on similarity with other pictures or blocks. Parts similar to other pictures or blocks can be encoded and decoded using signaling that was omitted in the current picture or block, which will be explained further below. Block-level predictions can be made.

[0067] Referring to Figure 8, the Reference picture is on the left and the Current picture is on the right. The Current picture, or a part of the Current picture, is predicted using its similarity to the Reference picture or a part of the Reference picture. If the rectangles shown with solid lines in the Current picture in Figure 8 represent the blocks currently being encoded and decoded, then the Current block is predicted from the rectangles shown with dotted lines in the Reference picture. In this process, there is information that indicates the block that the Current block should refer to (the Reference block). This information can be directly signaled, or it can be generated by a certain convention to reduce signaling overhead. The information that indicates the block that the Current block should refer to includes a motion vector. This is a vector that shows the relative position within the picture between the Current block and the Reference block. Referring to Figure 8, there is a part shown with a dotted line in the Reference picture, and the motion vector is the vector that shows how the Current block should move to reach the block that the Reference picture should refer to. In other words, the block that appears when the current block is moved according to the motion vector is the part indicated by the dotted line in the current picture in Figure 8, and the position of this part indicated by the dotted line is the same as the position of the reference block in the reference picture.

[0068] Furthermore, the information indicating the block that the current block should refer to includes information indicating a reference picture. This information includes a reference picture list and a reference picture index. The reference picture list is a list containing reference pictures, and a reference block can be used from the reference pictures included in the reference picture list. In other words, the current block can be predicted from the reference pictures included in the reference picture list. The reference picture index is an index used to indicate the reference picture to be used.

[0069] Figure 9 shows a motion vector signaling method according to one embodiment of the present disclosure. According to one embodiment of this disclosure, a motion vector (MV) is generated based on a motion vector predictor (MVP). For example, the motion vector predictor becomes the motion vector as follows.

[0070] MV=MVP. To give another example, for instance, the motion vector is based on the motion vector difference (MVD), as shown below. To show the motion vector accurately to the motion vector predictor, the motion vector difference (MVD) is added.

[0071] MV = MVP + MVD. In video coding, motion vector information determined by the encoder is transmitted to the decoder, which generates a motion vector from the received motion vector information to determine the prediction block. For example, the motion vector information includes information about the motion vector predictor and the motion vector difference. In this case, the components of the motion vector information may differ depending on the mode. For example, in merge mode, the motion vector information includes information about the motion vector predictor but does not include the motion vector difference. As another example, in AMVP (advanced motion vector prediction) mode, the motion vector information includes information about the motion vector predictor and the motion vector difference.

[0072] To determine, transmit, and receive information about motion vector predictors, the encoder and decoder generate MVP candidates in the same way. For example, the encoder and decoder generate the same MVP candidates in the same order. The encoder then transmits an index (mvp_lx_flag) indicating the MVP (motion vector predictor) determined from the generated MVP candidates to the decoder, and the decoder determines the MVP and MV based on this index (mvp_lx_flag). The index (mvp_lx_flag) includes the motion vector predictor index (mvp_l0_flag) of the 0th reference picture list (list 0) and the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1). The method for receiving the index (mvp_lx_flag) is explained in Figures 56 to 59.

[0073] The methods for generating MVP candidates include spatial candidates and temporal candidates. A spatial candidate is a motion vector for a block at a certain position from the current block. For example, it is a motion vector for a block or position that is adjacent to or not adjacent to the current block. A temporal candidate is a motion vector for a block in a different picture from the current block. Alternatively, an MVP candidate may include an affine motion vector, ATMVP, STMVP, a combination of the motion vectors mentioned above, the average vector of the motion vectors mentioned above, a zero motion vector, etc.

[0074] Furthermore, information indicating the reference picture is also transmitted from the encoder to the decoder. If the reference picture corresponding to the MVP candidate does not match the information indicating the reference picture, motion vector scaling is performed. Motion vector scaling is calculated based on the picture order count (POC) of the current picture, the POC of the reference picture in the current block, the POC of the MVP candidate's reference picture, and the MVP candidate itself.

[0075] Figure 10 shows a motion vector difference syntax according to one embodiment of the present disclosure.

[0076] Motion vector difference is coded with separate signature and absolute value values. In other words, the signature and absolute value of the motion vector difference have different syntax. The absolute value of the motion vector difference can be coded directly, or it can be coded with a flag indicating whether the absolute value is greater than N, as shown in Figure 10. If the absolute value is greater than N, the value of (absolute-N) is signaled together. In the example in Figure 10, the abs_mvd_greater0_flag is transmitted, which indicates whether the absolute value is greater than 0. If the abs_mvd_greater0_flag is shown and the absolute value is not greater than 0, then it is determined that the absolute value is 0. Also, if the abs_mvd_greater0_flag is shown and the absolute value is greater than 0, then there may be additional syntax. For example, there is a flag called abs_mvd_greater1_flag, which indicates whether the absolute value is greater than 1. If abs_mvd_greater1_flag is shown as not being greater than 1, then it is determined that the absolute value is 1. If abs_mvd_greater1_flag is shown as being greater than 1, then there may be additional syntax. For example, there is abs_mvd_minus2, which is the value (absolute value - 2). Because it was determined that the absolute value is greater than 1 (greater than or equal to 2) via the aforementioned abs_mvd_greater0_flag and abs_mvd_greater1_flag, it shows (absolute value - 2).The reason abs_mvd_minus2 is binarized to a variable length is to signal with fewer bits. For example, there are binarization methods that use a variable length, such as Exp-Golomb, truncated unary, and truncated Rice. Also, mvd_sign_flag is a flag that indicates the sign of the motion vector difference.

[0077] In this embodiment, the coding method was explained using motion vector difference, but information other than motion vector difference can also be separated with respect to sign and absolute value, and the absolute value can be coded as a flag indicating whether the absolute value is greater than a certain value or not, and as the value obtained by subtracting the aforementioned certain value from the absolute value. Also, in Figure 10, [0] and [1] represent component indices. For example, they represent x-component and y-component.

[0078] Figure 11 shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure. According to one embodiment of this disclosure, the resolution representing a motion vector or motion vector difference is diverse. In other words, the resolution to which a motion vector or motion vector difference is coded is diverse. For example, resolution is expressed based on pixels (pel). For example, a motion vector or motion vector difference is signaled in units such as 1 / 4 (quarter), 1 / 2 (half), 1 (integer), 2, and 4 pixels. For example, if you want to represent 16, it is coded as 64 in 1 / 4 units (1 / 4 * 64 = 16), as 1 in 1 units (1 * 16 = 16), and as 4 in 4 units (4 * 4 = 16). That is, the value is defined as follows.

[0079] valueDetermined=resolution*valuePerResolution

[0080] Here, valueDetermined is the value to be transmitted, which in this embodiment is either a motion vector or a motion vector difference. ValuePerResolution is the value that represents valueDetermined in units of resolution. In this case, if the value signaled by the motion vector or motion vector difference is not divisible by the resolution, an inaccurate value other than the motion vector or motion vector difference, which offers the best prediction performance (such as rounding), is sent. Using high reeolution reduces inaccuracy but uses more bits because the coded value is large, while using low reeolution increases inaccuracy but uses fewer bits because the coded value is small.

[0081] Furthermore, the resolution can be set differently for units such as blocks, CUs, and slices. Therefore, the resolution can be applied adaptively according to the unit.

[0082] The resolution is signaled from the encoder to the decoder. In this case, the signaling for the resolution is a signal that has been binarized to the variable length described above. In such a case, signaling with the index corresponding to the smallest value (the earliest value) reduces the signaling overhead.

[0083] As one example, the signaling index is matched in order from high resolution (detailed signaling) to low resolution.

[0084] Figure 11 shows the signaling for three resolutions. In this case, the three signalings are 0, 10, and 11, with each of the three signalings corresponding to resolution 1, resolution 2, and resolution 3, respectively. Signaling resolution 1 requires 1 bit, and signaling the remaining resolutions requires 2 bits, resulting in less signaling overhead when signaling resolution 1. In the example in Figure 11, resolutions 1, 2, and 3 are 1 / 4, 1, and 4pel, respectively.

[0085] In the following disclosure, motion vector resolution means the resolution of the motion vector difference.

[0086] Figure 12 shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure.

[0087] As explained in Figure 11, the number of bits required to signal a resolution can vary depending on the resolution, so the signaling method is changed according to the situation. For example, the signaling value for a given resolution may vary depending on the situation. For example, the signaling index and resolution are matched in a different order depending on the situation. For example, in a situation where there are resolutions corresponding to signaling 0, 10, 110, ... the order is resolution 1, resolution 2, resolution 3, ... respectively, while in other situations the order is not resolution 1, resolution 2, resolution 3, ... but something else. Furthermore, more than one situation can be defined. Referring to Figure 12, the resolutions corresponding to 0, 10, and 11 are resolution 1, resolution 2, and resolution 3 in Case 1, and resolutions 2, resolution 1, and resolution 3 in Case 2. In this case, there are two or more cases.

[0088] Figure 13 shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure.

[0089] As explained in Figure 12, motion vector resolution is signaled differently depending on the situation. For example, if resolutions of 1 / 4, 1, and 4pel exist, then in some situations the signaling described in Figure 11 is used, and in other situations the signaling shown in Figure 13(a) or Figure 13(b) is used. Only two cases exist, such as Figure 11, Figure 13(a), and Figure 13(b), or all three may exist. Therefore, in some situations, a resolution other than the highest resolution can be signaled with fewer bits.

[0090] Figure 14 shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure.

[0091] According to one embodiment of this disclosure, in the adaptive motion vector resolution described in Figure 11, the possible resolutions may vary depending on the situation. For example, the resolution value may change depending on the situation. In one embodiment, in one situation, resolution 1, resolution 2, resolution 3, resolution 4, ... may be used, and in another situation, resolution A, resolution B, resolution C, resolution D, ... may be used. Furthermore, {resolution 1, resolution 2, resolution 3, resolution 4, ...} and {resolution A, resolution B, resolution C, resolution D, ...} may have an intersection set rather than an empty set. In other words, a certain resolution value may be used in two or more situations, but the set of usable resolution values ​​may differ for each of these two or more situations. Also, the number of usable resolution values ​​may differ for each situation.

[0092] Referring to Figure 14, Case 1 uses resolutions 1, 2, and 3, while Case 2 uses resolutions A, B, and C. For example, resolutions 1, 2, and 3 are 1 / 4, 1, and 4pel. Also, for example, resolutions A, B, and C are 1 / 4, 1 / 2, and 1pel.

[0093] Figure 15 shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure.

[0094] Referring to Figure 15, the resolution signaling is differentiated depending on the motion vector candidate or motion vector predictor candidate. For example, this relates to whether the case described in Figures 12 to 14 is selected, or which candidate the signaled motion vector or motion vector predictor is. The method for differentiating the signaling is as shown in Figures 12 to 14.

[0095] For example, the situation can be defined differently depending on its position within the list of candidates. Alternatively, the situation can be defined differently depending on how the candidate was created.

[0096] The encoder or decoder generates a candidate list containing at least one MV candidate (motion vector candidate) or at least one MVP candidate (motion vector predictor candidate). MVPs that appear earlier in the candidate list for an MV candidate or MVP candidate tend to have higher accuracy, while MVPs that appear later tend to have lower accuracy. This may be because those appearing earlier in the candidate list are signaled with fewer bits, resulting in a design where MVPs appearing earlier are more accurate. In one embodiment of this disclosure, higher accuracy of the MVP indicates a smaller MVD (motion vector difference) value for a good predictive motion vector, while lower accuracy of the MVP indicates a larger MVD value for a good predictive motion vector. Therefore, if the accuracy of the MVP (Motion Vector Predictor) is low, it is possible to reduce the number of bits required to signal at low resolution and indicate the motion vector difference value (for example, the value that indicates the difference value based on resolution).

[0097] Because low resolution can be used when the accuracy of the MVP (motion vector predictor) is low based on this principle, according to one embodiment of the present disclosure, it is possible to promise that the resolution will be signaled with the fewest bits, rather than the highest resolution, depending on the MVP candidate. For example, if 1 / 4, 1, and 4pel are possible resolutions, then 1 or 4 will be signaled with the fewest bits (1-bit). Referring to Figure 15, Candidate 1 and Candidate 2 are signaled with the fewest bits for the high resolution of 1 / 4pel, while Candidate N, which is after Candidate 1 and Candidate 2, is signaled with the fewest bits for a resolution other than 1 / 4pel.

[0098] Figure 16 shows the signaling of adaptive motion vector resolution according to one embodiment of the present disclosure.

[0099] As explained in Figure 15, the motion vector resolution is varied depending on which candidate the determined motion vector or motion vector predictor is. The methods for varying the signaling are shown in Figures 12 to 14.

[0100] Referring to Figure 16, for some candidates, high resolution is signaled with the fewest bits, and for other candidates, a resolution that is not high resolution is signaled with the fewest bits. For example, a candidate that signals a resolution that is not high resolution with the fewest bits is an inaccurate candidate. For example, candidates that signal a resolution that is not high resolution with the fewest bits include temporal candidates, zero motion vectors, non-adjacent spatial candidates, and candidates depending on whether or not a refinement process exists.

[0101] A Temporal Candidate is a motion vector from another picture. A Zero motion vector is a motion vector where all vector components are 0. A Non-adjacent Spatial Candidate is a motion vector referenced from a position not adjacent to the current block. The Refinement process is the process of refining the motion vector predictor, for example, through template matching, bilateral matching, etc.

[0102] According to one embodiment of this disclosure, after refining the motion vector predictor, the motion vector difference is added. This is done to make the motion vector predictor more accurate and reduce the motion vector difference value. In such cases, the motion vector resolution signaling is made different for candidates that do not undergo the refinement process. For example, the resolution that is not the highest resolution is signaled with the fewest bits.

[0103] In other embodiments, a refinement process is performed after adding the motion vector difference to the motion vector predictor. During this refinement process, the motion vector resolution signaling is varied for each candidate. For example, the resolution that is not the highest resolution is signaled with the fewest bits. This is because, since the refinement process is performed after adding the motion vector difference, the error can be reduced through the refinement process even if the MV difference is not signaled with maximum accuracy (to minimize prediction error).

[0104] Another example is changing the motion vector resolution signaling if the selected candidate differs from other candidates by a certain amount.

[0105] As another example, the motion vector refinement process can be varied depending on which candidate the determined motion vector or motion vector predictor is. The motion vector refinement process is the process of finding a more accurate motion vector. For example, it is the process of finding a block that matches the current block according to a predetermined convention from a reference point (e.g., template matching or bilateral matching). The reference point is the position corresponding to the determined motion vector or motion vector predictor. In this case, the degree of movement from the reference point may differ according to the predetermined convention, and varying the motion vector refinement process means varying the degree of movement from the reference point. For example, for an accurate candidate, a detailed refinement process may be started, and for an inaccurate candidate, a less detailed refinement process may be started. Accurate and inaccurate candidates are determined by their position in the candidate list, or by the method used to generate them. The method used to generate candidates refers to the position from which the spatial candidate was brought. See Figure 16 for further details. Further, more detailed and less detailed refinements involve finding matching blocks by moving them gradually from a reference point, or by moving them significantly. If finding matching blocks by moving them significantly, an additional step is added where the best-matching block found by moving it significantly is then moved even further.

[0106] As another embodiment, the motion vector resolution signaling is changed based on the POC of the current picture and the POC of the motion vector or the reference picture of the motion vector predictor candidate. The method for changing the signaling is shown in Figures 12 to 14.

[0107] For example, if the difference between the current picture's POC (picture order count) and the reference picture of the motion vector or motion vector predictor candidate is large, the motion vector or motion vector predictor may be inaccurate, and a resolution other than high resolution will be signaled with the fewest bits.

[0108] As another example, motion vector resolution signaling is modified based on whether motion vector scaling should be performed. For instance, if the selected motion vector or motion vector predictor is a candidate for motion vector scaling, it signals a resolution other than high resolution with the fewest bits. Motion vector scaling should be performed if the reference picture for the current block is different from the reference picture of the referenced candidate.

[0109] Figure 17 shows an example of affine motion prediction according to one embodiment of the present disclosure.

[0110] Conventional prediction methods, as explained in Figure 8, predict based on a block simply moved to a position without rotation or scaling. They predict based on a reference block of the same size, shape, and angle as the current block. Figure 8 is simply a translation motion model. However, the content contained in actual video has more complex movements, and prediction performance can be further improved by predicting from a variety of patterns.

[0111] Referring to Figure 17, the currently predicted block is shown as a solid line in the current picture. The current block is predicted by referencing a block with a different shape, size, and angle, and this reference block is shown as a dotted line in the reference picture. The same position as the reference block within the picture is also shown as a dotted line in the current picture. In this case, the reference block is the block shown by the affine transformation of the current block. This allows for the representation of stretching (scaling), rotation, shearing, reflection, orthogonal projection, etc.

[0112] The number of parameters used to describe affine motion and affine transformation varies. Using more parameters allows for a wider variety of motions than using fewer parameters, but this can introduce overhead in signaling and calculations.

[0113] For example, the affine transformation can be represented by six parameters, or by three control point motion vectors.

[0114] Figure 18 shows an affine motion prediction according to one embodiment of the present disclosure.

[0115] As shown in Figure 17, complex motion can be represented using affine transformation, but to reduce the signaling overhead and computation, a simpler affine motion prediction or affine transformation is used. A simpler affine motion prediction can be performed by restricting the motion. When the motion is restricted, the shape that the current block will be transformed into in the reference block is limited.

[0116] Referring to Figure 18, affine motion prediction is performed using the control point motion vectors v0 and v1. Using two vectors, v0 and v1, is equivalent to using four parameters. The two vectors v0 and v1, or the four parameters, indicate which shape of reference block the current block predicts from. Using such a simple affine transformation, the rotation and scaling (zoom in / out) of a block can be shown. Referring to Figure 18, the current block, shown by the solid line, is predicted from the position shown by the dotted line in the reference picture. Each point (pixel) of the current block is mapped to other points via the affine transformation.

[0117] Figure 19 shows an equation representing the motion vector field according to one embodiment of the present disclosure. In Figure 18, the control point motion vector v0 is (v_0x, v_0y), which is the motion vector of the top-left corner control point. Similarly, the control point motion vector v1 is (v_1x, v_1y), which is the motion vector of the top-right corner control point. In this case, the motion vector (v_x, v_y) at position (x, y) is as shown in Figure 19. Therefore, the motion vector for each pixel position or a given position is estimated using the equation in Figure 19, which is based on v0 and v1.

[0118] Furthermore, in the formula in Figure 19, (x,y) represents the relative coordinates within the block. For example, (x,y) is the position when the left top position of the block is set to (0,0).

[0119] If v0 is the control point motion vector for position (x0, y0) on the picture, and v1 is the control point motion vector for position (x1, y1) on the picture, then to show the position (x, y) within the block using the same coordinates as v0 and v1, we change x and y to (x-x0, y-y0) respectively in the formula in Figure 19. Also, w (the width of the block) is (x1-x0).

[0120] Figure 20 shows an affine motion prediction according to one embodiment of the present disclosure.

[0121] According to one embodiment of this disclosure, affine motion is represented using a number of control point motion vectors or a number of parameters.

[0122] Referring to Figure 20, affine motion prediction is performed using the control point motion vectors v0, v1, and v2. Using the three vectors v0, v1, and v2 is equivalent to using six parameters. The three vectors v0, v1, and v2, or the six parameters, indicate which shape of reference block the current block predicts from. Referring to Figure 20, the current block shown by the solid line predicts from the position shown by the dotted line in the reference picture. Each point (pixel) of the current block is mapped to other points via the affine transformation.

[0123] Figure 21 shows an equation representing the motion vector field according to one embodiment of the present disclosure. In Figure 20, the control point motion vector v0 is (mv_0^x, mv_0^y) and is the motion vector of the top-left corner control point, the control point motion vector v1 is (mv_1^x, mv_1^y) and is the motion vector of the top-right corner control point, and the control point motion vector v2 is (mv_2^x, mv_2^y) and is the motion vector of the bottom-left corner control point. In this case, the motion vector (mv^x, mv^y) at the (x,y) position is as shown in Figure 21. Therefore, the motion vector for each pixel position or a certain position is estimated by the equation in Figure 21, and this is based on v0, v1, and v2.

[0124] Furthermore, in the formula in Figure 21, (x,y) are relative coordinates within the block. For example, (x,y) is the position when the left top position of the block is (0,0). Therefore, if v0 is the control point motion vector for position (x0,y0), v1 is the control point motion vector for position (x1,y1), and v2 is the control point motion vector for position (x2,y2), then to show (x,y) using the same coordinates as v0, v1, and v2, we change x and y to (x-x0,y-y0) respectively in the formula in Figure 21. Also, w (width of the block) is (x1-x0) and h (height of the block) is (y2-y0).

[0125] Figure 22 shows an affine motion prediction according to one embodiment of the present disclosure.

[0126] As mentioned above, a motion vector field exists, and a motion vector is calculated for each pixel. However, to simplify the process, an affine transformation is performed using a subblock-based approach, as shown in Figure 22. For example, the small rectangle in Figure 22(a) is a subblock. A representative motion vector is created for the subblock, and this representative motion vector is used for the pixels of that subblock. Alternatively, as shown in Figures 17, 18, and 20 to represent complex movements, the subblock can represent such movements and correspond to a reference block, or it can be simplified by applying only translation motion to the subblock. In Figure 22(a), v0, v1, and v2 are control point motion vectors.

[0127] In this case, the size of the subblock is M*N, where M and N are as shown in Figure 22(b). MvPre is the motion fraction accuracy. (v_0x,v_0y), (v_1x,v_1y), and (v_2x,v_2y) are the motion vectors of the top-left, top-right, and bottom-left control points, respectively. (v_2x,v_2y) is the motion vector of the current block's bottom-left control point, but for example, in the case of 4-parameters, it is the MV (motion vector) for bottom-left calculated by the formula in Figure 19.

[0128] Furthermore, when creating a representative motion vector for a subblock, the central sample position of the subblock is used to calculate the representative motion vector. In addition, when creating the subblock's motion vector, a motion vector with higher accuracy than the normal motion vector is used, and for this purpose, motion compensation interpolation filters are applied.

[0129] In another embodiment, the size of the subblock may not be variable but fixed to a specific size. For example, the size of the subblock may be fixed to a 4x4 size.

[0130] Figure 23 shows a mode of affine motion prediction according to one embodiment of the present disclosure.

[0131] According to one embodiment of this disclosure, an example of affine motion prediction is affine inter mode. There is a flag that indicates that it is affine inter mode. Referring to Figure 23, there are blocks at positions A, B, C, D, and E near v0 and v1, and the motion vectors corresponding to each block are vA, vB, vC, vD, and vE. Using these, a candidate list for motion vectors or motion vector predictors is created as follows.

[0132] {(v0,v1)|v0={vA,vB,vC}, v1={vD,vE}} In other words, a (v0,v1) pair is created using v0, selected from vA, vB, and vC, and v1, selected from vD and vE. At this time, the motion vector is scaled using the POC (picture order count) of the neighbor block's reference, the POC of the reference to the current CU (current coding unit; current block), and the POC of the current CU (current coding unit; current block). Once a candidate list is created with the motion vector pairs described above, it is possible to signal which of the candidate list was selected or has been selected. Also, if the candidate list is not sufficiently filled, it can be filled with candidates from other inter-prediction methods. For example, AMVP (Advanced Motion Vector Prediction) candidates can be used to fill it. Furthermore, instead of immediately using v0 and v1 selected from the candidate list as the control point motion vectors for affine motion prediction, it is possible to signal a difference to correct them, thereby creating better control point motion vectors. In other words, the decoder uses v0' and v1', created by adding the difference to v0 and v1 selected from the candidate list, as the control point motion vectors for affine motion prediction.

[0133] As one example, affine inter mode may be used for CUs (coding units) of a certain size or larger.

[0134] Figure 24 shows a mode of affine motion prediction according to one embodiment of the present disclosure.

[0135] According to one embodiment of this disclosure, an example of affine motion prediction is affine merge mode. There is a flag that indicates that affine merge mode is in use. In affine merge mode, if affine motion prediction is used around the current block, the control point motion vector of the current block is calculated from the motion vectors of those surrounding blocks. For example, when checking whether surrounding blocks have used affine motion prediction, candidate surrounding blocks are as shown in Figure 24(a). Furthermore, it is checked whether affine motion prediction was used in the order of A, B, C, D, and E, and once a block that has used affine motion prediction is found, the control point motion vector of the current block is calculated using the motion vector of that block or the motion vectors of the area around it. A, B, C, D, and E are left, above, above right, left bottom, and above left, respectively, as shown in Figure 24(a).

[0136] As one embodiment, if affine motion prediction is used for the block at position A as shown in Figure 24(b), v0 and v1 are calculated using the motion vectors of that block or its surroundings. The motion vectors of that block or its surroundings are v2, v3, and v4.

[0137] In one embodiment, the order of the surrounding blocks to be referenced is predetermined. However, the control point motion vector derived from a specific position does not always provide better performance. Therefore, in another embodiment, a signal is provided to indicate whether to reference a block at a certain position when deriving the control point motion vector. For example, the candidate positions for control point motion vector derivation are determined in the order of A, B, C, D, and E in Figure 24(a), and a signal is provided to indicate which of these to reference.

[0138] Another example is that when deriving control point motion vectors, accuracy can be improved by using blocks closest to each control point motion vector. For example, referring to Figure 24, when deriving v0, the left block is referenced, and when deriving v1, the above block is referenced. Alternatively, when deriving v0, A, D, or E may be referenced, and when deriving v1, B or C may be referenced.

[0139] Figure 25 shows an affine motion predictor derivation according to one embodiment of the present disclosure.

[0140] Control point motion vectors are necessary for affine motion prediction. Based on these control point motion vectors, a motion vector field, that is, a motion vector for a subblock or a certain position, is calculated. Control point motion vectors are also called seed vectors.

[0141] In this case, the control point MV (motion vector) is based on the predictor. For example, the predictor can be the control point MV. Alternatively, the control point MV may be calculated based on the predictor and the difference. More specifically, the control point MV is calculated by adding or subtracting the difference from the predictor.

[0142] In this process, the predictor for the control point MV is derived from the control point MVs (control point motion vectors) or MVs (motion vectors) of surrounding blocks that have undergone affine motion prediction (affine motion compensation (MC)). For example, if a block at a predetermined position is affine motion predicted, the control point MVs of that block are used to derive the predictor for the current block's affine motion compensation from the MVs. Referring to Figure 25, the predetermined positions are A0, A1, B0, B1, and B2. Alternatively, the predetermined positions include positions adjacent to and not adjacent to the current block. Furthermore, the control point MVs (control point motion vectors) or MVs (motion vectors) of the predetermined positions are referenced (spatial), and the temporal control point MVs or MVs of the predetermined positions are referenced.

[0143] A candidate for affine motion compensation (affine MC) can be created using the method shown in the example in Figure 25, and such a candidate may also be called an inherited candidate. Alternatively, such a candidate may also be called a merge candidate. Furthermore, in the method shown in Figure 25, when referencing pre-defined positions, they are referenced in a pre-defined order.

[0144] Figure 26 shows an affine motion predictor derivation according to one embodiment of the present disclosure.

[0145] Control point motion vectors are necessary for affine motion prediction. Based on these control point motion vectors, a motion vector field, that is, a motion vector for a subblock or a certain position, is calculated. Control point motion vectors are also called seed vectors.

[0146] In this case, the control point MV (motion vector) is based on the predictor. For example, the predictor can be the control point MV. Alternatively, the control point MV may be calculated based on the predictor and the difference. More specifically, the control point MV is calculated by adding or subtracting the difference from the predictor.

[0147] In this process, control point MVs are derived from surrounding MVs. This includes MVs (motion vectors) that are not affine-compensated (affine motion compensated). For example, when deriving each control point MV (control point motion vector) of a block, a pre-set MV is used as the predictor for each control point MV. For instance, the pre-set position is a part included in the block adjacent to that part.

[0148] Referring to Figure 26, control point MVs (motion vectors) mv0, mv1, and mv2 are determined. In this case, according to one embodiment of the present disclosure, the MVs (motion vectors) corresponding to the preset positions A, B, and C are used as predictors for mv0. The MVs (motion vectors) corresponding to the preset positions D and E are used as predictors for mv1. The MVs (motion vectors) corresponding to the preset positions F and G are used as predictors for mv2.

[0149] Furthermore, when determining the predictors for control point MV (control point motion vector) mv0, mv1, and mv2 using the embodiment shown in Figure 26, the order in which pre-set positions are referenced for each control point is predetermined. While there may be numerous pre-set positions referenced as predictors for each control point, the possible combinations of these pre-set positions may be predetermined.

[0150] A candidate for affine motion compensation (affine MC) can be created using the method shown in the embodiment in Figure 26, and such a candidate may also be called a constructed candidate. Alternatively, such a candidate may be called an inter candidate or a virtual candidate. Furthermore, in the method shown in Figure 41, when referencing pre-set positions, they are referenced in a pre-set order.

[0151] According to one embodiment of this disclosure, an affine motion compensation (affine MC) candidate list or an affine motion compensation (affine MC) control point MV candidate list can be generated using the embodiments described in Figures 23 to 26 or a combination thereof.

[0152] Figure 27 shows an affine motion predictor derivation according to one embodiment of the present disclosure.

[0153] As explained in Figures 24 and 25, the control point MV (motion vector) for the affine motion prediction of the current block is derived from the surrounding affine motion predicted blocks. The same method as in Figure 27 is used for this. In the equation in Figure 27, the top-left, top-right, and bottom-left MVs (motion vectors) or control point MVs (motion vectors) of the surrounding affine motion predicted blocks are (v_E0x,v_E0y), (v_E1x,v_E1y), and (v_E2x,v_E2y), respectively. Also, the coordinates of the top-left, top-right, and bottom-left of the surrounding affine motion predicted blocks are (x_E0,y_E0), (x_E1,y_E1), and (x_E2,y_E2), respectively. In this process, Figure 27 is used to calculate the predictor of the current block's control point MV (control point motion vector) or the control point MV itself, (v_0x,v_0y) and (v_1x,v_1y).

[0154] Figure 28 shows an affine motion predictor derivation according to one embodiment of the present disclosure.

[0155] As mentioned above, affine motion compensation requires a large number of control point MVs (motion vectors) or control point MV predictors. In this process, one control point MV or control point MV predictor induces other control point MVs or control point MV predictors.

[0156] For example, when two control point MVs (control point motion vectors) or two control point MV predictors (control point motion vector predictors) are created using the method described in the above diagram, other control point MVs (control point motion vectors) or other control point MV predictors (control point motion vector predictors) are generated based on them.

[0157] Referring to Figure 28, we see how to generate mv0, mv1, and mv2, which are top-left, top-right, and bottom-left control point MV predictors or control point MVs (control point motion vectors). In the figure, x and y represent the x-component and y-component, respectively, and the current block size is w*h.

[0158] Figure 29 shows a method for generating a control point motion vector according to one embodiment of the present disclosure.

[0159] According to one embodiment of this disclosure, a predictor for the control point MV (control point motion vector) is created to perform affine motion compensation (MC) on the current block, and the difference is added to it to determine the control point MV. According to one embodiment, the predictor for the control point MV can be created in the manner described in Figures 23 to 26. The difference is signaled from the encoder to the decoder.

[0160] Referring to Figure 29, a difference exists for each control point MV (control point motion vector). Furthermore, each difference for each control point MV is signaled. Figure 29(a) shows how to determine the control point MVs mv0 and mv1 in a 4-parameter model, and Figure 29(b) shows how to determine the control point MVs mv0, mv1, and mv2 in a 6-parameter model. The control point MVs are determined by adding the differences mvd0, mvd1, and mvd2 for each control point MV to the predictor.

[0161] In Figure 29, the bar at the top indicates the predictor of the control point MV (control point motion vector).

[0162] Figure 30 shows the method for determining the motion vector difference as described in Figure 29.

[0163] As one embodiment, the motion vector difference is signaled using the method described in Figure 10. The motion vector difference determined by the signaled method is lMvd in Figure 30. Furthermore, the same value as the signaled mvd (motion vector difference) shown in Figure 29, i.e., mvd0, mvd1, and mvd2, is lMvd in Figure 30. As explained in Figure 29, the signaled mvd (motion vector difference) is determined as the difference between it and the predictor of the control point MV (control point motion vector), and the determined difference is MvdL0 and MvdL1 in Figure 30. L0 represents reference list 0 (the 0th reference picture list), and L1 represents reference list 1 (the 1st reference picture list). compIdx is the component index, indicating the x, y component, etc.

[0164] Figure 31 shows a method for generating a control point motion vector according to one embodiment of the present disclosure.

[0165] According to one embodiment of this disclosure, a predictor for the control point MV (control point motion vector) is created to perform affine motion compensation (MC) on the current block, and the control point MV is determined by adding a difference to it. According to one embodiment, a predictor for the control point MV (control point motion vector) can be created in the manner described in Figures 23 to 26. The difference is signaled from the encoder to the decoder.

[0166] Referring to Figure 31, there is a predictor for the difference for each control point MV (control point motion vector). For example, the difference of one control point MV is used to determine the difference of other control point MVs. This is based on the similarity between the differences for the control point MVs. Because they are similar, once a predictor is determined, the difference with the predictor can be reduced. In this process, the difference predictor for the control point MV is signaled, and the difference between the control point MV and the difference predictor is signaled.

[0167] Figure 31(a) shows how to determine the control point MV (control point motion vector) mv0 and mv1 for a 4-parameter model, and Figure 31(b) shows how to determine the control point MV (control point motion vector) mv0, mv1, and mv2 for a 6-parameter model.

[0168] Referring to Figure 31, the difference for each control point MV is determined based on the difference (mvd0) of control point MV 0, which is mv0, and the control point MV is determined accordingly. mvd0, mvd1, and mvd2 shown in Figure 31 are signaled from the encoder to the decoder. Compared to the method described in Figure 29, the method in Figure 31 may result in different values ​​for signaled mvd1 and mvd2 even when using the same mv0, mv1, and mv2 and the same predictor as in Figure 29. If the differences between the predictor and the control point MVs mv0, mv1, and mv2 are similar, using the method in Figure 31 may result in lower absolute values ​​for mvd1 and mvd2 than using the method in Figure 29, thereby reducing the signaling overhead for mvd1 and mvd2. Referring to Figure 31, the difference between mv1 and the predictor is determined to be (mvd1 + mvd0), and the difference between mv2 and the predictor is determined to be (mvd2 + mvd0).

[0169] In Figure 31, the upper bar indicates the predictor of the control point MV.

[0170] Figure 32 shows the method for determining the motion vector difference as described in Figure 31.

[0171] In one embodiment, the motion vector difference is signaled using the method described in Figure 10 or Figure 33. The motion vector difference determined based on the signaled parameters is lMvd in Figure 32. Furthermore, the same values ​​as the signaled mvd shown in Figure 31, i.e., mvd0, mvd1, and mvd2, are the same as lMvd in Figure 32.

[0172] In Figure 32, MvdLX is the difference between each control point MV (control point motion vector) and the predictor. That is, (mv - mvp). In this case, as explained in Figure 31, for control point MV 0 (mv_0), the signaled motion vector difference is immediately used as the difference (MvdLX) for the control point MV. For the other control point MVs (mv_1, mv_2), the signaled motion vector difference (mvd1, mvd2 in Figure 31) and for control point MV 0 (mv_0), the signaled motion vector difference (mvd0 in Figure 31) are used to determine and use the MvdLX, which is the difference for the control point MV.

[0173] In Figure 32, LX represents the reference list (reference picture list X). compIdx is the component index, indicating components x, y, etc. cpIdx represents the control point index. cpIdx means 0, 1, or 0, 1, 2 as shown in Figure 31. The values ​​shown in Figures 10, 12, and 13 should take into account the resolution of the motion vector difference. For example, if the resolution is R, use the value of lMvd*R for lMvd in the drawing.

[0174] Figure 33 shows a motion vector difference syntax according to one embodiment of the present disclosure.

[0175] Referring to Figure 33, the motion vector difference is coded in a similar manner to that described in Figure 10. In this case, coding is performed separately using cpIdx and the control point index.

[0176] Figure 34 shows the structure of a higher-level signaling according to one embodiment of the present disclosure. According to one embodiment of the present disclosure, there is one or more higher-level signalings. A higher-level signaling refers to a signaling at a higher level. A higher level is a unit that contains a certain unit. For example, the higher levels of the current block or current coding unit include CTUs, slices, tiles, tile groups, pictures, sequences, etc. A higher-level signaling influences the lower levels of the corresponding higher level. For example, if the higher level is a sequence, it influences the CTUs, slices, tiles, tile groups, and picture units that are lower than the sequence. Here, "influence" means that the higher-level signaling influences the encoding or decoding of the lower levels.

[0177] Alternatively, higher-level signaling includes signaling that indicates whether a certain mode is available or not. Referring to Figure 34, the higher-level signaling includes sps_modeX_enabled_flag. In one embodiment, it is determined whether mode modeX is available or not based on sps_modeX_enabled_flag. For example, when sps_modeX_enabled_flag is a certain value, mode modeX cannot be used. However, when sps_modeX_enabled_flag is a different value, mode modeX can be used. Furthermore, when sps_modeX_enabled_flag is a different value, it is determined whether mode modeX should be used or not based on further signaling. For example, the aforementioned value may be 0, and the other value may be 1. However, it is not limited to this; the aforementioned value may be 1, and the other value may be 0.

[0178] According to one embodiment of this disclosure, there is a signaling that indicates whether affine motion compensation is available. For example, this signaling is a high-level signaling. Referring to Figure 34, this signaling is sps_affine_enabled_flag (affine-enabled flag). Referring to Figures 2 and 7, the signaling means a signal transmitted from the encoder to the decoder via the bitstream. The decoder purges the affine-enabled flag from the bitstream.

[0179] For example, if sps_affine_enabled_flag (affine-enabled flag) is 0, the syntax is restricted so that affine motion compensation is not used. Also, if sps_affine_enabled_flag (affine-enabled flag) is 0, inter_affine_flag (inter-affine flag) and cu_affine_type_flag (coding unit-affine type flag) do not exist.

[0180] For example, inter_affine_flag is a signaling that indicates whether or not affine motion compensation (MC) is used in a block. Referring to Figures 2 and 7, the signaling refers to the signal transmitted from the encoder to the decoder via the bitstream. The decoder purges inter_affine_flag from the bitstream. Furthermore, the cu_affine_type_flag (coding unit affine type flag) is a signaling that indicates which type of affine MC (affine motion compensation) will be used from the block. Here, the type indicates whether it is a 4-parameter affine model or a 6-parameter affine model. Also, if sps_affine_enabled_flag (affine enabled flag) is 1, then affine modelcomp (affine motion compensation) is available.

[0181] Affine motion compensation refers to motion compensation based on an affine model or motion compensation based on an affine model for inter prediction.

[0182] Furthermore, according to one embodiment of this disclosure, there exists a signaling that indicates which mode's specific type is available. For example, there exists a signaling that indicates whether a specific type of affine motion compensation is available. For example, this signaling is a higher-level signaling. Referring to Figure 34, this signaling is sps_affine_type_flag. The specific type refers to a 6-parameter affine model. For example, if sps_affine_type_flag is 0, the syntax is restricted so that the 6-parameter affine model is not used. Also, if sps_affine_type_flag is 0, then cu_affine_type_flag does not exist. Also, if sps_affine_type_flag is 1, then the 6-parameter affine model is used. If sps_affine_type_flag does not exist, its value is inferred to 0.

[0183] Furthermore, according to one embodiment of this disclosure, a signaling indicating which mode's specific type is available exists if there is a signaling indicating whether a certain mode is available. For example, if the signaling indicating which mode is available is 1, then the signaling indicating which mode's specific type is available is parsed. If the signaling indicating which mode is available is 0, then the signaling indicating which mode's specific type is available is not parsed. For example, the signaling indicating which mode is available includes sps_affine_enabled_flag (affine flag). The signaling indicating which mode's specific type is available includes sps_affine_type_flag (affine flag). Referring to Figure 34, if sps_affine_enabled_flag (affine flag) is 1, then sps_affine_type_flag is parsed. Additionally, if sps_affine_enabled_flag (affine-enabled flag) is 0, sps_affine_type_flag is not parsed, and its value is inferred (implied) to 0.

[0184] Furthermore, as mentioned above, adaptive motion vector resolution (AMVR) is used. The AMVR resolution set is used differently depending on the situation. For example, the AMVR resolution set is used differently depending on the prediction mode. For instance, the AMVR resolution set may differ when using regular interpretation like AMVP compared to when using affine MC (affine motion compensation). Also, AMVR applied to regular interpretation like AMVR is applied to the motion vector difference. Alternatively, AMVR applied to regular interpretation like AMVP (Advanced Motion Vector Prediction) is applied to the motion vector predictor. Furthermore, AMVR applied to affine MC (affine motion compensation) is applied to the control point motion vector or the control point motion vector difference.

[0185] Furthermore, according to one embodiment of this disclosure, there is a signaling that indicates whether AMVR is usable. This signaling is a high-level signaling. Referring to Figure 34, there is a sps_amvr_enabled_flag (AMVR enabled flag). Referring to Figures 2 and 7, the signaling means a signal transmitted from the encoder to the decoder via the bitstream. The decoder purges the AMVR enabled flag from the bitstream.

[0186] In one embodiment of this disclosure, the sps_amvr_enabled_flag (AMVR enabled flag) indicates whether adaptive motion vector difference resolution is used or not. Furthermore, in one embodiment of this disclosure, the sps_amvr_enabled_flag (AMVR enabled flag) indicates whether adaptive motion vector difference resolution is available or not. For example, according to one embodiment of this disclosure, if sps_amvr_enabled_flag (AMVR enabled flag) is 1, AMVR is available for motion vector coding. Also, according to one embodiment of this disclosure, if sps_amvr_enabled_flag (AMVR enabled flag) is 1, AMVR is used for motion vector coding. Furthermore, if sps_amvr_enabled_flag (AMVR enabled flag) is 1, there is further signaling to indicate which resolution is used. Also, if sps_amvr_enabled_flag (AMVR enabled flag) is 0, AMVR is not used for motion vector coding. Furthermore, if sps_amvr_enabled_flag (AMVR enabled flag) is 0, AMVR cannot be used for motion vector coding. According to one embodiment of this disclosure, AMVR corresponding to sps_amvr_enabled_flag (AMVR enabled flag) means that it is used for regular inter prediction. For example, AMVR corresponding to sps_amvr_enabled_flag (AMVR enabled flag) does not mean that it is used for affine MC (affine motion compensation). Also, the availability of affine MC (affine motion compensation) is indicated by inter_affine_flag (inter-affine flag). In other words, AMVR corresponding to sps_amvr_enabled_flag (AMVR enabled flag) means that it is used when inter_affine_flag (inter-affine flag) is 0, or it does not mean that it is used when inter_affine_flag (inter-affine flag) is 1.

[0187] According to one embodiment of this disclosure, there is a signaling that indicates whether AMVR is available for affine motion compensation (affine MC). This signaling is a high-level signaling. Referring to Figure 34, there is a signaling sps_affine_amvr_enabled_flag (affine AMVR enabled flag) that indicates whether AMVR is available for affine motion compensation (affine MC). Referring to Figures 2 and 7, the signaling means a signal transmitted from the encoder to the decoder via a bitstream. The decoder purges the affine AMVR enabled flag from the bitstream.

[0188] The sps_affine_amvr_enabled_flag (Affine AMVR enabled flag) indicates whether adaptive motion vector difference resolution is used for affine motion compensation. It also indicates whether adaptive motion vector difference resolution is available for affine motion compensation. According to one example, if sps_affine_amvr_enabled_flag (Affine AMVR enabled flag) is 1, AMVR is available for affine inter-mode motion vector coding. Furthermore, if sps_affine_amvr_enabled_flag (Affine AMVR enabled flag) is 0, AMVR is unavailable for affine inter-mode motion vector coding. If sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is 0, AMVR will not be used for affine inter-mode motion vector coding.

[0189] For example, if sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is 1, the AMVR corresponding to the one when inter_affine_flag (inter-affine flag) is 1 will be used. Also, if sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is 1, there is further signaling to indicate which resolution to use. If sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is 0, the AMVR corresponding to the one when inter_affine_flag (inter-affine flag) is 1 cannot be used.

[0190] Figure 35 shows the structure of a coding unit syntax according to one embodiment of the present disclosure.

[0191] As illustrated in Figure 34, there is further signaling to indicate the resolution based on the higher-level signaling using AMVR. Referring to Figure 34, the further signaling to indicate the resolution includes amvr_flag or amvr_precision_flag, which is information about the resolution of the motion vector difference.

[0192] According to one embodiment, there is a signaling mechanism indicated when amvr_flag is 0. Also, if amvr_flag is 1, amvr_precision_flag exists. Furthermore, if amvr_flag is 1, the resolution is determined based on amvr_precision_flag as well. For example, if amvr_flag is 0, the resolution is 1 / 4. If amvr_flag does not exist, the value of amvr_flag is inferred based on CuPredMode. For example, if CuPredMode is MODE_IBC, the value of amvr_flag is inferred to 1, and if CuPredMode is not MODE_IBC or CuPredMode (coding unit prediction mode) is MODE_INTER, the value of amvr_flag is inferred to 0.

[0193] Furthermore, if inter_affine_flag is 0 and amvr_precision_flag is 0, 1-pel resolution is used. Also, if inter_affine_flag is 1 and amvr_precision_flag is 0, 1 / 16-pel resolution is used. Also, if inter_affine_flag is 0 and amvr_precision_flag is 1, 4-pel resolution is used. Also, if inter_affine_flag is 1 and amvr_precision_flag is 1, 1-pel resolution is used.

[0194] If amvr_precision_flag is 0, infer its value to 0.

[0195] In one embodiment, the resolution is applied by the MvShift value. Furthermore, MvShift is determined by amvr_flag and amvr_precision_flag, which are information regarding the resolution of the motion vector difference. For example, if inter_affine_flag is 0, the MvShift value is determined as follows:

[0196] MvShift=(amvr_flag+amvr_precision_flag)<<1

[0197] Additionally, the Mvd (motion vector difference) value is shifted based on the MvShift value. For example, the Mvd (motion vector difference) is shifted as follows, and the AMVR resolution is applied accordingly.

[0198] MvdLX = MvdLX << (MvShift + 2)

[0199] As another example, if inter_affine_flag is 1, the MvShift value is determined as follows:

[0200] MvShift=amvr_precision_flag?(amvr_precision_flag<<1):(-(amvr_flag<<1))

[0201] Furthermore, the MvdCp (control point motion vector difference) value is shifted based on the MvShift value. MvdCp is either the control point motion vector difference or the control point motion vector. For example, MvdCp (control point motion vector difference) is shifted as follows, and the AMVR resolution is applied accordingly.

[0202] MvdCpLX = MvdCpLX << (MvShift + 2)

[0203] Additionally, Mvd or MvdCp are values ​​signaled by mvd_coding.

[0204] Referring to Figure 35, if CuPredMode is MODE_IBC, the amvr_flag, which contains information about the resolution of the motion vector difference, does not exist. However, if CuPredMode is MODE_INTER, the amvr_flag, which contains information about the resolution of the motion vector difference, does exist, and in this case, the amvr_flag can be parsed if certain conditions are met.

[0205] According to one embodiment of this disclosure, the parsing of syntax elements related to AMVR is determined based on a higher-level signaling value indicating whether AMVR is used. For example, if the higher-level signaling value indicating whether AMVR is used is 1, the syntax elements related to AMVR are parsed. If the higher-level signaling value indicating whether AMVR is used is 0, the syntax elements related to AMVR are not parsed. Referring to Figure 35, if CuPredMode is MODE_IBC and sps_amvr_enabled_flag (AMVR enabled flag) is 0, amvr_precision_flag, which is information about the resolution of the motion vector difference, is not parsed. Also, if CuPredMode is MODE_IBC and sps_amvr_enabled_flag (AMVR enabled flag) is 1, amvr_precision_flag, which is information about the resolution of the motion vector difference, is parsed. Further conditions for parsing are considered in this case.

[0206] For example, if at least one of the MvdLX (multiple motion vector differences) is not zero, amvr_precision_flag is parsed. MvdLX is the Mvd value for reference list LX. Also, Mvd (motion vector differences) are signaled via mvd_coding. LX includes L0 (0th reference picture list) and L1 (1st reference picture list). In addition, MvdLX has components corresponding to the x axis and y axis. For example, the x axis corresponds to the horizontal axis of the picture, and the y axis corresponds to the vertical axis of the picture. Referring to Figure 35, in MvdLX[x0][y0][0] and MvdLX[x0][y0][1], [0] and [1] indicate that they correspond to the x axis and y axis components, respectively. Also, if CuPredMode is MODE_IBC, only L0 is used. Referring to Figure 35, if 1) sps_amvr_enabled_flag is 1 and 2) MvdL0[x0][y0][0] or MvdL0[x0][y0][1] is not 0, then amvr_precision_flag is parsed. Also, if 1) sps_amvr_enabled_flag (AMVR enabled flag) is 0, or 2) both MvdL0[x0][y0][0] and MvdL0[x0][y0][1] are 0, then amvr_precision_flag is not parsed.

[0207] Furthermore, CuPredMode may not be MODE_IBC. In this case, referring to Figure 35, if sps_amvr_enabled_flag (AMVR enabled flag) is 1, inter_affine_flag (inter-affine flag) is 0, and at least one of the MvdLX (multiple motion vector difference) values ​​is non-zero, then amvr_flag is parsed. Here, amvr_flag is information about the resolution of the motion vector difference. As mentioned above, a sps_amvr_enabled_flag (AMVR enabled flag) of 1 indicates the use of adaptive motion vector difference resolution. Also, a inter_affine_flag (inter-affine flag) of 0 indicates that affine motion compensation is not used for the current block. Furthermore, as explained in Figure 34, the multiple motion vector differences for the current block are modified based on information about the resolution of the motion vector difference, such as amvr_flag. This condition is referred to as condition A. Furthermore, if sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is 1, inter_affine_flag (inter-affine flag) is 1, and at least one MvdCpLX (multiple control point motion vector difference) value is non-zero, amvr_flag is parsed. Here, amvr_flag is information about the resolution of the motion vector difference. As mentioned above, sps_affine_amvr_enabled_flag (affine AMVR enabled flag) being 1 indicates that an adaptive motion vector difference resolution is used for affine motion compensation. Also, inter_affine_flag (inter-affine flag) being 1 indicates that affine motion compensation is used for the current block. In addition, as explained in Figure 34, the multiple control point motion vector differences for the current block are modified based on information about the resolution of the motion vector difference, such as amvr_flag. This condition is referred to as condition B.

[0208] Furthermore, if either condition A or condition B is satisfied, amvr_flag is parsed. If neither condition A nor condition B is satisfied, amvr_flag, which contains information about the resolution of the motion vector difference, is not parsed. In other words, amvr_flag is not parsed if 1) sps_amvr_enabled_flag (AMVR enabled flag) is 0, inter_affine_flag (inter-affine flag) is 1, or MvdLX (multiple motion vector differences) is all 0, or 2) sps_affine_amvr_enabled_flag is 0, inter_affine_flag (inter-affine flag) is 0, or the MvdCpLX (multiple control point motion vector differences) value is all 0.

[0209] Furthermore, the parsing of amvr_precision_flag is determined based on the amvr_flag value. For example, if the amvr_flag value is 1, amvr_precision_flag will be parsed. If the amvr_flag value is 0, amvr_precision_flag will not be parsed.

[0210] Furthermore, MvdCpLX (Difference of Multiple Control Point Motion Vectors) represents the difference to the control point motion vector. MvdCpLX is also signaled via mvd_coding. LX includes L0 (0th reference picture list) and L1 (1st reference picture list). MvdCpLX also contains components corresponding to control point motion vectors 0, 1, 2, etc. For example, control point motion vectors 0, 1, 2, etc., are control point motion vectors corresponding to pre-set positions relative to the current block.

[0211] Referring to Figure 35, in MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][1][], and MvdCpLX[x0][y0][2][], [0], [1], and [2] indicate that they correspond to control point motion vectors 0, 1, and 2, respectively. The MvdCpLX (multiple control point motion vector difference) value corresponding to control point motion vector 0 is also used for other control point motion vectors. For example, control point motion vector 0 may be used as shown in Figure 47. In addition, MvdCpLX (multiple control point motion vector difference) has components corresponding to the x-axis and y-axis. For example, the x-axis corresponds to the horizontal axis of the picture, and the y-axis corresponds to the vertical axis of the picture. Referring to Figure 35, in MvdCpLX[x0][y0][][0] and MvdCpLX[x0][y0][][1], [0] and [1] indicate that they correspond to the x-axis and y-axis components, respectively.

[0212] Figure 36 shows the structure of a higher-level signaling according to one embodiment of the present disclosure. Higher-level signaling exists, as explained in Figures 34 and 35. For example, there are flags such as sps_affine_enabled_flag, sps_affine_amvr_enabled_flag, sps_amvr_enabled_flag, and sps_affine_type_flag.

[0213] According to one embodiment of the present disclosure, the higher-level signaling has a parsing dependency. For example, it determines whether or not to parsing other higher-level signalings based on which higher-level signaling value.

[0214] According to one embodiment of this disclosure, the availability of affine AMVR is determined based on whether affine MC (affine motion compensation) is available. For example, the availability of affine AMVR is determined based on higher-level signaling indicating whether affine MC is available. More specifically, the parsing of higher-level signaling indicating whether affine AMVR is available is determined based on higher-level signaling indicating whether affine MC is available.

[0215] As one example, if affine MC is usable, affine AMVR may also be usable. Conversely, if affine MC is not usable, affine AMVR may also be unusable.

[0216] More specifically, if the higher-level signaling indicating whether affine MC is available is 1, then affine AMVR is available. In this case, further signaling may exist. Also, if the higher-level signaling indicating whether affine MC is available is 0, then affine AMVR is unavailable. For example, if the higher-level signaling indicating whether affine MC is available is 1, then the higher-level signaling indicating whether affine AMVR is available is parsed. Also, if the higher-level signaling indicating whether affine MC is available is 0, then the higher-level signaling indicating whether affine AMVR is available is not parsed. Also, if there is no higher-level signaling indicating whether affine AMVR is available, its value is inferred. For example, inferred to 0. As another example, inferred based on the higher-level signaling indicating whether affine MC is available. As yet another example, inferred based on the higher-level signaling indicating whether AMVR is available.

[0217] In one embodiment, the affine AMVR is the AMVR used for the affine MC described in Figures 34 and 35. For example, the higher-level signaling indicating whether the affine MC is usable is sps_affine_enabled_flag. Also, the higher-level signaling indicating whether the affine AMVR is usable is sps_affine_amvr_enabled_flag.

[0218] Referring to Figure 36, if sps_affine_enabled_flag (affine flag) is 1, then sps_affine_amvr_enabled_flag is parsed. If sps_affine_enabled_flag (affine flag) is 0, then sps_affine_amvr_enabled_flag is not parsed. Also, if sps_affine_amvr_enabled_flag does not exist, its value is inferred to 0.

[0219] The embodiments of this disclosure are meaningful because affine AMVR is used when affine MC is used.

[0220] Figure 37 shows the structure of a higher-level signaling ring according to one embodiment of the present disclosure. Higher-level signaling exists, as explained in Figures 34 and 35. For example, there are flags such as sps_affine_enabled_flag (affine enabled flag), sps_affine_amvr_enabled_flag (affine AMVR enabled flag), sps_amvr_enabled_flag (AMVR enabled flag), and sps_affine_type_flag.

[0221] According to one embodiment of this disclosure, the higher-level signaling has a parsing dependency. For example, it determines whether or not to parsing other higher-level signalings based on which higher-level signaling value. According to one embodiment of this disclosure, the availability of affine AMVR is determined based on whether AMVR is available. For example, the availability of affine AMVR is determined based on higher-level signaling indicating whether AMVR is available. More specifically, the parsing feasibility of higher-level signaling indicating whether affine AMVR is available is determined based on higher-level signaling indicating whether AMVR is available.

[0222] As one example, if AMVR is available, affine AMVR may be available. Conversely, if AMVR is unavailable, affine AMVR may not be available.

[0223] More specifically, if the higher-level signaling indicating AMVR is available is 1, then affine AMVR is available. In this case, additional signaling may exist. Also, if the higher-level signaling indicating AMVR is available is 0, then affine AMVR is not available. For example, if the higher-level signaling indicating AMVR is available is 1, then the higher-level signaling indicating affine AMVR is available is parsed. Also, if the higher-level signaling indicating AMVR is available is 0, then the higher-level signaling indicating affine AMVR is available is not parsed. Furthermore, if there is no higher-level signaling indicating affine MC is available, its value is inferred. For example, inferred to 0. As another example, inferred based on the higher-level signaling indicating affine MC is available. As yet another example, inferred based on the higher-level signaling indicating AMVR is available.

[0224] In one embodiment, affine AMVR is an AMVR used for affine MC (affine motion compensation) as described in Figures 34 and 35. For example, the higher-level signaling indicating whether AMVR is available is sps_amvr_enabled_flag (AMVR enabled flag). Similarly, the higher-level signaling indicating whether affine AMVR is available is sps_affine_amvr_enabled_flag (affine AMVR enabled flag).

[0225] Referring to Figure 37(a), if sps_amvr_enabled_flag (AMVR enabled flag) is 1, then sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is parsed. If sps_amvr_enabled_flag (AMVR enabled flag) is 0, then sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is not parsed. Also, if sps_affine_amvr_enabled_flag (affine AMVR enabled flag) does not exist, its value is inferred to 0.

[0226] One embodiment of this disclosure is based on the fact that the efficiency of adaptive resolution may vary depending on the sequence.

[0227] Furthermore, the availability of affine AMVR is determined by considering both the availability of affine MC and AMVR. For example, the parsing of the higher-level signaling indicating whether affine AMVR is available is determined based on the higher-level signaling indicating whether affine MC is available and the higher-level signaling indicating whether AMVR is available. In one embodiment, if both the higher-level signaling indicating whether affine MC is available and the higher-level signaling indicating whether AMVR is available are 1, the higher-level signaling indicating whether affine AMVR is available is parsed. If either the higher-level signaling indicating whether affine MC is available or the higher-level signaling indicating whether AMVR is available is 0, the higher-level signaling indicating whether affine AMVR is available is not parsed. Also, if there is no higher-level signaling indicating whether affine MC is available, its value is inferred.

[0228] Referring to Figure 37(b), if both sps_affine_enabled_flag (affine-enabled flag) and sps_amvr_enabled_flag (AMVR-enabled flag) are 1, then sps_affine_amvr_enabled_flag (affine AMVR-enabled flag) is parsed. Also, if at least one of sps_affine_enabled_flag (affine-enabled flag) or sps_amvr_enabled_flag (AMVR-enabled flag) is 0, then sps_affine_amvr_enabled_flag (affine AMVR-enabled flag) is not parsed. Furthermore, if sps_affine_amvr_enabled_flag (affine AMVR-enabled flag) does not exist, its value is inferred to 0.

[0229] To explain Figure 37(b) in more detail, at line 3701, the availability of affine motion compensation is determined based on the sps_affine_enabled_flag (affine-enabled flag). As already explained in Figures 34 and 35, if sps_affine_enabled_flag (affine-enabled flag) is 1, it means that affine motion compensation is available. Conversely, if sps_affine_enabled_flag (affine-enabled flag) is 0, it means that affine motion compensation is unavailable.

[0230] If it is determined at line 3701 in Figure 37(b) that affine motion compensation will be used, then at line 3702, the use of adaptive motion vector difference resolution is determined based on sps_amvr_enabled_flag (AMVR enabled flag). As already explained in Figures 34 to 35, if sps_amvr_enabled_flag (AMVR enabled flag) is 1, it means that adaptive motion vector difference resolution will be used. If sps_amvr_enabled_flag (AMVR enabled flag) is 0, it means that adaptive motion vector difference resolution will not be used.

[0231] If it is determined in line 3701 that affine motion compensation will not be used, the availability of adaptive motion vector difference resolution will not be determined based on sps_amvr_enabled_flag (AMVR enabled flag). In other words, line 3702 in Figure 37(b) will not occur. More specifically, if affine motion compensation is not used, sps_affine_amvr_enabled_flag (affine AMVR enabled flag) will not be transmitted from the encoder to the decoder. In other words, the decoder will not receive sps_affine_amvr_enabled_flag (affine AMVR enabled flag), and sps_affine_amvr_enabled_flag (affine AMVR enabled flag) will not be purged. In this case, since sps_affine_amvr_enabled_flag (affine AMVR enabled flag) does not exist, it is implicitly set to 0. As mentioned above, if sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is 0, it indicates that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0232] If it is determined at line 3072 in Figure 37(b) that adaptive motion vector difference resolution should be used, then at line 3703, sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is purged from the bitstream, indicating whether adaptive motion vector difference resolution can be used for affine motion compensation.

[0233] If it is determined at line 3072 that adaptive motion vector difference resolution will not be used, the sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is not purged from the bitstream. In other words, line 3703 in Figure 37(b) does not occur. More specifically, if affine motion compensation is used and adaptive motion vector difference resolution is not used, the sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is not transmitted from the encoder to the decoder. In other words, the decoder does not receive the sps_affine_amvr_enabled_flag (affine AMVR enabled flag). The decoder does not purge the sps_affine_amvr_enabled_flag (affine AMVR enabled flag) from the bitstream. In this case, the sps_affine_amvr_enabled_flag (affine AMVR enabled flag) does not exist and is therefore implicitly set to 0. As mentioned above, if sps_affine_amvr_enabled_flag (affine AMVR enabled flag) is 0, it indicates that adaptive motion vector difference resolution cannot be used for affine motion compensation.

[0234] As shown in Figure 37(b), checking sps_affine_enabled_flag (affine-enabled flag) first and then sps_amvr_enabled_flag (AMVR-enabled flag) can reduce unnecessary processing steps and thus improve efficiency. For example, if sps_amvr_enabled_flag (AMVR-enabled flag) is checked first and then sps_affine_enabled_flag (affine-enabled flag) is checked next, then sps_affine_enabled_flag (affine-enabled flag) must be checked again in order to derive sps_affine_type_flag from the 7th row. However, by checking sps_affine_enabled_flag (affine-enabled flag) first and then sps_amvr_enabled_flag (AMVR-enabled flag) (AMVR-enabled flag), such unnecessary steps can be reduced.

[0235] Figure 38 shows the structure of a coding unit syntax according to one embodiment of the present disclosure.

[0236] As explained in Figure 35, whether or not syntax parsing related to AMVR is performed is determined by whether or not there is at least one non-zero value among MvdLX or MvdCpLX. However, the MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) used may differ depending on which reference list (reference picture list) is used and how many parameters the affine model uses. If the initial value of MvdLX or MvdCpLX is not 0, then since there are no MvdLX or MvdCpLX currently named in the block that are not 0, unnecessary AMVR-related syntax elements will be signaled, which may cause a mismatch between the encoder and decoder. Figures 38 and 39 illustrate how to avoid a mismatch between the encoder and decoder.

[0237] In one embodiment, inter_pred_idc (information about the reference picture list) indicates which reference list to use or the direction of the prediction. For example, inter_pred_idc (information about the reference picture list) is a value of PRED_L0, PRED_L1, or PRED_BI. If inter_pred_idc (information about the reference picture list) is PRED_L0, only reference list 0 (the 0th reference picture list) is used. If inter_pred_idc (information about the reference picture list) is PRED_L1, only reference list 1 (the 1st reference picture list) is used. If inter_pred_idc (information about the reference picture list) is PRED_BI, both reference list 0 (the 0th reference picture list) and reference list 1 (the 1st reference picture list) are used. If inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_L1, it is uni-prediction. Furthermore, if inter_pred_idc (information about the reference picture list) is PRED_BI, then it is a bi-prediction.

[0238] Furthermore, the system determines which affine model to use based on the MotionModelIdc value. It also determines which affine MC to use based on the MotionModelIdc value. For example, MotionModelIdc indicates translation motion, 4-parameter affine motion, or 6-parameter affine motion. For instance, if MotionModelIdc has translation motion values ​​of 0, 1, and 2, it indicates translation motion, 4-parameter affine motion, and 6-parameter affine motion, respectively. In one embodiment, MotionModelIdc is determined based on inter_affine_flag and cu_affine_type_flag. For example, if merge_flag is 0 (not merge mode), MotionModelIdc is determined based on inter_affine_flag and cu_affine_type_flag. For example, MotionModelIdc is (inter_affine_flag + cu_affine_type_flag). According to other examples, MotionModelIdc is determined by merge_subblock_flag. For example, if merge_flag is 1 (i.e., merge mode), MotionModelIdc is determined by merge_subblock_flag. For example, the MotionModelIdc value is set to the merge_subblock_flag value.

[0239] For example, if inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_BI, then the value corresponding to L0 is used in MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference). Therefore, when parsing syntax related to AMVR, MvdL0 or MvdCpL0 is only considered if inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_BI. In other words, if inter_pred_idc (information about the reference picture list) is PRED_L1, then MvdL0 (motion vector difference for the 0th reference picture list) or MvdCpL0 (control point motion vector difference for the 0th reference picture list) does not need to be considered.

[0240] Furthermore, if inter_pred_idc (information about the reference picture list) is PRED_L1 or PRED_BI, the value corresponding to L1 is used in MvdLX or MvdCPLX. Therefore, when parsing syntax related to AMVR, MvdL1 (motion vector difference for the first reference picture list) or MvdCpL1 (control point motion vector difference for the first reference picture list) is considered only if inter_pred_idc (information about the reference picture list) is PRED_L1 or PRED_BI. In other words, if inter_pred_idc (information about the reference picture list) is PRED_L0, MvdL1 or MvdCpL1 does not need to be considered.

[0241] Referring to Figure 38, MvdL0 and MvdCpL0 determine whether to parse AMVR-related syntax based on whether their value is non-zero, only if inter_pred_idc is not PRED_L1. In other words, if inter_pred_idc is PRED_L1, AMVR-related syntax will not be parsed if there is a non-zero value in MvdL0 or MvdCpL0.

[0242] Furthermore, if MotionModelIdc is 1, only MvdCpLX[x0][y0][0][] and MvdCpLX[x0][y0][1][] are considered from MvdCpLX[x0][y0][0][], MvdCpLXx0][y0][1][], and MvdCpLX[x0][y0][2][]. In other words, if MotionModelIdc is 1, MvdCpLX[x0][y0][2][] is not considered. For example, if MotionModelIdc is 1, whether or not there are non-zero values ​​in MvdCpLX[x0][y0][2][] does not affect whether AMVR-related syntax can be parsed.

[0243] Furthermore, if MotionModelIdc is 2, then MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][1][], and MvdCpLX[x0][y0][2][] are all considered. In other words, if MotionModelIdc is 2, then MvdCpLX[x0][y0][2][] is considered.

[0244] Furthermore, in the above embodiment, the MotionModelIdc values ​​of 1 and 2 can also be expressed as cu_affine_type_flag being 0 and 1. This is because it is possible to determine whether affine MC can be used or not. For example, the availability of affine MC can be determined by inter_affine_flag.

[0245] Referring to Figure 38, we consider whether there is a non-zero value in MvdCpLX[x0][y0][2][] only if MotionModelIdc is 2. If both MvdCpLX[x0][y0][0][] and MvdCpLX[x0][y0][1][] are 0, and there is at least one non-zero value in MvdCpLX[x0][y0][2][] (as in the embodiment described above, we may separate L0 and L1 here and consider only the value corresponding to one of L0 or L1), then unless MotionModelIdc is 2, we do not parse the syntax related to AMVR.

[0246] Figure 39 shows the setting of MVD base values ​​according to one embodiment of the present disclosure. As mentioned above, MvdLX (motion vector difference) or MvdCpLX (control point motion difference) is signaled by mvd_coding. In addition, the IMvd value is signaled by mvd_coding, and MvdLX or MvdCpLX is set to the IMvd value. Referring to Figure 39, if MotionModelIdc is 0, MvdLX is set by the IMvd value. If MotionModelIdc is not 0, MvdCpLX is set by the IMvd value. Furthermore, the refList value determines which of the LX actions to perform. Furthermore, as explained in Figure 10 or Figures 23, 34, and 35, there is an mvd_coding. The mvd_coding includes a step of parsing or determining abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, mvd_sign_flag, etc. The IMvd is determined by abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, mvd_sign_flag, etc.

[0247] Referring to Figure 39, IMvd is configured as follows. IMvd=abs_mvd_greater0_flag*(abs_mvd_minus2+2)*(1-2*mvd_sign_flag)

[0248] The base value of MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) is set to a preset value. According to one embodiment of this disclosure, the base value of MvdLX or MvdCpLX is set to 0. Alternatively, the base value of IMvd is set to a preset value. Alternatively, the base value of the associated syntax element is set so that the value of IMvd, MvdLX, or MvdCpLX is a preset value. The base value of the syntax element means the value to be inferred if the syntax element does not exist. The preset value is 0.

[0249] According to one embodiment of this disclosure, abs_mvd_greater0_flag indicates whether the absolute value of MVD is greater than 0. Also according to one embodiment of this disclosure, if abs_mvd_greater0_flag does not exist, its value is inferred to 0. In such cases, the IMvd value is set to 0. In such cases, the MvdLX or MvdCpLX value is also set to 0. Alternatively, according to one embodiment of this disclosure, if an IMvd, MvdLX, or MvdCpLX value is not set, it is set to a preset value. For example, it is set to 0.

[0250] Furthermore, according to one embodiment of this disclosure, if abs_mvd_greater0_flag does not exist, the corresponding IMvd, MvdLX, or MvdCpLX value is set to 0.

[0251] According to one embodiment of this disclosure, abs_mvd_greater1_flag indicates whether the absolute value of the MVD is greater than 1 or not. If abs_mvd_greater1_flag does not exist, its value is inferred to 0. Also, abs_mvd_minus2 + 2 indicates the absolute value of MVD. If the abs_mvd_minus2 value does not exist, infer its value as -1.

[0252] Also, mvd_sign_flag indicates the sign of MVD. If mvd_sign_flag is 0 or 1, it indicates that the corresponding MVD is a positive or negative value, respectively. If the mvd_sign_flag does not exist, infer its value as 0.

[0253] Figure 40 is a diagram showing the setting of the basic value of MVD according to an embodiment of the present disclosure. As described in FIG. 39 to prevent a missmatch between the encoder and the decoder by setting the initial values of MvdLX or MvdCpLX to 0, the values of MvdLX (multiple motion vector differences) or MvdCpLX (multiple control point motion vector differences) are initialized. At this time, the value to be initialized is 0. The embodiment of FIG. 40 will be described in more detail about this.

[0254] According to an embodiment of the present disclosure, the values of MvdLX (multiple motion vector differences) or MvdCpLX (multiple control point motion vector differences) are initialized to preset values. Also, the initialization position is after the position of parsing the syntax elements related to AMVR. The syntax elements related to the AMVR include amvr_flag, amvr_precision_flag, etc. in FIG. 40. amvr_flag or amvr_precision_flag is information regarding the resolution of the motion vector difference.

[0255] The resolution of MVD (motion vector difference) or MV (motion vector), or the resolution signaled by MVD or MV is determined by the syntax elements related to the AMVR. Also, in the embodiment of the present disclosure, the preset value to be initialized is 0.

[0256] By performing library diffractivity, the MvdLX and MvdCpLX values ​​between the encoder and decoder become the same when parsing AMVR-related syntax elements, thus preventing mismatches between the encoder and decoder. Additionally, by setting the initial value to 0, AMVR-related syntax elements are not unnecessarily included in the bitstream.

[0257] According to one embodiment of this disclosure, MvdLX is defined for a reference List (L0, L1, etc.), an x- or y-component, etc. MvdCpLX is defined for a reference List (L0, L1, etc.), an x- or y-component, control points 0, 1, 2, etc.

[0258] According to one embodiment of this disclosure, the values ​​of either MvdLX or MvdCpLX are initialized. Furthermore, the initialization takes place before mvd_coding is performed on the corresponding MvdLX or MvdCpLX. For example, even in a precision block that uses only L0, the MvdLX and MvdCpLX corresponding to L0 can be initialized. In order to initialize only those that are absolutely necessary, it is necessary to check the conditions, but by initializing without such distinction, the burden of checking the conditions can be reduced.

[0259] According to other embodiments of this disclosure, the values ​​corresponding to the unused MvdLX or MvdCpLX values ​​are initialized. Here, "unused" means not currently used in the block. For example, the values ​​corresponding to the currently unused reference list of MvdLX or MvdCpLX are initialized. For example, if L0 is not used, the values ​​corresponding to MvdL0 and MvdCpL0 are initialized. L0 is not used when inter_pred_idc is PRED_L1. Also, if L1 is not used, the values ​​corresponding to MvdL1 and MvdCpL1 are initialized. L1 is not used when inter_pred_idc is PRED_L0. L0 is used when inter_pred_idc is PRED_L0 or PRED_BI. Also, L1 is used when inter_pred_idc is PRED_L1 or PRED_BI. Referring to Figure 40, if inter_pred_idc is PRED_L1, then MvdL0[x0][y0][0], MvdL0[x0][y0][1], MvdCpL0[x0][y0][0][0], MvdCpL0[x0][y0][0][1], MvdCpL0[x0][y0][1][0], MvdCpL0[x0][y0][1][1], MvdCpL0[x0][y0][2][0], and MvdCpL0[x0][y0][2][1] are initialized. The value to be initialized is 0. Additionally, if inter_pred_idc is PRED_L0, initialize MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. The value to be initialized is 0.

[0260] MvdLX[x][y][compIdx] is the (x,y) position relative to the reference list LX and the motion vector difference relative to the component index compIdx. MvdCpLX[x][y][cpIdx][compIdx] is the motion vector difference relative to the reference list LX. Also, MvdCpLX[x][y][cpIdx][compIdx] is the motion vector difference relative to the position (x,y), the control point motion vector cpIdx, and the component index compIdx. Here, component refers to the x or y component.

[0261] Furthermore, according to one embodiment of this disclosure, the choice of whether to use MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) depends on whether affine motion compensation is available. For example, if affine motion compensation is available, MvdLX is initialized. If affine motion compensation is not available, MvdCpLX is initialized. For example, there is a signaling indicator that indicates whether affine motion compensation is available. Referring to Figure 40, inter_affine_flag is a signaling indicator that indicates whether affine motion compensation is available. For example, if inter_affine_flag is 1, affine motion compensation is used. Alternatively, MotionModelIdc is a signaling indicator that indicates whether affine motion compensation is available. For example, if MotionModelIdc is not 0, affine motion compensation is used.

[0262] Furthermore, according to one embodiment of this disclosure, the unused MvdLX or MvdCpLX parameters are used for affine motion compensation. For example, the unused MvdLX or MvdCpLX parameters may differ depending on whether a 4-parameter affine model or a 6-parameter affine model is used. For example, the unused cpIdx parameters in MvdCpLX[x][y][cpIdx][compIdx] may differ depending on which affine motion compensation is used. For example, when using a 4-parameter affine model, only a portion of MvdCpLX[x][y][cpIdx][compIdx] are used. Or, when not using a 6-parameter affine model, only a portion of MvdCpLX[x][y][cpIdx][compIdx] are used. Therefore, the unused MvdCpLX parameters are initialized to a pre-set value. In this case, the unused MvdCpLX corresponds to cpIdx, which is used in the 6-parameter affine model but not in the 4-parameter affine model. For example, if the 4-parameter affine model is used, the value of cpIdx in MvdCpLX[x][y][cpIdx][compIdx] that is 2 is not used, and it is initialized to a pre-set value. Also, as mentioned above, there is a signaling or parameter that indicates whether the 4-parameter affine model or the 6-parameter affine model is used. For example, MotionModelIdc or cu_affine_type_flag indicates whether the 4-parameter affine model or the 6-parameter affine model is used. A MotionModelIdc value of 1 or 2 indicates the use of the 4-parameter affine model and the 6-parameter affine model, respectively.Referring to Figure 40, if MotionModelIdc is 1, MvdCpL0[x0][y0][2][0], MvdCpL0[x0][y0][2][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1] are initialized to their pre-set values. Alternatively, instead of using the condition where MotionModelIdc is 1, the condition where MotionModelIdc is not 2 can be used. The aforementioned pre-set value is 0.

[0263] Furthermore, according to one embodiment of this disclosure, the unused MvdLX or MvdCpLX is determined based on mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list). For example, if mvd_l1_zero_flag is 1, MvdL1 and MvdCpL1 are initialized to a preset value. In addition, further conditions are considered. For example, the unused MvdLX or MvdCpLX is determined based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list). For example, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, MvdL1 and MvdCpL1 are initialized to a preset value. For example, mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) is a high-level signaling that the MVD value (e.g., MvdLX or MvdCpLX) for reference list L1 is 0. Signaling refers to a signal transmitted from the encoder to the decoder via the bitstream. The decoder purges mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) from the bitstream.

[0264] In other embodiments of this disclosure, all Mvd and MvdCp values ​​are initialized to pre-set values ​​before performing mvd_coding on a given block. In this case, mvd_coding refers to all mvd_coding for a given CU. Therefore, the problem described is solved by avoiding the need to parsing the mvd_coding syntax and initializing the determined Mvd or MvdCp value, thereby initializing all Mvd and MvdCP values.

[0265] Figure 41 shows the syntax structure related to AMVR according to one embodiment of the present disclosure. The embodiment in Figure 41 is based on the embodiment in Figure 38.

[0266] As explained in Figure 38, we check whether there is at least one non-zero value among MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference). In this case, the MvdLX or MvdCpLX to check is determined based on mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list). Therefore, the parsing of syntax related to AMVR is determined based on mvd_l1_zero_flag. As mentioned above, if mvd_l1_zero_flag is 1, it indicates that the MVD (motion vector difference) for reference list L1 is 0, so in this case, MvdL1 (motion vector difference for the first reference picture list) or MvdCpL1 (control point motion vector difference for the first reference picture list) is not considered. For example, if mvd_l1_zero_flag is 1, regardless of whether MvdL1 or MvdCpL1 is 0 or not, the syntax related to AMVR is parsed based on whether there is at least one value of 0 among MvdL0 (motion vector difference for the 0th reference picture list) or MvdCpL0 (control point motion vector difference for the 0th reference picture list). In a further embodiment, mvd_l1_zero_flag indicates that the MVD for reference list L1 is 0 only for the bi-prediction block. Therefore, the syntax related to AMVR is parsed based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list). For example, MvdLX or MvdCpLX is determined based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list) to determine whether there is at least one non-zero value. For example, based on mvd_l1_zero_flag and inter_pred_idc, it is not necessary to consider MvdL1 or MvdCpL1. For instance, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, then MvdL1 or MvdCpL1 is not considered.For example, if mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, then regardless of whether MvdL1 or MvdCpL1 is 0 or not, the syntax related to AMVR will be parsed based on whether there is at least one value of 0 among MvdL0 or MvdCpL0.

[0267] Referring to Figure 41, if mvd_l1_zero_flag is 1 and inter_pred_idc[x0][y0]==PRED_BI, then the operation will be performed regardless of whether or not there are non-zero values ​​among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. For example, if mvd_l1_zero_flag is 1 and inter_pred_idc[x0][y0]==PRED_BI, then even if there are no non-zero values ​​among MvdL0 or MvdCpL0, the syntax related to AMVR will not be parsed, even if there are non-zero values ​​among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1].

[0268] Additionally, if mvd_l1_zero_flag is 0, or if inter_pred_idc[x0][y0]!=PRED_BI, then consider whether there is a non-zero value among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. For example, if mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0]!=PRED_BI, then if there is a non-zero value among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], MvdCpL1[x0][y0][2][1], then the syntax related to AMVR will be parsed. For example, if mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0]!=PRED_BI, then if there is a non-zero value among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], MvdCpL1[x0][y0][2][1], then the syntax related to AMVR will be parsed even if both MvdL0 and MvdCpL0 are 0.

[0269] Figure 42 shows the structure of a syntax related to interpretation according to one embodiment of the present disclosure.

[0270] According to one embodiment of this disclosure, mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) is a signaling that the Mvd (motion vector difference) value for reference list L1 (the first reference picture list) is 0. Furthermore, this signaling is currently signaled at a higher level than the block. Therefore, based on the mvd_l1_zero_flag value, the Mvd value for reference list L1 can be 0 for many blocks. For example, if the mvd_l1_zero_flag value is 1, the Mvd value for reference list L1 is 0. Alternatively, the Mvd value for reference list L1 is 0 based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list). For example, if the mvd_l1_zero_flag value is 1 and inter_pred_idc is PRED_BI, the Mvd value for reference list L1 is 0. In this case, the Mvd value is MvdL1[x][y][compIdx]. Furthermore, the Mvd value does not represent the control point motion vector difference. In other words, the Mvd value does not represent the MvdCp value.

[0271] According to other embodiments of this disclosure, mvd_l1_zero_flag is a signaling that indicates the Mvd and MvdCp values ​​for reference list L1 are 0. This signaling is currently signaled at a higher level than the block. Therefore, based on the mvd_l1_zero_flag value, the Mvd and MvdCp values ​​for reference list L1 can be 0 for many blocks. For example, if the mvd_l1_zero_flag value is 1, the Mvd and MvdCp values ​​for reference list L1 are 0. Alternatively, the Mvd and MvdCp values ​​for reference list L1 can be 0 based on mvd_l1_zero_flag and inter_pred_idc. For example, if the mvd_l1_zero_flag value is 1 and inter_pred_idc is PRED_BI, the Mvd and MvdCp values ​​for reference list L1 are 0. In this case, the Mvd value is MvdL1[x][y][compIdx]. Also, the MvdCp value is MvdCpL1[x][y][cpIdx][compIdx].

[0272] Furthermore, an Mvd or MvdCp value of 0 means that the corresponding mvd_coding syntax structure will not be parsed. In other words, for example, if the mvd_l1_zero_flag value is 1, the mvd_coding syntax structure corresponding to MvdL1 or MvdCpL1 will not be parsed. Conversely, if the mvd_l1_zero_flag value is 0, the mvd_coding syntax structure corresponding to MvdL1 or MvdCpL1 will be parsed.

[0273] According to an embodiment of the present disclosure, if the MvdCp value is 0 based on mvd_l1_zero_flag, the signaling indicating MVP is not parsed. The signaling indicating MVP includes mvp_l1_flag. Also, according to the description of mvd_l1_zero_flag mentioned above, the mvd_l1_zero_flag signaling means that it is 0 for both Mvd and MvdCp. For example, if the condition indicating that the Mvd or MvdCp value for the reference list L1 is 0 is satisfied, the signaling indicating MVP is not parsed. In such a case, the signaling indicating MVP is inferred to a preset value. For example, if there is no signaling indicating MVP, its value is inferred to 0. Also, if the condition indicating that the Mvd or MvdCp value is 0 based on mvd_l1_zero_flag is not satisfied, the signaling indicating MVP is parsed. However, in such an embodiment, if the Mvd or MvdCp value is 0, there may be no freedom to select MVP, which may reduce the coding efficiency.

[0274] More specifically, if the condition indicating that the Mvd and MvdCp values for the reference list L1 are 0 is satisfied and affine MC is used, the signaling indicating MVP is not parsed. In such a case, the signaling indicating MVP is inferred to a preset value.

[0275] Referring to FIG. 42, if mvd_l1_zero_flag is 1 and the inter_pred_idc value is PRED_BI, mvp_l1_flag is not parsed. Also, in such a case, the mvp_l1_flag value is inferred to 0. Or, if mvd_l1_zero_flag is 0 and the inter_pred_idc value is not PRED_BI, mvp_l1_flag is parsed.

[0276] In this embodiment, determining whether a signaling indicating MVP can be parsed based on mvd_l1_zero_flag occurs when certain conditions are met. For example, certain conditions include the condition that general_merge_flag is 0. For example, general_merge_flag has the same meaning as merge_flag as described above. Also, certain conditions include conditions based on CuPerdMode. More specifically, certain conditions include the condition that CuPerdMode is not MODE_IBC. Also, certain conditions include the condition that CuPerdMode is MODE_INTER. If CuPerdMode is MODE_IBC, the prediction that references the current picture is used. Also, if CuPerdMode is MODE_IBC, there is a block vector or motion vector corresponding to that block. If CuPerdMode is MODE_INTER, the prediction that references a picture that is not the current picture is used. If CuPerdMode is MODE_INTER, there is a motion vector corresponding to that block. Therefore, according to one embodiment of this disclosure, if general_merge_flag is 0, CuPerdMode is not MODE_IBC, mvd_l1_zero_flag is 1, and inter_pred_idc is PRED_BI, then mvp_l1_flag is not parsed. Also, if mvp_l1_flag does not exist, its value is inferred to 0.

[0277] More specifically, if general_merge_flag is 0, CuPerdMode is not MODE_IBC, mvd_l1_zero_flag is 1, inter_pred_idc is PRED_BI, and affine MC is used, then mvp_l1_flag will not be parsed. Also, if mvp_l1_flag does not exist, its value will be inferred to 0.

[0278] Referring to Figure 42, sym_mvd_flag is the signaling that indicates a symmetric MVD. In the case of a symmetric MVD, other MVDs are determined based on one MVD. In the case of a symmetric MVD, other MVDs are determined based on an explicitly signaled MVD. For example, in the case of a symmetric MVD, the MVD for one reference list is determined based on the MVD for one reference list. For example, in the case of a symmetric MVD, the MVD for another reference list L1 is determined based on the MVD for reference list L0. When determining other MVDs based on one MVD, the value obtained by reversing the sign of the aforementioned MVD is determined as the other MVD.

[0279] Figure 43 shows the structure of a syntax related to interpretation according to one embodiment of the present disclosure.

[0280] The embodiment in Figure 43 is an embodiment for solving the problem described in Figure 42. According to one embodiment of this disclosure, if the Mvd (motion vector difference) or MvdCp (control point motion vector difference) value is 0 based on mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list), the signaling indicating MVP (motion vector predictor) is parsed. The signaling indicating MVP includes mvp_l1_flag (motion vector predictor index for the first reference picture list). Furthermore, according to the explanation of mvd_l1_zero_flag above, the mvd_l1_zero_flag signaling means that both Mvd and MvdCp are 0. For example, if the condition that the Mvd or MvdCp value for reference list L1 (first reference picture list) is 0 is satisfied, the signaling indicating MVP is parsed. Therefore, the signaling indicating MVP is not inferred. This provides the flexibility to select the MVP even if the Mvd or MvdCp value is 0 based on the mvd_l1_zero_flag. This improves coding efficiency. Furthermore, even if the conditions indicating that the Mvd or MvdCp value is 0 based on the mvd_l1_zero_flag are not met, the system will still parse the signaling indicating the MVP.

[0281] More specifically, if the conditions are met that the Mvd or MvdCp value for reference list L1 is 0, and affine MC is used, then the signaling indicating MVP is parsed.

[0282] Referring to line 4301 in Figure 43, information about the reference picture list (inter_pred_idc) is obtained for the current block. Referring to line 4302, if the information about the reference picture list (inter_pred_idc) indicates that it does not use only the 0th reference picture list (list 0), then at line 4303, the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1) is purged from the bitstream.

[0283] The mvd_l1_zero_flag (motion vector difference zero flag) is obtained from the bitstream. The mvd_l1_zero_flag indicates whether MvdLX (motion vector difference) and MvdCpLX (multiple control point motion vector difference) are set to 0 for the first reference picture list. Signaling refers to the signals transmitted from the encoder to the decoder via the bitstream. The decoder purges the mvd_l1_zero_flag (motion vector difference zero flag) from the bitstream.

[0284] If mvd_l1_zero_flag (motion vector difference zero flag) is 1 and the inter_pred_idc (information about the reference picture list) value is PRED_BI, then parsing mvp_l1_flag (motion vector predictor index) is performed. Here, PRED_BI indicates that both List 0 (0th reference picture list) and List 1 (1st reference picture list) are used. Alternatively, if mvd_l1_zero_flag (motion vector difference zero flag) is 0 and the inter_pred_idc (information about the reference picture list) value is not PRED_BI, then parsing mvp_l1_flag (motion vector predictor index) is performed. In other words, regardless of whether mvd_l1_zero_flag (motion vector difference zero flag) is 1 and inter_pred_idc (information about the reference picture list) indicates the use of both the 0th and 1st reference picture list, mvp_l1_flag (motion vector predictor index) is parsed.

[0285] In this embodiment, determining Mvd and MvdCp based on mvd_l1_zero_flag (motion vector difference zero flag of the first reference picture list) and parsing the signaling indicating MVP is possible when certain conditions are met. For example, the certain conditions include the condition that general_merge_flag is 0. For example, general_merge_flag has the same meaning as merge_flag described above. The certain conditions also include conditions based on CuPerdMode. More specifically, the certain conditions include the condition that CuPerdMode is not MODE_IBC. The certain conditions also include the condition that CuPerdMode is MODE_INTER. If CuPerdMode is MODE_IBC, the prediction that references the current picture is used. Also, if CuPerdMode is MODE_IBC, there is a block vector or motion vector corresponding to that block. If CuPerdMode is MODE_INTER, the prediction that references a picture that is not the current picture is used. If CuPerdMode is MODE_INTER, then a motion vector corresponding to that block exists.

[0286] Therefore, according to one embodiment of this disclosure, if general_merge_flag is 0, CuPerdMode is not MODE_IBC, mvd_l1_zero_flag is 1, and inter_pred_idc is PRED_BI, then mvp_l1_flag (motion vector predictor index for the first reference picture list) is parsed. Thus, mvp_l1_flag (motion vector predictor index for the first reference picture list) exists, and its value is not inferred.

[0287] More specifically, if general_merge_flag is 0, CuPerdMode is not MODE_IBC, mvd_l1_zero_flag (motion vector difference zero flag for the first reference picture list) is 1, inter_pred_idc (information about the reference picture list) is PRED_BI, and affine MC is used, then mvp_l1_flag is parsed. Also, if mvp_l1_flag (motion vector predictor index for the first reference picture list) exists, its value is not inferred.

[0288] The embodiments in Figures 43 and 40 may be implemented together. For example, after initializing Mvd or MvdCp, mvp_l1_flag is parsed. In this case, the initialization of Mvd or MvdCp is the initialization described in Figure 40. The parsing of mvp_l1_flag follows the description in Figure 43. For example, if the Mvd and MvdCp values ​​for reference list L1 are not 0 based on mvd_l1_zero_flag, and the MotionModelIdc value is 1, then the MvdCpL1 value for control point index 2 is initialized and mvp_l1_flag is parsed.

[0289] Figure 44 shows the structure of a syntax related to interpretation according to one embodiment of the present disclosure.

[0290] The embodiment in Figure 44 is an example designed to improve coding efficiency by eliminating the degree of freedom in selecting the MVP. Furthermore, the embodiment in Figure 44 represents the embodiment described in Figure 43 in a different way. Therefore, explanations that overlap with the embodiment in Figure 43 are omitted. In the embodiment shown in Figure 44, if the Mvd or MvdCp value is 0 based on mvd_l1_zero_flag, the signaling indicating MVP is parsed. The signaling indicating MVP includes mvp_l1_flag.

[0291] Figure 45 shows a syntax related to inter prediction according to one embodiment of the present disclosure.

[0292] According to one embodiment of this disclosure, the inter prediction method includes skip mpde, merge mode, inter mode, etc. According to one embodiment, in skip mode, the residual signal is not transmitted. In addition, in skip mode, a method for determining MV such as merge mode may be used. Whether or not to use skip mode is determined by the skip flag. Referring to Figure 33, whether or not to use skip mode is determined by the cu_skip_flag value.

[0293] According to one embodiment, merge mode does not use motion vector difference. Motion vectors are determined based on the motion candidate index. Whether or not merge mode can be used is determined by the merge flag. Referring to Figure 33, the merge_flag value determines whether or not merge mode can be used. Also, if skip mode is not used, merge mode can be used.

[0294] In Skip mode or merge mode, one or more types of candidate lists are selectively used. For example, merge candidate or subblock merge candidate can be used. Alternatively, merge candidate includes selection from signal neighboring, temporal candidate, etc. Also, merge candidate includes candidates that use motion vectors for the entire current block (CU: coding unit). That is, it includes candidates where the motion vectors for each subblock belonging to the current block are the same. Also, subblock merge candidate includes subblock-based temporal MV, affine merge candidate, etc. Also, subblock merge candidate includes candidates that can use different motion vectors for each subblock of the current block (CU). Affine merge candidate is a method created to determine the control point motion vector of affine motion prediction without using motion vector difference. Also, subblock merge candidate includes methods to determine motion vectors on a subblock basis in the current block. For example, in addition to the subblock-based temporal MV and affine merge candidate mentioned above, subblock merge candidate also includes planar MV, regression-based MV, STMVP, etc.

[0295] In one embodiment, the inter mode uses the motion vector difference. The motion vector predictor is determined based on the motion candidate index, and the motion vector is determined based on the motion vector predictor and the motion vector difference. Whether or not the inter mode can be used is determined by whether or not other modes can be used. In another embodiment, whether or not the inter mode can be used is determined by a flag. Figure 45 shows an example where the inter mode is used when other modes, skip mode and merge mode, are not used.

[0296] Inter mode includes AMVP mode, affine inter mode, etc. Inter mode is a mode in which the motion vector is determined based on the motion vector predictor and the motion vector difference. Affine inter mode is a method that uses the motion vector difference when determining the control point motion vector of affine motion prediction.

[0297] Referring to Figure 45, after determining whether to use skip mode or merge mode, it is decided whether to use a subblock merge candidate or a merge candidate. For example, if certain conditions are met, the merge_subblock_flag, which indicates whether a subblock merge candidate can be used, is parsed. The aforementioned conditions are related to block size. For example, they may be related to width, height, area, etc., or a combination of these may be used. Referring to Figure 45, for example, this is the condition when the width and height of the current block (CU) are greater than or equal to a certain value. When parsing merge_subblock_flag, its value is inferred to 0. If merge_subblock_flag is 1, a subblock merge candidate is used; if it is 0, a merge candidate is used. If a subblock merge candidate is used, the merge_subblock_Idx, which is the candidate index, is parsed; if a merge candidate is used, the merge_Idx, which is the candidate index, is parsed. In this case, if the maximum number of candidates in the candidate list is 1, parsing is not performed. If merge_subblock_Idx or merge_Idx is not parsed, it will be inferred to 0.

[0298] Figure 45 shows the coding_unit function, but the details regarding intra prediction are omitted, and Figure 45 shows the case where inter prediction is determined.

[0299] Figure 46 shows a triangle partitioning mode according to one embodiment of the present disclosure.

[0300] The triangle partitioning mode (TPM) referred to in this disclosure is also known by various names such as triangle partition mode, triangle prediction, triangle based prediction, triangle motion compensation, triangular prediction, triangle inter prediction, triangular merge mode, and triangle merge mode. Furthermore, TPM is included in geometric partitioning mode (GPM).

[0301] As shown in Figure 46, TPM divides a rectangular block into two triangles. However, GPM divides a block into two blocks in various ways. For example, GPM divides one rectangular block into two triangular blocks as shown in Figure 46. Also, GPM divides one rectangular block into one pentagonal block and one triangular block. Furthermore, GPM divides one rectangular block into two quadrilateral blocks. Here, a rectangle includes a square. In the following explanation, for convenience, we will use TPM, a simpler version of GPM, as the basis, but it should be interpreted as including GPM.

[0302] According to one embodiment of this disclosure, uni-prediction exists as a prediction method. Uni-prediction is a prediction method that uses one reference list. There are many such reference lists, but according to one embodiment, there are two, L0 and L1. When using uni-prediction, one reference list is used per block. Also, when using uni-prediction, one motion information is used to predict one pixel. In this disclosure, block means CU (coding unit) and PU (prediction unit). Also, in this disclosure, block means TU (transform unit).

[0303] In other embodiments of this disclosure, bi-prediction exists as a prediction method. Bi-prediction is a prediction method that uses multiple reference lists. Bi-prediction is a prediction method that uses two reference lists. For example, bi-prediction uses reference lists L0 and L1. When using bi-prediction, multiple reference lists are used in one block. For example, when using bi-prediction, two reference lists are used in one block. Also, when using bi-prediction, multiple motion information is used to predict a single pixel.

[0304] The aforementioned motion information includes a motion vector, a reference index, and a prediction list utilization flag.

[0305] The aforementioned reference list is a reference picture list.

[0306] In this disclosure, motion information corresponding to uni-prediction or bi-prediction is defined as a single motion information set.

[0307] According to one embodiment of this disclosure, when using TPM, a number of motion information sets are used. For example, when using TPM, two motion information sets are used. For example, when using TPM, at most two motion information sets are used. Furthermore, the way in which two motion information sets are applied within a block using TPM is based on location. For example, one motion information set is used for a predetermined location within a block using TPM, and another motion information set is used for other predetermined locations. Moreover, for other predetermined locations, both motion information sets are used together. For example, for other predetermined locations, prediction 3 is used for prediction based on prediction 1 based on one motion information set and prediction 2 based on another motion information set. For example, prediction 3 is the weighted sum of prediction 1 and prediction 2.

[0308] Referring to Figure 46, Partition 1 and Partition 2 schematically represent the aforementioned pre-set positions and other pre-set positions. When using TPM, one of the two splitting methods shown in Figure 46 is used. The two splitting methods include diagonal split and anti-diagonal split. The split divides the partition into two triangle-shaped partitions. As mentioned above, TPM is included in GPM. Since GPM has been explained above, a redundant explanation is omitted.

[0309] According to one embodiment of this disclosure, when using TPM, only uni-prediction is used for each partition. That is, one motion information is used for each partition. This is to reduce complexity such as memory access and computational complexity. Therefore, only two motion information pieces are used for the CU.

[0310] Furthermore, each motion information is determined from the candidate list. In one embodiment, the candidate list used by the TPM is based on the merge candidate list. In another embodiment, the candidate list used by the TPM is based on the AMVP candidate list. Therefore, a candidate index is signaled in order to use the TPM. Also, for blocks using the TPM, the TPM encodes, decodes, and parses candidate indices equal to the number of partitions, or at most equal to the number of partitions.

[0311] Furthermore, even if a block is predicted by the TPM based on a large amount of motion information, the transform and quantization are performed on the entire block.

[0312] Figure 47 shows a merge data syntax according to one embodiment of the present disclosure. According to one embodiment of the present disclosure, the merge data syntax includes signaling for various modes. These various modes include regular merge mode, MMVD (merge with MVD), subblock merge mode, CIIP (combined intra- and inter-prediction), TPM, etc. Regular merge mode is the same mode as merge mode in HEV. There is also signaling to indicate whether the various modes are used in a block. This signaling is either parsed as a syntax element or implicitly signaled. Referring to Figure 47, the signaling indicating whether to use regular merge mode, MMVD, subblock merge mode, CIIP, or TPM is regular_merge_flag, mmvd_merge_flag (or mmvd_flag), merge_subblock_flag, ciip_flag (or mh_intra_flag), and MergeTriangleFlag (or merge_triangle_flag), respectively.

[0313] According to one embodiment of this disclosure, when using merge mode, if it signals that all modes except one of the various modes will not be used, it is decided to use the mode in question. Also, when using merge mode, if it signals that at least one of the various modes except one will be used, it is decided not to use the mode in question. There is also a higher-level signaling that indicates whether a mode is available. The higher level is a unit that includes a block. Examples of higher levels include sequence, picture, slice, tile group, tile, CTU, etc. If the higher-level signaling that indicates whether a mode is available indicates that it is available, there is additional signaling that indicates whether to use that mode or not. If the higher-level signaling that indicates whether a mode is available indicates that it is unavailable, then that mode will not be used. For example, when using merge mode, if it signals that regular merge mode, MMVD, subblock merge mode, and CIIP will not be used, it is decided to use TPM. Furthermore, when using merge mode, if the system signals that at least one of the following will be used—regular merge mode, MMVD, subblock merge mode, or CIIP—it will decide not to use TPM. There is also a signaling mechanism to indicate whether merge mode will be used. For example, the signaling mechanism to indicate the use of merge mode is general_merge_flag or merge_flag. If merge mode is used, the system will parse the merge data syntax as shown in Figure 47.

[0314] Furthermore, the block size for which TPM can be used is limited. For example, TPM will be used if both the width and height are 8 or greater.

[0315] If TPM is used, the syntax elements associated with TPM are parsed. The syntax elements associated with TPM include signaling indicating the split method and signaling indicating the candidate index. The split method indicates the split direction. There are many (e.g., two) signaling indicators for the candidate index for a block using TPM. Referring to Figure 47, the signaling indicator for the split method is merge_triangle_split_dir. The signaling indicators for the candidate index are merge_triangle_Idx0 and merge_triangle_Idx1.

[0316] In this disclosure, the candidate indices for the TPM are denoted as m and n. For example, the candidate indices for Partition 1 and Partition 2 in Figure 46 are m and n, respectively. In one embodiment, m and n are determined based on the signaling that shows the candidate indices as described in Figure 47. In one embodiment of this disclosure, one of m and n is determined based on one of merge_triangle_Idx0 and merge_triangle_Idx1, and the other of m and n is determined based on both merge_triangle_Idx0 and merge_triangle_Idx1.

[0317] Alternatively, one of m and n is determined based on one of merge_triangle_Idx0 and merge_triangle_Idx1, and the remaining one of m and n is determined based on the remaining one of merge_triangle_Idx0 and merge_triangle_Idx1. More specifically, m is determined based on merge_triangle_Idx0, and n is determined based on merge_triangle_Idx0 (or m) and merge_triangle_Idx1. For example, m and n are determined as follows:

[0318] m=merge_triangle_Idx0 n=merge_triangle_Idx1+(merge_triangle_Idx1>=m)?1:0

[0319] According to one embodiment of this disclosure, m and n are not the same. This is because, in the TPM, if the two candidate indices are the same, that is, if the two motion information are the same, the partitioning effect cannot be obtained. Therefore, when signaling n in this signaling method, if n > m, the number of signaling bits is reduced. Of the total candidates, m is not n, so it can be excluded from signaling.

[0320] If the candidate list used in TPM is named mergeCanList, then mergeCanList[m] and mergeCanList[n] are used as motion information in TPM.

[0321] Figure 48 shows upper-level signaling according to one embodiment of the present disclosure. According to one embodiment of this disclosure, there are numerous higher-level signalings. Higher-level signaling is signaling transmitted in higher-level units. A higher-level unit includes one or more lower-level units. Higher-level signaling is signaling applied to one or more lower-level units. For example, a slice or sequence is a higher-level unit for CUs, PUs, TUs, etc. Conversely, a CU, PU, ​​or TU is a lower-level unit for a slice or sequence.

[0322] According to one embodiment of this disclosure, the higher-level signaling includes signaling indicating the maximum number of a candidate. For example, the higher-level signaling includes signaling indicating the maximum number of a merge candidate. For example, the higher-level signaling includes signaling indicating the maximum number of a candidate used in the TPM. The signaling indicating the maximum number of the merge candidate or the signaling indicating the maximum number of a candidate used in the TPM is signaled and parsed if inter-prediction is permitted. Whether inter-prediction is enforced is determined by the slice type. Slice types include I, P, B, etc. For example, if the slice type is I, inter-prediction is not permitted. For example, if the slice type is I, only intra-prediction or intra-block copy (IBC) is used. Also, if the slice type is P or B, inter-prediction is permitted. Also, if the slice type is P or B, intra-prediction, IBC, etc. are permitted. Also, if the slice type is P, at most one reference list is used to predict pixels. Furthermore, multiple reference lists are used to predict pixels whose slice type is B. For example, up to two reference lists are used to predict pixels whose slice type is B.

[0323] According to one embodiment of this disclosure, when signaling the maximum number, the signaling is based on a reference value. For example, (reference value - maximum number) is signaled. Therefore, the maximum number is derived based on the value parsed by the decoder and the reference value. For example, (base value - parsed value) is determined as the maximum number.

[0324] According to one embodiment, the reference value in the signaling indicating the maximum number of the merge candidate is 6.

[0325] According to one embodiment, the reference value in the signaling indicating the maximum number of candidate used in the TPM is the maximum number of merge candidate.

[0326] Referring to Figure 48, the signaling indicating the maximum number of the merge candidate is six_minus_max_num_merge_cand. Here, the merge candidate refers to the candidate for the merge mode vector prediction. For the sake of explanation, six_minus_max_num_merge_cand will be referred to as the first information. Referring to Figures 2 and 7, the signaling refers to the signal transmitted from the encoder to the decoder via the bitstream. six_minus_max_num_merge_cand (first information) is signaled on a sequence-by-sequence basis. The decoder purges six_minus_max_num_merge_cand (first information) from the bitstream.

[0327] Furthermore, the signaling indicating the maximum number of candidates used in the TPM is max_num_merge_cand_minus_max_num_triangle_cand. Referring to Figures 2 and 7, the signaling refers to the signals transmitted from the encoder to the decoder via the bitstream. The decoder purges max_num_merge_cand_minus_max_num_triangle_cand (third information) from the bitstream. max_num_merge_cand_minus_max_num_triangle_cand (third information) is information about the maximum number of merge mode candidates for the partitioned block.

[0328] Furthermore, the maximum number of merge candidates is MaxNumMergeCand (maximum number of merge candidates), and this value is based on six_minus_max_num_merge_cand (first information). Also, the maximum number of candidates used in TPM is MaxNumTriangleMergeCand, and this value is based on max_num_merge_cand_minus_max_num_triangle_cand. MaxNumMergeCand (maximum number of merge candidates) is used in merge mode, either when blocks are partitioned for movement compensation or when they are not partitioned. The above explanation is based on TPM, but the same method can be used to explain GPM.

[0329] According to one embodiment of this disclosure, there is a higher-level signaling that indicates whether TPM mode is available. Referring to Figure 48, the higher-level signaling that indicates whether TPM mode is available is sps_triangle_enabled_flag (second information). The information that indicates whether TPM mode is available is the same as the information that indicates whether block partitioning is possible for inter-prediction, as shown in Figure 46. Since GPM includes TPM, the information that indicates whether block partitioning is possible is the same as the information that indicates whether GPM mode is available. For inter-prediction means for motion compensation. In other words, sps_triangle_enabled_flag (second information) is information that indicates whether block partitioning is possible for inter-prediction. If the first information indicating whether block partitioning is possible is 1, it indicates that TPM or GPM is available. If the second information is 0, it indicates that TPM or GPM is unavailable. However, it is not limited to this; if the second information is 0, it indicates that TPM or GPM is available. If the second information is 1, it indicates that TPM or GPM is unavailable.

[0330] Referring to Figures 2 and 7, signaling refers to the signals transmitted from the encoder to the decoder via the bitstream. The decoder purges sps_triangle_enabled_flag (second information) from the bitstream.

[0331] According to one embodiment of this disclosure, a TPM is used only if there are candidates to be used in the TPM that are greater than or equal to the number of partitions in the TPM. For example, when a TPM is partitioned into two, the TPM is used only if there are two or more candidates to be used in the TPM. According to one embodiment, the candidates to be used in the TPM are based on merge candidates. Therefore, according to one embodiment of this disclosure, the TPM is used if the maximum number of merge candidates is 2 or greater. Therefore, if the maximum number of merge candidates is 2 or greater, the signaling related to the TPM is parsed. The signaling related to the TPM is a signaling that indicates the maximum number of candidates to be used in the TPM.

[0332] Referring to Figure 48, if sps_triangle_enabled_flag is 1 and MaxNumMergeCand is 2 or greater, then max_num_merge_cand_minus_max_num_triangle_cand is parsed. However, if sps_triangle_enabled_flag is 0 or MaxNumMergeCand is less than 2, then max_num_merge_cand_minus_max_num_triangle_cand is not parsed.

[0333] Figure 49 shows the maximum number of candidate used in a TPM according to one embodiment of the present disclosure.

[0334] Referring to Figure 49, the maximum candidate number used in TPM is MaxNumTriangleMergeCand. The signaling indicating the maximum candidate number used in TPM is max_num_merge_cand_minus_max_num_triangle_cand. The details explained in Figure 48 are omitted here.

[0335] According to one embodiment of this disclosure, the maximum number of a candidate used in the TPM is inculsively within a range from the number of partitions in the TPM to the reference value in the signaling that indicates the maximum number of a candidate used in the TPM. Therefore, if the number of partitions in the TPM is 2 and the reference value is the maximum number of a merge candidate, then, as shown in Figure 49, MaxNumTriangleMergeCand is inculsively within a range from 2 to MaxNumMergeCand.

[0336] According to one embodiment of this disclosure, if there is no signaling indicating the maximum number of a candidate used in the TPM, the signaling indicating the maximum number of a candidate used in the TPM is inferred, or the signaling indicating the maximum number of a candidate used in the TPM is inferred. For example, if there is no signaling indicating the maximum number of a candidate used in the TPM, the maximum number of a candidate used in the TPM is inferred to 0. Alternatively, if there is no signaling indicating the maximum number of a candidate used in the TPM, the signaling indicating the maximum number of a candidate used in the TPM is inferred to a reference value.

[0337] Furthermore, if there is no signaling indicating the maximum number of the candidate used by the TPM, the TPM will not be used. Also, if the maximum number of the candidate used by the TPM is smaller than the number of partitions in the TPM, the TPM will not be used. And if the maximum number of the candidate used by the TPM is 0, the TPM will not be used.

[0338] However, according to the examples in Figures 48 and 49, if the number of partitions in the TPM and the "reference value in the signaling indicating the maximum number of candidate used in the TPM" are the same, then there is only one possible value for the maximum number of candidate used in the TPM. However, according to the examples in Figures 48 and 49, even in such cases, the signaling indicating the maximum number of candidate used in the TPM is parsed, which is unnecessary. If MaxNumMergeCand is 2, then referring to Figure 49, there is only one possible value for MaxNumTriangleMergeCand: 2. However, referring to Figure 48, even in such cases, max_num_merge_cand_minus_max_num_triangle_cand is parsed.

[0339] Referring to Figure 49, MaxNumTriangleMergeCand is determined to be (MaxNumMergeCand - max_num_merge_cand_minus_max_num_triangle_cand).

[0340] Figure 50 shows the upper-level signaling for a TPM according to one embodiment of the present disclosure. According to one embodiment of the present disclosure, if the number of partitions in the TPM is the same as the "reference value in the signaling indicating the maximum number of candidates used in the TPM," the signaling indicating the maximum number of candidates used in the TPM is not parsed. Also, according to the embodiment described above, the number of partitions in the TPM is 2. Furthermore, the "reference value in the signaling indicating the maximum number of candidates used in the TPM" is the maximum number of the merge candidate. Therefore, if the maximum number of the merge candidate is 2, the signaling indicating the maximum number of candidates used in the TPM is not parsed.

[0341] Alternatively, if the "criterion value in the signaling indicating the maximum number of candidate used in the TPM" is less than or equal to the number of partitions in the TPM, the signaling indicating the maximum number of candidate used in the TPM will not be parsed. Therefore, if the maximum number of merge candidate is 2 or less, the signaling indicating the maximum number of candidate used in the TPM will not be parsed.

[0342] Referring to line 5001 in Figure 50, if MaxNumMergeCand (maximum number of merge candidates) is 2, or if MaxNumMergeCand (maximum number of merge candidates) is 2 or less, max_num_merge_cand_minus_max_num_triangle_cand (third information) is not parsed. Also, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) is parsed. Also, if sps_triangle_enabled_flag (second information) is 0, max_num_merge_cand_minus_max_num_triangle_cand (third information) is not parsed. Therefore, if sps_triangle_enabled_flag (second piece of information) is 0, or if MaxNumMergeCand (maximum number of merge candidates) is 2 or less, max_num_merge_cand_minus_max_num_triangle_cand (third piece of information) is not parsed.

[0343] Figure 51 shows the maximum candidate number used in a TPM according to one embodiment of the present disclosure. The embodiment in Figure 51 may also be carried out in conjunction with the embodiment in Figure 50. Furthermore, some of the above-mentioned aspects may be omitted from the explanation in these drawings.

[0344] According to one embodiment of this disclosure, if the "reference value in the signaling indicating the maximum number of candidates used in the TPM" is the number of partitions in the TPM, then the maximum number of candidates used in the TPM is inferred and set to the number of partitions in the TPM. Inferring and setting occurs when there is no signaling indicating the maximum number of candidates used in the TPM. According to the embodiment in Figure 50, if the "reference value in the signaling indicating the maximum number of candidates used in the TPM" is the number of partitions in the TPM, then the signaling indicating the maximum number of candidates used in the TPM is not parsed, and if there is a signaling indicating the maximum number of candidates used in the TPM, the value of the maximum number of candidates used in the TPM is inferred to the number. This is done if further conditions are met. The further condition is that the higher-level signaling indicating whether TPM mode is usable is 1.

[0345] Furthermore, although this embodiment describes inferring and setting the maximum number of the candidate used by the TPM, instead of inferring and setting the maximum number of the candidate used by the TPM, it is also possible to infer and set the signaling indicating the maximum number of the candidate used by the TPM so that the described maximum number value of the candidate used by the TPM is derived.

[0346] Referring to Figure 50, if sps_triangle_enabled_flag (second information) indicates 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) is received. In this case, referring to line 5101 in Figure 51, MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) is obtained using the clearly signaled max_num_merge_cand_minus_max_num_triangle_cand (third information). In short, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than or equal to 3, then MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) is obtained by subtracting the third information (max_num_merge_cand_minus_max_num_triangle_cand) from MaxNumMergeCand (maximum number of merge candidates).

[0347] Referring to line 5102 in Figure 51, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is 2, then MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) is set to 2. More specifically, as mentioned above, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) is received. Therefore, as in line 5102, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) is not received. In this case, MaxNumTriangleMergeCand (the maximum number of merge mode candidates for a partitioned block) is determined without max_num_merge_cand_minus_max_num_triangle_cand (third information).

[0348] Furthermore, referring to line 5103 in Figure 51, if sps_triangle_enabled_flag (second information) is not 0 or MaxNumMergeCand (maximum number of merge candidates) is not 2, MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) is inferred and set to 0. In this case, as described above, in line 5101, if sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than or equal to 3, max_num_merge_cand_minus_max_num_triangle_cand (third information) is signaled, so MaxNumTriangleMergeCand is inferred and set to 0 when the second information is 0 or the maximum number of merge candidates is 1. In short, if sps_triangle_enabled_flag (second information) is 0 or MaxNumMergeCand (maximum number of merge candidates) is 1, then MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) is set to 0.

[0349] MaxNumMergeCand (maximum number of merge candidates) and MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) are used for different purposes. For example, MaxNumMergeCand is used when a block is partitioned for movement compensation or not. However, MaxNumTriangleMergeCand is information used only when a block is partitioned. The number of candidates for a partitioned block that is in merge mode cannot exceed MaxNumTriangleMergeCand.

[0350] Other embodiments can also be used. In the embodiment shown in Figure 51, if MaxNumMergeCand is not 2, it is also greater than 2. In that case, the meaning of inferring and setting MaxNumTriangleMergeCand to 0 may become unclear. However, in such cases, there is a signaling that indicates the maximum number of the candidate used by the TPM, so it is not inferred, and there is no abnormality in operation. However, in this embodiment, the infer is performed in a way that makes sense.

[0351] If sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is 2 or greater, infer and set MaxNumTriangleMergeCand to 2 (or whatever MaxNumMergeCand is). Otherwise (i.e., sps_triangle_enabled_flag is 0 or MaxNumMergeCand is less than 2), infer and set MaxNumTriangleMergeCand to 0.

[0352] If sps_triangle_enabled_flag is 1 and MaxNumMergeCand is 2, then infer MaxNumTriangleMergeCand to 2. Otherwise, if sps_triangle_enabled_flag is 10, then infer MaxNumTriangleMergeCand to 0.

[0353] Therefore, in the above embodiment, according to the implementation shown in Figure 50, if MaxNumMergeCand is 0, 1, or 2, there is no signaling indicating the maximum number of the candidate used by the TPM. If MaxNumMergeCand is 0 or 1, the maximum number of the candidate used by the TPM is inferred and set to 0. If MaxNumMergeCand is 2, the maximum number of the candidate used by the TPM is inferred and set to 2.

[0354] Figure 52 shows a syntax element related to TPM according to one embodiment of the present disclosure.

[0355] As mentioned above, there is a maximum candidate number used by the TPM, and the number of partitions in the TPM is pre-configured. Furthermore, the candidate indices used by the TPM can differ from one another.

[0356] According to one embodiment of this disclosure, if the maximum number of a candidate used in the TPM is the same as the number of partitions in the TPM, a different signaling method is used than in other cases. For example, if the maximum number of a candidate used in the TPM is the same as the number of partitions in the TPM, a different candidate index signaling method is used than in other cases. This allows signaling to be performed with fewer bits. Also, if the maximum number of a candidate used in the TPM is less than or equal to the number of partitions in the TPM, a different signaling method is used than in other cases (in the case where it is less than the number of partitions in the TPM, it means that the TPM cannot be used).

[0357] If the TPM has 2 partitions, then two candidate indices are signaled. If the maximum number of candidates used in the TPM is 2, then there are only two possible candidate index combinations. These two combinations are when m and n are 0 and 1, and when m and n are 1 and 0. Therefore, the two candidate indices are signaled using only 1-bit signaling.

[0358] Referring to Figure 52, if MaxNumTriangleMergeCand is 2, a different candidate index signaling is performed than when it is not (or when MaxNumTriangleMergeCand is greater than 2 if TPM is used). Also, if MaxNumTriangleMergeCand is 2 or less, a different candidate index signaling is performed than when it is not (or when MaxNumTriangleMergeCand is greater than 2 if TPM is used). Another candidate index signaling method, referring to Figure 52, is merge_triangle_idx_indicator parsing. Another candidate index signaling method is a signaling method that does not parse merge_triangle_idx0 or merge_triangle_idx1. This is explained further in Figure 53.

[0359] Figure 53 shows the signaling of a TPM candidate index according to one embodiment of the present disclosure.

[0360] According to one embodiment of this disclosure, using the other index signaling described in Figure 52, the candidate index is determined based on the merge_triangle_idx_indicator. Also, using the other index signaling, it is possible that merge_triangle_idx0 or merge_triangle_idx1 does not exist.

[0361] According to one embodiment of this disclosure, if the maximum number of candidates used in the TPM is the same as the number of partitions in the TPM, the TPM candidate index is determined based on the merge_triangle_idx_indicator. This also applies to blocks that use the TPM.

[0362] More specifically, if MaxNumTriangleMergeCand is 2 (or if MaxNumTriangleMergeCand is 2 and MergeTriangleFlag is 1), the TPM candidate index is determined based on merge_triangle_idx_indicator. In this case, if merge_triangle_idx_indicator is 0, the TPM candidate indices m and n are set to 0 and 1 respectively, and if merge_triangle_idx_indicator is 1, the TPM candidate indices m and n are set to 1 and 0 respectively. Alternatively, m and n are inferred and set to merge_triangle_idx0 or merge_triangle_idx1, which are parsed values ​​(syntax elements) as described.

[0363] Referring to the method for setting m and n based on merge_triangle_idx0 and merge_triangle_idx1 explained in Figure 47, and referring to Figure 53, if merge_triangle_idx0 does not exist, and if MaxNumTriangleMergeCand is 2 (or MaxNumTriangleMergeCand is 2 and MergeTriangleFlag is 1), and merge_triangle_idx_indicator is 1, then the value of merge_triangle_idx0 is inferred to 1. Otherwise, the value of merge_triangle_idx0 is inferred to 0. Also, if merge_triangle_idx1 does not exist, then merge_triangle_idx1 is inferred to 0. Therefore, if MaxNumTriangleMergeCand is 2 and merge_triangle_idx_indicator is 0, then merge_triangle_idx0 and merge_triangle_idx1 are 0 and 0 respectively, and as a result m and n become 0 and 1 respectively. Furthermore, if MaxNumTriangleMergeCand is 2, then when merge_triangle_idx_indicator is 1, merge_triangle_idx0 and merge_triangle_idx1 are 1 and 0 respectively, and as a result m and n become 1 and 0 respectively.

[0364] Figure 54 shows the signaling of a TPM candidate index according to one embodiment of the present disclosure.

[0365] Figure 47 illustrates the method for determining TPM candidate indices, but the example in Figure 54 describes other determination and signaling methods. Explanations that overlap with those described above will be omitted. Also, m and n represent candidate indices as explained in Figure 47.

[0366] According to one embodiment of this disclosure, one of the merge_triangle_idx0 and merge_triangle_idx1 syntax elements is signaled with the smaller of m and n. The remaining one of merge_triangle_idx0 and merge_triangle_idx1 is signaled with a value based on the difference between m and n. In addition, a value indicating the relative magnitude of m and n is signaled.

[0367] For example, merge_triangle_idx0 is the smaller of m and n. Also, merge_triangle_idx1 is a value based on |m, n|, and merge_triangle_idx1 is (|m, n|-1). This is because m and n are not the same. Furthermore, the value that shows the relative magnitudes of m and n is merge_triangle_bigger in Figure 54.

[0368] Using this relationship, m and n are determined based on merge_triangle_idx0, merge_triangle_idx1, and merge_triangle_bigger. Referring to Figure 54, other actions are performed based on the merge_triangle_bigger value. For example, if merge_triangle_bigger is 0, then n is greater than m. In this case, m is merge_triangle_idx0, and n is (merge_triangle_idx1 + m + 1). Also, if merge_triangle_bigger is 1, then m is greater than n. In this case, n is merge_triangle_idx0, and m is (merge_triangle_idx1 + n + 1).

[0369] The method in Figure 54 has the advantage of reducing signaling overhead when the smaller of m and n is not zero (or is large) compared to the method in Figure 47. For example, if m and n are 3 and 4 respectively, the method in Figure 47 should signal merge_triangle_idx0 and merge_triangle_idx1 to 3 and 3 respectively. However, in the method in Figure 54, if m and n are 3 and 4 respectively, merge_triangle_idx0 and merge_triangle_idx1 should be signaled to 3 and 0 respectively (in addition to this, signaling indicating the greater-than / less-than relationship is required). Therefore, when using variable length signaling, the size of the values ​​to be encoded and decoded is reduced, allowing the use of smaller bits.

[0370] Figure 55 shows the signaling of a TPM candidate index according to one embodiment of the present disclosure.

[0371] Figure 47 illustrates the method for determining TPM candidate indices, but the example in Figure 55 describes other determination and signaling methods. Explanations that overlap with those described above will be omitted. Also, m and n represent candidate indices as explained in Figure 47.

[0372] According to one embodiment of this disclosure, one of the merge_triangle_idx0 and merge_triangle_idx1 syntax elements is signaled with the larger of m and n. The remaining one of merge_triangle_idx0 and merge_triangle_idx1 is signaled with a value based on the smaller of m and n. In addition, a value indicating the relative magnitudes of m and n is signaled.

[0373] For example, merge_triangle_idx0 is based on the larger of m and n. In one example, since m and n are not the same, the larger of m and n is 1 or greater. Therefore, considering that the larger of m and n will be excluded, signaling is done with smaller bits. For example, merge_triangle_idx0 is ((the larger of m and n)-1). In this case, the maximum value of merge_triangle_idx0 is (MaxNumTriangleMergeCand-1-1) (-1 because it is a value that starts from 0, and -1 because values ​​with a larger value of 0 can be excluded). The maximum value is used in binarization, and as the maximum value decreases, smaller bits may be used. For example, merge_triangle_idx1 is the smaller of m and n. Also, the maximum value of merge_triangle_idx1 is merge_triangle_idx0. Therefore, even smaller bits may be used than setting the maximum value to MaxNumTriangleMergeCand. Furthermore, if merge_triangle_idx0 is 0, that is, if the larger of m and n is 1, then the smaller of m and n is 0, and therefore no additional signaling occurs. For example, if merge_triangle_idx0 is 0, that is, if the larger of m and n is 1, then the smaller of m and n is determined to be 0. Also, if merge_triangle_idx0 is 0, that is, if the larger of m and n is 1, then merge_triangle_idx1 is inferred to 0 and determined to be 0. Referring to Figure 22, the parsing of merge_triangle_idx1 is determined based on merge_triangle_idx0. For example, if merge_triangle_idx0 is greater than 0, merge_triangle_idx1 is parsed, and if merge_triangle_idx0 is 0, merge_triangle_idx1 is not parsed.

[0374] Furthermore, the value that shows the relative magnitudes of m and n is merge_triangle_bigger in Figure 55.

[0375] Using this relationship, m and n are determined based on merge_triangle_idx0, merge_triangle_idx1, and merge_triangle_bigger. Referring to Figure 55, other actions are performed based on the merge_triangle_bigger value. For example, if merge_triangle_bigger is 0, m is greater than n. In this case, m is (merge_triangle_idx0 + 1), and n is merge_triangle_idx1. Also, if merge_triangle_bigger is 1, n is greater than m. In this case, n is (merge_triangle_idx0 + 1), and m is merge_triangle_idx1. Furthermore, if merge_triangle_idx1 does not exist, its value is inferred to 0.

[0376] The method in Figure 55 has the advantage of reducing signaling overhead depending on the values ​​of m and n compared to the method in Figure 47. For example, if m and n are 1 and 0 respectively, the method in Figure 47 should signal merge_triangle_idx0 and merge_triangle_idx1 to 1 and 0 respectively. However, in the method in Figure 55, if m and n are 1 and 0 respectively, merge_triangle_idx0 and merge_triangle_idx1 should be signaled to 0 and 0 respectively, but merge_triangle_idx1 is inferred without encoding or parsing (in addition to this, signaling indicating the greater-than / less-than relationship is required). Therefore, when using variable length signaling, the size of the values ​​to be encoded and decoded is reduced, allowing the use of smaller bits. Alternatively, if m and n are 2 and 1 respectively, the method in Figure 47 should signal merge_triangle_idx0 and merge_triangle_idx1 to 2 and 1 respectively, while the method in Figure 55 should signal merge_triangle_idx0 and merge_triangle_idx1 to 1 and 1 respectively. However, in this case, the method in Figure 55 signals with one less bit than when the maximum value is large, because the maximum value of merge_triangle_idx1 is (3-1-1)=1. For example, the method in Figure 55 has advantages when the difference between m and n is small, for example, when the difference is 1.

[0377] Furthermore, while it was explained that merge_triangle_idx0 is referenced in the syntax structure of Figure 55 to determine whether merge_triangle_idx1 can be parsed, the determination may also be based on the larger of m and n. That is, it is possible to distinguish between cases where the larger of m and n is 1 or greater, and cases where it is not. However, in this case, the parsing of merge_triangle_bigger should occur before the determination of whether merge_triangle_idx1 can be parsed.

[0378] Although the configuration has been described through specific embodiments, those skilled in the art should be able to modify and change it without departing from the spirit and scope of this disclosure. Therefore, anything that can be easily inferred by a person in the art to which this disclosure belongs from the detailed description and embodiments of this disclosure shall be interpreted as falling within the scope of the rights of this disclosure. [Explanation of symbols]

[0379] 100 Encoding device 110 Exchange section 120 Inverse quantization section 150 Prediction Section 200 Decoding Devices 210 Entropy Decoding Unit 230 Filtering section

Claims

1. In a method for processing video signals, The steps include: purging the AMVR (Adaptive Motion Vector Resolution) enabled flag (sps_amvr_enabled_flag) which indicates whether adaptive motion vector difference resolution is used from the bitstream; The steps include: purging the bitstream for an affine-enabled flag (sps_affine_enabled_flag) indicating whether affine motion compensation is available; A step of determining whether affine motion compensation is available based on the aforementioned affine enable flag (sps_affine_enabled_flag), If the aforementioned affine motion compensation is available, the step of determining whether or not adaptive motion vector difference resolution is used based on the AMVR enabled flag (sps_amvr_enabled_flag), If the adaptive motion vector difference resolution is used, the step is to purge the bitstream for an affine AMVR enabled flag (sps_affine_amvr_enabled_flag) indicating whether or not the adaptive motion vector difference resolution is available for affine motion compensation, A method for processing a video signal, characterized by including the following:

2. A method for processing a video signal according to claim 1, characterized in that one of the AMVR-enabled flag (sps_amvr_enabled_flag), the affine-enabled flag (sps_affine_enabled_flag), or the affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) is signaled in one of the Coding Tree Unit, slice, tile, tile group, picture, or sequence units.

3. The method for processing a video signal according to claim 1, characterized in that, when the affine motion compensation is available and the adaptive motion vector difference resolution is not used, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) implies that the adaptive motion vector difference resolution is unavailable for the affine motion compensation.

4. The method for processing a video signal according to claim 1, characterized in that, if the affine motion compensation is unavailable, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) implicitly indicates that adaptive motion vector difference resolution is unavailable for the affine motion compensation.

5. If the AMVR-enabled flag (sps_amvr_enabled_flag) indicates the use of adaptive motion vector difference resolution, the inter-affine flag (inter_affine_flag) obtained from the bitstream indicates that affine motion compensation is not used for the current block, and at least one of the multiple motion vector differences for the current block is not zero, then the step of purging information about the resolution of the motion vector difference from the bitstream, A method for processing a video signal according to claim 1, comprising the step of correcting the plurality of motion vector differences for the current block based on information regarding the resolution of the motion vector differences.

6. If the affine AMVR-enabled flag indicates that an adaptive motion vector difference resolution is available for the affine motion compensation, and the inter-affine flag obtained from the bitstream indicates the use of affine motion compensation for the current block, and at least one of the multiple control point motion vector differences for the current block is not zero, then the step of purging information regarding the resolution of the motion vector difference from the bitstream, A method for processing a video signal according to claim 1, comprising the step of correcting the plurality of control point motion vector differences for the current block based on information regarding the resolution of the motion vector differences.

7. The current step is to obtain information about the reference picture list (inter_pred_idc) for the block, If the information regarding the reference picture list (inter_pred_idc) indicates that it does not use only the 0th reference picture list (list 0), then the steps are to purge the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1) from the bitstream, The steps include generating candidate motion vector predictors, A step of obtaining a motion vector predictor from candidate motion vector predictors based on the motion vector predictor index, A method for processing a video signal according to claim 1, comprising the step of predicting the current block based on the motion vector predictor.

8. The process further includes obtaining a motion vector difference zero flag (mvd_l1_zero_flag) from the bitstream, which indicates whether the motion vector difference and the motion vector differences of multiple control points are set to zero for the first reference picture list. The step of purging the motion vector predictor index (mvp_l1_flag) is: A method for processing a video signal according to claim 7, comprising the step of purging the motion vector predictor index (mvp_l1_flag) regardless of whether the motion vector difference zero flag (mvd_l1_zero_flag) is 1 and the information regarding the reference picture list (inter_pred_idc) indicates the use of both the 0th reference picture list and the 1st reference picture list.

9. The first step is to purge the bitstream sequence by sequence for the first information (six_minus_max_num_merge_cand) regarding the maximum number of candidates for merge motion vector prediction, A step of obtaining the maximum number of merge candidates based on the first information, A step of purging second information from the bitstream to indicate whether or not the block can be partitioned for interpretation, A method for processing a video signal, comprising the step of purging from a bitstream a third information relating to the maximum number of merge mode candidates for a partitioned block, if the second information indicates 1 and the maximum number of merge candidates is greater than 2.

10. If the second information indicates 1 and the maximum number of merge candidates is greater than or equal to 3, the step of subtracting the third information from the maximum number of merge candidates to obtain the maximum number of merge mode candidates for the partitioned block, If the second information indicates 1 and the maximum number of merge candidates is 2, the step is to set the maximum number of merge mode candidates for the partitioned block to 2. A method for processing a video signal according to claim 9, further comprising the step of setting the maximum number of merge mode candidates for the partitioned block to 0 if the second information is 0 or the maximum number of merge candidates is 1.

11. A device for processing video signals includes a processor and memory. The processor, based on the instruction words stored in the memory, The AMVR (Adaptive Motion Vector Resolution) enabled flag (sps_amvr_enabled_flag) is purged from the bitstream to indicate whether adaptive motion vector difference resolution is used or not. The affine-enabled flag (sps_affine_enabled_flag) is purged from the bitstream to indicate whether or not affine motion compensation is available. Based on the aforementioned affine enable flag (sps_affine_enabled_flag), it is determined whether or not affine motion compensation is available. If the aforementioned affine motion compensation is available, it is determined whether or not adaptive motion vector difference resolution is used based on the AMVR enabled flag (sps_amvr_enabled_flag). A video signal processing apparatus characterized by purging an affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) from the bitstream, indicating whether or not the adaptive motion vector difference resolution is available for affine motion compensation, if the adaptive motion vector difference resolution is used.

12. The apparatus for processing a video signal according to claim 11, characterized in that one of the AMVR-enabled flag (sps_amvr_enabled_flag), the affine-enabled flag (sps_affine_enabled_flag), or the affine AMVR-enabled flag (sps_affine_amvr_enabled_flag) is signaled in one of the units of Coding Tree Unit, slice, tile, tile group, picture, or sequence.

13. The apparatus for processing a video signal according to claim 11, characterized in that, when the affine motion compensation is available and the adaptive motion vector difference resolution is not used, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) implies that the adaptive motion vector difference resolution is unavailable for the affine motion compensation.

14. The apparatus for processing a video signal according to claim 11, characterized in that, if the affine motion compensation is unavailable, the affine AMVR enabled flag (sps_affine_amvr_enabled_flag) implicitly indicates that adaptive motion vector difference resolution is unavailable for the affine motion compensation.

15. The processor, based on the instruction words stored in the memory, If the AMVR enabled flag (sps_amvr_enabled_flag) indicates the use of adaptive motion vector difference resolution, and the inter-affine flag (inter_affine_flag) obtained from the bitstream indicates that affine motion compensation is not being used for the current block, and at least one of the multiple motion vector differences for the current block is not zero, then purge the motion vector difference resolution information from the bitstream. The apparatus for processing a video signal according to claim 11, characterized in that it modifies the plurality of motion vector differences for the current block based on information regarding the resolution of the motion vector differences.

16. The processor, based on the instruction words stored in the memory, If the affine AMVR enabled flag indicates that an adaptive motion vector difference resolution is available for the affine motion compensation, and the inter-affine flag obtained from the bitstream indicates the use of affine motion compensation for the current block, and at least one of the multiple control point motion vector differences for the current block is not zero, then purge the bitstream for information regarding the resolution of the motion vector difference. The apparatus for processing a video signal according to claim 11, characterized in that it modifies the plurality of control point motion vector differences for the current block based on information regarding the resolution of the motion vector differences.

17. The processor, based on the instruction words stored in the memory, Currently, information about the reference picture list (inter_pred_idc) is obtained for the block, If the information regarding the aforementioned reference picture list (inter_pred_idc) indicates that it does not use only the 0th reference picture list (list 0), then the motion vector predictor index (mvp_l1_flag) of the 1st reference picture list (list 1) is purged from the bitstream. Generate candidate motion vector predictors, Based on the aforementioned motion vector predictor index, a motion vector predictor is obtained from the candidate motion vector predictors. The apparatus for processing a video signal according to claim 11, characterized in that it predicts the current block based on the motion vector predictor.

18. The processor, based on the instruction words stored in the memory, From the bitstream, obtain a motion vector difference zero flag (mvd_l1_zero_flag) that indicates whether the motion vector difference and the motion vector differences of multiple control points are set to zero for the first reference picture list. The video signal processing apparatus according to 17, characterized in that the motion vector difference zero flag (mvd_l1_zero_flag) is 1 and the motion vector predictor index (mvp_l1_flag) is purged regardless of whether the reference picture list information (inter_pred_idc) indicates the use of both the 0th reference picture list and the 1st reference picture list.

19. A device for processing video signals includes a processor and memory. The processor, based on the instruction words stored in the memory, The first piece of information (six_minus_max_num_merge_cand) regarding the maximum number of candidates for merge motion vector prediction is purged from the bitstream in sequence units. Based on the first information, obtain the maximum number of merge candidates, For inter-prediction, a second piece of information is purged from the bitstream to indicate whether or not the block can be partitioned. A device for processing a video signal, characterized by purging a third piece of information from a bitstream relating to the maximum number of merge mode candidates for a partitioned block, provided that the second piece of information indicates 1 and the maximum number of merge candidates is greater than 2.

20. The processor, based on the instruction words stored in the memory, If the second piece of information indicates 1 and the maximum number of merge candidates is greater than or equal to 3, the third piece of information is subtracted from the maximum number of merge candidates to obtain the maximum number of merge mode candidates for the partitioned block. If the second piece of information indicates 1 and the maximum number of merge candidates is 2, set the maximum number of merge mode candidates for the partitioned block to 2. The apparatus for processing a video signal according to claim 19, characterized in that if the second information is 0 or the maximum number of merge candidates is 1, the maximum number of merge mode candidates for the partitioned block is set to 0.

21. In a method for processing video signals, The steps include generating an AMVR (Adaptive Motion Vector Resolution) enabled flag (sps_amvr_enabled_flag) indicating whether or not adaptive motion vector difference resolution is used, The steps include generating an affine-enabled flag (sps_affine_enabled_flag) indicating whether or not affine motion compensation is available, A step of determining whether affine motion compensation is available based on the aforementioned affine enable flag (sps_affine_enabled_flag), If the aforementioned affine motion compensation is available, the step of determining whether or not adaptive motion vector difference resolution is used based on the AMVR enabled flag (sps_amvr_enabled_flag), If the adaptive motion vector difference resolution is used, the step of generating an affine AMVR enabled flag (sps_affine_amvr_enabled_flag) indicating whether or not the adaptive motion vector difference resolution is usable for the affine motion compensation, The steps include: generating a bitstream by entropy coding the AMVR-enabled flag (sps_amvr_enabled_flag), the affine-enabled flag (sps_affine_enabled_flag), and the AMVR-enabled flag (sps_amvr_enabled_flag); A method for processing a video signal, characterized by including the following:

22. The steps include generating first information (six_minus_max_num_merge_cand) regarding the maximum number of candidates for the merge motion vector prediction based on the maximum number of merge candidates, A step of generating second information indicating whether or not a block can be partitioned for interplanar prediction, If the second information indicates 1 and the maximum number of merge candidates is greater than 2, the step is to generate third information regarding the maximum number of merge mode candidates for the partitioned block. The steps include: entropy coding the first information (six_minus_max_num_merge_cand), the second information, and the third information to generate a bitstream in sequence units; A method for processing a video signal, characterized by including the following: