A method and apparatus for processing a video signal using adaptive motion vector resolution
Through adaptive motion vector resolution processing and affine motion compensation technology, the motion vector differential resolution is dynamically adjusted, which solves the problem of inefficient encoding efficiency in existing video signal processing and achieves more efficient video signal encoding.
Patent Information
- Application Number
- CN202080040478.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-17
- Filing Date
- 2020-05-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-05-04
AI Technical Summary
The existing video signal processing methods have insufficient encoding efficiency, and cannot effectively utilize the spatial, time and random correlation of the video signal, resulting in low encoding efficiency.
Adaptive motion vector resolution processing (AMVR) and affine motion compensation technology are used to analyze enable flags and inter prediction information, and the motion vector differential resolution is dynamically adjusted to improve coding efficiency.
It improves the encoding efficiency of video signals, enhances the flexibility and accuracy of the encoding process, and reduces the number of bits required for data transmission.
Smart Images

Figure CN113906750B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and apparatus for processing video signals, and more particularly, to a video signal processing method and apparatus for encoding and decoding video signals. Background Art
[0002] Compression coding refers to a series of signal processing techniques for transmitting digitized information over a communication line or storing information in a form suitable for a storage medium. The objects of compression coding include objects such as speech, video, and text, and in particular, the technique for performing compression coding on an image is called video compression. Considering spatial correlation, temporal correlation, and stochastic correlation, compression coding of video signals is performed by removing excessive information. However, with the latest development of various media and data transmission media, more efficient video signal processing methods and apparatuses are required. Summary of the Invention
[0003] Technical Problem
[0004] An object of the present disclosure is to increase the coding efficiency of video signals.
[0005] Technical Solution
[0006] A method for processing a video signal according to an embodiment of the present disclosure includes the following steps: parsing an Adaptive Motion Vector Resolution (AMVR) enable flag sps_amvr_enabled_flag indicating whether to use an adaptive motion vector difference resolution from a bitstream; parsing an affine enable flag sps_affine_enabled_flag indicating whether affine motion compensation is available from the bitstream; determining whether affine motion compensation is available based on the affine enable flag sps_affine_enabled_flag; when affine motion compensation is available, determining whether to use an adaptive motion vector difference resolution based on the AMVR enable flag sps_amvr_enabled_flag; and when using an adaptive motion vector difference resolution, parsing an affine AMVR enable flag sps_affine_amvr_enabled_flag indicating whether the adaptive motion vector difference resolution is available for affine motion compensation from the bitstream.
[0007] In a method for processing a video signal according to an embodiment of the present disclosure, one of the AMVR enable flag sps_amvr_enabled_flag, the affine enable flag sps_affine_enabled_flag, or the affine AMVR enable flag sps_affine_amvr_enabled_flag is signaled as one of a coding tree unit, a slice, a tile, a tile group, a picture, or a sequence unit.
[0008] In a method of processing a video signal according to an embodiment of the present disclosure, when affine motion compensation is available and adaptive motion vector difference resolution is not used, the affine AMVR enable flag sps_affine_amvr_enabled_flag infers that the adaptive motion vector difference resolution is not available for affine motion compensation.
[0009] In a method of processing a video signal according to an embodiment of the present disclosure, when affine motion compensation is not available, the affine AMVR enable flag sps_affine_amvr_enabled_flag infers that the adaptive motion vector difference resolution is not available for affine motion compensation.
[0010] The video signal processing method according to an embodiment of the present disclosure further includes, when the AMVR enable flag sps_amvr_enabled_flag indicates the use of adaptive motion vector difference resolution, the inter-frame affine flag inter_affine_flag obtained from the bitstream indicates that affine motion compensation is not used for the current block, and at least one of a plurality of motion vector differences for the current block is non-zero, parsing information about the resolution of the motion vector difference from the bitstream, and modifying the plurality of motion vector differences for the current block based on the information about the resolution of the motion vector difference.
[0011] The method of processing a video signal according to an embodiment of the present disclosure further includes, when the affine AMVR enable flag indicates that the adaptive motion vector difference resolution is available for affine motion compensation, the inter-frame affine flag inter_affine_flag obtained from the bitstream indicates that affine motion compensation is used for the current block, and at least one of a plurality of control point motion vector differences for the current block is non-zero, parsing information about the resolution of the motion vector difference from the bitstream, and modifying the plurality of control point motion vector differences for the current block based on the information about the resolution of the motion vector difference.
[0012] The video signal processing method according to an embodiment of the present disclosure further includes obtaining information inter_pred_idc about a reference picture list for the current block, when the information inter_pred_idc about the reference picture list indicates that not only the zero-th reference picture list is used, parsing a motion vector predictor index mvp_l1_flag of the first reference picture list from the bitstream, generating motion vector predictor candidates, obtaining a motion vector predictor from the motion vector predictor candidates based on the motion vector predictor index, and predicting the current block based on the motion vector predictor.
[0013] A method for processing a video signal according to an embodiment of the present disclosure further includes obtaining, from a bitstream, a motion vector difference zero flag mvd_l1_zero_flag indicating whether a motion vector difference and a plurality of control point motion vector differences are set to zero for a first reference picture list, wherein the step of parsing a motion vector predictor index mvp_l1_flag includes: the motion vector difference zero flag mvd_l1_zero_flag is 1, and the motion vector predictor index mvp_l1_flag is parsed regardless of whether the information inter_pred_idc regarding the reference picture list indicates the use of both a zero reference picture list and a first reference picture list.
[0014] A method for processing a video signal according to an embodiment of the present disclosure includes the following steps: parsing, on a sequence basis, from a bitstream, first information six_minus_max_num_merge_cand related to a maximum number of candidates for merging motion vector prediction, obtaining, based on the first information, the maximum number of merge candidates, parsing, from the bitstream, second information indicating whether a block is partitioned for inter prediction, and when the second information indicates 1 and the maximum number of merge candidates is greater than 2, parsing, from the bitstream, third information related to a maximum number of merge mode candidates for the partitioned block.
[0015] A method for processing a video signal according to an embodiment of the present invention further includes: when the second information indicates 1 and the maximum number of merge candidates is greater than or equal to 3, obtaining the maximum number of merge mode candidates for the partitioned block by subtracting the third information from the maximum number of merge candidates; when the second information indicates 1 and the maximum number of merge candidates is 2, setting the maximum number of merge mode candidates for the partitioned block to 2; and when the second information is 0 or the maximum number of merge candidates is 1, setting the maximum number of merge mode candidates for the partitioned block to 0.
[0016] An apparatus for processing a video signal according to an embodiment of the present invention includes a processor and a memory. Based on instructions stored in the memory, the processor parses an Adaptive Motion Vector Resolution (AMVR) enable flag sps_amvr_enabled_flag indicating whether to use an adaptive motion vector differential resolution from a bitstream, parses an affine enable flag sps_affine_enabled_flag indicating whether affine motion compensation is available from the bitstream, determines whether affine motion compensation is available based on the affine enable flag sps_affine_enabled_flag, when affine motion compensation is available, determines whether to use an adaptive motion vector differential resolution based on the AMVR enable flag sps_amvr_enabled_flag, and when using an adaptive motion vector differential resolution, parses an affine AMVR enable flag sps_affine_amvr_enabled_flag indicating whether the adaptive motion vector differential resolution is available for affine motion compensation from the bitstream.
[0017] In an apparatus for processing a video signal according to an embodiment of the present disclosure, one of the AMVR enable flag sps_amvr_enabled_flag, the affine enable flag sps_affine_enabled_flag, or the affine AMVR enable flag sps_affine_amvr_enabled_flag is signaled as one of a compile tree unit, a slice, a tile, a tile group, a picture, or a sequence unit.
[0018] In an apparatus for processing a video signal according to an embodiment of the present disclosure, when affine motion compensation is available and an adaptive motion vector differential resolution is not used, the affine AMVR enable flag sps_affine_amvr_enabled_flag infers that the adaptive motion vector differential resolution is not available for affine motion compensation.
[0019] In an apparatus for processing a video signal according to an embodiment of the present disclosure, when affine motion compensation is not available, the affine AMVR enable flag sps_affine_amvr_enabled_flag infers that the adaptive motion vector differential resolution is not available for affine motion compensation.
[0020] In an apparatus for processing a video signal according to an embodiment of the present disclosure, based on instructions stored in a memory, when an AMVR enable flag sps_amvr_enabled_flag indicates the use of an adaptive motion vector difference resolution, an inter-frame affine flag inter_affine_flag obtained from a bitstream indicates that affine motion compensation is not used for a current block, and at least one of a plurality of motion vector differences for the current block is non-zero, a processor parses information about a resolution of the motion vector difference from the bitstream and modifies the plurality of motion vector differences for the current block based on the information about the resolution of the motion vector difference.
[0021] In an apparatus for processing a video signal according to an embodiment of the present disclosure, based on instructions stored in a memory, when an affine AMVR enable flag indicates that an adaptive motion vector difference resolution is available for affine motion compensation, an inter-frame affine flag inter_affine_flag obtained from a bitstream indicates that affine motion compensation is used for the current block, and at least one of a plurality of control point motion vector differences for the current block is non-zero, a processor parses information about a resolution of the motion vector difference from the bitstream and modifies the plurality of control point motion vector differences for the current block based on the information about the resolution of the motion vector difference.
[0022] In an apparatus for processing a video signal according to an embodiment of the present disclosure, based on instructions stored in a memory, a processor obtains information inter_pred_idc about a reference picture list for a current block. When the information inter_pred_idc about the reference picture list indicates that not only the zero-th reference picture list list0 is used, a motion vector predictor index mvp_l1_flag of a first reference picture list list1 is parsed from a bitstream, a motion vector predictor candidate is generated, a motion vector predictor is obtained from the motion vector predictor candidate based on the motion vector predictor index, and the current block is predicted based on the motion vector predictor.
[0023] In an apparatus for processing a video signal according to an embodiment of the present invention, based on instructions stored in a memory, a processor obtains a motion vector difference zero flag mvd_l1_zero_flag from a bitstream that indicates whether motion vector differences and a plurality of control point motion vector differences are set to zero for a first reference picture list, and the motion vector difference zero flag mvd_l1_zero_flag is 1, and the motion vector predictor index mvp_l1_flag is parsed regardless of whether the information inter_pred_idc about the reference picture list indicates the use of both the zero-th reference picture list and the first reference picture list.
[0024] An apparatus for processing a video signal according to an embodiment of the present disclosure includes a processor and a memory. Based on instructions stored in the memory, the processor parses first information six_minus_max_num_merge_cand related to the maximum number of candidates for merging motion vector prediction in units of sequences from a bitstream, obtains the maximum number of merge candidates based on the first information, parses second information from the bitstream indicating whether a block is partitioned for inter prediction, and when the second information indicates 1 and the maximum number of merge candidates is greater than 2, parses third information related to the maximum number of merge mode candidates for the partitioned block from the bitstream.
[0025] In an apparatus for processing a video signal according to an embodiment of the present disclosure, based on instructions stored in the memory, when the second information indicates 1 and the maximum number of merge candidates is greater than or equal to 3, the processor obtains the maximum number of merge mode candidates for the partitioned block by subtracting the third information from the maximum number of merge candidates; when the second information indicates 1 and the maximum number of merge candidates is 2, sets the maximum number of merge mode candidates for the partitioned block to 2; and when the second information is 0 or the maximum number of merge candidates is 1, sets the maximum number of merge mode candidates for the partitioned block to 0.
[0026] A method for processing a video signal according to an embodiment of the present disclosure includes the following steps: generating an adaptive motion vector resolution (AMVR) enable flag sps_amvr_enabled_flag indicating whether to use an adaptive motion vector difference resolution, and generating an affine enable flag sps_affine_enabled_flag indicating whether affine motion compensation is available; based on the affine enable flag sps_affine_enabled_flag, determining whether affine motion compensation is available, and when affine motion compensation is available, determining whether to use an adaptive motion vector difference resolution based on the AMVR enable flag sps_amvr_enabled_flag, and when using an adaptive motion vector difference resolution, generating an affine AMVR enable flag sps_affine_amvr_enabled_flag indicating whether the adaptive motion vector difference resolution is available for affine motion compensation, and generating a bitstream by performing entropy coding on the AMVR enable flag sps_amvr_enabled_flag, the affine enable flag sps_affine_enabled_flag, and the AMVR enable flag sps_amvr_enabled_flag.
[0027] A method for processing a video signal according to an embodiment of the present disclosure further includes: generating first information six_minus_max_num_merge_cand related to the maximum number of candidates for merging motion vector prediction based on the maximum number of merge candidates, generating second information indicating whether a block can be partitioned for inter prediction, generating third information related to the maximum number of merge mode candidates for the partitioned block when the second information indicates 1 and the maximum number of merge candidates is greater than 2, and performing entropy coding on the first information six_minus_max_num_merge_cand, the second information, and the third information to generate a bitstream in units of sequences.
[0028] Beneficial effects
[0029] According to an embodiment of the present disclosure, the coding efficiency of a video signal can be increased. Description of the drawings
[0030] Figure 1 is a schematic block diagram of a video signal encoder device according to an embodiment of the present disclosure.
[0031] Figure 2 is a schematic block diagram of a video signal decoder device according to an embodiment of the present disclosure.
[0032] Figure 3 is a diagram illustrating an embodiment of the present disclosure for partitioning a coding unit.
[0033] Figure 4 is a diagram illustrating a method for hierarchically representing Figure 3 an embodiment of the segmentation structure.
[0034] Figure 5 is a diagram illustrating another embodiment of the present disclosure for partitioning a coding unit.
[0035] Figure 6 is a diagram illustrating a method for obtaining reference pixels for intra prediction.
[0036] Figure 7 is a diagram illustrating an embodiment of a prediction mode for intra prediction.
[0037] Figure 8 is a diagram illustrating inter prediction according to an embodiment of the present disclosure.
[0038] Figure 9 is a diagram illustrating a method for signaling a motion vector according to an embodiment of the present disclosure.
[0039] Figure 10 is a diagram illustrating a motion vector difference syntax according to an embodiment of the present disclosure.
[0040] Figure 11 FIG. is a diagram illustrating signaling for adaptive motion vector resolution according to an embodiment of the present disclosure.
[0041] Figure 12 FIG. is a diagram illustrating signaling for adaptive motion vector resolution according to an embodiment of the present disclosure.
[0042] Figure 13 FIG. is a diagram illustrating signaling for adaptive motion vector resolution according to an embodiment of the present disclosure.
[0043] Figure 14 FIG. is a diagram illustrating signaling for adaptive motion vector resolution according to an embodiment of the present disclosure.
[0044] Figure 15 FIG. is a diagram illustrating signaling for adaptive motion vector resolution according to an embodiment of the present disclosure.
[0045] Figure 16 FIG. is a diagram illustrating signaling for adaptive motion vector resolution according to an embodiment of the present disclosure.
[0046] Figure 17 FIG. is a diagram illustrating affine motion prediction according to an embodiment of the present disclosure.
[0047] Figure 18 FIG. is a diagram illustrating affine motion prediction according to an embodiment of the present disclosure.
[0048] Figure 19 FIG. is a diagram illustrating an expression of a motion vector field according to an embodiment of the present disclosure.
[0049] Figure 20 FIG. is a diagram illustrating affine motion prediction according to an embodiment of the present disclosure.
[0050] Figure 21 FIG. is a diagram illustrating an expression of a motion vector field according to an embodiment of the present disclosure.
[0051] Figure 22 FIG. is a diagram illustrating affine motion prediction according to an embodiment of the present disclosure.
[0052] Figure 23 FIG. is a diagram illustrating a mode of affine motion prediction according to an embodiment of the present disclosure.
[0053] Figure 24 FIG. is a diagram illustrating a mode of affine motion prediction according to an embodiment of the present disclosure.
[0054] Figure 25 FIG. is a diagram illustrating a derivation of an affine motion prediction sub according to an embodiment of the present disclosure.
[0055] Figure 26It is a diagram illustrating the derivation of an affine motion predictor according to an embodiment of the present disclosure.
[0056] Figure 27 It is a diagram illustrating the derivation of an affine motion predictor according to an embodiment of the present disclosure.
[0057] Figure 28 It is a diagram illustrating the derivation of an affine motion predictor according to an embodiment of the present disclosure.
[0058] Figure 29 It is a diagram illustrating a method for generating control point motion vectors according to an embodiment of the present disclosure.
[0059] Figure 30 It is a diagram illustrating by referring to Figure 29 a method for determining a motion vector difference described by the method.
[0060] Figure 31 It is a diagram illustrating a method for generating control point motion vectors according to an embodiment of the present disclosure.
[0061] Figure 32 It is a diagram illustrating by referring to Figure 31 a method for determining a motion vector difference described by the method.
[0062] Figure 33 It is a diagram illustrating the motion vector difference syntax according to an embodiment of the present disclosure.
[0063] Figure 34 It is a diagram illustrating the higher-level signaling structure according to an embodiment of the present disclosure.
[0064] Figure 35 It is a diagram illustrating the compilation unit syntax structure according to an embodiment of the present disclosure.
[0065] Figure 36 It is a diagram illustrating the higher-level signaling structure according to an embodiment of the present disclosure.
[0066] Figure 37 It is a diagram illustrating the higher-level signaling structure according to an embodiment of the present disclosure.
[0067] Figure 38 It is a diagram illustrating the compilation unit syntax structure according to an embodiment of the present disclosure.
[0068] Figure 39 It is a diagram illustrating the MVD default value setting according to an embodiment of the present disclosure.
[0069] Figure 40 It is a diagram illustrating the MVD default value setting according to an embodiment of the present disclosure.
[0070] Figure 41It is a diagram illustrating the AMVR-related syntax structure according to an embodiment of the present disclosure.
[0071] Figure 42 It is a diagram illustrating the inter-frame prediction-related syntax structure according to an embodiment of the present disclosure.
[0072] Figure 43 It is a diagram illustrating the inter-frame prediction-related syntax structure according to an embodiment of the present disclosure.
[0073] Figure 44 It is a diagram illustrating the inter-frame prediction-related syntax structure according to an embodiment of the present disclosure.
[0074] Figure 45 It is a diagram illustrating the inter-frame prediction-related syntax according to an embodiment of the present invention.
[0075] Figure 46 It is a diagram illustrating the triangular partition mode according to an embodiment of the present invention.
[0076] Figure 47 It is a diagram illustrating the merge data syntax according to an embodiment of the present invention.
[0077] Figure 48 It is a diagram illustrating the higher-level signaling according to an embodiment of the present invention.
[0078] Figure 49 It is a diagram illustrating the maximum number of candidates used in the TPM according to an embodiment of the present invention.
[0079] Figure 50 It is a diagram illustrating the higher-level signaling related to the TPM according to an embodiment of the present invention.
[0080] Figure 51 It is a diagram illustrating the maximum number of candidates used in the TPM according to an embodiment of the present invention.
[0081] Figure 52 It is a diagram of the TPM-related syntax elements according to an embodiment of the present invention.
[0082] Figure 53 It is a diagram illustrating the signaling of the TPM candidate index according to an embodiment of the present invention.
[0083] Figure 54 It is a diagram illustrating the signaling of the TPM candidate index according to an embodiment of the present invention.
[0084] Figure 55 It is a diagram illustrating the signaling of the TPM candidate index according to an embodiment of the present invention. Detailed Description
[0085] In view of the functions in the present invention, the terms used in this specification may be common terms widely used currently, but they may change according to the intentions, habits of those skilled in the art, or the emergence of new technologies. Additionally, in some cases, there may be terms arbitrarily selected by the applicant, and in such cases, their meanings are described in the corresponding description parts of the present invention. Therefore, the terms used in this specification should be interpreted based on the substantial meanings of the terms and the content throughout the specification.
[0086] In the present disclosure, the following terms can be interpreted based on the following criteria, and even terms not described can be interpreted for the following purposes. Compilation can be interpreted as encoding or decoding in some cases. Information is a term including all values, parameters, coefficients, elements, etc., and in some cases its meaning can be interpreted differently, and thus, the present disclosure is not limited thereto. "Unit" is used to refer to the basic unit of image (picture) processing or a specific position of a picture, and in some cases can be used interchangeably with terms such as "block", "partition", or "region". Additionally, in this specification, a unit can be used as a concept including all of a compilation unit, a prediction unit, and a transformation unit.
[0087] Figure 1 is a schematic block diagram of a video signal encoding device according to an embodiment of the present disclosure. Refer to Figure 1 , the encoding device 100 of the present disclosure mainly includes a transformation unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transformation unit 125, a filtering unit 130, and a prediction unit 150, as well as an entropy compilation unit 160.
[0088] The transformation unit 110 obtains transformation coefficient values by transforming the pixel values of the received video signal. For example, a discrete cosine transform (DCT) or a wavelet transform can be used. In particular, in the discrete cosine transform, the transformation is performed by dividing the input picture signal into blocks of a predetermined size. In the transformation, the compilation efficiency can vary according to the distribution and characteristics of the values in the transformation region.
[0089] The quantization unit 115 quantizes the transformation coefficient values output from the transformation unit 110. The inverse quantization unit 120 dequantizes the transformation coefficient values, and the inverse transformation unit 125 reconstructs the original pixel values using the dequantized transformation coefficient values.
[0090] The filtering unit 130 performs filtering calculations to improve the quality of the reconstructed picture. For example, it can include a deblocking filter and an adaptive loop filter. The filtered picture is output or stored in the decoded picture buffer 156 to be used as a reference picture.
[0091] To improve the compilation efficiency, instead of compiling the picture signal as it is, the following method is used: predict the picture via the prediction unit 150 by using the already-compiled area, and add the residual value between the original picture and the predicted picture to the predicted picture to obtain the reconstructed picture. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 predicts the current picture by using the reference pictures stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from the reconstructed area in the current picture and transmits the intra-compilation information to the entropy compilation unit 160. The inter prediction unit 154 may include a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains the motion vector value of the current area by referring to a specific reconstructed area. The motion estimation unit 154a transmits the position information (reference frame, motion vector, etc.) of the reference area to the entropy compilation unit 160 so that the position information can be included in the bitstream. The motion compensation unit 154b performs inter-frame motion compensation by using the motion vector value transmitted from the motion estimation unit 154a.
[0092] The entropy compilation unit 160 performs entropy compilation on the quantized transform coefficients, inter-compilation information, intra-compilation information, and reference area information input from the inter prediction unit 154 to generate a video signal bitstream. Here, in the entropy compilation unit 160, a variable length compilation (VLC) scheme, arithmetic compilation, etc. can be used. The variable length compilation (VLC) scheme transforms the input symbols into consecutive codewords, and the length of the codewords can be variable. For example, frequently occurring symbols are expressed as short codewords, and infrequently occurring symbols are expressed as long codewords. As a variable length compilation scheme, a context-based adaptive variable length compilation (CAVLC) scheme can be used. Arithmetic compilation transforms consecutive data symbols into a fraction, and in arithmetic compilation, the optimal fractional bits required to represent each symbol can be obtained. Context-based adaptive binary arithmetic compilation (CABAC) can be used as arithmetic compilation.
[0093] The generated bitstream is encapsulated using a network abstraction layer (NAL) unit as the basic unit. The NAL unit includes the compiled slices, and the slices are composed of an integer number of compilation tree units. To decode the bitstream in the video decoder, the bitstream should first be separated into NAL units, and then each separated NAL unit should be decoded.
[0094] Figure 2 is a schematic block diagram of a video signal decoding device 200 according to an embodiment of the present disclosure. Refer to Figure 2 , the decoding device 200 of the present disclosure includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.
[0095] The entropy decoding unit 210 performs entropy decoding on the video signal bitstream to extract the transform coefficients and motion information for each region. The inverse quantization unit 220 dequantizes the entropy decoded transform coefficients, and the inverse transform unit 225 reconstructs the original pixel values by using the dequantized transform coefficients.
[0096] Meanwhile, the filtering unit 230 improves the picture quality by performing filtering on the picture. In this filtering unit, a deblocking filter for reducing block distortion and / or an adaptive loop filter for removing distortion from the entire picture may be included. The filtered picture is output or stored in the decoded picture buffer 256 to be used as a reference picture for the next frame.
[0097] In addition, the prediction unit 250 of the present disclosure includes an intra prediction unit 252 and an inter prediction unit 254, and reconstructs a prediction picture by using the coding type decoded by the entropy decoding unit 210 described above, the transform coefficients of each region, the motion information, etc.
[0098] In this regard, the intra prediction unit 252 performs intra prediction based on the decoded samples in the current picture. The inter prediction unit 254 generates a prediction picture by using the reference pictures stored in the decoded picture buffer 256 and the motion information. The inter prediction unit 254 may be configured to further include a motion estimation unit 254a and a motion compensation unit 254b. The motion estimation unit 254a obtains a motion vector indicating the positional relationship between the current block and the reference block of the reference picture for compilation, and transmits the obtained motion vector to the motion compensation unit 254b.
[0099] The prediction sub output from the intra prediction unit 252 or the inter prediction unit 254 is added to the pixel values output from the inverse transform unit 225 to generate a reconstructed video frame.
[0100] Hereinafter, in the operations of the encoding device 100 and the decoding device 200, a method of dividing the compilation unit and the prediction unit with reference to Figures 3 to 5 will be described.
[0101] The compilation unit means a basic unit for processing a picture in the process of processing the video signal described above, for example, in the process of intra / inter prediction, transformation, quantization, and / or entropy compilation. The size of the compilation unit for compiling one picture may not be constant. The compilation unit may be rectangular, and one compilation unit may be divided into several compilation units.
[0102] Figure 3FIG. illustrates an embodiment of the present disclosure for dividing a compilation unit. For example, a compilation unit having a size of 2N×2N can be further divided into four compilation units having a size of N×N. This division of the compilation unit can be performed recursively, and not all compilation units need to be divided in the same form. However, for the convenience of the compilation and processing procedures, there may be limitations on the size of the largest compilation unit and / or the size of the smallest compilation unit.
[0103] For a compilation unit, information indicating whether the corresponding compilation unit is divided can be stored. Figure 4 FIG. illustrates an embodiment of a method of hierarchically representing Figure 3 the division structure of the compilation unit illustrated in. When a unit is divided, the value "1" can be assigned to the information, and when the unit is not divided, the value "0" can be assigned to it. As Figure 4 illustrated in, if the flag value indicating whether to divide is 1, the compilation unit corresponding to the corresponding node is further divided into four compilation units. If the flag value is 0, the compilation unit is not divided any further, and a processing procedure for the compilation unit can be executed.
[0104] The above structure of the compilation unit can be represented using a recursive tree structure. That is, a compilation unit that is divided into other compilation units with a picture or the largest-size compilation unit as the root has as many child nodes as the number of divided compilation units. Therefore, a compilation unit that is not divided any further becomes a leaf node. Assuming that a compilation unit can only be divided into a square shape, since a compilation unit can be divided into at most four other compilation units, the tree representing the compilation unit can be in the form of a quadtree.
[0105] In an encoder, the optimal size of the compilation unit is selected according to the characteristics of a video picture (e.g., resolution) or considering the compilation efficiency, and information about the optimal size of the compilation unit or information through which the optimal size of the compilation unit can be derived can be included in the bitstream. For example, the size of the largest compilation unit and the maximum depth of the tree can be defined. In the case of performing square division, since the height and width of the compilation unit are half of the height and width of the parent-node compilation unit, the size of the smallest compilation unit can be obtained using the above information. Alternatively, the size of the smallest compilation unit and the maximum depth of the tree can be predefined and used, and the size of the largest compilation unit can be derived and used through the size of the smallest compilation unit and the maximum depth of the tree. Since the size of the unit changes in multiples of 2 in square division, the actual size of the compilation unit is expressed as a logarithm to the base 2 to increase the transmission efficiency.
[0106] The decoder can obtain information indicating whether the current coding unit is split. If such information is obtained (transmitted) only under specific conditions, the efficiency can be increased. For example, since the condition under which the current coding unit can be split is that the size of the unit obtained by adding the current coding unit size at the current position is smaller than the size of the picture and the current unit size is greater than the preset minimum coding unit size, the information indicating whether the current coding unit is split can be obtained only in such a case.
[0107] If the above information indicates that the coding unit is split, the size of the split coding unit is changed to half of the current coding unit and is split into four square coding units based on the current processing position. The above processing can be repeated for each divided coding unit.
[0108] Figure 5 FIG. illustrates another embodiment of the present disclosure for splitting a coding unit. According to another embodiment of the present disclosure, the above quadtree-shaped coding unit can be further split into a binary tree structure that is horizontally split or vertically split. That is, square quadtree splitting can be first applied to the root coding unit, and additionally rectangular binary tree splitting can be applied to the leaf nodes of the quadtree. According to an embodiment, the binary tree splitting can be symmetric horizontal splitting or symmetric vertical splitting, but the present disclosure is not limited thereto.
[0109] In each split node of the binary tree, a flag indicating the split type (i.e., horizontal split or vertical split) can be additionally signaled. According to an embodiment, when the value of the flag is "0", it can indicate horizontal splitting, and when the value of the flag is "1", it can indicate vertical splitting.
[0110] However, the method for splitting a coding unit in the embodiments of the present disclosure is not limited to the above method, and asymmetric horizontal / vertical splitting, splitting into three rectangular coding units by a ternary tree, etc. can be applied thereto.
[0111] Picture prediction (motion compensation) for coding is performed on the coding unit that is no longer divided (i.e., the leaf node of the coding unit tree). The basic unit for performing such prediction is hereinafter referred to as a prediction unit or a prediction block.
[0112] Hereinafter, the term "unit" used in this specification can be used as a term to replace the prediction unit, which is the basic unit for performing prediction. However, the present disclosure is not limited thereto, and more generally, this term can be understood to include the concept of a coding unit.
[0113] To reconstruct the current unit for which decoding is performed, a decoded portion of the current picture or other pictures including the current unit can be used. A picture (slice) that is reconstructed using only the current picture, i.e., a picture (slice) that performs only intra prediction, is called an intra picture or I picture (slice), and a picture (slice) for which both intra prediction and inter prediction can be performed for reconstruction is called an inter picture (slice). Among the inter pictures (slices), a picture (slice) that uses at most one motion vector and reference index to predict each unit is called a predicted picture or P picture (slice), and a picture (slice) that uses at most two motion vectors and reference indexes to predict each unit among the inter pictures (slices) is called a bi-predicted picture or B picture (slice).
[0114] The intra prediction unit performs intra prediction for predicting the pixel value of a target unit based on the reconstructed region in the current picture. For example, the pixel value of the current unit can be predicted from the reconstructed pixels of the units located to the left and / or above the current unit centered on the current unit. In this case, the units located to the left of the current unit can include the left unit, the upper left unit, and the lower left unit adjacent to the current unit. In addition, the units located above the current unit can include the upper unit, the upper left unit, and the upper right unit adjacent to the current unit.
[0115] Meanwhile, the inter prediction unit performs inter prediction for predicting the pixel value of a target unit by using information about other reconstructed pictures than the current picture. In this case, the picture used for prediction is called a reference picture. An index indicating the reference picture including the corresponding reference region, motion vector information, etc. can be used to indicate which reference region will be used to predict the current unit during the inter prediction process.
[0116] Inter prediction can include L0 prediction, L1 prediction, and bi-prediction. L0 prediction means prediction using one reference picture included in L0 (the zero-th reference picture list), and L1 prediction means prediction using one reference picture included in L1 (the first reference picture list). For this, a set of motion information (e.g., motion vector and reference picture index) may be required. In the bi-prediction method, at most two reference regions can be used, and these two reference regions can exist in the same reference picture or in different pictures respectively. That is, in the bi-prediction method, at most two sets of motion information (e.g., motion vector and reference picture index) can be used, and these two motion vectors can correspond to the same reference picture index or different reference picture indexes. In this case, the reference pictures can be displayed (or output) before and after the current picture in terms of time.
[0117] A reference unit for a current unit can be obtained using a motion vector and a reference picture index. The reference unit exists in a reference picture having the reference picture index. In addition, the pixel value or interpolation value of the unit specified by the motion vector can be used as a predictor for the current unit. For motion prediction with pixel accuracy in units of sub-pixels (sub-pel), for example, an 8-tap interpolation filter can be used for the luminance signal, and a 4-tap interpolation filter can be used for the chrominance signal. However, the interpolation filter for motion prediction in units of sub-pixels is not limited to this. As described above, motion compensation for predicting the texture of the current unit from previously decoded pictures is performed using motion information.
[0118] Hereinafter, Figure 6 and Figure 7 the intra prediction method according to an embodiment of the present disclosure will be described in more detail. As described above, the intra prediction unit predicts the pixel value of the current unit by using adjacent pixels located to the left and / or above the current unit as reference pixels.
[0119] As Figure 6 illustrated, when the size of the current unit is N×N, up to (4N + 1) adjacent pixels located to the left and / or above the current unit can be used to set reference pixels. When at least some of the adjacent pixels to be used as reference pixels have not been reconstructed, the intra prediction unit can perform a reference sample filling process according to a preset rule to obtain reference pixels. In addition, the intra prediction unit can perform a reference sample filtering process to reduce errors in intra prediction. That is, reference pixels can be obtained by filtering adjacent pixels and / or pixels obtained by the reference sample filling process. The intra prediction unit uses the reference pixels obtained in this way to predict the pixel of the current unit.
[0120] Figure 7 An embodiment of a prediction mode for intra prediction is illustrated. For intra prediction, intra prediction mode information indicating the intra prediction direction can be signaled. When the current unit is an intra prediction unit, the video signal decoding device extracts the intra prediction mode information of the current unit from the bitstream. The intra prediction unit of the video signal decoding device performs intra prediction for the current unit based on the extracted intra prediction mode information.
[0121] According to an embodiment of the present disclosure, the intra prediction mode can include a total of 67 modes. Each intra prediction mode can be indicated by a preset index (i.e., intra mode index). For example, as Figure 7As illustrated, the intra mode index 0 may indicate the planar mode, the intra mode index 1 may indicate the DC mode, and the intra mode indices 2 to 66 may respectively indicate different directional modes (i.e., angular modes). The intra prediction unit determines the reference pixels and / or interpolated reference pixels to be used for intra prediction of the current unit based on the intra prediction mode information of the current unit. When the intra mode index indicates a specific directional mode, the reference pixels or interpolated reference pixels corresponding to the specific direction from the current pixel of the current unit are used for prediction of the current pixel. Thus, different sets of reference pixels and / or interpolated reference pixels may be used for intra prediction according to the intra prediction mode.
[0122] After performing intra prediction of the current unit using the reference pixels and the intra prediction mode information, the video signal decoding apparatus reconstructs the pixel value of the current unit by adding the residual signal of the current unit obtained from the inverse transform unit and the intra prediction sub of the current unit.
[0123] Figure 8 is a diagram illustrating inter prediction according to an embodiment of the present disclosure.
[0124] As described above, when encoding or decoding a current picture or block, it may be predicted from another picture or block. That is, the current picture or block may be encoded and decoded based on the similarity with other pictures or blocks. The current picture or block may be encoded and decoded using signaling that omits the parts similar to another picture or block from the current picture or block, which will be further described below. Prediction may be performed in units of blocks.
[0125] Reference Figure 8 , there is a reference picture on the left and a current picture on the right, and the similarity with the reference picture or a part of the reference picture may be used to predict the current picture or a part of the current picture. When in Figure 8 the rectangle indicated by the solid line in the current picture is the current block to be encoded and decoded, the current block may be predicted according to the rectangle indicated by the dotted line in the reference picture. In this case, there may be information indicating the block (reference block) to be referred to by the current block, which may be directly signaled or may be constructed by any convention to reduce signaling overhead. The information indicating the block to be referred to by the current block may include a motion vector. The motion vector may be a vector indicating the relative position between the current block and the reference block in the picture. Reference Figure 8 , there is a part indicated by the dotted line in the reference picture, and the vector indicating how to move the current block to the block to be referred to in the reference picture may be the motion vector. That is, the block that appears when the current block is moved according to the motion vector may be Figure 8 the part indicated by the dotted line in the current picture in
[0126] In addition, the information indicating the block to be referred to by the current block may include information indicating a reference picture. The information indicating a reference picture may include a list of reference pictures and a reference picture index. The list of reference pictures is a list including reference pictures, and a reference block in the reference pictures included in the list of reference pictures can be used. That is, the current block can be predicted based on the reference pictures included in the list of reference pictures. In addition, the reference picture index may be an index for indicating the reference picture to be used.
[0127] Figure 9 FIG. is a diagram illustrating a method of signaling a motion vector according to an embodiment of the present disclosure.
[0128] According to an embodiment of the present disclosure, a motion vector MV may be generated based on a motion vector predictor MVP. For example, the motion vector predictor may be a motion vector as illustrated below.
[0129] MV = MVP
[0130] As another example, the motion vector may be based on a motion vector difference (MVD) as follows. The motion vector difference MVD may be added to the motion vector predictor to represent an accurate motion vector.
[0131] MV = MVP + MVD
[0132] In addition, in video encoding, the encoder may transmit the determined motion vector information to the decoder, and the decoder may generate a motion vector and determine a prediction block based on the received motion vector information. For example, the motion vector information may include information about the motion vector predictor and the motion vector difference. In this case, the components of the motion vector information may vary according to the mode. For example, in the merge mode, the motion vector information may include information about the motion vector predictor but may not include the motion vector difference. As another example, in the advanced motion vector prediction (AMVP) mode, the motion vector information may include information about the motion vector predictor and may include the motion vector difference.
[0133] To determine, transmit, and receive information about motion vector predictors, the encoder and decoder may generate motion vector predictor (MVP) candidates in the same manner. For example, the encoder and decoder may generate the same MVP candidates in the same order. Also, the encoder may transmit an index mvp_lx_flag indicating the MVP (motion vector predictor) determined from the generated MVP candidates to the decoder, and the decoder may find the MVP (motion vector predictor) and MV determined based on this index mvp_lx_flag. The index mvp_lx_flag may include a motion vector predictor index mvp_l0_flag for the zero reference picture list list 0 and a motion vector predictor index mvp_l1_flag for the first reference picture list list 1. Refer to Figures 42 to 45 Describe the method of receiving the index mvp_lx_flag.
[0134] MVP candidates and MVP candidate generation methods may include spatial candidates, temporal candidates, etc. Spatial candidates may be the motion vectors of blocks located at a predetermined position from the current block. For example, spatial candidates may be the motion vectors corresponding to blocks or positions adjacent or non - adjacent to the current block. Temporal candidates may be the motion vectors of blocks in pictures different from the current picture. Alternatively, MVP candidates may include affine motion vectors, ATMVP, STMVP, combinations of the above motion vectors, average vectors of the above motion vectors, zero motion vectors, etc.
[0135] In addition, information indicating the above - mentioned reference pictures may also be transmitted from the encoder to the decoder. Also, when the reference picture corresponding to the MVP candidate does not correspond to the information indicating the reference picture, motion vector scaling may be performed. Motion vector scaling may be based on the picture order count (POC) of the current picture, the POC of the reference picture of the current block, the POC of the reference picture of the MVP candidate, and the calculation of the MVP candidate.
[0136] Figure 10 FIG. illustrates the motion vector difference syntax according to an embodiment of the present disclosure.
[0137] The motion vector difference may be coded by partitioning the sign and absolute value of the motion vector difference. That is, the sign and absolute value of the motion vector difference may be different syntaxes. Also, the absolute value of the motion vector difference may be directly coded, but may be coded as illustrated in Figure 10 in the case of including a flag indicating whether the absolute value is greater than N. If the absolute value is greater than N, the value (absolute value - N) may be signaled together. In Figure 10For example, abs_mvd_greater0_flag can be transmitted, and this flag can be a flag indicating whether the absolute value is greater than 0. If abs_mvd_greater0_flag indicates that the absolute value is not greater than 0, the absolute value can be determined to be 0. Additionally, if abs_mvd_greater0_flag indicates that the absolute value is greater than 0, there may be additional syntax. For example, there may be abs_mvd_greater1_flag, and this flag can be a flag indicating whether the absolute value is greater than 1. If abs_mvd_greater1_flag indicates that the absolute value is not greater than 1, the absolute value can be determined to be 1. If abs_mvd_greater1_flag indicates that the absolute value is greater than 1, there may be additional syntax. For example, there may be abs_mvd_minus2, whose value may be (absolute value - 2). It indicates (absolute value - 2) because it is determined by the above abs_mvd_greater0_flag and abs_mvd_greater1_flag that the absolute value is greater than 1 (more than 2). When abs_mvd_minus2 is binarized into a variable length, it is used to signal with fewer bits. For example, there are variable length binarization methods such as Exp-Golomb, truncated unary, and truncated Rice. In addition, mvd_sign_flag can be a flag indicating the sign of the motion vector difference.
[0138] Although the compilation method has been described by the motion vector difference in this embodiment, information other than the motion vector difference can also be partitioned regarding the sign and the absolute value, and the absolute value can be compiled with flags indicating whether the absolute value is greater than a certain value and the value obtained by subtracting a certain value from the absolute value. Additionally, Figure 10 The [0] and [1] in [] can represent component indices. For example, [0] and [1] can represent the x-component and the y-component.
[0139] Figure 11 is a diagram illustrating the signaling of adaptive motion vector resolution according to an embodiment of the present disclosure.
[0140] According to an embodiment of the present disclosure, the resolution indicating the motion vector or the motion vector difference may vary. In other words, the resolution at which the motion vector or the motion vector difference is encoded may vary. For example, the resolution may be expressed in pixels. For example, the motion vector or the motion vector difference may be signaled in units of 1 / 4 (quarter) pixel, 1 / 2 (half) pixel, 1 (integer) pixel, 2 pixels, 4 pixels, etc. When it is desired to represent 16, it may be encoded as 64 (1 / 4 * 64 = 16) for signaling in units of 1 / 4, encoded as 16 (1 * 16 = 16) for signaling in units of 1, and encoded as 4 (4 * 4 = 16) for signaling in units of 4. That is, the value may be determined as follows.
[0141] valueDetermined = resolution * valuePerResolution
[0142] Here, valueDetermined is the value to be transmitted and may be a motion vector or a motion vector difference in this embodiment. In addition, valuePerResolution may be the value obtained by expressing valueDetermined in units of [ / resolution].
[0143] In this case, if the value signaled as the motion vector or the motion vector difference is not divisible by the resolution, an inaccurate value may be transmitted, such as by rounding, rather than the motion vector or the motion vector difference with the best prediction performance. When a high resolution is used, the inaccuracy may be reduced, but since the encoded value is large, many bits may be used, and when a low resolution is used, the inaccuracy may increase, but since the encoded value is small, fewer bits may be used.
[0144] It is also possible to set the resolution differently for units such as blocks, CUs, slices, etc. Therefore, the resolution can be adaptively applied to suit the unit.
[0145] The resolution may be signaled from the encoder to the decoder. In this case, the signaling for the resolution may be the signaling binarized with the variable length described above. In this case, when signaling is performed with the index corresponding to the minimum value (the foremost value), the signaling overhead is reduced.
[0146] In one embodiment, the signaling index may be matched in the order from high resolution (detailed signaling) to low resolution.
[0147] Figure 11The figure illustrates signaling for three types of resolutions. In this case, the three signals can be 0, 10, and 11, and each of the three signals can correspond to resolution 1, resolution 2, and resolution 3. When signaling resolution 1, the signaling overhead is low because it requires 1 bit to signal resolution 1 and 2 bits to signal the remaining resolutions. In Figure 11 the example of
[0148] In the following disclosure, the motion vector resolution may refer to the resolution of the motion vector difference.
[0149] Figure 12 is a diagram illustrating the signaling of an adaptive motion vector resolution according to an embodiment of the present disclosure.
[0150] As described with reference to Figure 11 Since the number of bits required to signal the resolution may vary depending on the resolution, the signaling method can be changed according to the situation. For example, depending on the situation, the signaling value used to signal a certain resolution may be different. For example, the signaling index and the resolution may be matched in a different order depending on the situation. For example, the resolutions corresponding to signals 0, 10, 110,... can be resolution 1, resolution 2, resolution 3,... respectively in a certain situation, and can be in a different order from resolution 1, resolution 2, resolution 3,... respectively in another situation. It is also possible to define two or more situations.
[0151] Referring to Figure 12 , the resolutions corresponding to 0, 10, and 11 can be resolution 1, resolution 2, and resolution 3 respectively in case 1, and can be resolution 2, resolution 1, and resolution 3 respectively in case 2. In this case, there may be two or more situations.
[0152] Figure 13 is a diagram illustrating the signaling of an adaptive motion vector resolution according to an embodiment of the present disclosure.
[0153] As described with reference to Figure 12 , the motion vector resolution can be signaled differently depending on the situation. For example, assuming there are resolutions of 1 / 4, 1, and 4 pixels, in some cases, the signaling described with reference to Figure 11 can be used, and in some cases, the signaling illustrated in Figure 13 (a) or Figure 13 (b) can be used. There can be only Figure 11 , Figure 13 (a), and Figure 13(b) either two of the cases or all three cases. Thus, in some cases, it is possible to signal a resolution that is not the highest resolution with fewer bits.
[0154] Figure 14 FIG. is a diagram illustrating signaling of adaptive motion vector resolution according to an embodiment of the present disclosure.
[0155] According to an embodiment of the present disclosure, the possible resolutions may depend on the reference Figure 11 described cases in the adaptive motion vector resolution and change. For example, the resolution values may change depending on the case. In one embodiment, it is possible to use a resolution value among resolution 1, resolution 2, resolution 3, resolution 4,... in one case and a resolution value among resolution A, resolution B, resolution C, resolution D,... in another case. In addition, there may be a non-empty intersection between {resolution 1, resolution 2, resolution 3, resolution 4,...} and {resolution B, resolution C, resolution D,...}. That is, a certain resolution value can be used in two or more cases, and in two or more cases, the set of available resolution values may be different. In addition, the number of available resolution values for each case may be different.
[0156] Reference Figure 14 , resolution 1, resolution 2, and resolution 3 can be used in case 1, and resolution A, resolution B, and resolution C can be used in case 2. For example, resolution 1, resolution 2, and resolution 3 can be 1 / 4, 1, and 4 pixels. In addition, for example, resolution A, resolution B, and resolution C can be 1 / 4, 1 / 2, and 1 pixel.
[0157] Figure 15 FIG. is a diagram illustrating signaling of adaptive motion vector resolution according to an embodiment of the present disclosure.
[0158] Reference Figure 15 , the resolution signaling may be different according to the motion vector candidate or the motion vector predictor candidate. For example, it may be related to which case is selected in Figures 12 to 14 described, or which candidate the signaled motion vector or motion vector predictor is. The method of making the signaling different may follow Figures 12 to 14 's method.
[0159] For example, the cases can be defined differently depending on their position among the candidates. Alternatively, the cases can be defined differently depending on how the candidates are constructed.
[0160] An encoder or decoder may generate a candidate list including at least one MV candidate (motion vector candidate) or at least one MVP candidate (motion vector predictor candidate). There may be a tendency that the previous MVP (motion vector predictor) in the candidate list of MV candidates (motion vector candidates) or MVP candidates (motion vector predictor candidates) has high accuracy while the later MVP candidates (motion vector predictor candidates) in the candidate list have low accuracy. This may be designed such that the candidates located in the front of the candidate list are signaled with fewer bits and the previous MVP (motion vector predictor) in the candidate list has higher accuracy. In an embodiment of the present disclosure, if the accuracy of the MVP (motion vector predictor) is high, the motion vector difference (MVD) value representing a motion vector with good prediction performance may be small, and if the accuracy of the MVP (motion vector predictor) is low, the MVD (motion vector difference) value representing a motion vector with good prediction performance may be large. Therefore, when the accuracy of the MVP is low, the bits required to represent the motion vector difference (e.g., the value representing the difference based on the resolution) can be reduced by signaling at a low resolution.
[0161] Based on this principle, since a low resolution can be used when the accuracy of the MVP (motion vector predictor) is low, according to an embodiment of the present disclosure, it is possible to promise to signal a resolution that is not the highest resolution with the least number of bits according to the MVP candidate (motion vector predictor candidate). For example, when resolutions of 1 / 4, 1, or 4 pixels are possible, 1 or 4 can be signaled with the least number of bits (1 bit). Refer to Figure 15 , for candidate 1 and candidate 2, the high-resolution 1 / 4 pixel is signaled with the least number of bits, and for candidate N behind candidate 1 and candidate 2, a resolution that is not 1 / 4 pixel is signaled with the least number of bits.
[0162] Figure 16 is a diagram illustrating the signaling of adaptive motion vector resolution according to an embodiment of the present disclosure.
[0163] As described in reference Figure 15 , the motion vector resolution signaling may vary depending on which candidate the determined motion vector or motion vector predictor is. The method of varying the signaling may follow the method of Figures 12 to 14 .
[0164] Refer to Figure 16, for some candidates, the highest resolution is signaled using the fewest bits, while for some candidates, a resolution that is not the highest resolution is signaled using the fewest bits. For example, a candidate that signals a resolution that is not the highest resolution using the fewest bits may be an incorrect candidate. For example, a candidate that signals a resolution that is not the highest resolution using the fewest bits can be a temporal candidate, a zero motion vector, a non-adjacent spatial candidate, a candidate based on the presence or absence of a refinement process, etc.
[0165] A temporal candidate can be a motion vector from another picture. A zero motion vector can be a motion vector in which all vector components are 0. A non-adjacent spatial candidate can be a motion vector referenced from a position that is not adjacent to the current block. The refinement process can be a process of refining a motion vector predictor. For example, refinement can be performed by template matching, bilateral matching, etc.
[0166] According to an embodiment of the present disclosure, a motion vector difference can be added after refining the motion vector predictor. This can reduce the motion vector difference by making the motion vector predictor accurate. In this case, the motion vector resolution signaling can be different for candidates without a refinement process. For example, a resolution that is not the highest resolution can be signaled using the fewest bits.
[0167] According to another embodiment, after adding the motion vector difference to the motion vector predictor, a refinement process can be performed. In this case, the motion vector resolution signaling can be different for candidates with a refinement process. For example, a resolution that is not the highest resolution can be signaled using the fewest bits. This may be because, since the refinement process will be performed after adding the motion vector difference, even if the MV difference (motion vector difference) is not signaled as accurately as possible, the prediction error can be reduced through the refinement process (such that the prediction error is small).
[0168] As another example, when the selected candidate differs from another candidate by a specific value or more, the motion vector resolution signaling can be changed.
[0169] In another embodiment, depending on which candidate the determined motion vector or motion vector predictor is, the subsequent motion vector refinement process may be made different. The motion vector refinement process may be a process for finding a more accurate motion vector. For example, the motion vector refinement process may be a process of finding a block that matches the current block from a reference point according to a set convention (e.g., template matching or bilateral matching). The reference point may be a position corresponding to the determined motion vector or motion vector predictor. In this case, the degree of movement from the reference point may vary according to the set convention, and making the motion vector refinement process different may make the degree of movement from the reference point different. For example, for an accurate candidate, the motion vector refinement process may start from a detailed refinement process, and for an inaccurate candidate, from a less detailed refinement process. The accurate candidate and the inaccurate candidate may be determined according to the position in the candidate list or how the candidate is generated. How the candidate is generated may be related to the position that brings the spatial candidate. Additionally, reference may be made to Figure 16 for a description. Furthermore, the detailed refinement or the less detailed refinement may be finding a matching block when moving a little from the reference point or moving a lot from the reference point. Additionally, in the case of finding a matching block during a large movement, a process of adding to find by starting to move a little from the most matching block found during the large movement can be enabled.
[0170] In another embodiment, the motion vector resolution signaling may be changed based on the POC of the current picture and the POC of the reference picture of the motion vector or motion vector predictor candidate. The method of making the signaling different may follow Figures 12 to 14 the method.
[0171] For example, when the difference between the picture order count (POC) of the current picture and the POC of the reference picture of the motion vector or motion vector predictor candidate is large, the motion vector or motion vector predictor may be inaccurate, and the resolution that is not high resolution may be signaled using the fewest bits.
[0172] As another example, the motion vector resolution signaling may be changed based on whether motion vector scaling needs to be performed. For example, when the selected motion vector or motion vector predictor is a candidate for which motion vector scaling is performed, the resolution that is not high resolution may be signaled using the fewest bits. The case of performing motion vector scaling may be the case where the reference picture of the current block and the reference picture of the reference candidate are different.
[0173] Figure 17 is a diagram illustrating affine motion prediction according to an embodiment of the present disclosure.
[0174] In as referenced Figure 8In the described conventional prediction methods, the current block can be predicted from a block starting from the moving position without rotation or scaling. The prediction can be made from a reference block having the same size, shape, and angle as the current block. Figure 8 This is only a translational motion model. However, the content included in an actual image can have more complex motion, and the prediction performance can be further improved if the prediction is made from various shapes.
[0175] Reference Figure 17 , the currently predicted block is indicated by a solid line in the current picture. The current block can be predicted by referring to a block having a different shape, size, and angle from the current predicted block, and the reference block is indicated by a dashed line in the reference picture. The position identical to the position of the reference block in the picture is indicated by a dashed line in the current picture. In this case, the reference block can be a block represented by an affine transformation in the current block. In this way, stretching (scaling), rotation, shear, reflection, orthogonal projection, etc. can be represented.
[0176] The number of parameters representing the affine motion and the affine transformation can vary. If more parameters are used, more various motions can be represented than when fewer parameters are used, but there is a possibility that an overhead may occur in signaling or calculation, etc.
[0177] For example, the affine transformation can be represented by six parameters. Alternatively, the affine transformation can be represented by three control point motion vectors.
[0178] Figure 18 is a diagram illustrating affine motion prediction according to an embodiment of the present disclosure.
[0179] It is possible to represent complex motion using an affine transformation such as Figure 17 , but in order to reduce signaling overhead and calculation for this, a simpler affine motion prediction or affine transformation can be used. By restricting the motion, a simpler affine motion prediction can be performed. Restricting the movement can restrict the shape into which the current block is transformed into the reference block.
[0180] Reference Figure 18 , the affine motion prediction can be performed using the control point motion vectors of v0 and v1. Using two vectors v0 and v1 may be equivalent to using four parameters. With two vectors v0 and v1 or four parameters, the shape of the reference block from which the current block is predicted can be indicated. By using such a simple affine transformation, the motion of rotation and scaling (magnification / minification) of the block can be represented. Reference Figure 18 , it is possible to predict the current block indicated by the solid line from the position indicated by the dashed line in Figure 18 . Each point (pixel) of the current block can be mapped to another point by an affine transformation.
[0181] Figure 19It is a diagram showing an expression of a motion vector field according to an embodiment of the present disclosure. Figure 18 The control point motion vector v0 in Figure 19 can be (v_0x, v_0y) and can be the motion vector of the upper left control point. In addition, the control point motion vector v1 can be (v_1x, v_1y) and can be the motion vector of the upper right control point. In this case, the motion vector (v_x, v_y) at the (x, y) position can be as shown in Figure 19 . Therefore, the motion vector of each pixel position or a certain position can be estimated according to the expression of
[0182] . In addition, Figure 19 the (x, y) in the expression of
[0183] can be the relative coordinates within the block. For example, (x, y) can be the position when the upper left position of the block is (0, 0). Figure 19 If v0 is the control point motion vector at the position (x0, y0) on the picture and v1 is the control point motion vector at the position (x1, y1) on the picture, if it is intended to represent the position (x, y) within the block using the same coordinates as the positions of v0 and v1, it can be represented by changing x and y in the expression of
[0184] Figure 20 to (x - x0) and (y - y0) respectively. In addition, w (the width of the block) can be (x1 - x0).
[0185] According to an embodiment of the present disclosure, an affine motion can be represented using multiple control point motion vectors or multiple parameters.
[0186] Referring to Figure 20 , the control point motion vectors of v0, v1, and v2 can be used to perform affine motion prediction. Using the three vectors v0, v1, and v2 may be equivalent to using six parameters. The three vectors v0, v1, and v2 or the six parameters can indicate the shape of the reference block from which the current block is predicted. Referring to Figure 20 , the current block indicated by the solid line can be predicted from the position indicated by the dotted line in Figure 20 in the reference picture. Each point (pixel) of the current block can be mapped to another point through an affine transformation.
[0187] Figure 21 It is a diagram showing an expression of a motion vector field according to an embodiment of the present disclosure. In Figure 20Among them, the control point motion vector v0 can be (mv_0^x, mv_0^y) and can be the motion vector of the upper-left control point, the control point motion vector v1 can be (mv_1^x, mv_1^y) and can be the motion vector of the upper-right control point, and the control point motion vector v2 can be (mv_2^x, mv_2^y) and can be the motion vector of the lower-left control point. In this case, the motion vector (mv^x, mv^y) at the (x, y) position can be as shown in Figure 21 as illustrated. Therefore, the motion vector at each pixel position or a certain position can be estimated according to the expression in Figure 21 , which is based on v0, v1, and v2.
[0188] In addition, Figure 21 in the expression of (x, y) can be the relative coordinates within the block. For example, (x, y) can be the position when the upper-left position of the block is (0, 0). Therefore, if it is assumed that v0 is the control point motion vector at the position (x0, y0), v1 is the control point motion vector at the position (x1, y1), and v2 is the control point motion vector at the position (x2, y2) and it is intended to represent (x, y) using the same coordinates as the positions of v0, v1, and v2, it can be represented by changing x and y in the expression of Figure 21 to (x - x0) and (y - y0) respectively. In addition, w (the width of the block) can be (x1 - x0), and h (the height of the block) can be (y2 - y0).
[0189] Figure 22 is a diagram illustrating affine motion prediction according to an embodiment of the present disclosure.
[0190] As described above, there is a motion vector field, and the motion vector can be calculated for each pixel. However, for simplicity, the affine transformation can be performed based on the sub-blocks illustrated in Figure 22 . For example, Figure 22 a small rectangle in (a) is a sub-block, a representative motion vector of the sub-block can be constructed, and the representative motion vector can be used for the pixels of the sub-block. In addition, for complex motions represented in Figure 17 , 18 , 20, etc., the sub-block can correspond to the reference block by representing such motions, or it can be made simpler by applying only translational motion to the sub-block. In Figure 22 (a), v0, v1, and v2 can be control point motion vectors.
[0191] In this case, the size of the sub-block can be M*N, and M and N can be the same as Figure 22the same as that illustrated in (b). Also, MvPre may be motion vector fractional accuracy. Further, (v_0x, v_0y), (v_1x, v_1y), and (v_2x, v_2y) may be the motion vectors of the upper left, upper right, and lower left control points, respectively. (v_2x, v_2y) may be the motion vector of the lower left control point of the current block. For example, in the case of 4-parameters, it may be the MV (motion vector) of the lower left control point calculated by the expression of Figure 19 the lower left control point calculated by the expression of
[0192] Further, when constructing the representative motion vector of a sub-block, the center sample position of the sub-block can be used to calculate the representative motion vector. Further, when constructing the motion vector of a sub-block, a motion vector with higher accuracy than the normal motion vector can be used, and for this purpose, a motion compensation interpolation filter can be applied.
[0193] In another embodiment, the size of the sub-block is immutable and can be fixed to a specific size. For example, the sub-block size can be fixed to 4*4 size.
[0194] Figure 23 is a diagram illustrating a mode of affine motion prediction according to an embodiment of the present disclosure.
[0195] According to an embodiment of the present disclosure, there may be an affine inter-frame mode as an example of affine motion prediction. There may be a flag indicating that it is an affine inter-frame mode. Referring to Figure 23 there may be blocks at positions A, B, C, D, and E close to v0 and v1, and the motion vectors corresponding to these blocks may be called vA, vB, vC, vD, and vE, respectively. Using this, a candidate list can be constructed for the following motion vectors or motion vector predictors.
[0196] {(v0,v1)|v0={vA,vB,vC},v1={vD,vE}}
[0197] That is, the (v0, vl) pair can be constructed with v0 selected from vA, vB, and vC and vl selected from vD and vE. In this case, the motion vector can be scaled according to the picture order count (POC) of the reference of the neighboring block, the POC of the reference of the current CU (current coding unit; current block), and the POC of the current CU. When the candidate list is constructed with the same motion vector pairs as above, it is possible to signal which candidate in the candidate list is selected and whether it is selected. Further, if the candidate list is not sufficiently filled, the candidate list can be filled with other inter-frame prediction candidates. For example, advanced motion vector prediction (AMVP) candidates can be used for filling. In addition, instead of directly using v0 and v1 selected from the candidate list as the control point motion vectors for affine motion prediction, the difference for correction can be signaled, thereby enabling a better control point motion vector to be constructed. That is, in the decoder, v0' and v1' constructed by adding the difference to v0 and v1 selected from the candidate list can be used as the control point motion vectors for affine motion prediction.
[0198] In one embodiment, for a coding unit (CU) of a particular size or larger, an affine inter-frame mode can also be used.
[0199] Figure 24 FIG. is a diagram illustrating a mode of affine motion prediction according to an embodiment of the present disclosure.
[0200] According to an embodiment of the present disclosure, there may be an affine merge mode as an example of affine motion prediction. There may be a flag indicating that it is an affine merge mode. In the affine merge mode, when affine motion prediction is used around the current block, the control point motion vector of the current block can be calculated from the motion vectors of this block or around this block. For example, when checking whether a neighboring block uses affine motion prediction, the neighboring blocks that become candidates can be as Figure 24 illustrated in (a). Additionally, it is possible to check whether affine motion prediction is used in the order of A, B, C, D, and E, and when a block using affine motion prediction is found, the control point motion vector of the current block can be calculated using the motion vectors of the block or around the block. A, B, C, D, and E can be left, top, top - right, bottom - left, and top - left, respectively, as Figure 24 illustrated in (a).
[0201] In an embodiment, when the block at position A uses affine motion prediction as Figure 24 illustrated in (b), v0 and v1 can be calculated using the motion vectors of the block or around the block. The motion vectors around the block can be v2, v3, and v4.
[0202] In the previous embodiments, the order of neighboring blocks to be referred to is determined. However, the performance of the control point motion vectors derived from a specific position is not always better. Therefore, in another embodiment, it may be signaled which block located at which position is referred to for deriving the control point motion vectors. For example, in the order of A, B, C, D, E of Figure 24 (a), the candidate positions for deriving the control point motion vectors are determined, and the candidate positions to be referred to can be signaled.
[0203] In another embodiment, when deriving the control point motion vectors, the accuracy can be increased by obtaining from the neighboring blocks of each control point motion vector. For example, referring to Figure 24 , the left block can be referred to when deriving v0, and the upper block can be referred to when deriving v1. Alternatively, A, D, or E can be referred to when deriving v0, and B or C can be referred to when deriving v1.
[0204] Figure 25 is a diagram illustrating the derivation of an affine motion predictor according to an embodiment of the present disclosure.
[0205] Affine motion prediction may require control point motion vectors, and a motion vector field, i.e., the motion vectors of sub-blocks or a certain position, can be calculated based on the control point motion vectors. The control point motion vectors can be referred to as seed vectors.
[0206] In this case, the control point MV (control point motion vector) can be based on a predictor. For example, the predictor can be the control point MV (control point motion vector). As another example, the control point MV can be calculated based on the predictor and a difference. Specifically, the control point MV (control point motion vector) can be calculated by adding or subtracting the difference from the predictor.
[0207] In this case, during the process of constructing the predictor of the control point MV, it can be derived from the control point MV (control point motion vector) or MV (motion vector) of neighboring blocks for which affine motion prediction (affine motion compensation (MC)) is performed. For example, if the block corresponding to a preset position undergoes affine motion prediction, the predictor for affine motion compensation of the current block can be derived from the control point MV or MV of that block. Referring to Figure 25 , the preset positions can be A0, A1, B0, B1, and B2. Alternatively, the preset positions can include positions adjacent to the current block and positions not adjacent to the current block. In addition, the control point MV (control point motion vector) or the MV (motion vector) at the preset position (space) can be referred to, and the temporal control point MV or the MV at the preset position can be referred to.
[0208] Candidates for affine motion compensation (MC) can be in the same manner as Figure 25is constructed in the same manner as the embodiment, and this candidate may also be referred to as an inheritance candidate. Alternatively, such a candidate may be referred to as a merge candidate. Additionally, when referring to the preset positions in the method of Figure 25 they may be referred to in a preset order.
[0209] Figure 26 is a diagram illustrating the derivation of an affine motion predictor according to an embodiment of the present disclosure.
[0210] Affine motion prediction may require control point motion vectors, and the motion vector field, i.e., the motion vectors of sub-blocks or a certain position, may be calculated based on the control point motion vectors. The control point motion vectors may also be referred to as seed vectors.
[0211] In this case, the control point MV (control point motion vector) may be based on the predictor. For example, the predictor may be the control point MV (control point motion vector). As another example, the control point MV (control point motion vector) may be calculated based on the predictor and a difference. Specifically, the control point MV (control point motion vector) may be calculated by adding or subtracting the difference from the predictor.
[0212] In this case, it can be derived from neighboring MVs during the process of constructing the predictor of the control point MV. In this case, the neighboring MVs may include MVs (motion vectors) that have not undergone affine motion compensation (MC). For example, when deriving each control point MV (control point motion vector) of the current block, the MVs at the preset positions of each control point MV may be used as the predictor of the control point MV. For example, the preset position may be a part included in a block adjacent to a part of the preset position.
[0213] Referring to Figure 26 , the control point MVs (control point motion vectors) mv0, mv1, and mv2 can be determined. In this case, according to an embodiment of the present disclosure, the MVs (motion vectors) corresponding to the preset positions A, B, and C may be used as the predictor for mv0. Additionally, the MVs (motion vectors) corresponding to the preset positions D and E may be used as the predictor for mv1. The MVs (motion vectors) corresponding to the preset positions F and G may be used as the predictor for mv2.
[0214] Furthermore, when determining each predictor of the control point MVs (control point motion vectors) mv0, mv1, and mv2 according to the Figure 26 embodiment, the order of the preset positions referring to each control point position may be determined. Additionally, for each control point, there may be multiple preset positions referred to as predictors of the control point MV, and the possible combinations of the preset positions may be determined.
[0215] The candidates for affine MC (affine motion compensation) may be in the same manner asFigure 26 is constructed in the same manner as in the embodiments of, and this candidate may also be referred to as a constructed candidate. Alternatively, this candidate may be referred to as an inter-frame candidate or a virtual candidate. Additionally, when referring to Figure 41 the preset positions in the method of, the reference may be made in a preset order.
[0216] According to an embodiment of the present disclosure, using the embodiments described with reference to Figures 23 to 26 or combinations thereof, a candidate list for affine MC (affine motion compensation) or a candidate list for control point MV (control point motion vector) of affine MC (affine motion compensation) may be generated.
[0217] Figure 27 is a diagram illustrating the derivation of an affine motion predictor according to an embodiment of the present disclosure.
[0218] As described with reference to Figures 24 to 25 the control point MV (control point motion vector) for affine motion prediction of the current block may be derived from neighboring blocks that have undergone affine motion prediction. In this case, the same method as Figure 27 may be used. In Figure 27 the expressions, the MVs or control point MVs (motion vectors) of the upper left, upper right, and lower left of the neighboring blocks that have undergone affine motion prediction are (v_E0x, v_E0y), (v_E1x, v_E1y), and (v_E2x, v_E2y), respectively. Additionally, the coordinates of the upper left, upper right, and lower left of the neighboring blocks that have undergone affine motion prediction may be (x_E0, y_E0), (x_E1, y_E1), and (x_E2, y_E2), respectively. In this case, (v_0x, v_0y) and (v_1x, v_1y) as the predictor or control point MV (control point motion vector) of the control point MV (control point motion vector) of the current block may be calculated according to 27.
[0219] Figure 28 is a diagram illustrating the derivation of an affine motion predictor according to an embodiment of the present disclosure.
[0220] As described above, affine motion compensation may require multiple control motion MVs (control point motion vectors) or multiple control point MV predictors (control point motion vector predictors). In this case, another control motion MV (control point motion vector) or control point MV predictor (control point motion vector predictor) may be derived from a certain control motion MV (control point motion vector) or control point MV predictor (control point motion vector predictor).
[0221] For example, when constructing two control point MVs (control point motion vectors) or two control point MV predictors (control point motion vector predictors) by the method described in the previous figures, another control point MV (control point motion vector) or another control point MV predictor (control point motion vector predictor) can be generated based on this.
[0222] Reference Figure 28 , which illustrates a method of generating mv0, mv1, and mv2 of control point MVs (control point motion vectors) or control point MVs (control point motion vectors) as the upper left, upper right, and lower left. In the figure, x and y represent the x component and the y component respectively, and the current block size can be w*h.
[0223] Figure 29 is a diagram illustrating a method of generating a control point motion vector according to an embodiment of the present disclosure.
[0224] According to an embodiment of the present disclosure, the control point MV (control point motion vector) can be determined by constructing a predictor of the control point MV (control point motion vector) to facilitate performing affine MC (affine motion compensation) on the current block and adding a difference thereto. According to the embodiment, the predictor of the control point MV can be constructed by referring to Figures 23 to 26 the method described. The difference can be signaled from the encoder to the decoder.
[0225] Reference Figure 29 , there may be a difference in the control point MV (control point motion vector). In addition, the difference in the control point MV (control point motion vector) can be signaled separately. Figure 29 (a) illustrates a method of determining mv0 and mv1 of a control point MV as a 4-parameter model, and Figure 29 (b) illustrates a method of determining mv0, mv1, and mv2 of a control point MV as a 6-parameter model. The control point MV (control point motion vector) is determined by adding mvd0, mvd1, and mvd2, which are the differences of the control point MV (control point motion vector), to the predictor.
[0226] By Figure 29 The terms indicated by the overline in can be predictors of the control point MV (control point motion vector).
[0227] Figure 30 is a diagram illustrating a method of determining a motion vector difference by referring to Figure 29 the method described.
[0228] As an embodiment, the motion vector difference can be signaled by referring to Figure 10 the method described. And, the motion vector difference determined by the signaling method can be Figure 30 the lMvd of. Moreover, such as inFigure 29 The value of the motion vector difference mvd signaled in the middle, i.e., values such as mvd0, mvd1, mvd2 can be Figure 30 the lMvd in. As referenced Figure 29 As described, the signaled mvd (motion vector difference) can be determined as the difference from the predictor of the control point MV (control point motion vector), and the determined difference can be Figure 30 the MvdL0 and MvdL1 of. L0 can indicate reference list 0 (the zero - reference picture list), and L1 can indicate reference list 1 (the first reference picture list). compIdx is the component index and can indicate the x, y components, etc.
[0229] Figure 31 FIG. is a diagram illustrating a method of generating a control point motion vector according to an embodiment of the present disclosure.
[0230] According to an embodiment of the present disclosure, the control point MV (control point motion vector) can be determined by constructing a predictor of the control point MV to facilitate affine MC (affine motion compensation) for the current block and adding a difference thereto. According to an embodiment, the predictor of the control point MV (control point motion vector) can be constructed by referring to the Figures 23 to 26 method described. The difference can be signaled from the encoder to the decoder.
[0231] Referring to Figure 31 , there can be a predictor for the difference of each control point MV (control point motion vector). For example, the difference of another control point MV (control point motion vector) can be determined based on the difference of a certain control point MV (control point motion vector). This can be based on the similarity between the differences of the control point MV (control point motion vector). Because the differences are similar, if a predictor is determined, a small difference can be generated with the predictor. In this case, the difference predictor of the control point MV (control point motion vector) can be signaled, and the difference from the difference predictor of the control point MV (control point motion vector) can be signaled.
[0232] Figure 31 (a) illustrates a method of determining mv0 and mv1 of the control point MV (control point motion vector) as a 4 - parameter model, and Figure 31 (a) illustrates a method of determining mv0, mv1, and mv2 of the control point MV (control point motion vector) as a 6 - parameter model.
[0233] Referring to Figure 31 , for the difference of each control point MV, the difference of the control point MV and the control point MV are determined based on the difference mvd0 of mv0 as the control point MV 0. Figure 31 The mvd0, mvd1, and mvd2 illustrated in can be signaled from the encoder to the decoder. As referencedFigure 29 Compared with the described method, in Figure 31 's method, even if the same mv0, mv1, and mv2 and the same predictors as those in Figure 29 are used, the signaled values of mvd1 and mvd2 may be different. If the differences from the predictors of the control points MVmv0, mv1, and mv2 are similar, then when using Figure 31 's method, there is a possibility that the absolute values of mvd1 and mvd2 are less than those when using Figure 29 's method, and thus the signaling overhead of mvd1 and mvd2 can be reduced. Referring to Figure 31 , the difference from the predictor of mv1 can be determined as (mvd1 + mvd0), and the difference from the predictor of mv2 can be determined as (mvd2 + mvd0).
[0234] Through Figure 31 The terms indicated by the overline in may be the predictors of the control point MV.
[0235] Figure 32 is a diagram illustrating a method for determining the motion vector difference through the method described by referring to Figure 31 .
[0236] As an example, the motion vector difference can be signaled by referring to Figure 10 or Figure 33 's method. In addition, the motion vector difference determined based on the signaled parameters can be Figure 32 's lMvd. In addition, the mvd signaled in Figure 31 , that is, values such as mvd0, mvd1, mvd2 can be Figure 32 's lMvd.
[0237] Figure 32 's MvdLX can be the difference between each control point MV (control point motion vector) and the predictor. That is, it can be (mv - mvp). In this case, as described by referring to Figure 31 , for the control point MV 0mv_0, the signaled motion vector difference can be directly used for the difference MvdLX of the control point MV, and for other control point MVs (mv_1, mv_2), it can be determined and used as the difference MvdLX of the control point MV based on the signaled motion vector difference ( Figure 31 's mvd1 and mvd2) and the signaled motion vector difference for the control point MV 0mv_0 ( Figure 31 's mvd0).
[0238] Figure 32The LX in it can indicate a reference list X (reference picture list X). The compIdx is a component index and can indicate x, y components, etc. The cpIdx can indicate a control point index. The cpIdx can mean Figure 31 0, 1 or 0, 1, 2 illustrated in
[0239] It can be considered in Figure 10 , Figure 12 , Figure 13 etc. the resolution of the motion vector difference among the values illustrated. For example, when the resolution is R, the value of lMvd*R can be used for lMvd in the drawings.
[0240] Figure 33 is a diagram illustrating the motion vector difference syntax according to an embodiment of the present disclosure.
[0241] Refer to Figure 33 , and the motion vector difference can be compiled in a manner similar to the way described in reference Figure 10 . In this case, it can be compiled separately according to the control point index cpIdx.
[0242] Figure 34 is a diagram illustrating a higher-level signaling structure according to an embodiment of the present disclosure.
[0243] According to an embodiment of the present disclosure, there may be one or more higher-level signaling. The higher-level signaling may mean signaling at a higher level. The higher level can be a unit including any unit. For example, the higher level of the current block or the current compilation unit may include CTU, slice, tile, tile group, picture, sequence, etc. The higher-level signaling can affect the lower levels of the corresponding higher level. For example, if the higher level is a sequence, it may affect the CTU, slice, tile, tile group, and picture units that are the lower levels of the sequence. Here, the influence is that the higher-level signaling affects the encoding or decoding of the lower level.
[0244] In addition, the higher-level signaling can include signaling indicating which mode can be used. Refer to Figure 34, the higher-level signaling may include the sps_modeX_enabled_flag. According to an embodiment, it may be determined whether mode modeX can be used based on the sps_modeX_enabled_flag. For example, when the sps_modeX_enabled_flag is a certain value, mode modeX may not be used. In addition, when the sps_modeX_enabled_flag is any other value, mode modeX can also be used. Further, when the sps_modeX_enabled_flag is any other value, it can be determined whether to use mode modeX based on additional signaling. For example, a certain value may be 0, and the other value may be 1. However, the present disclosure is not limited thereto, and a certain value may be 1, and the other value may be 0.
[0245] According to an embodiment of the present disclosure, there may be signaling indicating whether affine motion compensation can be used. For example, this signaling may be higher-level signaling. Refer to Figure 34 , this signaling may be the sps_affine_enabled_flag (affine enable flag). Refer to Figure 2 and Figure 7 , the signaling may mean a signal transmitted from an encoder to a decoder through a bitstream. The decoder may parse the affine enable flag from the bitstream.
[0246] For example, if the sps_affine_enabled_flag (affine enable flag) is 0, the syntax may be restricted such that affine motion compensation is not used. In addition, when the sps_affine_enabled_flag (affine enable flag) is 0, the inter_affine_flag (inter-frame affine flag) and the cu_affine_type_flag (compiled unit affine type flag) may not exist.
[0247] For example, the inter_affine_flag (inter-frame affine flag) may be signaling indicating whether affine MC (affine motion compensation) is used in a block. Refer to Figure 2 and Figure 7 , the signaling may mean a signal transmitted from an encoder to a decoder through a bitstream. The decoder may parse the inter_affine_flag (inter-frame affine flag) from the bitstream.
[0248] In addition, the cu_affine_type_flag (compilation unit affine type flag) may be a signaling indicating which type of affine MC (affine motion compensation) is used in a block. Here, the type may indicate whether it is a 4-parameter affine model or a 6-parameter affine model. In addition, when the sps_affine_enabled_flag (affine enable flag) is 1, affine motion compensation can be used.
[0249] Affine motion compensation may mean motion compensation based on an affine model or motion compensation based on an affine model for inter prediction.
[0250] In addition, according to an embodiment of the present disclosure, there may be a specific type of signaling indicating which mode can be used. For example, there may be a signaling indicating whether a specific type of affine motion compensation can be used. For example, this signaling may be a higher-level signaling. Refer to Figure 34 , this signaling may be the sps_affine_type_flag. In addition, the specific type may mean a 6-parameter affine model. For example, if the sps_affine_type_flag is 0, the syntax may be restricted so that the 6-parameter affine model is not used. In addition, when the sps_affine_type_flag is 0, the cu_affine_type_flag may not exist. In addition, if the sps_affine_type_flag is 1, the 6-parameter affine model can be used. If the sps_affine_type_flag does not exist, it can be inferred that its value is equal to 0.
[0251] In addition, according to an embodiment of the present disclosure, when there is a signaling indicating that a certain mode can be used, there may be a specific type of signaling indicating that a certain mode can be used. For example, when the signaling value indicating whether a certain mode can be used is 1, the specific type of signaling indicating whether a certain mode can be used can be parsed. For example, when the signaling value indicating whether a certain mode can be used is 0, the specific type of signaling indicating whether a certain mode can be used cannot be parsed. For example, the signaling indicating whether a certain mode can be used may include the sps_affine_enabled_flag (affine enable flag). In addition, the specific type of signaling indicating that a certain mode can be used may include the sps_affine_type_flag (affine enable flag). Refer to Figure 34, when the sps_affine_enabled_flag (affine enable flag) is 1, the sps_affine_type_flag can be parsed. Additionally, when the sps_affine_enabled_flag (affine enable flag) is 0, the sps_affine_type_flag may not be parsed and its value can be inferred to be 0.
[0252] Furthermore, the Adaptive Motion Vector Resolution (AMVR) as described above can be used. The resolution set of AMVR can be used differently depending on the situation. For example, the resolution set of AMVR can be used differently according to the prediction mode. For example, the resolution set of AMVR may be different when using conventional inter prediction such as AMVP and when using affine MC (affine motion compensation). Additionally, the AMVR applied to conventional inter prediction such as AMVP can be applied to the motion vector difference. Alternatively, the AMVR applied to conventional inter prediction such as Advanced Motion Vector Prediction (AMVP) can be applied to the motion vector predictor. Moreover, the AMVR applied to affine MC (affine motion compensation) can be applied to the control point motion vector or the control point motion vector difference.
[0253] In addition, according to an embodiment of the present disclosure, there may be signaling indicating whether AMVR can be used. This signaling can be high-level signaling. Refer to Figure 34 , the sps_amvr_enabled_flag (AMVR enable flag) may exist. Refer to Figure 2 and Figure 7 , the signaling can mean a signal transmitted from an encoder to a decoder through a bitstream. The decoder can parse the AMVR enable flag from the bitstream.
[0254] The sps_amvr_enabled_flag (AMVR enable flag) according to an embodiment of the present disclosure may indicate whether adaptive motion vector differential resolution is used. In addition, the sps_amvr_enabled_flag (AMVR enable flag) according to an embodiment of the present disclosure may indicate whether adaptive motion vector differential resolution can be used. For example, according to an embodiment of the present disclosure, when the sps_amvr_enabled_flag (AMVR enable flag) is 1, AMVR can be used for motion vector compilation. In addition, according to an embodiment of the present disclosure, when the sps_amvr_enabled_flag (AMVR enable flag) is 1, AMVR can be used for motion vector compilation. In addition, when the sps_amvr_enabled_flag (AMVR enable flag) is 1, there may be additional signaling for indicating which resolution to use. In addition, when the sps_amvr_enabled_flag (AMVR enable flag) is 0, AMVR may not be used for motion vector compilation. In addition, when the sps_amvr_enabled_flag (AMVR enable flag) is 0, AMVR may not be able to be used for motion vector compilation. According to an embodiment of the present disclosure, the AMVR corresponding to the sps_amvr_enabled_flag (AMVR enable flag) may mean that it is used for regular inter-frame prediction. For example, the AMVR corresponding to the sps_amvr_enabled_flag (AMVR enable flag) may not mean that it is used for affine MC. In addition, whether affine MC (affine motion compensation) is used can be indicated by the inter_affine_flag (inter-frame affine flag). That is, the AMVR corresponding to the sps_amvr_enabled_flag (AMVR enable flag) means that it is used when the inter_affine_flag (inter-frame affine flag) is 0, or may not mean that it is used when the inter_affine_flag (inter-frame affine flag) is 1.
[0255] In addition, according to an embodiment of the present disclosure, there may be signaling indicating whether AMVR can be used for affine MC (affine motion compensation). This signaling may be higher-level signaling. Refer to Figure 34 , there may be a signaling sps_affine_amvr_enabled_flag (affine AMVR enable flag) indicating whether AMVR can be used for affine MC (affine motion compensation). Refer to Figure 2 and Figure 7 , the signaling may mean a signal transmitted from an encoder to a decoder through a bitstream. The decoder can parse the affine AMVR enable flag from the bitstream.
[0256] The sps_affine_amvr_enabled_flag can indicate whether adaptive motion vector differential resolution is used for affine motion compensation. Additionally, the sps_affine_amvr_enabled_flag can indicate whether adaptive motion vector differential resolution is available for affine motion compensation. According to an embodiment, when the sps_affine_amvr_enabled_flag is 1, it can be used for affine inter-frame mode motion vector compilation for AMVR. Additionally, when the sps_affine_amvr_enabled_flag is 1, AMVR can be used for affine inter-frame mode motion vector compilation. Additionally, when the sps_affine_amvr_enabled_flag is 0, it cannot be used for affine inter-frame mode motion vector compilation for AMVR. When the sps_affine_amvr_enabled_flag is 0, AMVR may not be used for affine inter-frame mode motion vector compilation.
[0257] For example, when the sps_affine_amvr_enabled_flag is 1, the AMVR corresponding to the case where the inter_affine_flag is 1 can be used. Additionally, when the sps_affine_amvr_enabled_flag is 1, there may be additional signaling for indicating which resolution to use. Additionally, when the sps_affine_amvr_enabled_flag is 0, the AMVR corresponding to the case where the inter_affine_flag is 1 may not be used for affine inter-frame mode motion vector compilation.
[0258] Figure 35 FIG. is a diagram illustrating the syntax structure of a coding unit according to an embodiment of the present disclosure.
[0259] As referred to Figure 34 As described, there may be additional signaling for indicating the resolution based on higher-level signaling using AMVR. Referring to Figure 34 , the additional signaling for indicating the resolution may include amvr_flag or amvr_precision_flag. The amvr_flag or amvr_precision_flag may be information about the resolution of the motion vector difference.
[0260] According to an embodiment, there may be signaling indicated when amvr_flag is 0. In addition, when amvr_flag is 1, there may be amvr_precision_flag. In addition, when amvr_flag is 1, the resolution may also be determined based on amvr_precision_flag. For example, if amvr_flag is 0, it may be 1 / 4 resolution. In addition, if amvr_flag does not exist, the amvr_flag value may be inferred based on CuPredMode. For example, when CuPredMode is MODE_IBC, the amvr_flag value can be inferred to be equal to 1, and when CuPredMode is not MODE_IBC or CuPredMode (compilation unit prediction mode) is MODE_INTER, the amvr_flag value can be inferred to be equal to 0.
[0261] In addition, when inter_affine_flag is 0 and amvr_precision_flag is 0, 1-pixel resolution can be used. In addition, when inter_affine_flag is 1 and amvr_precision_flag is 0, 1 / 16-pixel resolution can be used. In addition, when inter_affine_flag is 0 and amvr_precision_flag is 1, 4-pixel resolution can be used. In addition, when inter_affine_flag is 1 and amvr_precision_flag is 1, 1-pixel resolution can be used.
[0262] If amvr_precision_flag is 0, its value can be inferred to be equal to 0.
[0263] According to an embodiment, the resolution can be applied through the MvShift value. In addition, MvShift can be determined by amvr_flag and amvr_precision_flag, which are information about the resolution of the motion vector difference. For example, when inter_affine_flag is 0, the MvShift value can be determined as follows.
[0264] MvShift = (amvr_flag + amvr_precision_flag) << 1
[0265] In addition, the motion vector difference Mvd value can be shifted based on the MvShift value. For example, Mvd (motion vector difference) can be shifted as follows, and thus, the resolution of AMVR can be applied.
[0266] MvdLX = MvdLX << (MvShift + 2)
[0267] As another example, when the inter_affine_flag (inter-frame affine flag) is 1, the MvShift value can be determined as follows.
[0268] MvShift = amvr_precision_flag? (amvr_precision_flag << 1) : (-(amvr_flag << 1))
[0269] In addition, the control point motion vector difference MvdCp value can be shifted based on the MvShift value. MvdCP can be the control point motion vector difference or the control point motion vector. For example, the MvdCp (control point motion vector difference) is shifted as follows, and thus, the resolution of AMVR can be applied.
[0270] MvdCpLX = MvdCpLX << (MvShift + 2)
[0271] In addition, Mvd or MvdCp can be a value signaled by mvd_coding.
[0272] Reference Figure 35 , when CuPredMode is MODE_IBC, the amvr_flag as information about the resolution of the motion vector difference may not exist. In addition, when CuPredMode is MODE_INTER, the amvr_flag as information about the resolution of the motion vector difference may exist, and in this case, the amvr_flag can be parsed if a certain condition is met.
[0273] According to an embodiment of the present disclosure, it is possible to determine whether to parse AMVR-related syntax elements based on a higher-level signaling value indicating whether AMVR can be used. For example, when the higher-level signaling value indicating whether AMVR can be used is 1, the AMVR-related syntax elements can be parsed. In addition, when the higher-level signaling value indicating whether AMVR can be used is 0, the AMVR-related syntax elements may not be parsed. Reference Figure 35 , when CuPredMode is MODE_IBC and the sps_amvr_enabled_flag (AMVR enable flag) is 0, the amvr_precision_flag as information about the resolution of the motion vector difference may not be parsed. Additionally, when CuPredMode is MODE_IBC and the sps_amvr_enabled_flag (AMVR enable flag) is 1, the amvr_precision_flag as information about the resolution of the motion vector difference can be parsed. In this case, additional parsing conditions can be considered.
[0274] For example, when there is at least one non - zero value among the MvdLX (Multiple Motion Vector Differences) values, the amvr_precision_flag can be parsed. MvdLX can be the Mvd value of reference list LX. In addition, Mvd (Motion Vector Difference) can be signaled by mvd_coding. LX can include L0 (zero - th reference picture list) and L1 (first reference picture list). In addition, in MvdLX, there can be components corresponding to each of the x - axis and y - axis. For example, the x - axis can correspond to the horizontal axis of the picture, and the y - axis can correspond to the vertical axis of the picture. Refer to Figure 35 , it can be indicated that [0] and [1] in MvdLX[x0][y0][0] and MvdLX[x0][y0] are used for the x - axis and y - axis components respectively. In addition, when CuPredMode is MODE_IBC, only L0 can be used. Refer to Figure 35 , when 1) sps_amvr_enabled_flag is 1 and 2) MvdL0[x0][y0][0] or MvdL0[x0][y0][1] is not 0, the amvr_precision_flag can be parsed. In addition, when 1) sps_amvr_enabled_flag is 0 or 2) both MvdL0[x0][y0][0] and MvdL0[x0][y0][1] are 0, the amvr_precision_flag may not be parsed.
[0275] In addition, there may be cases where CuPredMode is not MODE_IBC. In this case, refer to Figure 35 , when sps_amvr_enabled_flag is 1, inter_affine_flag (inter - frame affine flag) is 0, and there is at least one non - zero value among the MvdLX (Multiple Motion Vector Differences) values, the amvr_flag can be parsed. Here, amvr_flag can be information about the resolution of the motion vector difference. In addition, as already described, sps_amvr_enabled_flag (AMVR enable flag) being 1 can indicate the use of adaptive motion vector differential resolution. In addition, inter_affine_flag (inter - frame affine flag) being 0 can indicate that affine motion compensation is not used for the current block. In addition, as referred to Figure 34 as described, the multiple motion vector differences of the current block can be corrected based on the information about the resolution of the motion vector difference such as amvr_flag. This condition can be referred to as condition A.
[0276] In addition, when the sps_affine_amvr_enabled_flag (Affine AMVR Enable Flag) is 1, the inter_affine_flag (Inter-frame Affine Flag) is 1, and there is at least one non-zero value in the MvdCpLX (Multiple Control Motion Vector Differences) value, the amvr_flag can be parsed. Here, the amvr_flag can be information about the resolution of the motion vector difference. In addition, as already described, the sps_affine_amvr_enabled_flag (Affine AMVR Enable Flag) being 1 can indicate that the adaptive motion vector differential resolution can be used for affine motion compensation. In addition, the inter_affine_flag (Inter-frame Affine Flag) being 1 can indicate that affine motion compensation is used for the current block. In addition, as referred to Figure 34 as described, the multiple control point motion vector differences of the current block can be corrected based on the information about the resolution of the motion vector difference such as amvr_flag. This condition can be referred to as Condition B.
[0277] In addition, if Condition A or Condition B is satisfied, the amvr_flag can be parsed. In addition, if Conditions A and B are not satisfied, the amvr_flag, which is information about the resolution of the motion vector difference, may not be parsed. That is, when 1) the sps_amvr_enabled_flag (AMVR Enable Flag) is 0, the inter_affine_flag (Inter-frame Affine Flag) is 1, or all of the MvdLX (Multiple Motion Vector Differences) are 0, and 2) the sps_affine_amvr_enabled_flag is 0, the inter_affine_flag (Inter-frame Affine Flag) is 0, or all of the MvdCpLX (Multiple Control Point Motion Vector Differences) values are 0, the amvr_flag may not be parsed.
[0278] It is also possible to determine whether to parse the amvr_precision_flag based on the amvr_flag value. For example, when the amvr_flag value is 1, the amvr_precision_flag can be parsed. In addition, when the amvr_flag value is 0, the amvr_precision_flag may not be parsed.
[0279] In addition, MvdCpLX (Multiple Control Motion Vector Differences) may mean to control the differences of point motion vectors. In addition, MvdCpLX (Multiple Control Motion Vector Differences) may be signaled by mvd_coding. LX may include L0 (Zero-th Reference Picture List) and L1 (First Reference Picture List). In addition, in MvdCpLX, there may be components corresponding to point motion vectors 0, 1, 2, etc. For example, point motion vectors 0, 1, 2, etc. may be point motion vectors based on preset positions corresponding to the current block.
[0280] Reference Figure 35 , it may indicate that [0], [1], and [2] in MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][1][], and MvdCpLX[x0][y0][2][] respectively correspond to point motion vectors 0, 1, 2. In addition, the MvdCpLX (Multiple Control Point Motion Vector Differences) value corresponding to point motion vector 0 may be used for other point motion vectors. For example, point motion vector 0 may be used as Figure 47 used in. In addition, in MvdCpLX (Multiple Control Point Motion Vector Differences), there may be components corresponding to the x-axis and y-axis respectively. For example, the x-axis may correspond to the horizontal axis of the picture, and the y-axis may correspond to the vertical axis of the picture. Reference Figure 35 , it may indicate that [0] and [1] in MvdCpLX[x0][y0][][0] and MvdCpLX[x0][y0][][1] are the x-axis and y-axis components respectively.
[0281] Figure 36 FIG. is a diagram illustrating a high-level signaling structure according to an embodiment of the present disclosure.
[0282] As described in reference Figures 34 to 35 , the high-level signaling may exist. For example, there may be sps_affine_enabled_flag (Affine Enable Flag), sps_affine_amvr_enabled_flag, sps_amvr_enabled_flag (AMVR Enable Flag), sps_affine_type_flag, etc.
[0283] According to an embodiment of the present disclosure, the above high-level signaling may have parsing dependencies. For example, it may be determined whether to parse other high-level signaling based on which high-level signaling value.
[0284] According to an embodiment of the present disclosure, it is possible to determine whether affine AMVR can be used based on whether affine MC (affine motion compensation) can be used. For example, it is possible to determine whether affine AMVR can be used based on a higher-level signaling indicating whether affine MC can be used. More specifically, it is possible to determine whether to parse a higher-level signaling indicating whether affine AMVR can be used based on a higher-level signaling indicating whether affine MC can be used.
[0285] In one embodiment, when affine MC can be used, affine AMVR can be used. In addition, when affine MC cannot be used, affine AMVR may not be available.
[0286] More specifically, if the higher-level signaling indicating whether affine MC can be used is 1, affine AMVR may be available. In this case, there may be additional signaling. In addition, when the higher-level signaling indicating whether affine MC can be used is 0, affine AMVR may not be available. For example, when the higher-level signaling indicating whether affine MC can be used is 1, the higher-level signaling indicating whether affine AMVR can be used can be parsed. In addition, when the higher-level signaling indicating whether affine MC can be used is 0, the higher-level signaling indicating whether affine AMVR can be used may not be parsed. In addition, when the higher-level signaling indicating whether affine AMVR can be used does not exist, its value can be inferred. For example, it can be inferred that the value is equal to 0. As another example, it can be inferred based on the higher-level signaling indicating whether affine MC can be used. As another example, it can be inferred based on the higher-level signaling indicating whether AMVR can be used.
[0287] According to an embodiment, affine AMVR can be an AMVR for the Figures 34 to 35 affine MC described above. For example, the higher-level signaling indicating whether affine MC can be used can be sps_affine_enabled_flag (affine enable flag). In addition, the higher-level signaling indicating whether affine AMVR can be used can be sps_affine_amvr_enabled_flag.
[0288] Referring to Figure 36 , when sps_affine_enabled_flag (affine enable flag) is 1, sps_affine_amvr_enabled_flag can be parsed. In addition, when sps_affine_enabled_flag (affine enable flag) is 0, sps_affine_amvr_enabled_flag may not be parsed. In addition, when sps_affine_amvr_enabled_flag does not exist, it can be inferred that its value is equal to 0.
[0289] This may be because, in embodiments of the present disclosure, affine AMVR may be meaningful when using affine MC.
[0290] Figure 37 FIG. is a diagram illustrating a high-level signaling structure according to an embodiment of the present disclosure.
[0291] As described in reference Figures 34 to 35 the high-level signaling may exist. For example, sps_affine_enabled_flag (affine enable flag), sps_affine_amvr_enabled_flag (affine AMVR enable flag), sps_amvr_enabled_flag (AMVR enable flag), sps_affine_type_flag, etc. may exist.
[0292] According to an embodiment of the present disclosure, the high-level signaling may have parsing dependencies. For example, whether to parse other high-level signaling may be determined based on a certain high-level signaling value.
[0293] According to an embodiment of the present disclosure, it is possible to determine whether affine AMVR can be used based on whether AMVR can be used. For example, it is possible to determine whether affine AMVR can be used based on high-level signaling indicating whether AMVR can be used. More specifically, it is possible to determine whether to parse high-level signaling indicating whether affine AMVR can be used based on high-level signaling indicating whether AMVR can be used.
[0294] In one embodiment, when AMVR can be used, affine AMVR can be used. In addition, when AMVR cannot be used, affine AMVR may not be available.
[0295] More specifically, when the high-level signaling indicating whether AMVR can be used is 1, affine AMVR can be used. In this case, there may be additional signaling. In addition, when the high-level signaling indicating whether AMVR can be used is 0, affine AMVR may not be available. For example, when the high-level signaling indicating whether AMVR can be used is 1, the high-level signaling indicating whether affine AMVR can be used can be parsed. In addition, when the high-level signaling indicating whether AMVR can be used is 0, the high-level signaling indicating whether affine AMVR can be used may not be parsed. In addition, if there is no high-level signaling indicating whether affine AMVR can be used, its value can be inferred. For example, it can be inferred that the value is equal to 0. As another example, it can be inferred based on high-level signaling indicating whether affine MC can be used. As another example, it can be inferred based on high-level signaling indicating whether AMVR can be used.
[0296] According to an embodiment, the affine AMVR may be an AMVR for the affine MC (affine motion compensation) described for reference. Figures 34 to 35 For example, the high-level signaling indicating whether the AMVR can be used may be the sps_amvr_enabled_flag (AMVR enable flag). In addition, the higher-level signaling indicating whether the affine AMVR can be used may be the sps_affine_amvr_enabled_flag (affine AMVR enable flag).
[0297] For reference Figure 37 (a), when the sps_amvr_enabled_flag (AMVR enable flag) is 1, the sps_affine_amvr_enabled_flag (affine AMVR enable flag) can be parsed. Additionally, when the sps_amvr_enabled_flag (AMVR enable flag) is 0, the sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be parsed. Furthermore, when the sps_affine_amvr_enabled_flag (affine AMVR enable flag) does not exist, its value can be inferred to be equal to 0.
[0298] This may be because, in the embodiments of the present disclosure, whether the adaptive resolution is effective may vary according to the sequence.
[0299] In addition, it can be determined whether the affine AMVR can be used by considering both whether the affine MC can be used and whether the AMVR can be used. For example, it can be determined whether to parse the higher-level signaling indicating whether the affine AMVR can be used based on the higher-level signaling indicating whether the affine MC can be used and the higher-level signaling indicating whether the AMVR can be used. According to an embodiment, when both the higher-level signaling indicating whether the affine MC can be used and the higher-level signaling indicating whether the AMVR can be used are 1, the higher-level signaling indicating whether the affine AMVR can be used can be parsed. Additionally, when the higher-level signaling indicating whether the affine MC can be used or the higher-level signaling indicating whether the AMVR can be used is 0, the higher-level signaling indicating whether the affine AMVR can be used may not be parsed. Furthermore, if the higher-level signaling indicating whether the affine AMVR can be used does not exist, its value can be inferred.
[0300] For reference Figure 37(b), when both sps_affine_enabled_flag (affine enable flag) and sps_amvr_enabled_flag (AMVR enable flag) are 1, sps_affine_amvr_enabled_flag (affine AMVR enable flag) can be parsed. Additionally, when at least one of sps_affine_enabled_flag (affine enable flag) and sps_amvr_enabled_flag (AMVR enable flag) is 0, sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be parsed. Moreover, when sps_affine_amvr_enabled_flag (affine AMVR enable flag) does not exist, it can be inferred that the value is equal to 0.
[0301] More specifically, referring to Figure 37 (b), it can be determined at line 3701 whether affine motion compensation can be used based on sps_affine_enabled_flag (affine enable flag). As described in the reference Figures 34 to 35 already, when sps_affine_enabled_flag (affine enable flag) is 1, it may mean that affine motion compensation can be used. Additionally, when sps_affine_enabled_flag (affine enable flag) is 0, it may mean that affine motion compensation cannot be used.
[0302] When it is determined at Figure 37 (b) line 3701 that affine motion compensation is used, it can be determined at line 3702 whether to use adaptive motion vector differential resolution based on sps_amvr_enabled_flag (AMVR enable flag). As described in the reference Figures 34 to 35 already, when sps_amvr_enabled_flag (AMVR enable flag) is 1, it may mean that adaptive motion vector differential resolution is used. When sps_amvr_enabled_flag (AMVR enable flag) is 0, it may mean that adaptive motion vector differential resolution is not used.
[0303] When it is determined at line 3701 that affine motion compensation is not used, it may not be determined whether to use adaptive motion vector differential resolution based on sps_amvr_enabled_flag (AMVR enable flag). That is, Figure 37(b)'s line 3702 may not be executed. Specifically, when affine motion compensation is not used, the sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be transmitted from the encoder to the decoder. That is, the decoder may not receive the sps_affine_amvr_enabled_flag (affine AMVR enable flag), and the sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be parsed by the decoder. In this case, since the sps_affine_amvr_enabled_flag (affine AMVR enable flag) does not exist, it can be inferred to be equal to 0. As already described, when the sps_affine_amvr_enabled_flag (affine AMVR enable flag) is 0, it may indicate that the adaptive motion vector differential resolution cannot be used for affine motion compensation.
[0304] When at Figure 37 it is determined at line 3702 of (b) to use the adaptive motion vector differential resolution, the sps_affine_amvr_enabled_flag indicating whether the adaptive motion vector differential resolution can be used for affine motion compensation can be parsed from the bitstream at line 3703.
[0305] When it is determined at line 3702 not to use the adaptive motion vector differential resolution, the sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be parsed from the bitstream. That is, Figure 37 (b)'s line 3703 may not be executed. More specifically, when affine motion compensation is used and the adaptive motion vector differential resolution is not used, the sps_affine_amvr_enabled_flag (affine AMVR enable flag) may not be transmitted from the encoder to the decoder. That is to say, the decoder may not receive the sps_affine_amvr_enabled_flag (affine AMVR enable flag). The decoder may not parse the sps_affine_amvr_enabled_flag (affine AMVR enable flag) from the bitstream. In this case, since the sps_affine_amvr_enabled_flag (affine AMVR enable flag) does not exist, it can be inferred to be equal to 0. As already described, when the sps_affine_amvr_enabled_flag (affine AMVR enable flag) is 0, it may indicate that the adaptive motion vector differential resolution cannot be used for affine motion compensation.
[0306] By first checking the sps_affine_enabled_flag (affine enable flag) and then checking the sps_amvr_enabled_flag (AMVR enable flag), as Figure 37 illustrated in (b), unnecessary processes can be reduced and efficiency can be increased. For example, when first checking the sps_amvr_enabled_flag (AMVR enable flag) and then checking the sps_affine_enabled_flag (affine enable flag), it may be necessary to check the sps_affine_enabled_flag (affine enable flag) again in order to derive the sps_affine_type_flag at line 7. However, by first checking the sps_affine_enabled_flag (affine enable flag) and then checking the sps_amvr_enabled_flag (AMVR enable flag), this unnecessary process can be reduced.
[0307] Figure 38 is a diagram illustrating the syntax structure of a compilation unit according to an embodiment of the present disclosure.
[0308] As referenced Figure 35 described, it is possible to determine whether to parse AMVR-related syntax based on whether there is at least one non-zero value among MvdLX or MvdCpLX. However, depending on which reference list (reference picture list) is used, how many parameters are used for the affine model, etc., MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) may vary. If the initial value of MvdLX or MvdCpLX is not 0, since the unused MvdLX or MvdCpLX in the current block is not 0, unnecessary AMVR-related syntax elements are signaled and a mismatch between the encoder and the decoder can occur. Figures 38 to 39 A method for not generating a mismatch between the encoder and the decoder can be described.
[0309] According to an embodiment, inter_pred_idc (information regarding the reference picture list) may indicate which reference list to use or what the prediction direction is. For example, inter_pred_idc (information regarding the reference picture list) may be a value of PRED_L0, PRED_L1, or PRED_BI. If inter_pred_idc (information regarding the reference picture list) is PRED_L0, only reference list 0 (the zero-th reference picture list) may be used. Further, when inter_pred_idc (information regarding the reference picture list) is PRED_L1, only reference list 1 (the first reference picture list) may be used. Further, when inter_pred_idc (information regarding the reference picture list) is PRED_BI, both reference list 0 (the zero-th reference picture list) and reference list 1 (the first reference picture list) may be used. When inter_pred_idc (information regarding the reference picture list) is PRED_L0 or PRED_L1, it may be unidirectional prediction. Further, when inter_pred_idc (information regarding the reference picture list) is PRED_BI, it may be bidirectional prediction.
[0310] It is also possible to determine the affine model to be used based on the MotionModelIdc value. It is also possible to determine whether to use affine MC based on the value of MotionModelIdc. For example, MotionModelIdc may indicate translational motion, 4-parameter affine motion, or 6-parameter affine motion. For example, when the MotionModelIdc value is 0, 1, and 2, they may indicate translational motion, 4-parameter affine motion, and 6-parameter affine motion, respectively. Further, according to an embodiment, MotionModelIdc may be determined based on inter_affine_flag (inter-frame affine flag) and cu_affine_type_flag. For example, when merge_flag is 0 (non-merge mode), MotionModelIdc may be determined based on inter_affine_flag (inter-frame affine flag) and cu_affine_type_flag. For example, MotionModelIdx may be (inter_affine_flag + cu_affine_type_flag). According to another embodiment, MotionModelIdc may be determined by merge_subblock_flag. For example, when merge_flag is 1 (merge mode), MotionModelIdc may be determined by merge_subblock_flag. For example, the MotionModelIdc value may be set to the merge_subblock_flag value.
[0311] For example, when inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_BI, the values corresponding to L0 in MvdLX (motion vector difference) or MvdCpLX (control point motion vector difference) can be used. Therefore, when parsing the AMVR-related syntax, only when inter_pred_idc (information about the reference picture list) is PRED_L0 or PRED_BI can MvdL0 or MvdCpL0 be considered. That is, when inter_pred_idc (information about the reference picture list) is PRED_L1, MvdL0 (motion vector difference of the zero-th reference picture list) or MvdCpL0 (control point motion vector difference of the zero-th reference picture list) can be not considered.
[0312] In addition, when inter_pred_idc (information about the reference picture list) is PRED_L1 or PRED_BI, the values corresponding to L1 in MvdLX or MvdCpLX can be used. Therefore, when parsing the AMVR-related syntax, only when inter_pred_idc (information about the reference picture list) is PRED_L1 or PRED_BI can MvdL1 (motion vector difference of the first reference picture list) or MvdCpL1 (control point motion vector difference of the first reference picture list) be considered. That is, when inter_pred_idc (information about the reference picture list) is PRED_L0, MvdL1 or MvdCpL1 can be not considered.
[0313] Reference Figure 38 , for MvdL0 and MvdCpL0, only when inter_pred_idc is not PRED_L1 can it be determined whether to parse the AMVR-related syntax according to whether their values are non-zero. That is, when inter_pred_idc is PRED_L1, even if there are non-zero values in MvdL0 or non-zero values in MvdCpL0, the AMVR-related syntax may not be parsed.
[0314] In addition, when MotionModelIdc is 1, it is possible to consider only MvdCpLX[x0][y0][0][] and MvdCpLX[x0][y0][1][] among MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][1][], and MvdCpLX[x0][y0][2][]. That is, when MotionModelIdc is 1, MvdCpLX[x0][y0][2][] can be not considered. For example, when MotionModelIdc is 1, whether there is a non-zero value in MvdCpLX[x0][y0][2][] may not affect whether to parse the AMVR-related syntax.
[0315] In addition, when MotionModelIdc is 2, it is possible to consider all of MvdCpLX[x0][y0][0][], MvdCpLX[x0][y0][1][], and MvdCpLX[x0][y0][2][]. That is, when MotionModelIdc is 2, MvdCpLX[x0][y0][2][] can be considered.
[0316] In addition, in the above embodiments, MotionModelIdc 1 or 2 can be expressed as cu_affine_type_flag being 0 or 1. This may be because it can determine whether to use affine MC. For example, it is possible to determine whether to use affine MC through inter_affine_flag (inter-frame affine flag).
[0317] Reference Figure 38 , it is only when MotionModelIdc is 2 that it is possible to consider whether there is a non-zero value in MvdCpLX[x0][y0][2][]. When MvdCpLX[x0][y0][0][] and MvdCpLX[x0][y0][1][] are both 0 and there is a non-zero value in MvdCpLX[x0][y0][2][] (according to the previous embodiments, here by dividing L0 and L1, it is possible to consider only one of L0 and L1), if MotionModelIdc is not 2, then it is possible not to parse the AMVR-related syntax.
[0318] Figure 39 is a diagram showing the MVD default value setting according to an embodiment of the present disclosure.
[0319] As described above, MvdLX (Motion Vector Difference) or MvdCpLX (Control Point Motion Vector Difference) can be signaled by mvd_coding. In addition, the lMvd value can be signaled by mvd_coding, and MvdLX or MvdCpLX can be set to the lMvd value. Refer to Figure 39 , when MotionModelIdc is 0, MvdLX can be set by the lMvd value. In addition, when MotionModelIdc is not 0, MvdCpLX can be set by the lMvd value. In addition, depending on the refList value, an operation corresponding to any one of the LXs to be performed can be determined.
[0320] In addition, there may be mvd_coding as described in reference Figure 10 or Figure 33 , Figure 34 and Figure 35 . In addition, mvd_coding may include steps of parsing or determining abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, mvd_sign_flag, etc. In addition, lMvd can be determined by abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, mvd_sign_flag, etc.
[0321] Refer to Figure 39 , lMvd can be set as follows.
[0322] lMvd = abs_mvd_greater0_flag * (abs_mvd_minus2 + 2) * (1 - 2 * mvd_sign_flag)
[0323] The default value of MvdLX (Motion Vector Difference) or MvdCpLX (Control Point Motion Vector Difference) can be set to a preset value. According to an embodiment of the present disclosure, the default value of MvdLX or MvdCpLX can be set to 0. Alternatively, the default value of lMvd can be set to a preset value. Alternatively, the default values of relevant syntax elements can be set such that the values of lMvd, MvdLX or MvdCpLX become preset values. The default value of a syntax element can mean the value to be inferred when the syntax element does not exist. The preset value can be 0.
[0324] According to an embodiment of the present disclosure, abs_mvd_greater0_flag may indicate whether the absolute value of MVD is greater than 0. In addition, according to an embodiment of the present disclosure, when abs_mvd_greater0_flag does not exist, it can be inferred that its value is equal to 0. In this case, the lMvd value may be set to 0. In addition, in this case, the MvdLX or MvdCpLX value may be set to 0.
[0325] Alternatively, according to an embodiment of the present disclosure, when the lMvd, MvdLX, or MvdCpLX value is not set, its value may be set to a preset value. For example, the value may be set to 0.
[0326] In addition, according to an embodiment of the present disclosure, when abs_mvd_greater0_flag does not exist, the corresponding lMvd, MvdLX, or MvdCpLX value may be set to 0.
[0327] In addition, abs_mvd_greater1_flag may indicate whether the absolute value of MVD is greater than 1. In addition, when abs_mvd_greater1_flag does not exist, it can be inferred that its value is equal to 0.
[0328] In addition, (abs_mvd_minus2 + 2) may indicate the absolute value of MVD. Additionally, when the value of abs_mvd_minus2 does not exist, it can be inferred to be equal to -1.
[0329] In addition, mvd_sign_flag may indicate the sign of MVD. When mvd_sign_flag is 0 and 1, it may indicate that the corresponding MVD has a positive value and a negative value, respectively. If mvd_sign_flag does not exist, it can be inferred that the value is equal to 0.
[0330] Figure 40 It is a diagram illustrating the MVD default value setting according to an embodiment of the present disclosure.
[0331] By setting the initial value of MvdLX or MvdCpLX to 0, the MvdLX (Multiple Motion Vector Differences) or MvdCpLX (Multiple Control Point Motion Vector Differences) value can be initialized, as described in Figure 39 so as to avoid a mismatch between the encoder and the decoder. Moreover, in this case, the initialization value may be 0. Figure 40 The embodiments of
[0332] According to an embodiment of the present disclosure, MvdLX (Multiple Motion Vector Differences) or MvdCpLX (Multiple Control Point Motion Vector Differences) values may be initialized to a preset value. In addition, the initialization position may be before the position of parsing AMVR-related syntax elements. AMVR-related syntax elements may include Figure 40 's amvr_flag, amvr_precision_flag, etc. The amvr_flag or amvr_precision_flag may be information about the resolution of the motion vector difference.
[0333] In addition, the resolution of MVD (Motion Vector Difference) or MV (Motion Vector) or the signaled resolution of MVD or MV may be determined by AMVR-related syntax elements. Additionally, the preset value for initialization in an embodiment of the present disclosure may be 0.
[0334] Since the MvdLX and MvdCpLX values between the encoder and the decoder may be the same when parsing AMVR-related syntax elements by performing the initialization, a mismatch between the encoder and the decoder may not occur. In addition, by performing the initialization set to the value 0, AMVR-related syntax elements may not be unnecessarily included in the bitstream.
[0335] According to an embodiment of the present disclosure, MvdLX may be defined for reference lists (L0, L1, etc.), x- or y-components, etc. In addition, MvdCpLX may be defined for reference lists (L0, L1, etc.), x or y components, control points 0, 1, or 2, etc.
[0336] According to an embodiment of the present disclosure, both MvdLX and MvdCpLX values can be initialized. In addition, the initialization position may be before the position of performing mvd_coding of the corresponding MvdLX or MvdCpLX. For example, even in a prediction block that only uses L0, it is possible to initialize MvdLX or MvdCpLX corresponding to L0. Conditional checks may be required to initialize only those that are necessary for initialization, but by performing the initialization indiscriminately in this way, the burden of conditional checks can be reduced.
[0337] According to another embodiment of the present disclosure, it is possible to initialize values corresponding to MvdLX or MvdCpLX values that are not used. Here, the values that are not used may mean that the values are not used in the current block. For example, it is possible to initialize the values corresponding to MvdLX or MvdCpLX that correspond to the reference list that is not currently used. For example, when L0 is not used, the values corresponding to MvdL0 or MvdCpL0 can be initialized. When inter_pred_idc is PRED_L1, L0 may not be used. In addition, when L1 is not used, the values corresponding to MvdL1 or MvdCpL1 can be initialized. When inter_pred_idc is PRED_L0, L1 may not be used. L0 can be used when inter_pred_idc is PRED_L0 or PRED_BI, and L1 can be used when inter_pred_idc is PRED_L1 or PRED_BI. Refer to Figure 40 , when inter_pred_idc is PRED_L1, MvdL0[x0][y0][0], MvdL0[x0][y0][1], MvdCpL0[x0][y0][0][0], MvdCpL0[x0][y0][0][1], MvdCpL0[x0][y0][1][0], MvdCpL0[x0][y0][1][1], MvdCpL0[x0][y0][2][0], and MvdCpL0[x0][y0][2][1] can be initialized. In addition, the initialization value can be 0. In addition, when inter_pred_idc is PRED_L0, MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1] can be initialized. In addition, the initialization value may be 0.
[0338] MvdLX[x][y][compIdx] can be the motion vector difference at the (x, y) position of the reference list LX and the component index compIdx. MvdCpLX[x][y][cpIdx][compIdx] can be the motion vector difference of the reference list LX. In addition, MvdCpLX[x][y][cpIdx][compIdx] can be the motion vector difference at the position (x, y), the control point motion vector index cpIdx, and the component index compIdx. Here, the component can indicate the x or y component.
[0339] In addition, according to an embodiment of the present disclosure, which of MvdLX (Motion Vector Difference) or MvdCpLX (Control Point Motion Vector Difference) is not used may depend on whether affine motion compensation is used. For example, when affine motion compensation is used, MvdLX may be initialized. In addition, when affine motion compensation is not used, MvdCpLX may be initialized. For example, there may be signaling indicating whether affine motion compensation is used. Refer to Figure 40 , inter_affine_flag (inter-frame affine flag) may be signaling indicating whether affine motion compensation is used. For example, when inter_affine_flag (inter-frame affine flag) is 1, affine motion compensation may be used. Alternatively, MotionModelIdc may be signaling indicating whether affine motion compensation is used. For example, when MotionModelIdc is not 0, affine motion compensation may be used.
[0340] In addition, according to an embodiment of the present disclosure, which of MvdLX or MvdCpLX is not used may be related to which affine motion model is used. For example, depending on whether a 4-parameter affine model or a 6-parameter affine model is used, the unused MvdLX or MvdCpLX may be different. For example, depending on which affine motion model is used, the MvdCpLX for cpIdx not used for MvdCpLX[x][y][cpIdx][compIdx] may be different. For example, when a 4-parameter affine model is used, only a part of MvdCpLX[x][y][cpIdx][compIdx] can be used. Alternatively, when a 6-parameter affine model is not used, only a part of MvdCpLX[x][y][cpIdx][compIdx] can be used. Therefore, the unused MvdCpLX can be initialized to a preset value. In this case, the unused MvdCpLX may correspond to the cpIdx used in the 6-parameter affine model and not used in the 4-parameter affine model. For example, when a 4-parameter affine model is used, the value of MvdCpLX[x][y][cpIdx][compIdx] where cpIdx is 2 may not be used and can be initialized to a preset value. In addition, as described above, there may be signaling or a parameter indicating whether a 4-parameter affine model or a 6-parameter affine model is used. For example, it can be known whether a 4-parameter affine model or a 6-parameter affine model is used through MotionModelIdc or cu_affine_type_flag. MotionModelIdc values of 1 and 2 may indicate using a 4-parameter affine model and a 6-parameter affine model, respectively. Refer to Figure, when MotionModelIdc is 1, MvdCpL0[x0][y0][2][0], MvdCpL0[x0][y0][2][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1] can be initialized to preset values. Additionally, in this case, the condition when MotionModelIdc is not 2 can be used instead of the condition when MotionModelIdc is 1. The preset value can be 0.
[0341] Furthermore, according to an embodiment of the present disclosure, which of MvdLX or MvdCpLX is not used can be based on the value of mvd_l1_zero_flag (motion vector difference zero flag of the first reference picture list). For example, when mvd_l1_zero_flag is 1, MvdL1 and MvdCpL1 can be initialized to preset values. Additionally, in this case, additional conditions can be considered. For example, which of MvdLX or MvdCpLX is not used can be determined based on mvd_l1_zero_flag and inter_pred_idc (information regarding the reference picture list). For example, when mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, MvdL1 and MvdCpL1 can be initialized to preset values. For example, mvd_l1_zero_flag (motion vector difference zero flag of the first reference picture list) can be a higher-level signaling, which can indicate that the MVD value (e.g., MvdLX or MvdCpLX) of reference list L1 is 0. Signaling can mean a signal transmitted from the encoder to the decoder through the bitstream. The decoder can parse mvd_l1_zero_flag (motion vector difference zero flag of the first reference picture list) from the bitstream.
[0342] According to another embodiment of the present disclosure, all Mvd and MvdCp can be initialized to preset values before performing mvd_coding on a certain block. In this case, mvd_coding may mean all mvd_coding of a certain CU. Therefore, it is possible to prevent parsing the mvd_coding syntax and initializing the determined Mvd or MvdCp values, and the above problems can be solved by initializing all Mvd and MvdCp values.
[0343] is a diagram illustrating an AMVR-related syntax structure according to an embodiment of the present disclosure.
[0344] The embodiment of can be based on the embodiment of.
[0345] As described in reference it is possible to check whether there is at least one non-zero value in MvdLX (Motion Vector Difference) and MvdCpLX (Motion Vector Difference of Control Points). In this case, it is possible to determine MvdLX or MvdCpLX to be checked based on mvd_l1_zero_flag (Motion Vector Difference Zero Flag for the First Reference Picture List). Therefore, it is possible to determine whether to parse the AMVR-related syntax based on mvd_l1_zero_flag. As described above, when mvd_l1_zero_flag is 1, it can indicate that the MVD (Motion Vector Difference) of reference list L1 is 0, and thus, in this case, MvdL1 (Motion Vector Difference of the First Reference Picture List) or MvdCpL1 (Motion Vector Difference of Control Points of the First Reference Picture List) may not be considered. For example, when mvd_l1_zero_flag is 1, regardless of whether MvdL1 or MvdCpL1 is 0, it is possible to parse the AMVR-related syntax list based on whether there is at least one value equal to 0 in MvdL0 (Motion Vector Difference of the Zero-th Reference Picture List) or MvdCpL0 (Motion Vector Difference of Control Points of the Zero-th Reference Picture List). According to an additional embodiment, mvd_l1_zero_flag indicating that the MVD for reference list L1 is 0 can be used only for blocks that are bi-directionally predicted. Therefore, it is possible to parse the AMVR-related syntax based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list). For example, it is possible to determine MvdLX or MvdCpLX for determining whether there is at least one non-zero value based on mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list). For example, MvdL1 or MvdCpL1 may not be considered based on mvd_l1_zero_flag and inter_pred_idc. For example, when mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, MvdL1 or MvdCpL1 may not be considered. For example, when mvd_l1_zero_flag is 1 and inter_pred_idc is PRED_BI, regardless of whether MvdL1 or MvdCpL1 is 0, it is possible to parse the AMVR-related syntax based on whether there is at least one 0 value in MvdL0 or MvdCpL0.
[0346] Reference , when mvd_l1_zero_flag is 1 and inter_pred_idc[x0][y0] = PRED_BI, operations can be performed regardless of whether there are non-zero values among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. For example, when mvd_l1_zero_flag is 1 and inter_pred_idc[x0][y0] = PRED_BI, if there are no non-zero values in MvdL0 or MvdCpL0, even if there are non-zero values among MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1], the AMVR-related syntax may not be parsed.
[0347] In addition, when mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0] != PRED_BI, whether there are non-zero values in MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1] can be considered. For example, when mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0] != PRED_BI, if there are non-zero values in MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], MvdCpL1[x0][y0][2][1], then the AMVR-related syntax can be parsed. For example, when mvd_l1_zero_flag is 0 or inter_pred_idc[x0][y0] != PRED_BI, if there are non-zero values in MvdL1[x0][y0][0], MvdL1[x0][y0][1], MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1], even if both MvdL0 and MvdCpL0 are 0, the AMVR-related syntax can be parsed.
[0348] FIG. is a diagram illustrating an inter prediction-related syntax structure according to an embodiment of the present disclosure.
[0349] According to an embodiment of the present disclosure, the mvd_l1_zero_flag (motion vector difference zero flag of the first reference picture list) may be a signaling indicating that the Mvd (motion vector difference) value of the reference list L1 (first reference picture list) is 0. In addition, this signaling may be signaled at a level higher than the current block. Therefore, based on the mvd_l1_zero_flag value, the Mvd values of the reference list L1 in multiple blocks may be 0. For example, when the mvd_l1_zero_flag value is 1, the Mvd value of the reference list L1 may be 0. Alternatively, based on the mvd_l1_zero_flag and inter_pred_idc (information about the reference picture list), the Mvd value of the reference list L1 may be 0. For example, when the mvd_l1_zero_flag is 1 and the inter_pred_idc is PRED_BI, the Mvd value of the reference list L1 may be 0. In this case, the Mvd value may be MvdL1[x][y][compIdx]. In addition, in this case, the Mvd value may not mean the control point motion vector difference. That is, in this case, the Mvd value may not mean the MvdCp value.
[0350] According to another embodiment of the present disclosure, the mvd_l1_zero_flag may be a signaling indicating that the Mvd and MvdCp values of the reference list L1 are 0. In addition, this signaling may be signaled at a level higher than the current block. Therefore, the Mvd and MvdCp values of the reference list L1 in multiple blocks may be 0. For example, when the mvd_l1_zero_flag value is 1, based on the mvd_l1_zero_flag value, the Mvd and MvdCp values of the reference list L1 may be 0. Alternatively, based on the mvd_l1_zero_flag and inter_pred_idc, the Mvd and MvdCp values of the reference list L1 may be 0. For example, when the mvd_l1_zero_flag is 1 and the inter_pred_idc is PRED_BI, the Mvd and MvdCp values of the reference list L1 may be 0. In this case, the Mvd value may be MvdL1[x][y][compIdx]. In addition, the MvdCp value may be MvdCpL1[x][y][cpIdx][compIdx].
[0351] Alternatively, a value of 0 for Mvd or MvdCp may mean that the corresponding mvd_coding syntax structure is not parsed. That is, for example, when the value of mvd_l1_zero_flag is 1, the mvd_coding syntax structure corresponding to MvdL1 or MvdCpL1 may not be parsed. Additionally, when the value of mvd_l1_zero_flag is 0, the mvd_coding syntax structure corresponding to MvdL1 or MvdCpL1 may be parsed.
[0352] According to an embodiment of the present disclosure, when the Mvd or MvdCp value based on mvd_l1_zero_flag is 0, the signaling indicating the MVP may not be parsed. The signaling indicating the MVP may include mvp_l1_flag. Additionally, according to the above description of mvd_l1_zero_flag, the signaling of mvd_l1_zero_flag may mean indicating that it is 0 for both Mvd and MvdCp. For example, when the condition that the Mvd or MvdCp value indicating reference list L1 is 0 is satisfied, the signaling indicating the MVP may not be parsed. In this case, the signaling indicating the MVP may be inferred as a preset value. For example, when the signaling indicating the MVP does not exist, its value may be inferred to be equal to 0. Additionally, when the condition indicating that the Mvd or MvdCp value is 0 based on mvd_l1_zero_flag is not satisfied, the signaling indicating the MVP may be parsed. However, in this embodiment, when the Mvd or MvdCp value is 0, the freedom to select the MVP may be lost, and thus the coding efficiency may be reduced.
[0353] More specifically, when the condition that the Mvd or MvdCp value indicating reference list L1 is 0 is satisfied and affine MC is used, the signaling indicating the MVP may not be parsed. In this case, the signaling indicating the MVP may be inferred as a preset value.
[0354] Refer to , when mvd_l1_zero_flag is 1 and the inter_pred_idc value is PRED_BI, mvp_l1_flag may not be parsed. Additionally, in this case, the value of mvp_l1_flag may be inferred to be equal to 0. Alternatively, when mvd_l1_zero_flag is 0 or the inter_pred_idc value is not PRED_BI, mvp_l1_flag may be parsed.
[0355] In this embodiment, it is possible to determine whether to parse the signaling indicating the MVP based on mvd_l1_zero_flag when specific conditions are met. For example, the specific conditions may include the condition where general_merge_flag is 0. For example, general_merge_flag may have the same meaning as merge_flag described above. In addition, the specific conditions may include conditions based on CuPredMode. More specifically, the specific conditions may include the condition where CuPredMode is not MODE_IBC. Alternatively, the specific conditions may include the condition where CuPredMode is MODE_INTER. When CuPredMode is MODE_IBC, prediction using the current picture as a reference can be used. In addition, when CuPredMode is MODE_IBC, there may be a block vector or a motion vector corresponding to the block. When CuPredMode is MODE_INTER, prediction using a picture other than the current picture as a reference can be used. When CuPredMode is MODE_INTER, there may be a motion vector corresponding to the block.
[0356] Therefore, according to an embodiment of the present disclosure, when general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag is 1, and inter_pred_idc is PRED_BI, mvp_l1_flag may not be parsed. In addition, when mvp_l1_flag does not exist, its value can be inferred to be equal to 0.
[0357] More specifically, when general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag is 1, inter_pred_idc is PRED_BI, and affine MC is used, mvp_l1_flag may not be parsed. In addition, when mvp_l1_flag does not exist, its value can be inferred to be equal to 0.
[0358] Reference , the sym_mvd_flag may be a signaling indicating a symmetric MVD. In the case of symmetric MVD, one MVD can be determined based on a certain MVD. In the case of symmetric MVD, the other MVD can be determined based on the explicitly signaled MVD. For example, in the case of symmetric MVD, the MVD of one reference list can be used to determine the MVD of another reference list. For example, in the case of symmetric MVD, the MVD of reference list L0 can be used to determine the MVD of reference list L1. When determining one MVD based on a certain MVD, the value obtained by inverting the sign of a certain MVD can be determined as the other MVD.
[0359] is a diagram of an inter-frame prediction related syntax structure according to an embodiment of the present disclosure.
[0360] An embodiment of can be used to solve the reference described problem.
[0361] According to an embodiment of the present disclosure, when the Mvd (Motion Vector Difference) or MvdCp (Control Point Motion Vector Difference) value based on mvd_l1_zero_flag (Motion Vector Difference Zero Flag for the First Reference Picture List) is 0, the signaling indicating the MVP (Motion Vector Predictor) can be parsed. The signaling indicating the MVP may include mvp_l1_flag (Motion Vector Predictor Index for the First Reference Picture List). In addition, according to the above description of mvd_l1_zero_flag, the mvd_l1_zero_flag signaling may mean indicating 0 for both Mvd and MvdCp. For example, when the condition indicating that the Mvd or MvdCp value of reference list L1 (the first reference picture list) is 0 is satisfied, the signaling indicating the MVP can be parsed. Therefore, it may not be necessary to infer the signaling indicating the MVP. Therefore, even when the Mvd or MvdCp value based on mvd_l1_zero_flag is 0, the freedom to select the MVP can be maintained. Therefore, the compilation efficiency can be improved. In addition, even when the condition indicating that the Mvd or MvdCp value is 0 is not satisfied based on mvd_l1_zero_flag, the signaling indicating the MVP can be parsed.
[0362] More specifically, when the condition indicating that the Mvd or MvdCp value of reference list L1 is 0 is satisfied and affine MC is used, the signaling indicating the MVP can be parsed.
[0363] Reference At line 4301, information inter_pred_idc about the reference picture list for the current block can be obtained. Referring to line 4302, when the information inter_pred_idc about the reference picture list indicates that it is not only using the zero-th reference picture list list 0, the motion vector predictor index mvp_l1_flag of the first reference picture list list 1 can be parsed from the bitstream at line 4303.
[0364] The mvd_l1_zero_flag (motion vector difference zero flag) can be obtained from the bitstream. The mvd_l1_zero_flag (motion vector difference zero flag) can indicate whether MvdLX (motion vector difference) and MvdCpLX (multiple control point motion vector differences) are set to 0 for the first reference picture list. Signaling can mean the signal transmitted from the encoder to the decoder through the bitstream. The decoder can parse the mvd_l1_zero_flag (motion vector difference zero flag) from the bitstream.
[0365] When the mvd_l1_zero_flag (motion vector difference zero flag) is 1 and the inter_pred_idc (information about the reference picture list) is PRED_BI, the mvp_l1_flag (motion vector predictor index) can be parsed. Here, PRED_BI can indicate the use of both list 0 (the zero-th reference picture list) and list 1 (the first reference picture list). Alternatively, when the mvd_l1_zero_flag (motion vector difference zero flag) is 0 or the inter_pred_idc (information about the reference picture list) is not PRED_BI, the mvp_l1_flag (motion vector predictor index) can be parsed. That is, when the mvd_l1_zero_flag (motion vector difference zero flag) is 1, the mvp_l1_flag (motion vector predictor index) can be parsed regardless of whether the inter_pred_idc (information about the reference picture list) indicates that both the zero-th reference picture list and the first reference picture list are used.
[0366] In this embodiment, when specific conditions are met, signaling for determining Mvd and MvdCp and parsing the indication of MVP based on mvd_l1_zero_flag (motion vector difference zero flag of the first reference picture list) can occur. For example, the specific conditions can include the condition where general_merge_flag is 0. For example, general_merge_flag can have the same meaning as the above-mentioned merge_flag. In addition, the specific conditions can include conditions based on CuPredMode. More specifically, the specific conditions can include the condition where CuPredMode is not MODE_IBC. Alternatively, the specific conditions can include the condition where CuPredMode is MODE_INTER. When CuPredMode is MODE_IBC, prediction using the current picture as a reference can be used. In addition, when CuPredMode is MODE_IBC, there may be a block vector or a motion vector corresponding to the block. If CuPredMode is MODE_INTER, prediction using a picture other than the current picture as a reference can be used. When CuPredMode is MODE_INTER, there may be a motion vector corresponding to the block.
[0367] Therefore, according to an embodiment of the present disclosure, when general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag is 1, and inter_pred_idc is PRED_BI, mvp_l1_flag can be parsed. Therefore, mvp_l1_flag (motion vector predictor index of the first reference picture list) exists, and its value may not be inferred.
[0368] More specifically, when general_merge_flag is 0, CuPredMode is not MODE_IBC, mvd_l1_zero_flag (motion vector difference zero flag of the first reference picture list) is 1, inter_pred_idc (information about the reference picture list) is PRED_BI, and affine MC is used, mvp_l1_flag can be parsed. In addition, mvp_l1_flag (motion vector predictor index of the first reference picture list) exists, and its value may not be inferred.
[0369] Embodiments of and can also be implemented together. For example, mvp_l1_flag can be parsed after initializing Mvd or MvdCp. In this case, the initialization of Mvd or MvdCp can be a reference Description initialization. Additionally, the mvp_l1_flag parsing can follow the description. For example, when the Mvd and MvdCp values of reference list L1 based on mvd_l1_zero_flag are not 0, if the MotionModelIdc value is 1, the MvdCpL1 value of control point index 2 can be initialized and the mvp_l1_flag can be parsed.
[0370] is a diagram illustrating an inter-frame prediction related syntax structure according to an embodiment of the present disclosure.
[0371] The embodiment of can be an embodiment that increases compilation efficiency by not removing the degrees of freedom in selecting the MVP. Additionally, the embodiment of can be an embodiment obtained by describing the reference
[0372] in the embodiment of, when the Mvd or MvdCp value is indicated as 0 based on mvd_l1_zero_flag, the signaling indicating the MVP can be parsed. The signaling indicating the MVP can include the mvp_l1_flag.
[0373] is a diagram illustrating an inter-frame prediction related syntax according to an embodiment of the present disclosure.
[0374] According to an embodiment of the present disclosure, the inter-frame prediction method can include a skip mode, a merge mode, an inter-frame mode, etc. According to the embodiment, the residual signal may not be transmitted in the skip mode. Additionally, an MV determination method such as the merge mode can be used in the skip mode. Whether to use the skip mode can be determined according to the skip flag. Referring to , whether to use the skip mode can be determined according to the value of the cu_skip_flag.
[0375] According to the embodiment, the motion vector difference may not be used in the merge mode. The motion vector can be determined based on the motion candidate index. Whether to use the merge mode can be determined according to the merge flag. Referring to , whether to use the merge mode can be determined according to the merge_flag value. Additionally, the merge mode can be used when the skip mode is not used.
[0376] In skip mode or merge mode, one or more candidate list types can be selectively used. For example, merge candidates or sub-block merge candidates can be used. In addition, merge candidates can include spatially adjacent candidates, temporal candidates, etc. In addition, merge candidates can include candidates using the motion vector for the entire current block (CU; coding unit). That is, the motion vectors of each sub-block belonging to the current block can include the same candidates. In addition, sub-block merge candidates can include temporal MVs based on sub-blocks, affine merge candidates, etc. In addition, sub-block merge candidates can include candidates that can use different motion vectors for each sub-block of the current block (CU). Affine merge candidates can be a method constructed by determining the control point motion vector of affine motion prediction without using the motion vector difference when determining the control point motion vector. In addition, sub-block merge candidates can include methods for determining motion vectors in units of sub-blocks in the current block. For example, in addition to the above-mentioned temporal MVs based on sub-blocks and affine merge candidates, sub-block merge candidates can include planar MVs, regression-based MVs, STMVP, etc.
[0377] According to an embodiment, the motion vector difference can be used in the inter-frame mode. The motion vector predictor can be determined based on the motion candidate index, and the motion vector can be determined based on the difference between the motion vector predictor and the motion vector difference. Whether to use the inter-frame mode can be determined according to whether other modes are used. In another embodiment, whether to use the inter-frame mode can be determined by a flag. An example of using the inter-frame mode when the skip mode and the merge mode, which are other modes, are not used is illustrated.
[0378] The inter-frame mode can include the AMVP mode, the affine inter-frame mode, etc. The inter-frame mode can be a mode for determining a motion vector based on a motion vector predictor and a motion vector difference. The affine inter-frame mode can be a method of using the motion vector difference when determining the control point motion vector of affine motion prediction.
[0379] Reference , it is possible to determine whether to use sub-block merge candidates or merge candidates after determining the skip mode or the merge mode. For example, when a specific condition is satisfied, the merge_subblock_flag indicating whether to use sub-block merge candidates can be parsed. In addition, the specific condition can be a condition related to the block size. For example, it can be a condition related to the width, height, area, etc., and combinations of these can be used. Reference , for example, it can be a condition when the width and height of the current block (CU) are greater than or equal to a specific value. When parsing merge_subblock_flag, its value can be inferred to be equal to 0. If merge_subblock_flag is 1, sub-block merge candidates can be used, and if merge_subblock_flag is 0, merge candidates can be used. When using sub-block merge candidates, merge_subblock_idx as the candidate index can be parsed, and when using merge candidates, merge_idx as the candidate index can be parsed. In this case, when the maximum number of the candidate list is 1, the parsing may not be performed. When merge_subblock_idx or merge_idx is not parsed, it can be inferred to be equal to 0.
[0380] Illustrate the coding_unit function, where the content related to intra prediction can be omitted, and Illustrate the case of determining inter prediction.
[0381] It is a diagram illustrating the triangle partitioning mode according to an embodiment of the present disclosure.
[0382] The triangle partitioning mode (TPM) mentioned in the present disclosure can be referred to by various names, such as triangle partition mode, triangle prediction, triangle-based prediction, triangle motion compensation, triangle prediction, triangle inter prediction, triangle merge mode, and triangle combination mode. In addition, TPM can be included in the geometric partitioning mode (GPM).
[0383] As illustrated, TPM can be a method of dividing a rectangular block into two triangles. However, GPM can divide a block into two blocks in various ways. For example, GPM can divide a rectangular block into two triangular blocks, as illustrated. In addition, GPM can divide a rectangular block into a pentagonal block and a triangular block. In addition, GPM can divide a rectangular block into two quadrilateral blocks. Here, the rectangle can include a square. Hereinafter, for the convenience of description, the description is based on TPM, which is a simple version of GPM, but it should be understood to include GPM.
[0384] According to an embodiment of the present disclosure, there may be unidirectional prediction as a prediction method. Unidirectional prediction may be a prediction method using one reference list. There may be multiple reference lists, and according to an embodiment, there may be two reference lists of L0 and L1. When using unidirectional prediction, one reference list can be used in one block. In addition, when using unidirectional prediction, one motion information can be used to predict one pixel. In the present disclosure, a block may refer to a coding unit (CU) or a prediction unit (PU). In addition, in the present disclosure, a block may refer to a transform unit (TU).
[0385] According to another embodiment of the present disclosure, there may be bidirectional prediction as a method for prediction. Bidirectional prediction may be a prediction method using multiple reference lists. In an embodiment, bidirectional prediction may be a prediction method using two reference lists. For example, bidirectional prediction may use L0 and L1 reference lists. When using bidirectional prediction, multiple reference lists can be used in one block. For example, when using bidirectional prediction, two reference lists can be used in one block. In addition, when using bidirectional prediction, multiple motion information can be used to predict one pixel.
[0386] Motion information may include a motion vector, a reference index, and a prediction list usage flag.
[0387] The reference list may be a reference picture list.
[0388] In the present disclosure, the motion information corresponding to unidirectional prediction or bidirectional prediction may be defined as one set of motion information.
[0389] According to an embodiment of the present disclosure, multiple sets of motion information can be used when using TPM. For example, when using TPM, two sets of motion information can be used. For example, when using TPM, up to two sets of motion information can be used. In addition, the method of applying two sets of motion information within a block using TPM may be based on position. For example, within a block using TPM, one set of motion information for a preset position can be used and another set of motion information for another preset position can be used. Additionally, for another preset position, two sets of motion information can be used together. For example, for another preset position, prediction 3 based on prediction 1 from one set of motion information and prediction 2 from another set of motion information can be used for prediction. For example, prediction 3 may be a weighted sum of prediction 1 and prediction 2.
[0390] Reference , partition 1 and partition 2 may schematically represent the preset position and other preset positions. When using TPM, one of two splitting methods can be used, such as As shown in the figure. The two splitting methods may include diagonal splitting and anti-diagonal splitting. The block may also be divided into two triangular-shaped partitions by splitting. As described above, the TPM may be included in the GPM. Since the GPM has been described, redundant descriptions will be omitted.
[0391] According to an embodiment of the present disclosure, when using the TPM, it is possible to use only unidirectional prediction for each partition. That is, it is possible to use one motion information for each partition. This may be to reduce memory access and complexity, such as computational complexity. Therefore, it is possible to use only two motion information for each CU.
[0392] It is also possible to determine each motion information from the candidate list. According to an embodiment, the candidate list for the TPM may be based on the merge candidate list. In another embodiment, the candidate list for the TPM may be based on the AMVP candidate list. Therefore, it is possible to signal the candidate index for easy use of the TPM. In addition, for a block using the TPM, it is possible to encode, decode, and parse as many candidate indexes as the number of partitions or the maximum number of partitions in the TPM.
[0393] In addition, even if a block is predicted based on multiple motion information by the TPM, it is possible to perform transformation and quantization on the entire block.
[0394] FIG. is a diagram illustrating merge data syntax according to an embodiment of the present disclosure.
[0395] According to an embodiment of the present disclosure, the merge data syntax may include signaling related to various modes. The various modes may include a regular merge mode, merge with MVD (MMVD), sub-block merge mode, combined intra and inter prediction (CIIP), TPM, etc. The regular merge mode may be the same mode as the merge mode in HEVC. In addition, there may be signaling indicating whether various modes are used in a block. In addition, these signals may be parsed as syntax elements or may be implicitly signaled. Refer to , the signals indicating whether to use the regular merge mode, MMVD, sub-block merge mode, CIIP, and TPM may be regular_merge_flag, mmvd_merge_flag (or mmvd_flag), merge_subblock_flag, ciip_flag (or mh_intra_flag), MergeTriangleFlag (or merge_triangle_flag), respectively.
[0396] According to an embodiment of the present disclosure, when using the merge mode, if it is signaled that all modes except a certain mode among various modes are not used, then it can be determined that the certain mode is used. In addition, when using the merge mode, if it is signaled that at least one mode among the modes except a certain mode among various modes is used, then it can be determined that the certain mode is not used. In addition, there may be a higher-level signaling indicating whether a mode can be used. The higher level may be a unit including blocks. The higher level may be a sequence, a picture, a slice, a tile group, a tile, a CTU, etc. If the higher-level signaling indicating whether a mode can be used indicates that it can be used, then there may be additional signaling indicating whether the mode is used, and the mode may or may not be used. If the higher-level signaling indicating whether a mode can be used indicates that it cannot be used, then the mode may not be used. For example, when using the merge mode, if it is signaled that the normal merge mode, MMVD, the sub-block merge mode, and CIIP are not used, then it can be determined that TPM is used. In addition, when using the merge mode, if it is signaled that at least one of the normal merge mode, MMVD, the sub-block merge mode, and CIIP is used, then it can be determined that TPM is not used. In addition, there may be signaling indicating whether the merge mode is used. For example, the signaling indicating whether the merge mode is used may be general_merge_flag or merge_flag. If the merge mode is used, then the merge data syntax as illustrated in can be parsed.
[0397] In addition, the block size available for TPM may be limited. For example, when both the width and height are 8 or more, TPM can be used.
[0398] If TPM is used, then the syntax elements related to TPM can be parsed. The syntax elements related to TPM may include signaling indicating the splitting method and signaling indicating the candidate index. The splitting method may mean the splitting direction. For a block using TPM, there may be multiple (e.g., two) signaling indicating the candidate index. Referring to , the signaling indicating the splitting method may be merge_triangle_split_dir. In addition, the signaling indicating the candidate index may be merge_triangle_idx0 and merge_triangle_idx1.
[0399] In the present disclosure, the candidate index for TPM may be m and n. For example, the candidate indices of partition 1 and partition 2 of Determined by signaling of the described candidate index. According to an embodiment of the present disclosure, one of m and n can be determined based on one of merge_triangle_idx0 and merge_triangle_idx1, and the other of m and n can be determined based on both merge_triangle_idx0 and merge_triangle_idx1.
[0400] Alternatively, one of m and n can be determined based on one of merge_triangle_idx0 and merge_triangle_idx1, and the other of m and n can be determined based on the other of merge_triangle_idx0 and merge_triangle_idx1.
[0401] More specifically, m can be determined based on merge_triangle_idx0, and n can be determined based on merge_triangle_idx0 (or m) and merge_triangle_idx1. For example, m and n can be determined as follows.
[0402] m = merge_triangle_idx0
[0403] n = merge_triangle_idx1 + (merge_triangle_idx1 >= m)? 1 : 0
[0404] According to an embodiment of the present disclosure, m and n can be different. This is because, in the TPM, when two candidate indexes are the same, that is, when two motion information is the same, the effect of partitioning may not be obtained. Therefore, the above signaling method can be used to reduce the number of signaling bits when signaling n in the case of n > m. Since m will not be n among all candidates, it can be excluded from the signaling.
[0405] If the candidate list used in the TPM is mergeCandList, mergeCandList[m] and mergeCandList[n] can be used as motion information in the TPM.
[0406] It is a diagram showing a high-level signaling according to an embodiment of the present disclosure.
[0407] According to embodiments of the present disclosure, there may be multiple higher-level signaling. The higher-level signaling may be signaling transmitted in a higher-level unit. The higher-level unit may include one or more lower-level units. The higher-level signaling may be signaling applied to one or more lower-level units. For example, a slice or a sequence may be a higher-level unit of a CU, a PU, a TU, etc. Conversely, a CU, a PU, or a TU may be a lower-level unit of a slice or a sequence.
[0408] According to embodiments of the present disclosure, the higher-level signaling may include signaling indicating a maximum number of candidates. For example, the higher-level signaling may include signaling indicating a maximum number of merge candidates. For example, the higher-level signaling may include signaling indicating a maximum number of candidates used in the TPM. When inter-frame prediction is allowed, signaling indicating a maximum number of merge candidates or signaling indicating a maximum number of candidates used in the TPM may be signaled and parsed. Whether inter-frame prediction is allowed may be determined by a slice type. As the slice type, there may be I, P, B, etc. For example, when the slice type is I, inter-frame prediction may not be allowed. For example, when the slice type is I, only intra-frame prediction or intra-block copy (IBC) may be used. In addition, when the slice type is P or B, inter-frame prediction may be allowed. In addition, when the slice type is P or B, intra-frame prediction, IBC, etc. may be allowed. In addition, when the slice type is P, at most one reference list may be used to predict pixels. In addition, when the slice type is B, multiple reference lists may be used to predict pixels. For example, if the slice type is B, two reference lists may be used to predict pixels.
[0409] According to embodiments of the present disclosure, when signaling the maximum number, it may be signaled based on a reference value. For example, (reference value - maximum number) may be signaled. Therefore, the maximum number may be derived based on the value obtained by parsing by the decoder and the reference value. For example, (reference value - value obtained by parsing) may be determined as the maximum number.
[0410] According to an embodiment, the reference value in the signaling indicating the maximum number of merge candidates may be 6.
[0411] According to an embodiment, the reference value in the signaling indicating the maximum number of candidates used in the TPM may be the maximum number of merge candidates.
[0412] Reference , the signaling indicating the maximum number of merge candidates may be six_minus_max_num_merge_cand. Here, the merge candidate may mean a candidate for merging motion vector prediction. Hereinafter, for convenience of description, six_minus_max_num_merge_cand will also be referred to as the first information. Reference and A signaling may refer to a signal transmitted from an encoder to a decoder through a bitstream. six_minus_max_num_merge_cand (the first information) may be signaled in units of a sequence. The decoder may parse six_minus_max_num_merge_cand (the first information) from the bitstream.
[0413] In addition, the signaling indicating the maximum number of candidates used in the TPM may be max_num_merge_cand_minus_max_num_triangle_cand. Refer to and A signaling may refer to a signal transmitted from an encoder to a decoder through a bitstream. The decoder may parse max_num_merge_cand_minus_max_num_triangle_cand (the third information) from the bitstream. max_num_merge_cand_minus_max_num_triangle_cand (the third information) may be information related to the maximum number of merge mode candidates for a partitioned block.
[0414] In addition, the maximum number of merge candidates may be MaxNumMergeCand (the maximum number of merge candidates), and this value may be based on six_minus_max_num_merge_cand (the first information). In addition, the maximum number of candidates used in the TPM may be MaxNumTriangleMergeCand, and this value may be based on max_num_merge_cand_minus_max_num_triangle_cand. MaxNumMergeCand (the maximum number of merge candidates) may be used for the merge mode and is information that can be used when partitioning or not partitioning a block for motion compensation. Above, it has been described based on the TPM, but the GPM may also be described in the same way.
[0415] According to an embodiment of the present disclosure, there may be a higher-level signaling indicating whether the TPM mode can be used. Refer to The higher-level signaling indicating whether the TPM mode can be used may be sps_triangle_enabled_flag (the second information). The information indicating whether the TPM mode can be used may be related to the information indicating whether a block can be as The information for partitioning for inter prediction shown therein is the same. Since GPM includes TPM, the information indicating whether a block can be partitioned can be the same as the information indicating whether the GPM mode is used. Performing inter prediction can indicate performing motion compensation. That is, sps_triangle_enabled_flag (the second information) can be the information indicating whether a block can be partitioned for inter prediction. When the second information indicating whether a block can be partitioned is 1, it can indicate that TPM or GPM can be used. Additionally, when the second information is 0, it can indicate that TPM or GPM cannot be used. However, the present disclosure is not limited thereto, and when the second information is 0, it can indicate that TPM or GPM can be used. Additionally, when the second information is 1, it can indicate that TPM or GPM cannot be used.
[0416] Reference and Figure 7 , signaling can mean a signal transmitted from an encoder to a decoder through a bitstream. The decoder can parse sps_triangle_enabled_flag (the second information) from the bitstream.
[0417] According to an embodiment of the present disclosure, TPM can be used only when there can be candidates used in TPM that are greater than or equal to the number of partitions of TPM. For example, when TPM is partitioned into two partitions, TPM can be used only when there can be two or more candidates used in TPM. According to an embodiment, the candidates used in TPM can be based on merge candidates. Thus, according to an embodiment of the present disclosure, when the maximum number of merge candidates is 2 or more, TPM can be used. Thus, when the maximum number of merge candidates is 2 or more, signaling related to TPM can be parsed. The signaling related to TPM can be a signaling indicating the maximum number of candidates used in TPM.
[0418] Reference Figure 48 , when sps_triangle_enabled_flag is 1 and MaxNumMergeCand is 2 or greater, max_num_merge_cand_minus_max_num_triangle_cand can be parsed. Additionally, when sps_triangle_enabled_flag is 0 or MaxNumMergeCand is less than 2, max_num_merge_cand_minus_max_num_triangle_cand cannot be parsed.
[0419] Figure 49 is a diagram illustrating the maximum number of candidates used in TPM according to an embodiment of the present disclosure.
[0420] ReferenceFigure 49 , the maximum number of candidates used in TPM can be MaxNumTriangleMergeCand. In addition, the signaling indicating the maximum number of candidates used in TPM can be max_num_merge_cand_minus_max_num_triangle_cand. In addition, the content referred to Figure 48 can be omitted.
[0421] According to an embodiment of the present disclosure, the maximum number of candidates used in TPM can exist within the range from "the number of partitions of TPM" to "the reference value in the signaling indicating the maximum number of candidates used in TPM", including the endpoints. Therefore, when the number of partitions of TPM is 2 and the reference value is the maximum number of merge candidates, MaxNumTriangleMergeCand can exist within the range from 2 to MaxNumMergeCand (including 2 and MaxNumMergeCand), as Figure 49 illustrated.
[0422] According to an embodiment of the present disclosure, when the signaling indicating the maximum number of candidates used in TPM does not exist, it is possible to infer the signaling indicating the maximum number of candidates used in TPM or infer the maximum number of candidates used in TPM. For example, when the signaling indicating the maximum number of candidates used in TPM does not exist, it can be inferred that the maximum number of candidates used in TPM is equal to 0. Alternatively, when the signaling indicating the maximum number of candidates used in TPM does not exist, the signaling indicating the maximum number of candidates used in TPM can be inferred as the reference value.
[0423] In addition, when the signaling indicating the maximum number of candidates used in TPM does not exist, it is possible not to use TPM. Alternatively, when the maximum number of candidates used in TPM is less than the number of partitions in TPM, it is possible not to use TPM. Alternatively, when the maximum number of candidates used in TPM is 0, it is possible not to use TPM.
[0424] However, according to Figures 48 to 49 the embodiment, when the number of partitions of TPM is the same as "the reference value in the signaling indicating the maximum number of candidates used in TPM", there may be only one possible value as the maximum number of candidates used in TPM. However, according to Figures 48 to 49 the embodiment, even in this case, it may be possible to parse the signaling indicating the maximum number of candidates used in TPM, which may be unnecessary. When MaxNumMergeCand is 2, referring to Figure 49 , the possible value for MaxNumTriangleMergeCand can be only 2. However, referring toFigure 48 , even in this case, max_num_merge_cand_minus_max_num_triangle_cand can be parsed.
[0425] Reference Figure 49 , MaxNumTriangleMergeCand can be determined as (MaxNumMergeCand - max_num_merge_cand_minus_max_num_triangle_cand).
[0426] Figure 50 FIG. is a diagram illustrating higher-level signaling related to TPM according to an embodiment of the present disclosure.
[0427] According to an embodiment of the present disclosure, when the number of partitions of the TPM is the same as the "reference value in the signaling indicating the maximum number of candidates used in the TPM", the signaling indicating the maximum number of candidates used in the TPM may not be parsed. In addition, according to the above embodiment, the number of partitions of the TPM may be 2. In addition, the "reference value in the signaling indicating the maximum number of candidates used in the TPM" may be the maximum number of merge candidates. Therefore, when the maximum number of merge candidates is 2, the signaling indicating the maximum number of candidates used in the TPM may not be parsed.
[0428] Alternatively, when the "reference value in the signaling indicating the maximum number of candidates used in the TPM" is less than or equal to the number of partitions in the TPM, the signaling indicating the maximum number of candidates used in the TPM may not be parsed. Therefore, when the maximum number of merge candidates is 2 or less, the signaling indicating the maximum number of candidates used in the TPM may not be parsed.
[0429] Reference Figure 50For line 5001, when MaxNumMergeCand (the maximum number of merge candidates) is 2 or MaxNumMergeCand (the maximum number of merge candidates) is 2 or less, max_num_merge_cand_minus_max_num_triangle_cand (the third information) may not be parsed. Additionally, when sps_triangle_enabled_flag (the second information) is 1 and MaxNumMergeCand (the maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (the third information) can be parsed. Further, when sps_triangle_enabled_flag (the second information) is 0, max_num_merge_cand_minus_max_num_triangle_cand (the third information) may not be parsed. Therefore, when sps_triangle_enabled_flag (the second information) is 0 or MaxNumMergeCand (the maximum number of merge candidates) is 2 or less, max_num_merge_cand_minus_max_num_triangle_cand (the third information) may not be parsed.
[0430] Figure 51 is a diagram showing the maximum number of candidates used in the TPM according to an embodiment of the present disclosure.
[0431] Figure 51 The embodiment of Figure 50 can be implemented together with the case of the embodiment of . Additionally, the description of the above can be omitted in this diagram.
[0432] According to an embodiment of the present disclosure, when "the reference value in the signaling indicating the maximum number of candidates used in the TPM" is the number of partitions in the TPM, the maximum number of candidates used in the TPM can be inferred and set to the number of partitions of the TPM. Additionally, the inference and setting can be performed in the case where the signaling indicating the maximum number of candidates used in the TPM does not exist. According to Figure 50 the embodiment of , when "the reference value in the signaling indicating the maximum number of candidates used in the TPM" is the number of partitions in the TPM, the signaling indicating the maximum number of candidates used in the TPM may not be parsed, and when the signaling indicating the maximum number of candidates used in the TPM does not exist, the value of the maximum number of candidates used in the TPM can be inferred as the number of partitions in the TPM. The inference and setting can also be performed when additional conditions are met. The additional condition can be a condition where the higher-level signaling indicating whether the TPM mode can be used is 1.
[0433] In addition, in this embodiment, although the maximum number of candidates used in the TPM has been described as being inferred and set, it is also possible to infer and set the signaling indicating the maximum number of candidates used in the TPM, such that the maximum number of candidates used in the described TPM is derived, rather than inferring and setting the maximum number of candidates used in the TPM.
[0434] Reference Figure 50 , when the sps_triangle_enabled_flag (second information) indicates 1 and MaxNumMergeCand (maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (third information) can be received. In this case, referring to Figure 51 line 5101 of
[0435] Reference Figure 51For line 5102, when sps_triangle_enabled_flag (the second piece of information) is 1 and MaxNumMergeCand (the maximum number of merge candidates) is 2, MaxNumTriangleMergeCand (the maximum number of merge mode candidates for partitioned blocks) can be set to 2. More specifically, as already described, when sps_triangle_enabled_flag (the second piece of information) indicates 1 and MaxNumMergeCand (the maximum number of merge candidates) is greater than 2, max_num_merge_cand_minus_max_num_triangle_cand (the third piece of information) can be received, and thus, as at line 5102, when sps_triangle_enabled_flag (the second piece of information) is 1 and MaxNumMergeCand (the maximum number of merge candidates) is 2, max_num_merge_cand_minus_max_num_triangle_cand (the third piece of information) may not be received. In this case, MaxNumTriangleMergeCand (the maximum number of merge mode candidates for partitioned blocks) can be determined without max_num_merge_cand_minus_max_num_triangle_cand (the third piece of information).
[0436] In addition, referring to Figure 51 line 5103, when sps_triangle_enabled_flag (the second piece of information) is 0 or MaxNumMergeCand (the maximum number of merge candidates) is not 2, MaxNumTriangleMergeCand (the maximum number of merge mode candidates for partitioned blocks) can be inferred and set to 0. In this case, as already described with reference to line 5101, when sps_triangle_enabled_flag (the second piece of information) indicates 1 and MaxNumMergeCand (the maximum number of merge candidates) is greater than or equal to 3, max_num_merge_cand_minus_max_num_triangle_cand (the third piece of information) can be signaled, and thus the case where MaxNumMergeCand is inferred and set to 0 can be the case where the second piece of information is 0 or the maximum number of merge candidates is 1. In summary, when sps_triangle_enabled_flag (the second piece of information) is 0 or MaxNumMergeCand (the maximum number of merge candidates) is 1, MaxNumTriangleMergeCand (the maximum number of merge mode candidates for partitioned blocks) can be set to 0.
[0437] MaxNumMergeCand (maximum number of merge candidates) and MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) can be used for different purposes. For example, MaxNumMergeCand (maximum number of merge candidates) can be used when blocks are partitioned or not partitioned for motion compensation. However, MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks) is information that can be used when blocks are partitioned. The number of candidates for partitioned blocks in merge mode cannot exceed MaxNumTriangleMergeCand (maximum number of merge mode candidates for partitioned blocks).
[0438] Another embodiment can also be used. In Figure 51 the embodiment, cases where MaxNumMergeCand is not 2 include cases where MaxNumMergeCand is greater than 2, and in such cases, the meaning of inferring and setting MaxNumTriangleMergeCand to 0 may not be clear, but in such cases, since there is signaling indicating the maximum number of candidates used in the TPM, no inference is performed, and thus, there is no operational problem. However, in this embodiment, inference can be performed using the meaning.
[0439] If sps_triangle_enabled_flag (second information) is 1 and MaxNumMergeCand (maximum number of merge candidates) is 2 or more, it is possible to infer and set MaxNumTriangleMergeCand to 2 (or MaxNumMergeCand). Otherwise (i.e., when sps_triangle_enabled_flag is 0 or MaxNumMergeCand is less than 2), it is possible to infer and set MaxNumTriangleMergeCand to 0.
[0440] Alternatively, if sps_triangle_enabled_flag is 1 and MaxNumMergeCand is 2, it is possible to infer and set MaxNumTriangleMergeCand to 2. Otherwise, if sps_triangle_enabled_flag is 0, it is possible to infer and set MaxNumTriangleMergeCand to 0.
[0441] Therefore, according to the Figure 50In an embodiment, when MaxNumMergeCand is 0, 1, or 2, there may be no signaling indicating the maximum number of candidates used in the TPM. When MaxNumMergeCand is 0 or 1, the maximum number of candidates used in the TPM can be inferred and set to 0. When MaxNumMergeCand is 2, the maximum number of candidates used in the TPM can be inferred and set to 2.
[0442] Figure 52 FIG. is a diagram illustrating syntax elements related to the TPM according to an embodiment of the present disclosure.
[0443] As described above, the maximum number of candidates used in the TPM can exist, and the number of partitions of the TPM can be preset. In addition, the candidate indexes used in the TPM may be different.
[0444] According to an embodiment of the present disclosure, when the maximum number of candidates used in the TPM is the same as the number of partitions in the TPM, signaling different from the case where they are not the same can be performed. For example, when the maximum number of candidates used in the TPM is the same as the number of partitions in the TPM, signaling different from the case where they are not the same can be performed. Therefore, signaling can be sent with fewer bits. Alternatively, when the maximum number of candidates used in the TPM is less than or equal to the number of partitions of the TPM, signaling different from the case where the maximum number of candidates is greater than the number of partitions can be performed (among these, when the number of partitions is less than the number of partitions of the TPM, it may be a case where the TPM cannot be used).
[0445] When the TPM has two partitions, two candidate indexes can be signaled. If the maximum number of candidates used in the TPM is 2, there may be only two possible combinations of candidate indexes. These two combinations can be the combination where m and n are 0 and 1 respectively, and the combination where m and n are 1 and 0 respectively. Therefore, the two candidate indexes can be signaled with only 1-bit signaling.
[0446] Refer to Figure 52 , when MaxNumTriangleMergeCand is 2, candidate index signaling different from the case where it is not (otherwise, when using the TPM, the case where MaxNumTriangleMergeCand is greater than 2) can be performed. Alternatively, when MaxNumTriangleMergeCand is 2 or less, candidate index signaling different from the case where it is not (otherwise, when using the TPM, the case where MaxNumTriangleMergeCand is greater than 2) can be performed. When referring to Figure 52When different candidate index signaling can be merge_triangle_idx_indicator parsing. Different candidate index signaling can be a signaling method that does not parse merge_triangle_idx0 or merge_triangle_idx1. This will be described with reference to Figure 53 in further detail.
[0447] Figure 53 is a diagram illustrating the signaling of TPM candidate indexes according to an embodiment of the present disclosure.
[0448] According to an embodiment of the present disclosure, when using different index signaling from that described in the reference Figure 52 candidate indexes can be determined based on merge_triangle_idx_indicator. In addition, when using different index signaling, merge_triangle_idx0 or merge_triangle_idx1 may not exist.
[0449] According to an embodiment of the present disclosure, when the maximum number of candidates used in the TPM is the same as the number of partitions in the TPM, the TPM candidate indexes can be determined based on merge_triangle_idx_indicator. In addition, this may be the case for the blocks using the TPM.
[0450] More specifically, when MaxNumTriangleMergeCand is 2 (or when MaxNumTriangleMergeCand is 2 and MergeTriangleFlag is 1), the TPM candidate indexes can be determined based on merge_triangle_idx_indicator. In this case, if merge_triangle_idx_indicator is 0, m and n as the TPM candidate indexes can be set to 0 and 1 respectively, and if merge_triangle_idx_indicator is 1, m and n as the TPM candidate indexes can be set to 1 and 0 respectively. Alternatively, merge_triangle_idx0 or merge_triangle_idx1, which can be parsed to have the same values (syntax elements) as those described for m and n, can be inferred and set.
[0451] Referring to Figure 47 the method of setting m and n based on merge_triangle_idx0 and merge_triangle_idx1 described in Figure 53, when merge_triangle_idx0 does not exist, if MaxNumTriangleMergeCand is 2 and merge_triangle_idx_indicator is 1 (or if MaxNumTriangleMergeCand is 2 and MergeTriangleFlag is 1), the value of merge_triangle_idx0 can be inferred to be equal to 1. Additionally, otherwise, the value of merge_triangle_idx0 may be inferred to be equal to 0. Additionally, when merge_triangle_idx1 does not exist, merge_triangle_idx1 can be inferred to be equal to 0. Therefore, if MaxNumTriangleMergeCand is 2, when merge_triangle_idx_indicator is 0, merge_triangle_idx0 and merge_triangle_idx1 are 0 and 0 respectively, and accordingly m and n can be 0 and 1 respectively. Additionally, if MaxNumTriangleMergeCand is 2, when merge_triangle_idx_indicator is 1, merge_triangle_idx0 and merge_triangle_idx1 are 1 and 0 respectively, and accordingly, m and n can be 1 and 0 respectively.
[0452] Figure 54 FIG. is a diagram illustrating signaling of TPM candidate indices according to an embodiment of the present disclosure.
[0453] has been referred to Figure 47 described a method for determining TPM candidate indices, but in Figure 54 embodiments of, another determination method and a signaling method are described. Descriptions redundant with the above can be omitted. Additionally, m and n can represent candidate indices as described in reference to Figure 47
[0454] According to an embodiment of the present disclosure, the smaller value among m and n can be signaled to a preset syntax element among merge_triangle_idx0 and merge_triangle_idx1. Additionally, a value based on the difference between m and n can be signaled to the other one of merge_triangle_idx0 and merge_triangle_idx1. Additionally, a value indicating the magnitude relationship between m and n can be signaled.
[0455] For example, merge_triangle_idx0 can be the smaller value of m and n. Additionally, merge_triangle_idx1 can be based on the value of |m - n|. merge_triangle_idx1 can be (|m - n| - 1). This is because m and n may be different. Additionally, the value representing the size relationship between m and n can be Figure 54 merge_triangle_bigger.
[0456] Using this relationship, m and n can be determined based on merge_triangle_idx0, merge_triangle_idx1, and merge_triangle_bigger. Referring to Figure 54 , another operation can be performed based on the merge_triangle_bigger value. For example, when merge_triangle_bigger is 0, n may be greater than m. In this case, m can be merge_triangle_idx0. Additionally, n can be (merge_triangle_idx1 + m + 1). Additionally, when merge_triangle_bigger is 1, m may be greater than n. In this case, n can be merge_triangle_idx0. Additionally, m can be (merge_triangle_idx1 + n + 1).
[0457] Compared with the Figure 47 method in Figure 54 the advantage of the method in Figure 47 is that when the smaller value of m and n is not 0 (or greater), signaling overhead can be reduced. For example, when m and n are 3 and 4 respectively, in the Figure 54 method, merge_triangle_idx0 and merge_triangle_idx1 may need to be signaled as 3 and 3 respectively. However, in the
[0458] Figure 55 is a diagram illustrating the signaling of TPM candidate indices according to an embodiment of the present disclosure.
[0459] A method for determining TPM candidate indices has been described with reference to Figure 47 but in Figure 55Another determination method and signaling method will be described in the embodiments. Descriptions redundant with the above description may be omitted. Further, m and n may represent candidate indices as described in reference Figure 47 as described.
[0460] According to an embodiment of the present disclosure, a value based on the larger value among m and n may be signaled to a preset syntax element in merge_triangle_idx0 and merge_triangle_idx1. Further, a value based on the smaller value among m and n may be signaled to the other of merge_triangle_idx0 and merge_triangle_idx1. Further, a value indicating the magnitude relationship between m and n can be signaled.
[0461] For example, merge_triangle_idx0 may be based on the larger value among m and n. According to the embodiment, since m and n are not equal, the larger value among m and n will be greater than or equal to 1. Thus, considering that the larger value among m and n excludes 0, it can be signaled with fewer bits. For example, merge_triangle_idx0 may be ((the larger value among m and n) - 1). In this case, the maximum value of merge_triangle_idx0 may be (MaxNumTriangleMergeCand - 1 - 1) (-1 because it is a value starting from 0, and -1 because the larger value of 0 can be excluded). The maximum value can be used for binarization, and when the maximum value decreases, there may be a case of using fewer bits. Further, merge_triangle_idx1 may be the smaller value among m and n. Further, the maximum value of merge_triangle_idx1 may be merge_triangle_idx0. Thus, there may be a case of using fewer bits than when setting the maximum value to MaxNumTriangleMergeCand. Additionally, when merge_triangle_idx0 is 0, i.e., when the larger value among m and n is 1, the smaller value among m and n is 0, and thus there may be no additional signaling. For example, when merge_triangle_idx0 is 0, i.e., when the larger value among m and n is 1, it can be determined that the smaller value among m and n is 0. Alternatively, when merge_triangle_idx0 is 0, i.e., when the larger value among m and n is 1, merge_triangle_idx1 may be inferred to be equal to and determined to be 0. Refer to Figure 22, it is possible to determine whether to parse merge_triangle_idx1 based on merge_triangle_idx0. For example, when merge_triangle_idx0 is greater than 0, merge_triangle_idx1 can be parsed, and when merge_triangle_idx0 is 0, merge_triangle_idx1 may not be parsed.
[0462] In addition, the value representing the size relationship between m and n can be Figure 55 merge_triangle_bigger.
[0463] Using this relationship, m and n can be determined based on merge_triangle_idx0, merge_triangle_idx1, and merge_triangle_bigger. Refer to Figure 55 , another operation can be performed based on the value of merge_triangle_bigger. For example, when merge_triangle_bigger is 0, m may be greater than n. In this case, m can be (merge_triangle_idx0 + 1). In addition, n can be merge_triangle_idx1. In addition, when merge_triangle_bigger is 1, n may be greater than m. In this case, n can be (merge_triangle_idx0 + 1). In addition, m can be merge_triangle_idx1. In addition, when merge_triangle_idx1 does not exist, its value can be inferred to be equal to 0.
[0464] Compared with Figure 47 the method in Figure 55 , the advantage of the method in Figure 47 is that the signaling overhead can be reduced according to the m and n values. For example, when m and n are 1 and 0 respectively, in Figure 55 the method inFigure 47 In the method, it may be necessary to separately signal merge_triangle_idx0 and merge_triangle_idx1 with values of 2 and 1 respectively, and in Figure 55 the method, separately signal merge_triangle_idx0 and merge_triangle_idx1 with values of 1 and 1 respectively. However, in this case, in Figure 55 the method, since the maximum value of merge_triangle_idx1 is (3 – 1 – 1) = 1, 1 can be signaled with fewer bits compared to when the maximum value is larger. For example, Figure 55 the method may be a method that has an advantage when the difference between m and n is small, for example, when the difference is 1.
[0465] Reference to merge_triangle_idx0 has also been described to facilitate determining whether to parse Figure 55 merge_triangle_idx1 in the syntax structure, but it is also possible to determine whether to parse merge_triangle_idx1 based on the larger value among m and n. That is, it can be classified into cases where the larger value among m and n is 1 or greater and cases where it is not 1. However, in this case, merge_triangle_bigger parsing may need to occur before determining whether to parse merge_triangle_idx1.
[0466] In the above description, the configuration has been described through specific examples, but those skilled in the art can make modifications and variations without departing from the spirit and scope of the present disclosure. Therefore, what can be easily inferred by those skilled in the technical field to which the present disclosure pertains from the detailed description and examples of the present disclosure is construed as falling within the scope of the rights of the present disclosure.
Claims
1. A method for processing a video signal, comprising: parsing from a bitstream an Adaptive Motion Vector Resolution (AMVR) enable flag indicating whether adaptive motion vector differential resolution is enabled; parsing from the bitstream an affine enable flag indicating whether affine motion compensation is enabled; when the AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled and the affine enable flag indicates that the affine motion compensation is enabled, parsing from the bitstream an affine AMVR enable flag indicating whether the adaptive motion vector differential resolution is enabled for the affine motion compensation; parsing from the bitstream information about a reference picture list for a current block; parsing from the bitstream a motion vector difference zero flag, the motion vector difference zero flag indicating whether, for reference picture list 1, the motion vector difference and a plurality of control point motion vector differences are set to zero; when the information about the reference picture list indicates that a reference picture list other than reference picture list 0 is available, parsing from the bitstream a motion vector predictor index for reference picture list 1; wherein the motion vector predictor index is parsed regardless of the motion vector difference zero flag.
2. The method according to claim 1, Among them, wherein at least one of the AMVR enable flag, the affine enable flag, or the affine AMVR enable flag is signaled as one of a coding tree unit, a slice, a tile, a tile group, a picture, or a sequence unit.
3. The method according to claim 1, Among them, when the affine motion compensation is enabled and the adaptive motion vector differential resolution is not enabled, the value of the affine AMVR enable flag is inferred to be a value indicating that the adaptive motion vector differential resolution is not enabled for the affine motion compensation.
4. The method according to claim 1, Among them, when the affine motion compensation is not enabled, the value of the affine AMVR enable flag is inferred to be a value indicating that the adaptive motion vector differential resolution is not enabled for the affine motion compensation.
5. The method according to claim 1, the method further comprising: when the AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled, an inter-frame affine flag obtained from the bitstream indicates that the affine motion compensation is not used for the current block, and at least one of a plurality of motion vector differences for the current block is non-zero, parsing from the bitstream information about the resolution of the motion vector difference; and modifying the plurality of motion vector differences for the current block based on the information about the resolution of the motion vector difference.
6. The method according to claim 1, the method further comprising: when the affine AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled for the affine motion compensation, an inter-frame affine flag obtained from the bitstream indicates that affine motion compensation is used for the current block, and at least one of a plurality of control point motion vector differences for the current block is non-zero, parsing from the bitstream information about the resolution of the motion vector difference; and Modify the multiple control point motion vector differences for the current block based on information about the resolution of the motion vector differences.
7. A decoding apparatus for processing a video signal, the decoding apparatus comprising: a processor, wherein the processor is configured to: parse from a bitstream an Adaptive Motion Vector Resolution (AMVR) enable flag indicating whether adaptive motion vector differential resolution is enabled; parse from the bitstream an affine enable flag indicating whether affine motion compensation is enabled; when the AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled and the affine enable flag indicates that the affine motion compensation is enabled, parse from the bitstream an affine AMVR enable flag indicating whether the adaptive motion vector differential resolution is enabled for the affine motion compensation; parse from the bitstream information about a reference picture list for a current block; parse from the bitstream a motion vector difference zero flag, the motion vector difference zero flag indicating whether a motion vector difference and multiple control point motion vector differences are set to zero for reference picture list 1; when the information about the reference picture list indicates that a reference picture list other than reference picture list 0 is available, parse from the bitstream a motion vector predictor index for the reference picture list 1; wherein the motion vector predictor index is parsed regardless of the motion vector difference zero flag.
8. The decoding apparatus according to claim 7, Among them, at least one of the AMVR enable flag, the affine enable flag, or the affine AMVR enable flag is signaled as one of a compile tree unit, a slice, a tile, a tile group, a picture, or a sequence unit.
9. The decoding apparatus according to claim 7, Among them, when the affine motion compensation is enabled and the adaptive motion vector differential resolution is not enabled, the value of the affine AMVR enable flag is inferred to be a value indicating that the adaptive motion vector differential resolution is not enabled for the affine motion compensation.
10. The decoding apparatus according to claim 7, Among them, when the affine motion compensation is not enabled, the value of the affine AMVR enable flag is inferred to be a value indicating that the adaptive motion vector differential resolution is not enabled for the affine motion compensation.
11. The decoding apparatus according to claim 7, Among them, the processor is configured to: when the AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled, an inter-frame affine flag obtained from the bitstream indicates that the affine motion compensation is not used for the current block, and at least one of the multiple motion vector differences for the current block is non-zero, parse from the bitstream information about the resolution of the motion vector differences; and modify the multiple motion vector differences for the current block based on the information about the resolution of the motion vector differences.
12. The decoding apparatus according to claim 7, Among them, the processor is configured to: When the affine AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled for the affine motion compensation, the inter-frame affine flag obtained from the bitstream indicates the use of the affine motion compensation for the current block, and at least one of the motion vector differences of multiple control points for the current block is non-zero, parse information about the resolution of the motion vector difference from the bitstream; and Modify the motion vector differences of multiple control points for the current block based on the information about the resolution of the motion vector difference.
13. An encoding apparatus for processing a video signal, comprising: A processor, wherein the processor is configured to: Obtain an adaptive motion vector resolution (AMVR) enable flag indicating whether the adaptive motion vector differential resolution is enabled; Obtain an affine enable flag indicating whether the affine motion compensation is enabled; When the AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled, and the affine enable flag indicates that the affine motion compensation is enabled, obtain an affine AMVR enable flag indicating whether the adaptive motion vector differential resolution is enabled for the affine motion compensation, Obtain information about a reference picture list for the current block; Obtain a motion vector difference zero flag, the motion vector difference zero flag indicating whether the motion vector difference and the motion vector differences of multiple control points are set to zero for reference picture list 1; When the information about the reference picture list indicates that a reference picture list other than reference picture list 0 is available, obtain a motion vector predictor index of the reference picture list 1, wherein the motion vector predictor index is obtained regardless of the motion vector difference zero flag, Obtain a bitstream including at least one of the AMVR enable flag, the affine enable flag, the affine AMVR enable flag, the information about the reference picture, the motion vector difference zero flag, or the motion vector predictor index.
14. The encoding apparatus according to claim 13, Among them, At least one of the AMVR enable flag, the affine enable flag, or the affine AMVR enable flag is signaled as one of a coding tree unit, a slice, a tile, a tile group, a picture, or a sequence unit.
15. A method for obtaining a bitstream, the method comprising: Obtain an adaptive motion vector resolution (AMVR) enable flag indicating whether the adaptive motion vector differential resolution is enabled; Obtain an affine enable flag indicating whether the affine motion compensation is enabled; When the AMVR enable flag indicates that the adaptive motion vector differential resolution is enabled, and the affine enable flag indicates that the affine motion compensation is enabled, obtain an affine AMVR enable flag indicating whether the adaptive motion vector differential resolution is enabled for the affine motion compensation, Obtain information about a reference picture list for the current block; Obtain a motion vector difference zero flag, the motion vector difference zero flag indicating whether the motion vector difference and the motion vector differences of multiple control points are set to zero for reference picture list 1; When the information about the reference picture list indicates that a reference picture list other than reference picture list 0 is available, obtain the motion vector predictor index of the reference picture list 1, wherein the motion vector predictor index is obtained regardless of the motion vector difference zero flag, Obtain a bitstream including at least one of the AMVR enable flag, the affine enable flag, the affine AMVR enable flag, the information about the reference picture, the motion vector difference zero flag, or the motion vector predictor index.