Method and apparatus for processing video signals using affine prediction

JP7923860B2Active Publication Date: 2026-09-18GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025069754
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-08-03
Filing Date
2025-04-21
Publication Date
2026-09-18
Estimated Expiration
2038-08-03

AI Technical Summary

Benefits of technology

【0014】 本発明は、アフィン予測(affine prediction)を適応的に行う方法を提供することにより、アフィン予測の性能を向上させることができ、アフィン予測の複雑度を減少させることにより、より効率的なコーディングを行うことができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007923860000020
    Figure 0007923860000020
  • Figure 0007923860000021
    Figure 0007923860000021
  • Figure 0007923860000022
    Figure 0007923860000022
Patent Text Reader

Abstract

To provide a method for decoding a video signal including a current block on the basis of an affine motion prediction mode (AF mode).SOLUTION: A method for decoding a video signal includes the steps of: checking whether an AF mode is applied to a current block, the AF mode representing a motion prediction mode that uses an affine motion model; checking whether an AF4 mode is used when the AF mode is applied to the current block, the AF4 mode representing a mode in which a motion vector is predicted using four parameters constituting the affine motion model; generating a motion vector predictor using four parameters when the AF4 mode is used and generating a motion vector predictor using six parameters constituting the affine motion model when the AF4 mode is not used; and obtaining the motion vector of the current block on the basis of the motion vector predictor.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for encoding / decoding a video signal, and more specifically, to a method and apparatus for adaptively performing affine prediction. [Background Art]

[0002] Compression coding refers to a series of signal processing techniques for transmitting digitized information via a communication line or storing the information in a form suitable for a storage medium. Media such as moving pictures, still images and audio are targets of compression coding, and in particular, a technique of performing compression coding targeting images is referred to as video image compression.

[0003] Next-generation video content has the characteristics of high spatial resolution, high frame rate, and high dimensionality of scene representation. Processing such content will bring about a significant increase in terms of memory storage, memory access rate, and processing power.

[0004] Therefore, it is necessary to design coding tools for more efficiently processing next-generation video content. [Summary of the Invention] [Problem to be Solved by the Invention]

[0005] The present invention proposes a method for more efficiently encoding and decoding a video signal.

[0006] Furthermore, the present invention proposes a method for encoding or decoding that takes into account both AF4 mode, which is a four-parameter affine prediction mode that uses four parameters, and AF6 mode, which is a six-parameter affine prediction mode that uses six parameters.

[0007] Furthermore, the present invention proposes a method for adaptively determining (or selecting) the optimal coding mode based on the block size, using at least one of AF4 mode or AF6 mode.

[0008] Furthermore, the present invention proposes a method for adaptively determining (or selecting) the optimal coding mode based on at least one of AF4 mode or AF6 mode, depending on whether adjacent blocks have been coded by affine prediction. [Means for solving the problem]

[0009] In order to solve the aforementioned technical challenges,

[0010] This invention provides a method for adaptively performing affine prediction based on block size.

[0011] Furthermore, the present invention provides a method for adaptively performing affine prediction based on whether or not adjacent blocks have been coded by affine prediction.

[0012] Furthermore, the present invention provides a method for adaptively determining (or selecting) the optimal coding mode based on at least one of AF4 mode or AF6 mode.

[0013] Furthermore, the present invention provides a method for adaptively performing affine prediction based on whether or not at least one pre-defined condition is satisfied, in which case the pre-defined condition may include at least one of the following: block size, number of pixels in a block, block width, block height, and whether or not adjacent blocks have been coded by affine prediction. [Effects of the Invention]

[0014] This invention provides a method for adaptively performing affine prediction, thereby improving the performance of affine prediction and reducing the complexity of affine prediction, which in turn enables more efficient coding. [Brief explanation of the drawing]

[0015] [Figure 1] A schematic block diagram of an encoder used to encode a video signal is shown as an embodiment to which the present invention is applied. [Figure 2] As an embodiment to which the present invention is applied, a schematic block diagram of a decoder that decodes a video signal is shown. [Figure 3] This figure illustrates a QT (QuadTree, hereinafter referred to as "QT") block partitioning structure as an embodiment to which the present invention can be applied. [Figure 4] This diagram illustrates a BT (Binary Tree, hereinafter referred to as "BT") block partitioning structure as an embodiment to which the present invention can be applied. [Figure 5] This figure illustrates a TT (Ternary Tree, hereinafter referred to as "TT") block partitioning structure as an embodiment to which the present invention can be applied. [Figure 6] This diagram illustrates an AT (Asymmetric Tree, hereinafter referred to as "AT") block partitioning structure as an embodiment to which the present invention can be applied. [Figure 7]It is a diagram for explaining an inter prediction mode as an embodiment to which the present invention is applied. [Figure 8] It is a diagram for explaining an affine motion model as an embodiment to which the present invention is applied. [Figure 9] It is a diagram for explaining an affine motion prediction method using a control point motion vector as an embodiment to which the present invention is applied. [Figure 10] It is a flowchart for explaining a process of processing a video signal including a current block by using an Affine prediction mode as an embodiment to which the present invention is applied. [Figure 11] As embodiment (1-1) to which the present invention is applied, it shows a flowchart for adaptively determining an optimal coding mode based on at least one of an AF4 mode or an AF6 mode. [Figure 12] As embodiment (1-2) to which the present invention is applied, it shows a flowchart for performing adaptive decoding based on the AF4 mode or the AF6 mode. [Figure 13] As embodiment (1-3) to which the present invention is applied, it shows a syntax structure for performing decoding based on the AF4 mode or the AF6 mode. [Figure 14] As embodiment (2-1) to which the present invention is applied, it shows a flowchart for adaptively determining an optimal coding mode from motion vector prediction modes including the AF4 mode or the AF6 mode based on condition A. [Figure 15] As embodiment (2-2) to which the present invention is applied, it shows a flowchart for performing adaptive decoding according to the AF4 mode or the AF6 mode based on condition A. [Figure 16] As embodiment (2-3) to which the present invention is applied, it shows a syntax structure for performing decoding according to the AF4 mode or the AF6 mode based on condition A. [Figure 17] As embodiment (3-1) to which the present invention is applied, there is shown a flowchart for adaptively determining an optimal coding mode from motion vector prediction modes including AF4 mode or AF6 mode based on at least one of condition B or condition C. [Figure 18] As embodiment (3-2) to which the present invention is applied, there is shown a flowchart for performing adaptive decoding in accordance with AF4 mode or AF6 mode based on at least one of condition B or condition C. [Figure 19] As embodiment (3-3) to which the present invention is applied, there is shown a syntax structure for performing decoding in accordance with AF4 mode or AF6 mode based on at least one of condition B or condition C. [Figure 20] As embodiment (4-1) to which the present invention is applied, there is shown a flowchart for adaptively determining an optimal coding mode from motion vector prediction modes including AF4 mode or AF6 mode based on the coding mode of an adjacent block. [Figure 21] As embodiment (4-2) to which the present invention is applied, there is shown a flowchart for performing adaptive decoding in accordance with AF4 mode or AF6 mode based on the coding mode of an adjacent block. [Figure 22] As embodiment (4-3) to which the present invention is applied, there is shown a syntax structure for performing decoding in accordance with AF4 mode or AF6 mode based on the coding mode of an adjacent block. [Figure 23] As embodiment (5-1) to which the present invention is applied, there is shown a flowchart for adaptively determining an optimal coding mode from motion vector prediction modes including AF4 mode or AF6 mode based on at least one of condition A, condition B or condition C. [Figure 24]As an embodiment (5-2) to which the present invention is applied, a flowchart is shown which adaptively decodes according to AF4 mode or AF6 mode based on at least one of condition A, condition B, or condition C. [Figure 25] As an embodiment (5-3) to which the present invention is applied, a syntax structure is shown that performs decoding according to AF4 mode or AF6 mode based on at least one of condition A, condition B, or condition C. [Figure 26] As an embodiment (6-1) to which the present invention is applied, a flowchart is shown for adaptively determining the optimal coding mode from among motion vector prediction modes, including AF4 mode or AF6 mode, based on condition A or at least one of the coding modes of adjacent blocks. [Figure 27] As an embodiment (6-2) to which the present invention is applied, a flowchart is shown which adaptively decodes according to AF4 mode or AF6 mode based on condition A or at least one of the coding modes of adjacent blocks. [Figure 28] As an embodiment (6-3) to which the present invention is applied, a syntax structure is shown that decodes according to AF4 mode or AF6 mode based on condition A or at least one of the coding modes of adjacent blocks. [Figure 29] An embodiment to which the present invention is applied is shown, illustrating a flowchart for generating a motion vector predictor based on at least one of AF4 mode or AF6 mode. [Figure 30] As an embodiment to which the present invention is applied, a flowchart for generating a motion vector predictor based on AF4_flag and AF6_flag is shown. [Figure 31] As an embodiment to which the present invention is applied, a flowchart is shown which adaptively decodes according to AF4 mode or AF6 mode based on whether or not adjacent blocks are coded in AF mode. [Figure 32] As an embodiment to which the present invention is applied, a syntax for adaptively decoding based on AF4_flag and AF6_flag is shown. [Figure 33] As an embodiment to which the present invention is applied, a syntax is shown that adaptively decodes according to AF4 mode or AF6 mode based on whether an adjacent block is coded in AF mode or not. [Figure 34] This shows a video coding system to which the present invention is applied. [Figure 35] This document illustrates a content streaming system to which the present invention is applied. [Best Mode for Carrying Out the Invention]

[0016] The present invention provides a method for decoding a video signal including a current block based on an affine motion prediction mode (AF mode), comprising the steps of: checking whether the AF mode is applied to the current block, wherein the AF mode represents a motion prediction mode that utilizes an affine motion model; checking whether AF4 mode is used if the AF mode is applied to the current block, wherein the AF4 mode represents a mode that predicts motion vectors using four parameters constituting the affine motion model; generating a motion vector predictor using the four parameters if AF4 mode is used, and generating a motion vector predictor using six parameters constituting the affine motion model if AF4 mode is not used; and obtaining the motion vector of the current block based on the motion vector predictor.

[0017] In the present invention, the method further includes the step of obtaining an affine flag from the video signal, wherein the affine flag indicates whether or not the AF mode is applied to the current block, and whether or not the AF mode is applied to the current block is determined based on the affine flag.

[0018] In the present invention, the method further includes the step of obtaining an affine parameter flag from the video signal when the AF mode is applied to the current block by the affine flag, wherein the affine parameter flag indicates whether the motion vector predictor is generated using the four parameters or the six parameters.

[0019] In the present invention, the affine flag and the affine parameter flag are defined at at least one level of slice, maximum coding unit, coding unit, or prediction unit.

[0020] In the present invention, the method further includes a step of checking whether the size of the current block satisfies a previously set condition, wherein the previously set condition indicates whether at least one of the number of pixels in the current block, the width and / or height of the current block is greater than a previously set threshold, and if the size of the current block satisfies the previously set condition, a step of checking whether the AF mode is applied to the current block is performed.

[0021] In the present invention, if the size of the current block does not satisfy the previously set conditions, the current block is decoded based on a coding mode other than the AF mode.

[0022] In the present invention, the method further includes a step of checking whether the AF mode has been applied to an adjacent block if the AF mode has been applied to the current block, wherein if the AF mode has been applied to the adjacent block, a motion vector predictor is generated using the four parameters, and if the AF mode has not been applied to the adjacent block, the method includes a step of checking whether the AF4 mode is used.

[0023] The present invention provides an interpretation unit for decoding a video signal including a current block based on an affine motion prediction mode (AF mode), wherein the unit checks whether the AF mode is applied to the current block, checks whether the AF4 mode is used if the AF mode is applied to the current block, generates a motion vector predictor using four parameters if the AF4 mode is used, generates a motion vector predictor using six parameters constituting an affine motion model if the AF4 mode is not used, and obtains the motion vector of the current block based on the motion vector predictor, wherein the AF mode indicates a motion prediction mode that uses the affine motion model, and the AF4 mode indicates a mode that predicts a motion vector using four parameters constituting the affine motion model.

[0024] In the present invention, the apparatus further includes a parsing unit that parses an affine flag from the video signal, wherein the affine flag indicates whether or not the AF mode is applied to the current block, and whether or not the AF mode is applied to the current block is confirmed based on the affine flag.

[0025] In the present invention, the apparatus includes a parsing unit that acquires an affine parameter flag from the video signal when the AF mode is applied to the current block by the affine flag, wherein the affine parameter flag indicates whether the motion vector predictor is generated using the four parameters or the six parameters.

[0026] In the present invention, the apparatus includes an inter-prediction unit that checks whether the size of the current block satisfies a previously set condition, wherein the previously set condition indicates whether at least one of the number of pixels in the current block, the width and / or height of the current block is greater than a previously set threshold, and if the size of the current block satisfies the previously set condition, a step is performed to check whether the AF mode is applied to the current block.

[0027] In the present invention, the apparatus includes an inter-prediction unit that checks whether the AF mode has been applied to an adjacent block when the AF mode has been applied to the current block, and if the AF mode has been applied to the adjacent block, a motion vector predictor is generated using the four parameters, and if the AF mode has not been applied to the adjacent block, the apparatus performs a step of checking whether the AF4 mode is used. [Modes for carrying out the invention]

[0028] The configuration and operation of embodiments of the present invention will be described below with reference to the attached drawings. The configuration and operation of the present invention described in the drawings are described as one embodiment, and this does not limit the technical idea, core configuration, and operation of the present invention.

[0029] Furthermore, while the terminology used in this invention has been selected as widely used and general terms as possible, in certain cases, the applicant may use terms of their own choosing for explanation. In such cases, the meaning will be clearly described in the detailed explanation of the relevant section, so it is important to clarify that the explanation should not be based solely on the names of the terms used in the description of this invention, but rather the meaning of those terms should also be understood.

[0030] Furthermore, while the terminology used in this invention is general terminology selected to describe the invention, if other terms with similar meanings exist, they can be substituted for more appropriate analysis. For example, terms such as signal, data, sample, picture, frame, and block can be appropriately substituted and analyzed in each coding process. Similarly, terms such as partitioning, decomposition, splitting, and division can also be appropriately substituted and analyzed in each coding process.

[0031] Figure 1 shows a schematic block diagram of an encoder used to encode a video signal, as an embodiment to which the present invention is applied.

[0032] As shown in Figure 1, the encoder 100 is composed of an image splitting unit 110, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, a filtering unit 160, a decoded picture buffer (DPB) 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy encoding unit 190.

[0033] The image splitting unit 110 splits the input image (or picture, frame) input to the encoder 100 into one or more processing units. For example, the processing units may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In this case, the splitting is performed using at least one of the following methods: QT (Quad Tree), BT (Binary Tree), TT (Ternary Tree), or AT (Asymmetric Tree).

[0034] However, the terms used herein are for convenience in explaining the present invention, and the present invention is not limited to the definitions of those terms. Furthermore, for convenience in this specification, the term coding unit is used to refer to a unit used in the process of encoding or decoding a video signal, but the present invention is not limited thereto and can be appropriately analyzed depending on the content of the invention.

[0035] The encoder 100 generates a residual signal by subtracting the prediction signal output from the inter-prediction unit 180 or intra-prediction unit 185 from the input image signal, and the generated residual signal is transmitted to the conversion unit 120.

[0036] The conversion unit 120 applies a conversion technique to the residual signal to generate a transform coefficient. The conversion process may be applied to pixel blocks of the same size and square shape, or to blocks of a variable size that are not square.

[0037] The quantization unit 130 quantizes the conversion coefficients and transmits them to the entropy encoding unit 190, which then entropy encodes the quantized signal and outputs it to a bitstream.

[0038] The quantized signal output from the quantization unit 130 may be used to generate a prediction signal. For example, the quantized signal can be reconstructed by applying inverse quantization and inverse transformation to the inverse quantization unit 140 and inverse transformation unit 150 within the loop. The reconstructed residual signal is added to the prediction signal output from the inter-prediction unit 180 or intra-prediction unit 185 to generate a reconstructed signal.

[0039] On the other hand, during the compression process described above, adjacent blocks may be quantized using different quantization parameters, potentially causing degradation where block boundaries become visible. This phenomenon is called blocking artifacts, and it is an important factor in evaluating image quality. A filtering process may be performed to reduce such degradation. Such a filtering process can eliminate blocking artifacts and improve image quality by reducing errors in the current picture.

[0040] The filtering unit 160 applies filtering to the restored signal and outputs it to the playback device or transmits it to the decoding picture buffer 170. The filtered signal transmitted to the decoding picture buffer 170 can be used as a reference picture in the inter-prediction unit 180. By using the filtered picture as a reference picture in the inter-screen prediction mode in this way, not only image quality but also encoding efficiency is improved.

[0041] The decoded picture buffer 170 stores the filtered picture for use as a reference picture in the interpretation unit 180.

[0042] The interpretation unit 180 performs temporal and / or spatial predictions by referring to the reconstructed picture to remove temporal and / or spatial overlaps. Here, the reference picture used for prediction is a transformed signal that has undergone quantization and dequantization in block units during encoding / decoding in the past, and therefore may contain blocking artifacts and ringing artifacts.

[0043] Therefore, the interpretation unit 180 can interpolate the signals between pixels in sub-pixel units by applying a low-pass filter to resolve such signal discontinuities and performance degradation due to quantization. Here, a sub-pixel refers to a virtual pixel generated by applying the interpolation filter, and an integer pixel refers to an actual pixel present in the restored picture. Linear interpolation, bilinear interpolation, Wiener filters, etc., may be applied as interpolation methods.

[0044] The interpolation filter is applied to the reconstructed picture to improve the accuracy of the prediction. For example, the interpretation unit 180 can generate interpolated pixels by applying the interpolation filter to integer pixels, and then use the interpolated block composed of these interpolated pixels as a prediction block to perform predictions.

[0045] The intra-prediction unit 185 can predict the current block by referring to samples surrounding the block currently to be encoded. The intra-prediction unit 185 performs the following process to perform intra-prediction. First, it prepares the reference samples necessary to generate the prediction signal. Then, it generates the prediction signal using the prepared reference samples. Subsequently, it encodes the prediction mode. Here, the reference samples are prepared by reference sample padding and / or reference sample filtering. Since the reference samples have gone through the prediction and reconstruction process, quantization errors may exist. Therefore, in order to reduce such errors, the reference sample filtering process is performed for each prediction mode used in intra-prediction.

[0046] The prediction signal generated by the inter-prediction unit 180 or the intra-prediction unit 185 is used to generate a restoration signal or to generate a residual signal.

[0047] Figure 2 shows a schematic block diagram of a decoder that decodes a video signal, as an embodiment to which the present invention is applied.

[0048] As shown in Figure 2, the decoder 200 is composed of a parsing unit (not shown), an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 230, a filtering unit 240, a decoded picture buffer (DPB) 250, an inter-prediction unit 260, an intra-prediction unit 265, and a reconstruction unit (not shown).

[0049] The decoder 200 receives the signal output from the encoder 100 in Figure 1 and parses or acquires the syntax elements via a parsing unit (not shown). The parsed or acquired signal is then entropy-decoded by the entropy-decode unit 210.

[0050] In the inverse quantization unit 220, the transformation coefficient is obtained from the entropy-decoded signal using the quantization step size information.

[0051] In the inverse conversion unit 230, the conversion coefficients are inversely converted to obtain a residual signal.

[0052] The reconstruction unit (not shown) generates a reconstructed signal by adding the acquired residual signal to the prediction signal output from the inter-prediction unit 260 or the intra-prediction unit 265.

[0053] The filtering unit 240 applies filtering to the reconstructed signal and outputs it to the playback device or transmits it to the decoding picture buffer unit 250. The filtered signal transmitted to the decoding picture buffer unit 250 can be used as a reference picture in the interpretation unit 260.

[0054] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the encoder 100 can also be similarly applied to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the decoder, respectively.

[0055] The reconstructed video signal output via the decoder 200 can be played back by a playback device.

[0056] Figure 3 is a diagram illustrating a QT (QuadTree, hereinafter referred to as "QT") block division structure as an embodiment to which the present invention can be applied.

[0057] In video coding, a single block can be divided using a Quad Tree (QT) base. A subblock divided by QT can then be recursively divided further using QT. Leaf blocks that cannot be further divided using QT can be divided using at least one of the following methods: Binary Tree (BT), Ternary Tree (TT), or Asymmetric Tree (AT). BT has two forms of division: horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). TT also has two forms of division: horizontal TT (2N×1 / 2N, 2N×N, 2N×1 / 2N) and vertical TT (1 / 2N×2N, N×2N, 1 / 2N×2N). AT has four forms of partitioning: horizontal-up AT (2N × 1 / 2N, 2N × 3 / 2N), horizontal-down AT (2N × 3 / 2N, 2N × 1 / 2N), vertical-left AT (1 / 2N × 2N, 3 / 2N × 2N), and vertical-right AT (3 / 2N × 2N, 1 / 2N × 2N). Each of BT, TT, and AT can be further partitioned recursively using BT, TT, and AT.

[0058] Figure 3 shows an example of QT partitioning. Block A can be divided into four subblocks (A0, A1, A2, A3) by QT. Subblock A1 can be further divided into four subblocks (B0, B1, B2, B3) by QT.

[0059] Figure 4 is a diagram illustrating a BT (Binary Tree, hereinafter referred to as "BT") block partitioning structure as an embodiment to which the present invention can be applied.

[0060] Figure 4 shows an example of BT partitioning. Block B3, which cannot be further partitioned by QT, can be partitioned by vertical BT(C0, C1) or horizontal BT(D0, D1). Each subblock, like block C0, can be recursively further partitioned in the form of horizontal BT(E0, E1) or vertical BT(F0, F1).

[0061] Figure 5 is a diagram illustrating a TT (Ternary Tree, hereinafter referred to as "TT") block partitioning structure as an embodiment to which the present invention can be applied.

[0062] Figure 5 shows an example of TT partitioning. Block B3, which cannot be further partitioned by QT, can be partitioned by vertical TT(C0, C1, C2) or horizontal TT(D0, D1, D2). Each subblock, like block C1, can be recursively further partitioned in the form of horizontal TT(E0, E1, E2) or vertical TT(F0, F1, F2).

[0063] Figure 6 is a diagram illustrating an AT (Asymmetric Tree, hereinafter referred to as "AT") block partitioning structure as an embodiment to which the present invention can be applied.

[0064] Figure 6 shows an example of AT partitioning. Block B3, which cannot be further partitioned by QT, can be partitioned by vertical AT(C0, C1) or horizontal AT(D0, D1). Each subblock, like block C1, can be further partitioned recursively in the form of horizontal AT(E0, E1) or vertical TT(F0, F1).

[0065] On the other hand, BT, TT, and AT partitioning can all be used together for partitioning. For example, a subblock partitioned by BT can be partitioned by TT or AT. Similarly, a subblock partitioned by TT can be partitioned by BT or AT. A subblock partitioned by AT can be partitioned by BT or TT. For example, after horizontal BT partitioning, each subblock can be partitioned by vertical BT, or after vertical BT partitioning, each subblock can be partitioned by horizontal BT. The two partitioning methods described above differ in their order of partitioning, but the final partitioned shape is the same.

[0066] Furthermore, once a block is divided, the order in which the block is searched can be defined in various ways. Generally, searching a block from left to right and from top to bottom means either the order in which to decide whether or not to perform additional block divisions on each divided subblock, or, if the block cannot be divided any further, the encoding order of each subblock, or the search order when referencing information from other adjacent blocks within a subblock.

[0067] Figure 7 is a diagram illustrating an interprediction mode as an embodiment to which the present invention is applied.

[0068] Interpretation mode

[0069] In the interpretation mode to which the invention is applied, merge mode, AMVP (Advanced Motion Vector Prediction) mode, or affine prediction mode (hereinafter referred to as "AF mode") is used to reduce the amount of motion information.

[0070] 1) Merge mode

[0071] Merge mode refers to a method of deriving motion parameters (or information) from spatially or temporally adjacent blocks.

[0072] In merge mode, the available candidate set consists of spatially adjacent candidates, temporal candidates, and generated candidates.

[0073] As shown in Figure 7(a), the availability of each spatial candidate block is determined in the order of {A1, B1, B0, A0, B2}. At that time, if a candidate block is encoded in intra-prediction mode and no motion information exists, or if a candidate block is currently located outside the picture (or slice), that candidate block is unavailable.

[0074] After determining the effectiveness of spatial candidates, spatial merge candidates can be constructed by removing unnecessary candidate blocks from the candidate blocks of the current processing block. For example, if the candidate block of the current prediction block is the first prediction block within the same coding block, that candidate block can be removed, as can candidate blocks that have the same motion information.

[0075] Once the spatial merge candidate construction is complete, the temporal merge candidate construction process proceeds in the order of {T0, T1}.

[0076] In the temporal candidate configuration, if the right bottom block (T0) of the collocated blocks in the reference picture is available, that block is configured as a temporal merge candidate. A collocated block refers to a block that is located in the position corresponding to the currently processed block in the selected reference picture. If this is not the case, the block located in the center of the collocated blocks (T1) is configured as a temporal merge candidate.

[0077] The maximum number of merge candidates can be specified in the slice header. If the number of merge candidates is greater than the maximum, the remaining number of spatial and temporal candidates is maintained. Otherwise, the number of merge candidates is increased by combining the candidates added up to date to generate additional merge candidates (i.e., combined bi-predictive merging candidates) until the number of candidates reaches the maximum.

[0078] In the encoder, a merge candidate list is constructed using the method described above, and motion estimation is performed to signal the decoder with the candidate block information selected from the merge candidate list as a merge index (for example, merge_idx[x0][y0]'). Figure 7(b) illustrates the case where block B1 is selected from the merge candidate list, in which case "Index 1" is signaled to the decoder as the merge index.

[0079] In the decoder, similar to the encoder, a merge candidate list is constructed, and motion information for the current block is derived from the motion information of the candidate block corresponding to the merge index received from the encoder in the merge candidate list. Then, the decoder generates a predicted block for the currently processed block based on the derived motion information.

[0080] 2) AMVP (Advanced Motion Vector Prediction) mode

[0081] AMVP mode refers to a method for deriving motion vector predictions from surrounding blocks. Therefore, horizontal and vertical motion vector difference (MVD), reference index, and interpretation mode are signaled to the decoder. The horizontal and vertical motion vector values ​​are calculated using the derived motion vector predictions and the motion vector difference (MVD) provided by the encoder.

[0082] In other words, the encoder constructs a list of motion vector prediction candidate values ​​and signals the decoder with a motion reference flag (i.e., candidate block information) (e.g., 'mvp_lX_flag[x0][y0]') selected from the motion vector prediction candidate value list by performing motion estimation. The decoder constructs a list of motion vector prediction candidate values ​​in the same way as the encoder and derives the motion vector prediction value of the currently processed block using the motion information of the candidate block indicated by the motion reference flag received from the encoder in the motion vector prediction candidate value list. The decoder then obtains the motion vector value for the currently processed block using the derived motion vector prediction value and the motion vector difference value sent from the encoder. Finally, the decoder generates a predicted block for the currently processed block based on the derived motion information (i.e., motion compensation).

[0083] In AMVP mode, two spatial motion candidates are selected from the five available candidates mentioned above, as shown in Figure 7. The first spatial motion candidate is selected from the {A0, A1} set located on the left, and the second spatial motion candidate is selected from the {B0, B1, B2} set located higher up. Here, if the reference index of adjacent candidate blocks is not the same as the currently predicted block, the motion vector is scaled.

[0084] If two candidates are selected from the search results for spatial motion candidates, the candidate construction is completed. If there are fewer than two candidates, temporal motion candidates are added.

[0085] The decoder (e.g., the interpretation unit) decodes motion parameters for the processing block (e.g., the prediction unit).

[0086] For example, if a processing block utilizes merge mode, the decoder can decode the merge index signaled by the encoder. Then, it can derive the motion parameters of the current processing block from the motion parameters of the candidate blocks indicated in the merge index.

[0087] Furthermore, if AMVP mode is applied to the processing block, the decoder can decode the horizontal and vertical motion vector difference (MVD), reference index, and interpretation mode signaled from the encoder. Then, it can derive the motion vector prediction value from the motion parameters of the candidate block indicated by the motion reference flag, and use the motion vector prediction value and the received motion vector difference value to derive the motion vector value of the current processing block.

[0088] The decoder uses the decoded motion parameters (or information) to perform motion compensation for the prediction unit.

[0089] In other words, the encoder / decoder performs motion compensation, using the decoded motion parameters to predict the image of the current unit from a previously decoded picture.

[0090] 3) AF Mode (Affine Mode)

[0091] The AF mode refers to a motion prediction mode that uses an affine motion model, and may include at least one of an affine merge mode or an affine inter mode. The affine inter mode may include at least one of an AF4 mode or an AF6 mode. Here, the AF4 mode refers to a four-parameter affine prediction mode that uses four parameters, and the AF6 mode refers to a six-parameter affine prediction mode that uses six parameters.

[0092] However, in this invention, for the sake of explanation, we will refer to them as AF4 mode or AF6 mode, but this does not necessarily need to be defined as a separate prediction mode. The AF4 mode or AF6 mode can be understood simply as being distinguished by whether or not it uses four parameters or six parameters.

[0093] The aforementioned AF modes will be explained in more detail in Figures 8 to 10.

[0094] Figure 8 is a diagram illustrating an affine motion model as an embodiment to which the present invention is applied.

[0095] Common image coding techniques use translation motion models to represent the movement of coding blocks. Here, a translation motion model describes a translation-based prediction method for each block; that is, the motion information of a coding block is represented using a single motion vector. However, the optimal motion vector for each pixel within an actual coding block may differ. If the optimal motion vector could be determined for each pixel or sub-block with less information, coding efficiency could be improved.

[0096] Therefore, in order to improve the performance of interpretation, the present invention proposes not only a translated block-based prediction method, but also an interpretation-based image processing method that reflects diverse motions in the image.

[0097] Furthermore, the present invention proposes an affine motion prediction method that performs encoding / decoding using an affine motion model. The affine motion model represents a prediction method that induces motion vectors on a pixel-by-pixel or sub-block-by-subblock basis using motion vectors of control points. In this specification, the affine motion prediction mode that uses the affine motion model is referred to as AF mode (Affine Mode).

[0098] Furthermore, the present invention provides a method for adaptively performing affine prediction based on block size.

[0099] Furthermore, the present invention provides a method for adaptively performing affine prediction based on whether adjacent blocks have been coded by affine prediction.

[0100] Furthermore, the present invention provides a method for adaptively determining (or selecting) the optimal coding mode based on at least one of AF4 mode or AF6 mode. Here, AF4 mode refers to a four-parameter affine prediction mode that utilizes four parameters, and AF6 mode refers to a six-parameter affine prediction mode that utilizes six parameters.

[0101] As shown in Figure 8, various methods are used to represent image distortion as motion information, and in particular, the affine motion model can represent the four types of motion shown in Figure 8.

[0102] For example, affine motion models can model not only image translation, scaling, rotation, and shear, but also any induced image distortion.

[0103] Affine motion models can be represented in various ways, but this invention proposes a method that utilizes motion information at specific reference points (or reference pixels / samples) of a block to display (or identify) distortion, and uses this to perform interpretation. Here, the reference points are called control points (CP) (or control pixels, control samples), and the motion vectors at such reference points are called control point motion vectors (CPMV). The degree of distortion that can be represented changes depending on the number of such control points.

[0104] The affine motion model is expressed using six parameters (a, b, c, d, e, f) as shown in Equation 1 below.

[0105]

number

[0106] Here, (x, y) indicates the position of the top-left pixel of the coding block. And v x and v y These represent the motion vectors at (x, y), respectively.

[0107] Figure 9 illustrates an affine motion prediction method using a control point motion vector as an embodiment to which the present invention is applied.

[0108] As shown in Figure 9(a), the upper left control point (CP0) 902 (hereinafter referred to as the first control point), the upper right control point (CP1) 903 (hereinafter referred to as the second control point), and the lower left control point (CP2) 904 (hereinafter referred to as the third control point) of block 901 can each have independent motion information. These are represented as CP0, CP1, and CP2, respectively. However, this is merely one embodiment of the present invention, and the present invention is not limited thereto. For example, various control points can be defined, such as a lower right control point, a center control point, and other position-specific control points for subblocks.

[0109] In one embodiment of the present invention, at least one of the first to third control points may be a pixel included in the current block. Alternatively, as another example, at least one of the first to third control points may be a pixel adjacent to the current block that is not included in the current block.

[0110] By utilizing the motion information of one or more of the aforementioned control points, pixel-by-pixel or sub-block-by-subblock motion information for block 901 can be derived.

[0111] For example, an affine motion model using the motion vectors of the upper left control point 902, the upper right control point 903, and the lower left control point 904 of block 901 is defined as shown in Equation 2 below.

[0112]

number

[0113] Here, JPEG0007923860000003.jpg19150 shows the motion vector of the upper left control point 902. JPEG0007923860000004.jpg24150 is the motion vector of the upper right control point 903. When JPEG0007923860000005.jpg25150 is the motion vector of the lower left control point 904, JPEG0007923860000006.jpg21150, JPEG0007923860000007.jpg20150, It can be defined as JPEG0007923860000008.jpg20150. Then, in equation 2, w currently represents the width of block 901, and h currently represents the height of block 901. And, JPEG0007923860000009.jpg19150 shows the motion vector at the {x, y} position.

[0114] In this invention, an affine motion model can be defined that represents three types of motion that can be expressed by an affine motion model: translation, scaling, and rotation. In this specification, this will be referred to as a simplified affine motion model (or similarity affine motion model).

[0115] The aforementioned simplified affine motion model can be expressed using four parameters (a, b, c, d) as shown in the following equation 3.

[0116]

number

[0117] Here, {v x , v yThe vectors {x, y} represent the motion vectors at the {x, y} positions, respectively. An affine motion model that utilizes these four parameters is called AF4. The present invention is not limited to this; when six parameters are used, it is called AF6, and the embodiments described above can be applied identically.

[0118] As shown in Figure 9(b), The motion vector of control point 1001 on the upper left side of the current block in JPEG0007923860000011.jpg23150, When JPEG0007923860000012.jpg20150 is the motion vector of the upper right control point 1002, JPEG0007923860000013.jpg23150, It can be defined as JPEG0007923860000014.jpg22150. Here, the affine motion model of AF4 can be defined as shown in the following equation 4.

[0119]

number

[0120] In equation 4, w represents the current block width, and h represents the current block height. JPEG0007923860000016.jpg19150 shows the motion vectors at the {x,y} positions.

[0121] The encoder or decoder can determine (or induce) the motion vector of each pixel position using the motion vectors of the control points (for example, the motion vectors of the upper left control point 1001 and the upper right control point 1002).

[0122] In the present invention, the set of motion vectors determined by affine motion prediction can be defined as an affine motion vector field. The affine motion vector field is determined using at least one of the formulas 1 to 4.

[0123] During the encoding / decoding process, motion vectors based on affine motion prediction can be determined on a pixel-by-pixel basis or on a predefined (or pre-configured) block (or subblock) basis. For example, if determined on a pixel basis, the motion vector is derived based on each pixel within a block; if determined on a subblock basis, the motion vector is derived based on each subblock within the current block. As another example, if determined on a subblock basis, the motion vector of that subblock is derived based on the upper-left pixel or the center pixel.

[0124] In the following description of the present invention, for the sake of convenience, the explanation will mainly focus on the case where motion vectors are determined in 4x4 block units by affine motion prediction. However, the present invention is not limited to this, and can be applied in pixel units or block units of other sizes.

[0125] On the other hand, let's assume that the current block size is 16 × 16, as shown in Figure 9(b). The encoder or decoder uses the motion vectors of the upper left control point 1001 and the upper right control point 1002 of the current block to determine the motion vector in 4 × 4 sub-block units. Then, the motion vector of each sub-block is determined based on the central pixel value of that sub-block.

[0126] In Figure 9(b), the arrows displayed in the center of each subblock indicate the motion vectors obtained by the affine motion model.

[0127] Affine motion prediction can be used in affine merge mode (hereinafter referred to as "AF merge mode") and affine inter mode (hereinafter referred to as "AF inter mode"). AF merge mode is a method that does not encode the motion vector difference, similar to skip mode or merge mode, but instead induces two control point motion vectors and then encodes or decodes them. AF inter mode is a method that determines the control point motion vector predictor and the control point motion vector, and then encodes or decodes the control point motion vector difference (CPMVD) corresponding to the difference between them. In this case, two control point motion vector difference values ​​are transmitted in AF4 mode, and three control point motion vector difference values ​​are transmitted in AF6 mode.

[0128] Here, AF4 mode has the advantage of being able to represent control point motion vectors (CPMVs) with fewer bits because it transmits fewer motion vector difference values ​​compared to AF6 mode, while AF6 mode has the advantage of being able to reduce the number of bits used for residual coding because it transmits three CPMVDs, enabling superior predictor generation.

[0129] Therefore, the present invention proposes a method for considering both AF4 mode and AF6 mode (or simultaneously) in the AF intermode.

[0130] Figure 10 is a flowchart illustrating the process of processing a video signal including the current block using an affine prediction mode (hereinafter referred to as "AF mode") as an embodiment to which the present invention is applied.

[0131] This invention provides a method for processing a video signal, including the current block, using AF mode.

[0132] First, the video signal processing device generates a candidate list of motion vector pairs using the motion vectors of pixels or blocks adjacent to at least two control points of the current block (S1010). Here, the control points represent the corner pixels of the current block, and the motion vector pair represents the motion vectors of the upper left corner pixel and the upper right corner pixel of the current block.

[0133] In one embodiment, the control point includes at least two of the upper-left corner pixel, upper-right corner pixel, lower-left corner pixel, or lower-right corner pixel of the current block, and the candidate list consists of pixels or blocks adjacent to the upper-left corner pixel, the upper-right corner pixel, and the lower-left corner pixel.

[0134] In one embodiment, the candidate list can be generated based on the motion vectors of the diagonally adjacent pixels (A), the upper adjacent pixels (B), and the left adjacent pixels (C) of the upper left corner pixel; the motion vectors of the upper adjacent pixels (D) and the diagonally adjacent pixels (E) of the upper right corner pixel; and the motion vectors of the left adjacent pixel (F) and the diagonally adjacent pixels (G) of the lower left corner pixel.

[0135] In one embodiment, the method may further include the step of adding an AMVP candidate list to the candidate list if the number of motion vector pairs in the candidate list is less than two.

[0136] In one embodiment, when the current block is of size N × 4, the control point motion vector of the current block is determined to be a motion vector induced based on the central positions of the left subblock and the right subblock within the current block, and when the current block is of size 4 × N, the control point motion vector of the current block is determined to be a motion vector induced based on the central positions of the upper subblock and the lower subblock within the current block.

[0137] In one embodiment, when the current block is of size N×4, the control point motion vector of the left subblock within the current block is determined by the average value of the first control point motion vector and the third control point motion vector, and the control point motion vector of the right subblock is determined by the average value of the second control point motion vector and the fourth control point motion vector. When the current block is of size 4×N, the control point motion vector of the upper subblock within the current block is determined by the average value of the first control point motion vector and the second control point motion vector, and the control point motion vector of the lower subblock is determined by the average value of the third control point motion vector and the fourth control point motion vector.

[0138] In another embodiment, the method can signal a predictive mode or flag information indicating whether or not the AF mode is performed.

[0139] In this case, the video signal processing device can receive the prediction mode or flag information, perform the AF mode according to the prediction mode or flag information, and guide the motion vector according to the AF mode. Here, the AF mode is characterized by indicating a mode in which the motion vector is guided on a pixel or subblock basis using the control point motion vector of the current block.

[0140] On the other hand, the video signal processing device determines a final candidate list of a predetermined number of motion vector pairs based on the divergence value of the motion vector pair (S1020). Here, the final candidate list is determined in ascending order of divergence value, where the divergence value represents a value indicating the similarity of the directions of the motion vectors.

[0141] The video signal processing device determines the control point motion vector of the current block based on the rate-distortion cost from the final candidate list (S1030).

[0142] The video signal processing device generates a motion vector predictor for the current block based on the control point motion vector (S1040).

[0143] Figure 11 shows a flowchart (1-1) of an embodiment to which the present invention is applied, which adaptively determines the optimal coding mode based on at least one of AF4 mode or AF6 mode.

[0144] The video signal processing device makes a prediction based on at least one of skip mode, merge mode, or intermode (S1110). Here, merge mode may include not only a general merge mode but also the aforementioned AF merge mode, and intermode may include not only a general intermode but also the aforementioned AF intermode.

[0145] The video signal processing device performs motion vector prediction based on at least one of AF4 mode or AF6 mode (S1120). Here, the order of steps S1110 and S1120 is not restricted.

[0146] The video signal processing device compares the results of step S1120 to determine the optimal coding mode among the modes (S1130). Here, the results of step S1120 are compared based on the rate-distortion cost.

[0147] Thereafter, the video signal processing device generates a motion vector predictor for the current block based on the optimal coding mode, and subtracts the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.

[0148] From this point forward, the encoding / decoding processes described in Figures 1 and 2 above will be applied identically.

[0149] Figure 12 shows a flowchart illustrating an embodiment (1-2) to which the present invention is applied, in which adaptive decoding is performed based on AF4 mode or AF6 mode.

[0150] The decoder receives a bitstream (S1210). The bitstream contains information about the coding mode of the current block in the video signal.

[0151] The decoder checks whether the coding mode of the current block is AF mode (S1220). Here, AF mode means an affine motion prediction mode that utilizes an affine motion model, and may include, for example, at least one of affine merge mode or affine inter mode, and the affine inter mode may include at least one of AF4 mode or AF6 mode.

[0152] Here, step S1220 is confirmed by an affine flag indicating whether or not AF mode is performed. For example, the affine flag can be expressed as affine_flag. If affine_flag=1, it indicates that AF mode is performed for the current block, and if affine_flag=0, it indicates that AF mode is not performed for the current block.

[0153] If AF mode is not currently applied to a block, the decoder decodes (i.e., predicts motion vectors) using a coding mode other than AF mode (S1230). For example, skip mode, merge mode, or inter-mode may be used.

[0154] If AF mode is applied to the current block, the decoder checks whether AF4 mode is applied to the current block (S1240).

[0155] Here, step S1240 can be confirmed by an affine parameter flag indicating whether or not AF4 mode is performed (or whether or not affine motion prediction is performed using four parameters). For example, the affine parameter flag is expressed as affine_param_flag. If affine_param_flag=0, motion vector prediction is performed by AF4 mode (S1250), and if affine_param_flag=1, motion vector prediction is performed by AF6 mode (S1260), but the present invention is not limited thereto.

[0156] For example, the affine parameter flag includes at least one of AF4_flag and AF6_flag.

[0157] The AF4_flag flag indicates whether or not AF4 mode is currently being applied to the block. If AF4_flag=1, AF4 mode is currently being applied to the block; if AF4_flag=0, AF4 mode is not being applied to the block. Here, the application of AF4 mode means that motion vector prediction is performed using an affine motion model expressed by four parameters.

[0158] The AF6_flag flag indicates whether or not AF6 mode is being applied to the current block. If AF6_flag=1, AF6 mode is being applied to the current block; if AF4_flag=0, AF6 mode is not being applied to the current block. Here, the application of AF6 mode means that motion vector prediction is performed using an affine motion model expressed by four parameters.

[0159] The affine flag and the affine parameter flag can be defined at at least one level of slice, maximum coding unit, coding unit, or prediction unit.

[0160] For example, at least one of AF_flag, AF4_flag, and AF6_flag may be defined at the slice level, and may also be defined at the block level or the prediction unit level.

[0161] Figure 13 shows a syntax structure that performs decoding based on AF4 mode or AF6 mode as an embodiment (1-3) to which the present invention is applied.

[0162] The decoder retrieves the merge_flag to check whether merge mode is applied to the current block (S1310).

[0163] If merge mode is not applied to the current block, the decoder can obtain affine_flag (S1320). Here, affine_flag indicates whether or not AF mode is performed.

[0164] If affine_flag=1, that is, if AF mode is performed for the current block, the decoder can obtain affine_param_flag (S1330). Here, affine_param_flag indicates whether AF4 mode is performed (or whether affine motion prediction is performed using four parameters).

[0165] When affine_param_flag=0, that is, when motion vector prediction is performed by AF4 mode, the decoder can obtain two motion vector difference values, mvd_CP0 and mvd_CP1 (S1340). Here, mvd_CP0 represents the motion vector difference value with respect to control point 0, and mvd_CP1 represents the motion vector difference value with respect to control point 1.

[0166] Then, when affine_param_flag=1, that is, when motion vector prediction is performed in AF6 mode, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S1350).

[0167] Figure 14 shows a flowchart (2-1) of an embodiment to which the present invention is applied, which adaptively determines the optimal coding mode from among motion vector prediction modes, including AF4 mode or AF6 mode, based on condition A.

[0168] The encoder makes predictions based on at least one of skip mode, merge mode, or intermode (S1410).

[0169] The encoder checks whether condition A is satisfied for the current block in order to determine the optimal coding mode for motion vector prediction (S1420).

[0170] Here, condition A refers to a condition on the block size. For example, the embodiment shown in Table 1 can be applied.

[0171] [Table 1]

[0172] In Example 1 of Table 1, condition A indicates whether the number of pixels (pixNum) of the current block is greater than the threshold (TH1). Here, the threshold has values ​​such as 64, 128, 256, 512, 1024, ..., for example, TH1=64 means that the block size is 4×16, 8×8, or 16×4, and TH1=128 means that the block size is 32×4, 16×8, 8×16, or 4×32.

[0173] In Example 2, this indicates whether both the width and height of the current block are greater than the threshold (TH1).

[0174] In Example 3, this indicates whether the width of the current block is greater than the threshold (TH1), or whether the height of the current block is greater than the threshold (TH1).

[0175] If condition A is satisfied, the encoder performs motion vector prediction based on at least one of AF4 mode or AF6 mode (S1430).

[0176] The encoder compares the results of steps S1410 and S1430 to determine the optimal coding mode from among the motion vector prediction modes, including AF4 mode or AF6 mode (S1440).

[0177] On the other hand, if condition A is not satisfied, the encoder determines the optimal coding mode from among the modes other than AF mode (S1440).

[0178] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and obtain a motion vector difference value by subtracting the motion vector predictor from the motion vector of the current block.

[0179] From this point forward, the encoding / decoding processes described in Figures 1 and 2 can be applied identically.

[0180] Figure 15 shows a flowchart illustrating an embodiment (2-2) to which the present invention is applied, in which adaptive decoding is performed using AF4 mode or AF6 mode based on condition A.

[0181] The decoder receives a bitstream (S1510). The bitstream includes a video signal and information regarding the coding mode of the current block.

[0182] The decoder checks whether condition A is satisfied for the current block in order to determine the optimal coding mode for motion vector prediction (S1520). Here, condition A refers to a condition on the block size. For example, the embodiments in Table 1 can be applied.

[0183] If condition A is satisfied, the decoder checks whether the coding mode of the current block is AF mode (S1530). Here, AF mode means an affine motion prediction mode using an affine motion model, and embodiments described herein can be applied.

[0184] Here, step S1530 can be confirmed by an affine flag indicating whether or not AF mode is performed. For example, the affine flag is expressed as affine_flag. If affine_flag=1, it indicates that AF mode is performed for the current block, and if affine_flag=0, it indicates that AF mode is not performed for the current block.

[0185] If condition A is not satisfied, or if AF mode is not performed on the current block, the decoder may decode (i.e., predict motion vectors) using a coding mode other than AF mode (S1540). For example, skip mode, merge mode, or inter mode may be used.

[0186] When AF mode is applied to the current block, the decoder checks whether AF4 mode is applied to the current block (S1550).

[0187] Here, step S1550 can be confirmed by an affine parameter flag that indicates whether AF4 mode is performed (or whether affine motion prediction is performed using four parameters). For example, the affine parameter flag can be expressed as affine_param_flag. If affine_param_flag=0, motion vector prediction is performed by AF4 mode (S1560), and if affine_param_flag=1, motion vector prediction is performed by AF6 mode (S1570), but the present invention is not limited thereto.

[0188] Figure 16 shows a syntax structure that performs decoding in AF4 mode or AF6 mode based on condition A, as an embodiment (2-3) to which the present invention is applied.

[0189] The decoder retrieves the merge_flag to check whether merge mode is applied to the current block (S1610).

[0190] If merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S1620). Here, condition A refers to a condition on the block size. For example, the embodiments in Table 1 can be applied.

[0191] If condition A is satisfied, the decoder can obtain affine_flag (S1620). Here, affine_flag indicates whether or not AF mode is performed.

[0192] If affine_flag=1, that is, if AF mode is performed on the current block, the decoder can obtain affine_param_flag (S1630). Here, affine_param_flag indicates whether AF4 mode is performed (or whether affine motion prediction is performed using four parameters).

[0193] When affine_param_flag=0, that is, when motion vector prediction is performed by AF4 mode, the decoder can obtain two motion vector difference values, mvd_CP0 and mvd_CP1 (S1640). Here, mvd_CP0 represents the motion vector difference value with respect to control point 0, and mvd_CP1 represents the motion vector difference value with respect to control point 1.

[0194] Then, when affine_param_flag=1, that is, when motion vector prediction is performed by AF6 mode, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S1650).

[0195] Figure 17 shows a flowchart (3-1) of an embodiment to which the present invention is applied, which adaptively determines the optimal coding mode from among motion vector prediction modes, including AF4 mode or AF6 mode, based on at least one of condition B or condition C.

[0196] The present invention provides a method for adaptively selecting between AF4 mode and AF6 mode based on the size of the current block.

[0197] For example, AF6 mode transmits one additional motion vector difference value compared to AF4 mode, making it effective for relatively large blocks. Therefore, if the current block size is smaller than (or equal to) the already set size, encoding is performed considering only AF4 mode; if the current block size is larger than or equal to (or greater than) the already set size, encoding is performed considering only AF6 mode.

[0198] On the other hand, in areas where it is not clearly determined that only one of AF4 mode or AF6 mode is superior, both AF4 mode and AF6 mode can be considered, and only the optimal mode can be signaled.

[0199] As shown in Figure 17, the encoder makes predictions based on at least one of skip mode, merge mode, or inter mode (S1710).

[0200] The encoder checks whether condition B is satisfied for the current block (S1720). Here, condition B refers to a condition on the block size. For example, the embodiment shown in Table 2 below can be applied.

[0201] [Table 2]

[0202] In Example 1 of Table 2, condition B indicates whether the number of pixels (pixNum) of the current block is less than the threshold (TH2). Here, the threshold has values ​​such as 64, 128, 256, 512, 1024, ..., for example, TH2=64 means that the block size is 4×16, 8×8, or 16×4, and TH2=128 means that the block size is 32×4, 16×8, 8×16, or 4×32.

[0203] In Example 2, condition B indicates whether both the width and height of the current block are less than the threshold (TH2).

[0204] In Example 3, condition B indicates whether the current block's width is less than the threshold (TH2) or whether the current block's height is less than the threshold (TH2).

[0205] If condition B is satisfied, the encoder performs motion vector prediction based on the AF4 mode (S1730).

[0206] If condition B is not satisfied, the encoder checks whether condition C is satisfied for the current block (S1740). Here, condition C refers to a condition on the block size. For example, the embodiment shown in Table 3 can be applied.

[0207] [Table 3]

[0208] In Example 1 of Table 3, condition A indicates whether the number of pixels (pixNum) of the current block is greater than or equal to the threshold (TH3). Here, the threshold has values ​​such as 64, 128, 256, 512, 1024, ..., for example, TH3=64 means that the block size is 4×16, 8×8, or 16×4, and TH3=128 means that the block size is 32×4, 16×8, 8×16, or 4×32.

[0209] In Example 2, this indicates whether the width and height of the current block are both greater than or equal to the threshold (TH3).

[0210] In Example 3, this indicates whether the width of the current block is greater than or equal to the threshold (TH1), or whether the height of the current block is greater than or equal to the threshold (TH1).

[0211] If the above condition C is satisfied, the encoder performs motion vector prediction based on the AF6 mode (S1760).

[0212] If condition C is not satisfied, the encoder can perform motion vector prediction based on AF4 mode and AF6 mode (S1750).

[0213] On the other hand, in the conditions B and C, the thresholds (TH2) and (TH3) can be determined to satisfy the following equation 5.

[0214] [Number 5] TH_2 ≤ TH_3

[0215] The encoder compares the results of steps S1710, S1730, S1750, and S1760 to determine the optimal coding mode from among the motion vector prediction modes, including AF4 mode or AF6 mode (S1770).

[0216] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and obtain a motion vector difference value by subtracting the motion vector predictor from the motion vector of the current block.

[0217] From this point forward, the encoding / decoding processes described in Figures 1 and 2 can be applied identically.

[0218] Figure 18 shows a flowchart illustrating an embodiment (3-2) to which the present invention is applied, in which adaptive decoding is performed by AF4 mode or AF6 mode based on at least one of condition B or condition C.

[0219] The decoder checks whether the current coding mode of the block is AF mode (S1810). Here, AF mode means an affine motion prediction mode using an affine motion model, and embodiments described herein can be applied, so redundant explanations are omitted.

[0220] When AF mode is performed on the current block, the decoder checks whether condition B is satisfied for the current block (S1820). Here, condition B refers to a condition on the block size. For example, the embodiments in Table 2 can be applied, and redundant explanations are omitted.

[0221] If condition B is satisfied, the decoder performs motion vector prediction based on the AF4 mode (S1830).

[0222] If condition B is not satisfied, the decoder checks whether condition C is satisfied for the current block (S1840). Here, condition C refers to a condition for the block size. For example, the embodiments in Table 3 can be applied, and redundant explanations are omitted.

[0223] On the other hand, in the conditions B and C, the thresholds (TH2) and (TH3) can be determined to satisfy the formula 5.

[0224] If condition C is satisfied, the decoder performs motion vector prediction based on the AF6 mode (S1860).

[0225] If condition C is not satisfied, the decoder checks whether AF4 mode is applied to the current block (S1850).

[0226] Here, step S1850 can be confirmed by an affine parameter flag indicating whether or not AF4 mode is performed (or whether or not affine motion prediction is performed using the four parameters).

[0227] For example, the affine parameter flag can be represented by affine_param_flag. When affine_param_flag=0, motion vector prediction is performed by AF4 mode (S1830), and when affine_param_flag=1, motion vector prediction is performed by AF6 mode (S1860), but the present invention is not limited thereto.

[0228] On the other hand, if AF mode is not performed for the current block, the decoder performs decoding (i.e., motion vector prediction) using a coding mode other than AF mode (S1870). For example, skip mode, merge mode, or inter mode can be used.

[0229] Figure 19 shows a syntax structure that performs decoding by AF4 mode or AF6 mode based on at least one of condition B or condition C, as an embodiment (3-3) to which the present invention is applied.

[0230] The decoder retrieves the merge_flag to check whether merge mode is applied to the current block (S1910).

[0231] If merge mode is not applied to the current block, the decoder obtains affine_flag (S1920). Here, affine_flag indicates whether or not AF mode is performed.

[0232] When affine_flag=1, that is, when AF mode is performed on the current block, the decoder checks whether condition B is satisfied (S1620). Here, condition B refers to a condition on the block size. For example, the embodiment shown in Table 2 can be applied.

[0233] If condition B is satisfied, the decoder sets affine_param_flag to 0 (S1930). Here, affine_param_flag indicates whether AF4 mode is performed (or whether affine motion prediction is performed using the four parameters). affine_param_flag=0 means that motion vector prediction is performed using AF4 mode.

[0234] If condition B is not satisfied but condition C is satisfied, the decoder sets affine_param_flag to 1 (S1940). Here, affine_param_flag=1 means that motion vector prediction is performed using AF6 mode.

[0235] On the other hand, if neither condition B nor condition C is satisfied, the decoder can obtain affine_param_flag (S1950).

[0236] When affine_param_flag=0, the decoder can obtain two motion vector difference values, mvd_CP0 and mvd_CP1 (S1960).

[0237] When affine_param_flag=1, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S1970).

[0238] Figure 20 shows a flowchart (4-1) of an embodiment to which the present invention is applied, which adaptively determines the optimal coding mode from among motion vector prediction modes, including AF4 mode or AF6 mode, based on the coding mode of the adjacent block.

[0239] The encoder makes predictions based on at least one of skip mode, merge mode, or intermode (S2010).

[0240] The encoder checks whether the adjacent block is coded in AF mode (S2020). Here, whether the adjacent block is coded in AF mode can be expressed by isNeighborAffine(). For example, isNeighborAffine()=0 means that the adjacent block is not coded in AF mode, and isNeighborAffine()=1 means that the adjacent block is coded in AF mode.

[0241] If the adjacent block is not coded in AF mode, the encoder performs motion vector prediction based on AF4 mode (S2030).

[0242] If the adjacent block is coded in AF mode, the encoder performs motion vector prediction based on AF4 mode and also performs motion vector prediction based on AF6 mode (S2040).

[0243] The encoder compares the results of steps S2030 and S2040 to determine the optimal coding mode from among the motion vector prediction modes, including AF4 mode or AF6 mode (S2050).

[0244] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and obtain a motion vector difference value by subtracting the motion vector predictor from the motion vector of the current block.

[0245] From this point forward, the encoding / decoding processes described in Figures 1 and 2 can be applied identically.

[0246] Figure 21 shows a flowchart illustrating an embodiment (4-2) to which the present invention is applied, in which adaptive decoding is performed using AF4 mode or AF6 mode based on the coding mode of adjacent blocks.

[0247] The decoder can receive a bitstream (S2110). The bitstream contains information about the coding mode of the current block in the video signal.

[0248] The decoder checks whether the coding mode of the current block is AF mode (S2120).

[0249] If AF mode is not performed for the current block, the decoder performs decoding (i.e., motion vector prediction) using a coding mode other than AF mode (S2170). For example, skip mode, merge mode, or inter mode may be used.

[0250] If AF mode is currently applied to a block, the decoder checks whether the adjacent block has been coded in AF mode (S2130). Here, whether the adjacent block has been coded in AF mode can be expressed by isNeighborAffine(). For example, isNeighborAffine()=0 means that the adjacent block has not been coded in AF mode, and isNeighborAffine()=1 means that the adjacent block has been coded in AF mode.

[0251] If the adjacent block is coded in AF mode, the decoder performs motion vector prediction based on AF4 mode (S2140).

[0252] If the adjacent block is not coded in AF mode, the decoder checks whether AF4 mode is applied to the current block (S2150).

[0253] Here, the step S2150 can be confirmed by an affine parameter flag indicating whether AF4 mode is performed (or whether affine motion prediction is performed using four parameters). For example, the affine parameter flag can be expressed as affine_param_flag. When affine_param_flag=0, motion vector prediction is performed in AF4 mode (S2140), and when affine_param_flag=1, motion vector prediction is performed in AF6 mode (S2160).

[0254] FIG. 22 shows, as an embodiment (4-3) to which the present invention is applied, a syntax structure for performing decoding in AF4 mode or AF6 mode based on the coding mode of an adjacent block.

[0255] The decoder acquires merge_flag and checks whether merge mode is applied to the current block (S2210).

[0256] When merge mode is not applied to the current block, the decoder can acquire affine_flag (S2220). Here, affine_flag indicates whether AF mode is performed.

[0257] When affine_flag=1, that is, when AF mode is performed on the current block, the decoder checks whether an adjacent block is coded in AF mode (S2230).

[0258] When an adjacent block is coded in AF mode, the decoder can acquire affine_param_flag (S2230). Here, affine_param_flag indicates whether AF4 mode is performed (or whether affine motion prediction is performed using four parameters).

[0259] If an adjacent block is not coded in AF mode, the decoder sets affine_param_flag to 0 (S2240).

[0260] When affine_param_flag=0, that is, when motion vector prediction is performed by AF4 mode, the decoder can obtain two motion vector difference values, mvd_CP0 and mvd_CP1 (S2250).

[0261] Then, when affine_param_flag=1, that is, when motion vector prediction is performed in AF6 mode, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S2260).

[0262] Figure 23 shows a flowchart (5-1) of an embodiment to which the present invention is applied, which adaptively determines the optimal coding mode from among motion vector prediction modes, including AF4 mode or AF6 mode, based on at least one of condition A, condition B, or condition C.

[0263] This invention presents an embodiment that combines Embodiment 2 and Embodiment 3. Figure 23 illustrates an example where all conditions A, B, and C are considered, and the order of the conditions can be changed.

[0264] As shown in Figure 23, the encoder makes predictions based on at least one of skip mode, merge mode, or intermode (S2310).

[0265] The encoder checks whether condition A is satisfied for the current block (S2320). Here, condition A refers to a condition for the block size, and the embodiments in Table 1 can be applied.

[0266] If condition A is satisfied, the encoder determines the optimal coding mode from among the modes excluding AF mode (S2380).

[0267] On the other hand, if condition A is not satisfied, the encoder checks whether condition B is satisfied for the current block (S2330). Here, condition B refers to a condition for the block size, and the embodiments in Table 2 can be applied.

[0268] If condition B is satisfied, the encoder performs motion vector prediction based on the AF4 mode (S2340).

[0269] If condition B is not satisfied, the encoder checks whether condition C is satisfied for the current block (S2350). Here, condition C refers to a condition for the block size, and the embodiments in Table 3 can be applied.

[0270] If the above condition C is satisfied, the encoder performs motion vector prediction based on the AF6 mode (S2370).

[0271] If condition C is not satisfied, the encoder performs motion vector prediction based on AF4 mode and also performs motion vector prediction based on AF6 mode (S2360).

[0272] On the other hand, in the conditions B and C, the thresholds (TH2) and (TH3) can be determined to satisfy the formula 5.

[0273] The encoder compares the results of steps S2310, S2340, S2360, and S2370 to determine the optimal coding mode (S2380).

[0274] Subsequently, the encoder generates a motion vector predictor for the current block based on the optimal coding mode, and can subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.

[0275] Subsequently, the encoding / decoding processes described in FIG. 1 and FIG. 2 can be applied identically.

[0276] FIG. 24 is a flowchart showing adaptive decoding performed in AF4 mode or AF6 mode based on at least one of condition A, condition B or condition C, as embodiment (5-2) to which the present invention is applied.

[0277] The decoder checks whether condition A is satisfied for the current block (S2410). Here, the condition A refers to a condition for block size. For example, the embodiment of Table 1 can be applied.

[0278] If the condition A is satisfied, the decoder checks whether the coding mode of the current block is AF mode (S2420). Here, the AF mode refers to an affine motion prediction mode using an affine motion model, the embodiment described in the present specification can be applied, and overlapping description is omitted.

[0279] If the condition A is not satisfied, or AF mode is not applied to the current block, the decoder performs decoding (i.e., motion vector prediction) according to a coding mode other than the AF mode (S2480). For example, skip mode, merge mode or inter mode can be used.

[0280] When AF mode is performed on the current block, the decoder checks whether condition B is satisfied for the current block (S2430). Here, condition B refers to a condition on the block size. For example, the embodiments in Table 2 can be applied, and redundant explanations are omitted.

[0281] If condition B is satisfied, the decoder performs motion vector prediction based on the AF4 mode (S2440).

[0282] If condition B is not satisfied, the decoder checks whether condition C is satisfied for the current block (S2450). Here, condition C refers to a condition on the block size. For example, the embodiments in Table 3 can be applied, and redundant explanations are omitted.

[0283] On the other hand, in the conditions B and C, the thresholds TH2 and TH3 can be determined to satisfy the formula 5.

[0284] If the above condition C is satisfied, the decoder performs motion vector prediction based on the AF6 mode (S2470).

[0285] If condition C is not satisfied, the decoder checks whether AF4 mode is applied to the current block (S2460).

[0286] Here, step S2460 can be confirmed by an affine parameter flag indicating whether or not AF4 mode is performed (or whether or not affine motion prediction is performed using the four parameters).

[0287] For example, the affine parameter flag can be represented by affine_param_flag. When affine_param_flag=0, motion vector prediction is performed by AF4 mode (S2440), and when affine_param_flag=1, motion vector prediction is performed by AF6 mode (S2470), but the present invention is not limited thereto.

[0288] Figure 25 shows a syntax structure that performs decoding by AF4 mode or AF6 mode based on at least one of condition A, condition B, or condition C, as an embodiment (5-3) to which the present invention is applied.

[0289] The decoder retrieves the merge_flag to check whether merge mode is applied to the current block (S2510).

[0290] If merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S2520). Here, condition A refers to a condition on the block size. For example, the embodiments in Table 1 can be applied.

[0291] If condition A is satisfied, the decoder can obtain affine_flag (S2520). Here, affine_flag indicates whether or not AF mode is performed.

[0292] When affine_flag=1, that is, when AF mode is performed on the current block, the decoder checks whether condition B is satisfied (S2530). Here, condition B refers to a condition on the block size. For example, the embodiment shown in Table 2 can be applied.

[0293] If condition B is satisfied, the decoder sets affine_param_flag to 0 (S2540). Here, affine_param_flag indicates whether AF4 mode is performed (or whether affine motion prediction is performed using the four parameters). affine_param_flag=0 means that motion vector prediction is performed using AF4 mode.

[0294] If condition B is not satisfied but condition C is satisfied, the decoder sets affine_param_flag to 1 (S2550). Here, affine_param_flag=1 means that motion vector prediction is performed using AF6 mode.

[0295] On the other hand, if neither condition B nor condition C is satisfied, the decoder can obtain affine_param_flag (S2560).

[0296] When affine_param_flag=0, the decoder can obtain two motion vector difference values, mvd_CP0 and mvd_CP1 (S2570).

[0297] When affine_param_flag=1, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S2580).

[0298] Figure 26 shows a flowchart (6-1) of an embodiment to which the present invention is applied, which adaptively determines the optimal coding mode from among motion vector prediction modes, including AF4 mode or AF6 mode, based on condition A or at least one of the coding modes of adjacent blocks.

[0299] The encoder can make predictions based on at least one of skip mode, merge mode, or intermode (S2610).

[0300] The encoder checks whether condition A is satisfied for the current block (S2620). Here, condition A refers to a condition for the block size, and the embodiments in Table 1 can be applied.

[0301] If condition A is satisfied, the encoder determines the optimal coding mode from among the modes excluding AF mode (S2660).

[0302] On the other hand, if condition A is not satisfied, the encoder checks whether the adjacent block is coded in AF mode (S2630). Here, whether or not the adjacent block is coded in AF mode can be expressed by isNeighborAffine(). For example, isNeighborAffine()=0 means that the adjacent block is not coded in AF mode, and isNeighborAffine()=1 means that the adjacent block is coded in AF mode.

[0303] If the adjacent block is not coded in AF mode, the encoder performs motion vector prediction based on AF4 mode (S2640).

[0304] If the adjacent block is coded in AF mode, the encoder performs motion vector prediction based on AF4 mode and also performs motion vector prediction based on AF6 mode (S2650).

[0305] The encoder compares the results of steps S2610, S2640, and S2650 to determine the optimal coding mode (S2660).

[0306] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and obtain a motion vector difference value by subtracting the motion vector predictor from the motion vector of the current block.

[0307] From this point forward, the encoding / decoding processes described in Figures 1 and 2 can be applied identically.

[0308] Figure 27 shows a flowchart illustrating an embodiment (6-2) to which the present invention is applied, in which adaptive decoding is performed by AF4 mode or AF6 mode based on at least one of condition A or the coding mode of an adjacent block.

[0309] The decoder can receive a bitstream (S2710). The bitstream contains information about the coding mode of the current block in the video signal.

[0310] The decoder checks whether condition A is satisfied for the current block in order to determine the optimal coding mode for motion vector prediction (S2720). Here, condition A refers to a condition on the block size. For example, the embodiments in Table 1 can be applied.

[0311] If condition A is satisfied, the decoder checks whether the coding mode of the current block is AF mode (S2730).

[0312] Hereafter, steps S2730 through S2780 can be applied to the content explained in steps S2120 through S2170 in Figure 21, and redundant explanations will be omitted.

[0313] Figure 28 shows a syntax structure (6-3) to which the present invention is applied, which is decoded by AF4 mode or AF6 mode based on condition A or at least one of the coding modes of adjacent blocks.

[0314] The decoder retrieves the merge_flag to check whether merge mode is applied to the current block (S2810).

[0315] If merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S2820). Here, condition A refers to a condition on the block size. For example, the embodiments in Table 1 can be applied.

[0316] If condition A is satisfied, the decoder can obtain affine_flag (S2820). Here, affine_flag indicates whether or not AF mode is performed.

[0317] In the following steps S2830 to S2860, the content explained in steps S2230 to S2260 in Figure 22 can be applied, and redundant explanations will be omitted.

[0318] Figure 29 shows a flowchart illustrating an embodiment to which the present invention is applied, illustrating the generation of a motion vector predictor based on at least one of AF4 mode or AF6 mode.

[0319] The decoder checks whether the AF mode is currently applied to the block (S2910). Here, the AF mode refers to a motion prediction mode that uses an affine motion model.

[0320] For example, the decoder can obtain an affine flag from the video signal, and whether or not the AF mode is applied to the current block can be determined based on the affine flag.

[0321] The decoder checks whether the AF4 mode is used when the AF mode is applied to the current block (S2920). Here, the AF4 mode is a mode that predicts the motion vector using the four parameters that constitute the affine motion model.

[0322] For example, if the AF mode is applied to the current block by the affine flag, the decoder may obtain an affine parameter flag from the video signal, which indicates whether the motion vector predictor is generated using the four parameters or the six parameters.

[0323] Here, the affine flag and the affine parameter flag can be defined at at least one level of slice, maximum coding unit, coding unit, or prediction unit.

[0324] When AF4 mode is used, the decoder generates a motion vector predictor using the four parameters, and when AF4 mode is not used, it generates a motion vector predictor using the six parameters that constitute the affine motion model (S2930).

[0325] The decoder can obtain the motion vector of the current block based on the motion vector predictor (S2940).

[0326] In one embodiment, the decoder can check whether the size of the current block satisfies a previously set condition. The previously set condition indicates whether at least one of the number of pixels in the current block, the width of the current block, and / or the height is greater than a previously set threshold.

[0327] For example, if the size of the current block satisfies the already set conditions, the decoder can check whether the AF mode is applied to the current block.

[0328] Conversely, if the size of the current block does not satisfy the previously set conditions, the current block can be decoded based on a coding mode other than the AF mode.

[0329] In one embodiment, the decoder can check whether the AF mode has been applied to an adjacent block when the AF mode is applied to the current block.

[0330] If AF mode is applied to the adjacent block, the motion vector predictor is generated using the four parameters; if AF mode is not applied to the adjacent block, the decoder performs a step to check whether AF4 mode is used.

[0331] Figure 30 shows a flowchart illustrating an embodiment to which the present invention is applied, in which a motion vector predictor is generated based on AF4_flag and AF6_flag.

[0332] The decoder obtains at least one of AF4_flag and AF6_flag from the video signal (S3010). Here, AF4_flag indicates whether AF4 mode is performed for the current block, and AF6_flag indicates whether AF6 mode is performed for the current block.

[0333] Here, at least one of AF4_flag and AF6_flag may be defined at the slice level, and may also be defined at the block level or the prediction unit level. However, the present invention is not limited thereto, and at least one of AF4_flag and AF6_flag may be defined at at least one level of slice, maximum coding unit, coding unit, or prediction unit.

[0334] The decoder checks the values ​​of AF4_flag and AF6_flag (S3020).

[0335] If AF4_flag=1, AF4 mode is applied to the current block; if AF4_flag=0, AF4 mode is not applied to the current block. Here, the application of AF4 mode means that motion vector prediction is performed using an affine motion model expressed by four parameters.

[0336] If AF6_flag=1, AF6 mode is applied to the current block; if AF4_flag=0, AF6 mode is not applied to the current block. Here, applying AF6 mode means that motion vector prediction is performed using an affine motion model expressed by four parameters.

[0337] If AF4_flag=0 and AF6_flag=0, the decoder performs motion vector prediction using a mode other than AF4 mode and AF6 mode (S3030).

[0338] If AF4_flag=1 and AF6_flag=0, the decoder performs motion vector prediction in AF4 mode (S3040).

[0339] If AF4_flag=0 and AF6_flag=1, the decoder performs motion vector prediction in AF6 mode (S3050).

[0340] If AF4_flag=1 and AF6_flag=1, the decoder performs motion vector prediction in AF4 mode or AF6 mode (S3060).

[0341] Figure 31 shows a flowchart illustrating an embodiment to which the present invention is applied, in which adaptive decoding is performed using either AF4 mode or AF6 mode based on whether or not adjacent blocks are coded in AF mode.

[0342] The decoder checks whether AF mode is currently applied to the block (S3110).

[0343] The decoder checks whether the adjacent block has been coded in AF mode if the AF mode is applied to the current block (S3120).

[0344] The decoder can obtain at least one of AF4_flag or AF6_flag if the adjacent block is coded in AF mode (S3130).

[0345] The decoder generates a motion vector predictor using four or six parameters based on at least one of AF4_flag or AF6_flag (S3140). For example, if AF4_flag=1, the decoder can perform motion vector prediction in AF4 mode, and if AF6_flag=1, the decoder can perform motion vector prediction in AF6 mode.

[0346] The decoder can obtain the motion vector of the current block based on the motion vector predictor (S3150).

[0347] Figure 32 shows a syntax for adaptive decoding based on AF4_flag and AF6_flag as an embodiment to which the present invention is applied.

[0348] The decoder can obtain AF4_flag and AF6_flag at the slice level (S3010). Here, AF4_flag indicates whether AF4 mode is currently applied to the block, and AF6_flag indicates whether AF6 mode is currently applied to the block. The AF4_flag can be represented as affine_4_flag, and the AF6_flag can be represented as affine_6_flag.

[0349] The decoder can adaptively decode based on AF4_flag and AF6_flag at the block level or prediction unit level.

[0350] If affine_4_flag is not 0 or affine_6_flag is not 0 (i.e., affine_4_flag=0 && affine_6_flag=0), the decoder can obtain the affine flag (S3220). The affine flag indicates whether or not AF mode is performed.

[0351] When AF mode is activated, the decoder can adaptively decode based on the AF4_flag and AF6_flag values.

[0352] If affine_4_flag=1 and affine_6_flag=0, the decoder can set affine_param_flag to 0. In other words, affine_param_flag=0 means that AF4 mode is performed.

[0353] If affine_4_flag=0 and affine_6_flag=1, the decoder can set affine_param_flag to 1. That is, affine_param_flag=1 means that AF6 mode will be performed.

[0354] If affine_4_flag=1 and affine_6_flag=1, the decoder can parse or retrieve affine_param_flag. Here, the decoder can decode in AF4 mode or AF6 mode depending on the value of affine_param_flag at the block level or prediction unit level.

[0355] Furthermore, the syntax structure can be applied to the embodiments described above, and redundant explanations will be omitted.

[0356] Figure 33 shows a syntax that adaptively decodes an adjacent block using either AF4 mode or AF6 mode, based on whether or not the adjacent block is coded in AF mode, as an embodiment to which the present invention is applied.

[0357] In this embodiment, the content that overlaps with Figure 32 can be explained using the previously described explanation, and only the differing parts will be explained.

[0358] If affine_4_flag=1 and affine_6_flag=1, the decoder can check whether the adjacent block was coded in AF mode.

[0359] If an adjacent block is coded in AF mode, the decoder parses or retrieves affine_param_flag (S3310). Here, the decoder can decode in AF4 mode or AF6 mode based on the affine_param_flag value at the block level or prediction unit level.

[0360] Conversely, if an adjacent block is not coded in AF mode, the decoder can set affine_param_flag to 0. That is, affine_param_flag=0 means that AF4 mode will be performed.

[0361] Figure 34 shows a video coding system to which the present invention is applied.

[0362] A video coding system includes a source device and a receiving device. The source device transmits encoded video / image information or data to the receiving device via a digital storage medium or network in file or streaming format.

[0363] The source device includes a video source, an encoding apparatus, and a transmitter. The receiving device includes a receiver, a decoding apparatus, and a renderer. The encoding apparatus may also be called a video / image encoding apparatus, and the decoding apparatus may also be called a video / image decoding apparatus. The transmitter may be included in the encoding apparatus. The receiver may be included in the decoding apparatus. The renderer may include a display unit, which may consist of a separate device or external component.

[0364] A video source can acquire video / images through video / image capture, synthesis, or generation processes. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices include, for example, one or more cameras, or a video / image archive containing previously captured video / images. Video / image generation devices include, for example, computers, tablets, and smartphones, which can generate video / images (electronically). For example, a computer can generate virtual video / images, in which case the process of generating the associated data can replace the video / image capture process.

[0365] An encoding device encodes the input video / image. Encoding involves a series of steps, including prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) is output in bitstream format.

[0366] The transmitting unit transmits encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium includes a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit includes elements for generating media files according to a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit extracts the bitstream and transmits it to a decoding device.

[0367] The decoding device decodes video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of the encoding device.

[0368] The renderer renders the decoded video / image. The rendered video / image is displayed via the display unit.

[0369] Figure 35 shows a content streaming system to which the present invention is applied.

[0370] As shown in Figure 35, the content streaming system to which the present invention is applied broadly includes an encoding server, a streaming server, a web server, a media storage facility, user equipment, and multimedia input devices.

[0371] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting it to the streaming server. In another example, if the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.

[0372] The bitstream is generated by an encoding method or bitstream generation method to which the present invention is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0373] The streaming server transmits multimedia data to the user's device based on user requests via the web server, and the web server acts as an intermediary to inform the user of available services. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. The content streaming system may also include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0374] The streaming server receives content from a media storage facility and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0375] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs (head-mounted displays)), digital TVs, desktop computers, and digital signage.

[0376] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0377] As described above, the embodiments described in the present invention can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in Figures 1, 2, 34, and 35 can be implemented on a computer, processor, microprocessor, controller, or chip.

[0378] Furthermore, decoders and encoders to which the present invention applies can include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and can be used to process video signals and data signals.

[0379] Furthermore, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored on a computer-readable recording medium. Multimedia data having the data structure according to the present invention can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices on which computer-readable data is stored. The computer-readable storage medium can include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media implemented in the form of a carrier wave (for example, transmission over the Internet). Furthermore, the bitstream generated by the encoding method can be stored on a computer-readable recording medium or transmitted over a wireless communication network. [Industrial applicability]

[0380] The preferred embodiments of the present invention described above are disclosed for illustrative purposes only, and those skilled in the art will be able to improve, modify, substitute, or add various other embodiments within the technical concept and technical scope of the present invention disclosed in the claims appended below.

Claims

1. In a method for decoding a video signal including the current block based on affine motion prediction mode (AF mode), A step of obtaining a merge flag from the video signal, wherein the merge flag indicates whether a merge mode is applied to the current block, and the set of candidates available in the merge mode consists of spatially adjacent candidates, temporal candidates, and generated candidates. When the merge mode is not applied to the current block, the motion vectors in the motion vector prediction candidate list include spatial motion candidates, or include spatial motion candidates and temporal motion candidates, and the step of obtaining an affine flag from the video signal based on whether the width and height of the current block are equal to or greater than 16, wherein the affine flag indicates whether the AF mode is applied to the current block, the AF mode represents a motion prediction mode using an affine motion model, the affine motion model induces motion vectors on a pixel-by-pixel or subblock-by-subblock basis using control point motion vectors, and the control points are two or more of the upper left corner pixel, upper right corner pixel, and lower left corner pixel of the current block, The steps include obtaining an affine parameter flag indicating whether four or six parameters are used for the affine motion model, based on the application of the AF mode to the current block, A step of obtaining a motion vector predictor based on the use of the four parameters or the six parameters for the affine motion model, The steps include obtaining a predicted sample for the current block based on the motion vector predictor, The steps include obtaining a residual sample for the current block, A step of restoring the current block based on the predicted sample and the residual sample, The steps include filtering the restored current block, A method in which the affine flag and the affine parameter flag are obtained based on the merge flag indicating that the merge mode is not applied to the current block.

2. The method according to claim 1, wherein the affine flag and the affine parameter flag are defined at the coding unit level.

3. The method according to claim 1, wherein, based on the fact that the width and height of the current block are less than 16, the current block is decoded based on a coding mode other than the AF mode.

4. In a method for encoding a video signal including the current block based on affine motion prediction mode (AF mode), A merge flag is generated indicating whether the merge mode is applied to the current block, and the set of candidates available in the merge mode consists of spatially adjacent candidates, temporal candidates, and generated candidates, in steps, When the merge mode is not applied to the current block, the motion vectors in the motion vector prediction candidate list include spatial motion candidates, or include spatial motion candidates and temporal motion candidates, and an affine flag is generated based on whether the width and height of the current block are equal to or greater than 16, wherein the affine flag indicates whether the AF mode is applied to the current block, the AF mode represents a motion prediction mode using an affine motion model, the affine motion model induces motion vectors on a pixel-by-pixel or subblock-by-subblock basis using control point motion vectors, and the control points are two or more of the upper left corner pixel, upper right corner pixel, and lower left corner pixel of the current block, The steps include generating an affine parameter flag indicating whether four or six parameters are used for the affine motion model, based on the application of the AF mode to the current block, A step of obtaining a motion vector predictor based on the use of the four parameters or the six parameters for the affine motion model, The steps include generating a prediction sample for the current block based on the motion vector predictor, The steps include generating a residual sample for the current block based on the predicted sample, The steps include performing a transformation, quantization, and entropy encoding on the said residual sample, Based on the predicted sample and the residual sample, the current block is restored. A method in which the affine flag and the affine parameter flag are generated based on the merge flag indicating that the merge mode is not applied to the current block.

5. In a method for transmitting data for a video signal, A step of generating a bitstream for the video signal, wherein the bitstream is generated by an encoding method, The step of transmitting the data, which includes the bitstream, The aforementioned encoding method is A merge flag is generated indicating whether the merge mode is currently applied to the block, and the set of candidates available in the merge mode consists of spatially adjacent candidates, temporal candidates, and generated candidates, in steps, When the merge mode is not applied to the current block, the motion vectors in the motion vector prediction candidate list include spatial motion candidates, or include spatial motion candidates and temporal motion candidates, and an affine flag is generated based on whether the width and height of the current block are equal to or greater than 16, wherein the affine flag indicates whether an affine motion prediction mode (AF mode) is applied to the current block, the AF mode represents a motion prediction mode using an affine motion model, the affine motion model induces motion vectors on a pixel-by-pixel or subblock-by-subblock basis using control point motion vectors, and the control points are two or more of the upper left corner pixel, upper right corner pixel, and lower left corner pixel of the current block, The steps include generating an affine parameter flag indicating whether four or six parameters are used for the affine motion model, based on the application of the AF mode to the current block, A step of obtaining a motion vector predictor based on the use of the four parameters or the six parameters for the affine motion model, The steps include generating a prediction sample for the current block based on the motion vector predictor, The steps include generating a residual sample for the current block based on the predicted sample, The steps include performing a transformation, quantization, and entropy encoding on the said residual sample, Based on the predicted sample and the residual sample, the current block is restored. A method in which the affine flag and the affine parameter flag are generated based on the merge flag indicating that the merge mode is not applied to the current block.

Citation Information

Patent Citations

  • Method and apparatus for global motion compensation in video coding system

    WO2017087751A1

  • Method and apparatus for affine merge mode prediction for video coding system

    WO2017118409A1

  • Prediction image generation device, moving image decoding device, and moving image encoding device

    WO2017130696A1

  • Affine prediction for video coding

    WO2017156705A1

  • Method and apparatus of video coding with affine motion compensation

    WO2017157259A1