Method and apparatus for processing a video signal using affine prediction

By adaptively selecting the four-parameter or six-parameter affine prediction mode, based on the block size and the encoding mode of adjacent blocks, the encoding efficiency problem when processing high-resolution and high-frame rate videos in the prior art is solved, and more efficient video signal processing is achieved.

JP7673160B2Active Publication Date: 2025-05-08GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Patent Information

Application Number
JP2023198969
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-08-03
Filing Date
2023-11-24
Publication Date
2025-05-08
Estimated Expiration
2038-08-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently encode and decode next-generation video content, especially when processing high-resolution, high frame rate and high-dimensional scene representations, resulting in a significant increase in storage, access rate and processing capabilities.

Method used

A method for adaptive affine prediction is proposed, and the best encoding mode is selected based on the block size and whether affine prediction is used in neighboring blocks. The method includes adaptively selecting the four-parameter affine prediction mode (AF4) or the six-parameter affine prediction mode (AF6) and applying affine prediction when specific conditions are met.

Benefits of technology

Through adaptive affine prediction, the performance of video encoding is improved, encoding complexity is reduced, and more efficient video signal processing is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007673160000020
    Figure 0007673160000020
  • Figure 0007673160000021
    Figure 0007673160000021
  • Figure 0007673160000022
    Figure 0007673160000022
Patent Text Reader

Abstract

To provide a method for decoding a video signal including a current block on the basis of an affine motion prediction mode (AF mode).SOLUTION: A method for decoding a video signal includes the steps of: checking whether an AF mode is applied to a current block, the AF mode representing a motion prediction mode that uses an affine motion model; checking whether an AF4 mode is used when the AF mode is applied to the current block, the AF4 mode representing a mode in which a motion vector is predicted using four parameters constituting the affine motion model; generating a motion vector predictor using four parameters when the AF4 mode is used and generating a motion vector predictor using six parameters constituting the affine motion model when the AF4 mode is not used; and obtaining the motion vector of the current block on the basis of the motion vector predictor.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method and apparatus for encoding / decoding a video signal, and more particularly to a method and apparatus for adaptively performing affine prediction. [Background technology]

[0002] Compression coding refers to a series of signal processing techniques for transmitting digitized information over communication lines or storing it in a suitable form on a storage medium. Media such as pictures, images, and sounds are the targets of compression coding, and the technology that compresses and codes images in particular is called video image compression.

[0003] Next-generation video content is characterized by high spatial resolution, high frame rate, and high dimensionality of scene representation. Processing such content will bring about a significant increase in memory storage, memory access rate, and processing power.

[0004] Therefore, there is a need to design coding tools to more efficiently process next generation video content. Summary of the Invention [Problem to be solved by the invention]

[0005] The present invention proposes a more efficient method for encoding and decoding video signals.

[0006] In addition, the present invention proposes a method of encoding or decoding taking into consideration both AF4 mode, which is a four parameter affine prediction mode, and AF6 mode, which is a six parameter affine prediction mode.

[0007] The present invention also proposes a method for adaptively determining (or selecting) an optimal coding mode based on at least one of AF4 mode or AF6 mode based on a block size.

[0008] The present invention also proposes a method for adaptively determining (or selecting) an optimal coding mode based on at least one of AF4 mode or AF6 mode depending on whether the neighboring blocks are coded by affine prediction. [Means for solving the problem]

[0009] In order to solve the above-mentioned technical problems,

[0010] The present invention provides a method for adaptively performing affine prediction based on block size.

[0011] The present invention also provides a method for adaptively performing affine prediction based on whether a neighboring block is coded using affine prediction.

[0012] The present invention also provides a method for adaptively determining (or selecting) an optimal coding mode based on at least one of AF4 mode or AF6 mode.

[0013] The present invention also provides a method for adaptively performing affine prediction based on whether at least one pre-set condition is satisfied, in which the pre-set condition may include at least one of a block size, a number of pixels in a block, a block width, a block height, and whether an adjacent block is coded using affine prediction. Effect of the Invention

[0014] The present invention provides a method for adaptively performing affine prediction, thereby improving the performance of affine prediction, and by reducing the complexity of affine prediction, more efficient coding can be performed. [Brief description of the drawings]

[0015] [Figure 1] 1 shows a schematic block diagram of an encoder for encoding a video signal as an embodiment to which the present invention is applied. [Diagram 2] 1 shows a schematic block diagram of a decoder in which a video signal is decoded as an embodiment to which the present invention is applied. [Diagram 3] FIG. 1 is a diagram for explaining a QT (QuadTree, hereinafter referred to as "QT") block division structure as an embodiment to which the present invention can be applied. [Figure 4] 1 is a diagram for explaining a BT (Binary Tree, hereinafter referred to as "BT") block division structure as an embodiment to which the present invention can be applied. FIG. [Diagram 5] 1 is a diagram for explaining a TT (Ternary Tree, hereinafter referred to as "TT") block division structure as an embodiment to which the present invention can be applied. FIG. [Figure 6] 1 is a diagram for explaining an AT (Asymmetric Tree, hereinafter referred to as "AT") block division structure as an embodiment to which the present invention can be applied. FIG. [Figure 7]FIG. 1 is a diagram for explaining an inter prediction mode as an embodiment to which the present invention is applied. [Figure 8] FIG. 1 is a diagram for explaining an affine motion model as an embodiment to which the present invention is applied. [Figure 9] 1 is a diagram illustrating an affine motion prediction method using a control point motion vector as an embodiment to which the present invention is applied. [Figure 10] 1 is a flowchart illustrating a process of processing a video signal including a current block using an affine prediction mode as an embodiment to which the present invention is applied. [Figure 11] As an embodiment (1-1) to which the present invention is applied, a flowchart for adaptively determining an optimal coding mode based on at least one of the AF4 mode and the AF6 mode is shown. [Figure 12] As an embodiment (1-2) to which the present invention is applied, a flowchart for adaptively decoding based on the AF4 mode or the AF6 mode is shown. [Figure 13] As an embodiment (1-3) to which the present invention is applied, a syntax structure for performing decoding based on the AF4 mode or AF6 mode is shown. [Figure 14] As an embodiment (2-1) to which the present invention is applied, a flowchart for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode based on condition A will be shown. [Figure 15] As an embodiment (2-2) to which the present invention is applied, a flowchart showing adaptive decoding in accordance with AF4 mode or AF6 mode based on condition A will be shown. [Figure 16] As an embodiment (2-3) to which the present invention is applied, a syntax structure for performing decoding in accordance with AF4 mode or AF6 mode based on condition A will be shown. [Figure 17] As an embodiment (3-1) to which the present invention is applied, a flowchart is shown for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode based on at least one of condition B (condition B) or condition C (condition C). [Figure 18] As an embodiment (3-2) to which the present invention is applied, a flowchart is shown in which decoding is adaptively performed in accordance with AF4 mode or AF6 mode based on at least one of condition B or condition C. [Figure 19] As an embodiment (3-3) to which the present invention is applied, a syntax structure for performing decoding according to AF4 mode or AF6 mode based on at least one of condition B or condition C is shown. [Figure 20] As an embodiment (4-1) to which the present invention is applied, a flowchart is shown in which an optimal coding mode is adaptively determined from among motion vector prediction modes including AF4 mode or AF6 mode based on the coding mode of an adjacent block. [Figure 21] As an embodiment (4-2) to which the present invention is applied, a flowchart is shown in which decoding is adaptively performed in accordance with the AF4 mode or AF6 mode based on the coding mode of an adjacent block. [Figure 22] As an embodiment (4-3) to which the present invention is applied, a syntax structure for performing decoding according to AF4 mode or AF6 mode based on the coding mode of an adjacent block is shown. [Figure 23] As an embodiment (5-1) to which the present invention is applied, a flowchart is shown for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode and AF6 mode, based on at least one of condition A, condition B, and condition C. [Figure 24]As an embodiment (5-2) to which the present invention is applied, a flowchart is shown in which decoding is adaptively performed in accordance with AF4 mode or AF6 mode based on at least one of condition A, condition B, and condition C. [Diagram 25] As an embodiment (5-3) to which the present invention is applied, a syntax structure is shown in which decoding is performed according to AF4 mode or AF6 mode based on at least one of condition A, condition B, or condition C. [Figure 26] As an embodiment (6-1) to which the present invention is applied, a flowchart is shown for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode based on at least one of condition A or the coding mode of an adjacent block. [Figure 27] As an embodiment (6-2) to which the present invention is applied, a flowchart is shown in which adaptive decoding is performed in accordance with AF4 mode or AF6 mode based on at least one of condition A or the coding mode of an adjacent block. [Figure 28] As an embodiment (6-3) to which the present invention is applied, a syntax structure for performing decoding according to AF4 mode or AF6 mode based on at least one of condition A or the coding mode of an adjacent block is shown. [Figure 29] As an embodiment to which the present invention is applied, a flowchart for generating a motion vector predictor based on at least one of AF4 mode and AF6 mode is shown. [Diagram 30] As an embodiment to which the present invention is applied, a flowchart for generating a motion vector predictor based on AF4_flag and AF6_flag is shown. [Diagram 31] As an embodiment to which the present invention is applied, a flowchart will be shown in which adaptive decoding is performed in AF4 mode or AF6 mode depending on whether adjacent blocks are coded in AF mode. [Diagram 32] As an embodiment to which the present invention is applied, a syntax for adaptively performing decoding based on AF4_flag and AF6_flag will be shown. [Diagram 33] As an embodiment to which the present invention is applied, a syntax for adaptively decoding in AF4 mode or AF6 mode based on whether a neighboring block is coded in AF mode will be described. [Diagram 34] 1 shows a video coding system to which the present invention is applied; [Diagram 35] 1 shows a content streaming system to which the present invention is applied. BEST MODE FOR CARRYING OUT THEINVENTION

[0016] The present invention provides a method for decoding a video signal including a current block based on an affine motion prediction mode (AF mode), the method comprising the steps of: determining whether the AF mode is applied to the current block, where the AF mode indicates a motion prediction mode using an affine motion model; determining whether an AF4 mode is used if the AF mode is applied to the current block, where the AF4 mode indicates a mode of predicting a motion vector using four parameters constituting the affine motion model; generating a motion vector predictor using the four parameters if the AF4 mode is used, and generating a motion vector predictor using six parameters constituting the affine motion model if the AF4 mode is not used; and obtaining a motion vector of the current block based on the motion vector predictor.

[0017] In the present invention, the method further includes a step of acquiring an affine flag from the video signal, the affine flag indicating whether the AF mode is applied to the current block, and whether the AF mode is applied to the current block is confirmed based on the affine flag.

[0018] In the present invention, the method further includes a step of obtaining an affine parameter flag from the video signal when the AF mode is applied to the current block according to the affine flag, the affine parameter flag indicating whether the motion vector predictor is generated using the four parameters or the six parameters.

[0019] In the present invention, the affine flag and the affine parameter flag are defined at least at one level of a slice, a maximum coding unit, a coding unit, or a prediction unit.

[0020] In the present invention, the method further includes a step of checking whether a size of the current block satisfies a pre-set condition, the pre-set condition indicating whether at least one of a number of pixels in the current block, a width and / or a height of the current block is greater than a pre-set threshold, and if the size of the current block satisfies the pre-set condition, a step of checking whether the AF mode is applied to the current block is performed.

[0021] In the present invention, if the size of the current block does not satisfy a pre-set condition, the current block is decoded based on a coding mode other than the AF mode.

[0022] In the present invention, the method further includes a step of checking whether an AF mode is applied to an adjacent block when the AF mode is applied to the current block, and if the AF mode is applied to the adjacent block, a motion vector predictor is generated using the four parameters, and if the AF mode is not applied to the adjacent block, a step of checking whether the AF4 mode is used is performed.

[0023] The present invention provides an apparatus for decoding a video signal including a current block based on an affine motion prediction mode (AF mode), the apparatus including: a) determining whether the AF mode is applied to the current block; b) determining whether an AF4 mode is used if the AF mode is applied to the current block; c) generating a motion vector predictor using four parameters if the AF4 mode is used; and d) generating a motion vector predictor using six parameters constituting an affine motion model if the AF4 mode is not used; and d) obtaining a motion vector of the current block based on the motion vector predictor, the AF mode indicating a motion prediction mode using the affine motion model; and d) the AF4 mode indicating a mode of predicting a motion vector using four parameters constituting the affine motion model.

[0024] In the present invention, the device further includes a parsing unit that parses an affine flag from the video signal, the affine flag indicating whether the AF mode is applied to the current block, and whether the AF mode is applied to the current block is confirmed based on the affine flag.

[0025] In the present invention, the device includes a parsing unit that acquires an affine parameter flag from the video signal when the AF mode is applied to the current block according to the affine flag, and the affine parameter flag indicates whether the motion vector predictor is generated using the four parameters or the six parameters.

[0026] In the present invention, the device includes an inter prediction unit that checks whether a size of the current block satisfies a pre-set condition, the pre-set condition indicating whether at least one of a number of pixels in the current block, a width and / or a height of the current block is greater than a pre-set threshold, and if the size of the current block satisfies the pre-set condition, a step of checking whether the AF mode is applied to the current block is performed.

[0027] In the present invention, the device includes the inter prediction unit which, when the AF mode is applied to the current block, checks whether the AF mode is applied to a neighboring block, and, when the AF mode is applied to the neighboring block, a motion vector predictor is generated using the four parameters, and, when the AF mode is not applied to the neighboring block, performs a step of checking whether the AF4 mode is used. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] Hereinafter, the configuration and operation of an embodiment of the present invention will be described with reference to the accompanying drawings. The configuration and operation of the present invention described with reference to the drawings will be described as one embodiment, and the technical idea of ​​the present invention and its core configuration and operation are not limited thereby.

[0029] In addition, the terms used in the present invention are generally used as widely as possible, but in specific cases, terms arbitrarily selected by the applicant are used for explanation. In such cases, the meanings are clearly described in the detailed description of the relevant part, so it is clarified that the terms used in the description of the present invention should not be analyzed simply based on the names of the terms, but should be analyzed by understanding the meanings of the relevant terms.

[0030] In addition, the terms used in the present invention are general terms selected to describe the invention, but if there are other terms with similar meanings, they can be substituted for more appropriate analysis. For example, in the case of signal, data, sample, picture, frame, block, etc., they can be appropriately substituted and analyzed in each coding process. In addition, in the case of partitioning, decomposition, splitting, division, etc., they can be appropriately substituted and analyzed in each coding process.

[0031] FIG. 1 is a schematic block diagram of an encoder for encoding a video signal as an embodiment to which the present invention is applied.

[0032] As shown in FIG. 1, the encoder 100 includes an image division unit 110, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, a filtering unit 160, a decoded picture buffer (DPB) 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190.

[0033] The image division unit 110 divides an input image (or picture, frame) input to the encoder 100 into one or more processing units. For example, the processing units may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In this case, the division is performed by at least one of a QuadTree (QT), a Binary Tree (BT), a Ternary Tree (TT), and an Asymmetric Tree (AT).

[0034] However, the above terms are used merely for the convenience of explanation of the present invention, and the present invention is not limited to the definition of the terms. In addition, for the convenience of explanation, the term "coding unit" is used in this specification as a unit used in the process of encoding or decoding a video signal, but the present invention is not limited thereto, and can be appropriately analyzed according to the contents of the invention.

[0035] The encoder 100 subtracts a prediction signal output from the inter prediction unit 180 or the intra prediction unit 185 from the input image signal to generate a residual signal, and the generated residual signal is transmitted to the conversion unit 120.

[0036] The transform unit 120 applies a transform technique to the residual signal to generate transform coefficients. The transform process may be applied to a pixel block having a uniform size, such as a square, or may be applied to a block of a variable size, such as a non-square.

[0037] The quantization unit 130 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 190, and the entropy encoding unit 190 performs entropy coding on the quantized signal and outputs the result as a bitstream.

[0038] The quantized signal output from the quantizer 130 may be used to generate a prediction signal. For example, the quantized signal may be subjected to inverse quantization and inverse transformation by the inverse quantizer 140 and inverse transformer 150 in a loop to reconstruct a residual signal. A reconstructed signal is generated by adding the reconstructed residual signal to a prediction signal output from the inter prediction unit 180 or intra prediction unit 185.

[0039] Meanwhile, in the above-mentioned compression process, adjacent blocks may be quantized with different quantization parameters, so that block boundaries may become visible. This phenomenon is called blocking artifacts, which is one of the important factors for evaluating image quality. In order to reduce this artifact, a filtering process may be performed. This filtering process can remove the blocking artifacts and reduce errors for the current picture, thereby improving image quality.

[0040] The filtering unit 160 applies filtering to the reconstructed signal and outputs it to a playback device or transmits it to a decoded picture buffer 170. The filtered signal transmitted to the decoded picture buffer 170 can be used as a reference picture in the inter prediction unit 180. In this way, by using the filtered picture as a reference picture in the inter prediction mode, not only the image quality but also the coding efficiency is improved.

[0041] The decoded picture buffer 170 stores the filtered pictures for use as reference pictures by the inter predictor 180 .

[0042] The inter prediction unit 180 performs temporal prediction and / or spatial prediction to remove temporal redundancy and / or spatial redundancy by referring to a reconstructed picture. Here, the reference picture used for prediction is a transformed signal that has been quantized and dequantized in units of blocks during previous encoding / decoding, and thus may have blocking artifacts and ringing artifacts.

[0043] Therefore, in order to solve the performance degradation due to the signal discontinuity and quantization, the inter prediction unit 180 may apply a low pass filter to interpolate signals between pixels in sub-pixel units. Here, a sub-pixel refers to a virtual pixel generated by applying an interpolation filter, and an integer pixel refers to an actual pixel present in a restored picture. As an interpolation method, linear interpolation, bi-linear interpolation, a Wiener filter, etc. may be applied.

[0044] The interpolation filter is applied to a reconstructed picture to improve the accuracy of prediction. For example, the inter prediction unit 180 may apply an interpolation filter to integer pixels to generate interpolated pixels, and perform prediction using an interpolated block composed of the interpolated pixels as a prediction block.

[0045] The intra prediction unit 185 can predict a current block by referring to samples in the vicinity of a block to be currently encoded. The intra prediction unit 185 performs the following process to perform intra prediction. First, reference samples necessary for generating a prediction signal are prepared. Then, a prediction signal is generated using the prepared reference samples. Then, a prediction mode is encoded. Here, the reference samples are prepared by reference sample padding and / or reference sample filtering. Since the reference samples have undergone prediction and restoration processes, there is a possibility that quantization errors may exist. Therefore, in order to reduce such errors, a reference sample filtering process is performed for each prediction mode used in intra prediction.

[0046] A prediction signal generated by the inter prediction unit 180 or the intra prediction unit 185 is used to generate a reconstructed signal or is used to generate a residual signal.

[0047] FIG. 2 is a schematic block diagram of a decoder for decoding a video signal as an embodiment to which the present invention is applied.

[0048] As shown in FIG. 2, the decoder 200 includes a parsing unit (not shown), an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, a filtering unit 240, a decoded picture buffer (DPB) 250, an inter prediction unit 260, an intra prediction unit 265, and a restoration unit (not shown).

[0049] The decoder 200 receives the signal output from the encoder 100 of Fig. 1, and parses or obtains syntax elements through a parsing unit (not shown). The parsed or obtained signal is entropy decoded by an entropy decoding unit 210.

[0050] The inverse quantization unit 220 obtains transform coefficients from the entropy decoded signal using the quantization step size information.

[0051] In the inverse transform unit 230, the transform coefficients are inversely transformed to obtain a residual signal.

[0052] A reconstruction unit (not shown) generates a reconstructed signal by adding the acquired residual signal to a prediction signal output from the inter prediction unit 260 or the intra prediction unit 265.

[0053] The filtering unit 240 applies filtering to the reconstructed signal and outputs the signal to a playback device or transmits the signal to the decoded picture buffer unit 250. The filtered signal transmitted to the decoded picture buffer unit 250 can be used as a reference picture in the inter prediction unit 260.

[0054] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the encoder 100 can be similarly applied to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the decoder, respectively.

[0055] The reconstructed video signal outputted through the decoder 200 can be reproduced by a reproduction device.

[0056] FIG. 3 is a diagram for explaining a QT (QuadTree, hereinafter referred to as "QT") block division structure as an embodiment to which the present invention can be applied.

[0057] In video coding, a block can be divided on a quadtree (QT) basis. A subblock divided by QT can be further divided recursively using QT. A leaf block that is not further divided by QT can be divided by at least one of binary tree (BT), ternary tree (TT) and asymmetric tree (AT). BT has two types of division, horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). TT has two types of division, horizontal TT (2N×1 / 2N, 2N×N, 2N×1 / 2N) and vertical TT (1 / 2N×2N, N×2N, 1 / 2N×2N). AT has four forms of partitioning: horizontal-up AT (2N×1 / 2N, 2N×3 / 2N), horizontal-down AT (2N×3 / 2N, 2N×1 / 2N), vertical-left AT (1 / 2N×2N, 3 / 2N×2N), vertical-right AT (3 / 2N×2N, 1 / 2N×2N). Each BT, TT, and AT can be further partitioned recursively using BT, TT, and AT.

[0058] Figure 3 shows an example of QT division. Block A can be divided into four sub-blocks (A0, A1, A2, A3) by QT. Sub-block A1 can be further divided into four sub-blocks (B0, B1, B2, B3) by QT.

[0059] FIG. 4 is a diagram for explaining a BT (Binary Tree, hereinafter referred to as "BT") block division structure as an embodiment to which the present invention can be applied.

[0060] Figure 4 shows an example of BT division. Block B3, which is not further divided by QT, can be divided by vertical BT (C0, C1) or horizontal BT (D0, D1). Each subblock, such as block C0, can be further divided recursively, such as in the form of horizontal BT (E0, E1) or vertical BT (F0, F1).

[0061] FIG. 5 is a diagram for explaining a TT (Ternary Tree, hereinafter referred to as "TT") block division structure as an embodiment to which the present invention can be applied.

[0062] Figure 5 shows an example of TT partitioning. Block B3, which is not further divided by QT, can be partitioned by vertical TT (C0, C1, C2) or horizontal TT (D0, D1, D2). Each subblock, such as block C1, can be further partitioned recursively, such as in the form of horizontal TT (E0, E1, E2) or vertical TT (F0, F1, F2).

[0063] FIG. 6 is a diagram for explaining an AT (Asymmetric Tree, hereinafter referred to as "AT") block division structure as an embodiment to which the present invention can be applied.

[0064] Figure 6 shows an example of AT partitioning. Block B3, which is not further divided by QT, can be partitioned by vertical AT(C0, C1) or horizontal AT(D0, D1). Each subblock, such as block C1, can be further partitioned recursively, such as in the form of horizontal AT(E0, E1) or vertical TT(F0, F1).

[0065] Meanwhile, BT, TT, and AT division can be used together for division. For example, sub-blocks divided by BT can be divided by TT or AT. Also, sub-blocks divided by TT can be divided by BT or AT. Sub-blocks divided by AT can be divided by BT or TT. For example, after horizontal BT division, each sub-block can be divided by vertical BT, or after vertical BT division, each sub-block can be divided by horizontal BT. The above two division methods have different division orders, but the final divided shapes are the same.

[0066] In addition, when a block is divided, the order of searching the block can be defined in various ways. In general, searching a block from left to right and from top to bottom means an order of determining whether to perform additional block division of each divided sub-block, means an encoding order of each sub-block when the block is not further divided, or means a search order when referring to information of other neighboring blocks in a sub-block.

[0067] FIG. 7 is a diagram for explaining an inter prediction mode as an embodiment to which the present invention is applied.

[0068] Inter Prediction Mode

[0069] In the inter prediction mode to which the present invention is applied, a merge mode, an Advanced Motion Vector Prediction (AMVP) mode, or an affine prediction mode (hereinafter referred to as "AF mode") is used to reduce the amount of motion information.

[0070] 1) Merge mode

[0071] Merge mode refers to a method of deriving motion parameters (or information) from spatially or temporally neighboring blocks.

[0072] The set of candidates available in merge mode consists of spatial neighbor candidates, temporal candidates, and generated candidates.

[0073] As shown in (a) of Figure 7, whether each spatial candidate block is available is determined according to the order of {A1, B1, B0, A0, B2}. At that time, if the candidate block is encoded in intra prediction mode and there is no motion information, or if the candidate block is located outside the current picture (or slice), the candidate block is unavailable.

[0074] After determining the validity of the spatial candidates, the spatial merge candidates can be constructed by removing unnecessary candidate blocks from the candidate blocks of the current processing block. For example, if the candidate block of the current prediction block is the first prediction block in the same coding block, the candidate block can be removed, and candidate blocks having the same motion information can also be removed.

[0075] Once the construction of spatial merge candidates is completed, the construction process of temporal merge candidates is carried out according to the order {T0, T1}.

[0076] In the temporal candidate configuration, if the right bottom block (T0) of the collocated block of the reference picture is available, the block is configured as the temporal merge candidate. The collocated block means a block that exists at a position corresponding to the currently processed block in the selected reference picture. On the other hand, if not, the center block (T1) of the collocated blocks is configured as the temporal merge candidate.

[0077] The maximum number of merging candidates can be specified in the slice header. If the number of merging candidates is greater than the maximum, a smaller number of spatial and temporal candidates are maintained. Otherwise, the number of merging candidates is increased by combining the currently added candidates to generate additional merging candidates (i.e., combined bi-predictive merging candidates) until the number of candidates reaches the maximum number.

[0078] In the encoder, a merge candidate list is constructed as described above, and information on a candidate block selected from the merge candidate list by performing motion estimation is signaled to the decoder as a merge index (e.g., merge_idx[x0][y0]'). Figure 7(b) illustrates an example in which block B1 is selected from the merge candidate list, and in this case, "Index 1" is signaled to the decoder as the merge index.

[0079] The decoder generates a merge candidate list in the same manner as the encoder, derives motion information for the current block from motion information of a candidate block corresponding to a merge index received from the encoder in the merge candidate list, and generates a prediction block for the current block based on the derived motion information.

[0080] 2) AMVP (Advanced Motion Vector Prediction) mode

[0081] The AMVP mode refers to a method of deriving a motion vector prediction value from neighboring blocks. Thus, horizontal and vertical motion vector difference (MVD), reference index and inter prediction mode are signaled to the decoder. The horizontal and vertical motion vector values ​​are calculated using the deriving motion vector prediction value and the motion vector difference (MVD) provided by the encoder.

[0082] That is, the encoder configures a motion vector predictor candidate list, and performs motion estimation to signal a motion reference flag (i.e., candidate block information) (e.g., mvp_lX_flag[x0][y0]') selected from the motion vector predictor candidate list to the decoder. The decoder configures a motion vector predictor candidate list in the same manner as the encoder, and derives a motion vector predictor value for a currently processed block using motion information of a candidate block indicated by a motion reference flag received from the encoder in the motion vector predictor candidate list. The decoder then obtains a motion vector value for a currently processed block using the derived motion vector predictor value and a motion vector difference value transmitted from the encoder. The decoder then generates a predictive block for a currently processed block based on the derived motion information (i.e., motion compensation).

[0083] For AMVP mode, two spatial motion candidates are selected from the five available candidates previously described in Figure 7. The first spatial motion candidate is selected from the set {A0, A1} located on the left side, and the second spatial motion candidate is selected from the set {B0, B1, B2} located on the top side. Here, the motion vectors are scaled if the reference index of the neighboring candidate block is not the same as the current prediction block.

[0084] If the number of candidates selected from the search result of spatial motion candidates is two, the candidate configuration is terminated, but if the number is less than two, a temporal motion candidate is added.

[0085] A decoder (eg, an inter predictor) decodes motion parameters for a processing block (eg, a prediction unit).

[0086] For example, if the processing block utilizes a merge mode, the decoder can decode a merge index signaled from the encoder and derive the motion parameters of the current processing block from the motion parameters of the candidate block indicated in the merge index.

[0087] Also, when the processing block is applied with the AMVP mode, the decoder can decode horizontal and vertical motion vector difference (MVD), reference index, and inter prediction mode signaled from the encoder, derive a motion vector prediction value from motion parameters of a candidate block indicated by a motion reference flag, and derive a motion vector value of a current processing block using the motion vector prediction value and the received motion vector difference value.

[0088] The decoder uses the decoded motion parameters (or information) to perform motion compensation for the prediction unit.

[0089] That is, in the encoder / decoder, the decoded motion parameters are used to perform motion compensation, which predicts the image of the current unit from previously decoded pictures.

[0090] 3) AF mode (Affine Mode)

[0091] The AF mode refers to a motion prediction mode using an affine motion model, and may include at least one of an affine merge mode or an affine inter mode. The affine inter mode may include at least one of an AF4 mode or an AF6 mode. Here, the AF4 mode indicates an affine prediction mode using four parameters, and the AF6 mode indicates an affine prediction mode using six parameters.

[0092] However, in the present invention, for convenience of explanation, it is expressed as AF4 mode or AF6 mode, but this does not necessarily need to be defined as a separate prediction mode, and the AF4 mode or AF6 mode can be understood as being distinguished according to whether it uses only four parameters or six parameters.

[0093] The AF modes will be explained in more detail with reference to FIGS.

[0094] FIG. 8 is a diagram for explaining an affine motion model as an embodiment to which the present invention is applied.

[0095] A typical image coding technique uses a translation motion model to represent the motion of a coding block. Here, the translation motion model refers to a translated block-based prediction method. That is, the motion information of a coding block is represented using one motion vector. However, the optimal motion vector for each pixel in an actual coding block may differ. If the optimal motion vector can be determined for each pixel or sub-block with only a small amount of information, the coding efficiency can be improved.

[0096] Therefore, in order to improve the performance of inter prediction, the present invention proposes not only a translated block-based prediction method but also an inter prediction-based image processing method that reflects various motions of an image.

[0097] The present invention also proposes an affine motion prediction method for performing encoding / decoding using an affine motion model. The affine motion model is a prediction method for deriving a motion vector in pixel units or subblock units using a motion vector of a control point. In this specification, an affine motion prediction mode using the affine motion model is referred to as an AF mode.

[0098] The present invention also provides a method for adaptively performing affine prediction based on block size.

[0099] The present invention also provides a method for adaptively performing affine prediction based on whether a neighboring block is coded using affine prediction.

[0100] The present invention also provides a method for adaptively determining (or selecting) an optimal coding mode based on at least one of an AF4 mode or an AF6 mode, where the AF4 mode indicates a four parameter affine prediction mode and the AF6 mode indicates a six parameter affine prediction mode.

[0101] As shown in FIG. 8, various methods are used to express image distortion as motion information, and in particular, the affine motion model can express the four motions shown in FIG.

[0102] For example, an affine motion model can model image translate, image scale, image rotate, image shear, as well as any induced image distortion.

[0103] Affine motion models are expressed in various ways, among which the present invention proposes a method of displaying (or identifying) distortion using motion information at a specific reference point (or reference pixel / sample) of a block, and performing inter prediction using the distortion. Here, the reference point is called a control point (CP) (or control pixel, control sample), and the motion vector at such a reference point is called a control point motion vector (CPMV). The degree of distortion that can be expressed varies depending on the number of such control points.

[0104] The affine motion model is expressed using six parameters (a, b, c, d, e, and f) as shown in the following Equation 1.

[0105]

number

[0106] where (x, y) indicates the position of the upper left pixel of the coding block. x and v y indicate the motion vector at (x, y), respectively.

[0107] FIG. 9 is a diagram for explaining an affine motion prediction method using a control point motion vector as an embodiment to which the present invention is applied.

[0108] As shown in (a) of FIG. 9, an upper left control point (CP0) 902 (hereinafter referred to as a first control point), an upper right control point (CP1) 903 (hereinafter referred to as a second control point), and a lower left control point (CP2) 904 (hereinafter referred to as a third control point) of a current block 901 may have independent motion information. These are represented as CP0, CP1, and CP2, respectively. However, this corresponds to an embodiment of the present invention, and the present invention is not limited thereto. For example, various control points such as a lower right control point, a center control point, and other control points according to the positions of sub-blocks may be defined.

[0109] In one embodiment of the present invention, at least one of the first control point to the third control point may be a pixel included in the current block, or, as another example, at least one of the first control point to the third control point may be a pixel adjacent to the current block but not included in the current block.

[0110] Using the control point motion information of one or more of the control points, motion information for each pixel or each sub-block of the current block 901 can be derived.

[0111] For example, an affine motion model using the motion vectors of an upper left control point 902, an upper right control point 903, and a lower left control point 904 of a current block 901 is defined as in Equation 2 below.

[0112]

number

[0113] Where: JPEG0007673160000003.jpg1297 is the motion vector of the upper left control point 902, JPEG0007673160000004.jpg1486, the motion vector of the upper right control point 903, If JPEG0007673160000005.jpg1377 is the motion vector of the lower left control point 904, JPEG0007673160000006.jpg1390, JPEG0007673160000007.jpg1290, JPEG0007673160000008.jpg1397. In Equation 2, w represents the width of the current block 901, and h represents the height of the current block 901. JPEG0007673160000009.jpg12100 shows the motion vector for position {x, y}.

[0114] In the present invention, an affine motion model that expresses three types of motion, namely translation, scale, and rotation, among the motions that can be expressed by the affine motion model, can be defined. In this specification, this is called a simplified affine motion model or a similarity affine motion model.

[0115] The simplified affine motion model is expressed using four parameters (a, b, c, d) as shown in the following Equation 3.

[0116]

number

[0117] Here, {v x , v y} indicates a motion vector at the {x, y} position. An affine motion model using four parameters in this way is called AF4. The present invention is not limited to this, and when six parameters are used, it is called AF6, and the above-mentioned embodiment can be applied in the same way.

[0118] As shown in FIG. 9(b), JPEG0007673160000011.jpg1278 is the motion vector of the upper left control point 1001 of the current block, If JPEG0007673160000012.jpg14100 is the motion vector of the upper right control point 1002, JPEG0007673160000013.jpg15100, JPEG0007673160000014.jpg16107. Here, the affine motion model of AF4 can be defined as Equation 4 below.

[0119]

number

[0120] In Equation 4, w represents the width of the current block, and h represents the height of the current block. JPEG0007673160000016.jpg13105 shows the motion vectors for each {x,y} position.

[0121] An encoder or decoder can use the control point motion vectors (eg, the motion vectors of the top-left control point 1001 and the top-right control point 1002) to determine (or derive) a motion vector for each pixel position.

[0122] In the present invention, a set of motion vectors determined by affine motion prediction may be defined as an affine motion vector field, which is determined using at least one of Equations 1 to 4.

[0123] In the encoding / decoding process, a motion vector according to affine motion prediction may be determined in pixel units or in predefined (or preset) block (or sub-block) units. For example, when determined in pixel units, a motion vector is derived based on each pixel in a block, and when determined in sub-block units, a motion vector is derived based on each sub-block unit in a current block. As another example, when determined in sub-block units, a motion vector of the sub-block is derived based on the upper left pixel or center pixel.

[0124] For the sake of convenience, in the following description of the present invention, we will focus on the case where motion vectors are determined in 4x4 block units by affine motion prediction, but the present invention is not limited to this and can be applied in pixel units or block units of other sizes.

[0125] Meanwhile, assume that the size of the current block is 16x16 as shown in Figure 9(b). The encoder or decoder determines a motion vector in units of 4x4 sub-blocks using the motion vectors of the upper left control point 1001 and the upper right control point 1002 of the current block. Then, the motion vector of each sub-block is determined based on the central pixel value of the sub-block.

[0126] In FIG. 9(b), the arrows displayed in the center of each subblock indicate the motion vectors obtained by the affine motion model.

[0127] Affine motion prediction can be used as an affine merge mode (hereinafter referred to as "AF merge mode") and an affine inter mode (hereinafter referred to as "AF inter mode"). The AF merge mode is a method of deriving two control point motion vectors and encoding or decoding them without encoding a motion vector difference, similar to the skip mode or merge mode. The AF inter mode is a method of determining a control point motion vector predictor and a control point motion vector, and then encoding or decoding a control point motion vector difference (CPMVD) corresponding to the difference. In this case, in the case of AF4 mode, two control point motion vector difference values ​​are transmitted, and in the case of AF6 mode, three control point motion vector difference values ​​are transmitted.

[0128] Here, the AF4 mode has the advantage of being able to express a control point motion vector (CPMV) with fewer bits since it transmits fewer motion vector differential values ​​than the AF6 mode, and the AF6 mode has the advantage of being able to reduce bits for residual coding since it transmits three CPMVDs, enabling excellent predictor generation.

[0129] Therefore, the present invention proposes a method that takes into account both the AF4 mode and the AF6 mode (or simultaneously) in the AF inter mode.

[0130] FIG. 10 is a flowchart illustrating a process of processing a video signal including a current block using an affine prediction mode (hereinafter, referred to as AF mode) according to an embodiment of the present invention.

[0131] The present invention provides a method for processing a video signal including a current block using an AF mode.

[0132] First, the video signal processing device generates a candidate list of motion vector pairs using motion vectors of pixels or blocks adjacent to at least two control points of a current block (S1010), where the control points refer to corner pixels of the current block, and the motion vector pairs indicate the motion vectors of the upper left corner pixel and the upper right corner pixel of the current block.

[0133] In one embodiment, the control points include at least two of the upper left corner pixel, the upper right corner pixel, the lower left corner pixel, or the lower right corner pixel of the current block, and the candidate list is composed of pixels or blocks adjacent to the upper left corner pixel, the upper right corner pixel, and the lower left corner pixel.

[0134] In one embodiment, the candidate list may be generated based on motion vectors of the diagonal neighboring pixel (A), the upper neighboring pixel (B), and the left neighboring pixel (C) of the upper left corner pixel, the motion vectors of the upper neighboring pixel (D) and the diagonal neighboring pixel (E) of the upper right corner pixel, and the motion vectors of the left neighboring pixel (F) and the diagonal neighboring pixel (G) of the lower left corner pixel.

[0135] In one embodiment, the method may further comprise the step of adding an AMVP candidate list to the candidate list if the candidate list has less than two motion vector pairs.

[0136] In one embodiment, when the current block has a size of N×4, the control point motion vector of the current block is determined to be a motion vector derived based on a center position of a left sub-block and a right sub-block within the current block, and when the current block has a size of 4×N, the control point motion vector of the current block is determined to be a motion vector derived based on a center position of an upper sub-block and a lower sub-block within the current block.

[0137] In one embodiment, when the current block has a size of N×4, the control point motion vector of a left sub-block in the current block is determined by an average value of the first control point motion vector and the third control point motion vector, and the control point motion vector of a right sub-block is determined by an average value of the second control point motion vector and the fourth control point motion vector. When the current block has a size of 4×N, the control point motion vector of an upper sub-block in the current block is determined by an average value of the first control point motion vector and the second control point motion vector, and the control point motion vector of a lower sub-block is determined by an average value of the third control point motion vector and the fourth control point motion vector.

[0138] In another embodiment, the method may signal a predicted mode or flag information indicating whether the AF mode is performed.

[0139] In this case, the video signal processing device may receive the prediction mode or flag information, perform the AF mode according to the prediction mode or the flag information, and derive a motion vector according to the AF mode, wherein the AF mode indicates a mode for deriving a motion vector in pixel or sub-block units using a control point motion vector of the current block.

[0140] Meanwhile, the video signal processing device determines a final candidate list of a preset number of motion vector pairs based on divergence values ​​of the motion vector pairs (S1020), where the final candidate list is determined in ascending order of divergence value, and the divergence value indicates a similarity in the direction of the motion vectors.

[0141] The video signal processing apparatus determines a control point motion vector of the current block from the final candidate list based on a rate-distortion cost (S1030).

[0142] The video signal processing device generates a motion vector predictor for the current block based on the control point motion vector (S1040).

[0143] FIG. 11 shows a flowchart for adaptively determining an optimal coding mode based on at least one of the AF4 mode and the AF6 mode as an embodiment (1-1) to which the present invention is applied.

[0144] The video signal processing device performs prediction based on at least one of a skip mode, a merge mode, or an inter mode (S1110), where the merge mode may include not only a general merge mode but also the above-mentioned AF merge mode, and the inter mode may include not only a general inter mode but also the above-mentioned AF inter mode.

[0145] The video signal processing device performs motion vector prediction based on at least one of an AF4 mode and an AF6 mode (S1120), where the order of steps S1110 and S1120 is not limited.

[0146] The video signal processing device compares the results of step S1120 to determine an optimal coding mode among the modes (S1130), where the results of step S1120 are compared based on a rate-distortion cost.

[0147] Thereafter, the video signal processing device generates a motion vector predictor for the current block based on the optimal coding mode, and subtracts the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.

[0148] Thereafter, the encoding / decoding process described above with reference to FIGS. 1 and 2 is applied in the same manner.

[0149] FIG. 12 shows a flowchart of adaptive decoding based on the AF4 mode or the AF6 mode as an embodiment (1-2) to which the present invention is applied.

[0150] The decoder receives a bitstream (S1210), the bitstream including information about the coding mode of a current block in a video signal.

[0151] The decoder determines whether the coding mode of the current block is an AF mode (S1220). Here, the AF mode refers to an affine motion prediction mode using an affine motion model, and may include at least one of an affine merge mode or an affine inter mode, and the affine inter mode may include at least one of an AF4 mode or an AF6 mode.

[0152] Here, step S1220 is checked based on an affine flag indicating whether or not the AF mode is performed. For example, the affine flag may be expressed as affine_flag. When affine_flag=1, it indicates that the AF mode is performed for the current block, and when affine_flag=0, it indicates that the AF mode is not performed for the current block.

[0153] If the AF mode is not performed for the current block, the decoder performs decoding (i.e., motion vector prediction) according to a coding mode other than the AF mode (S1230). For example, a skip mode, a merge mode, or an inter mode may be used.

[0154] If the AF mode is to be performed on the current block, the decoder checks whether the AF4 mode is to be applied to the current block (S1240).

[0155] Here, the S1240 step can be confirmed by an affine parameter flag indicating whether or not AF4 mode is performed (or whether or not affine motion prediction is performed using four parameters). For example, the affine parameter flag is expressed as affine_param_flag. When affine_param_flag=0, it means that motion vector prediction is performed using AF4 mode (S1250), and when affine_param_flag=1, it means that motion vector prediction is performed using AF6 mode (S1260), but the present invention is not limited thereto.

[0156] For example, the affine parameter flags include at least one of AF4_flag and AF6_flag.

[0157] AF4_flag indicates whether or not AF4 mode is performed for the current block. If AF4_flag=1, AF4 mode is performed for the current block, and if AF4_flag=0, AF4 mode is not performed for the current block. Here, performing AF4 mode means that motion vector prediction is performed using an affine motion model expressed by four parameters.

[0158] AF6_flag indicates whether or not the AF6 mode is performed for the current block. If AF6_flag=1, the AF6 mode is performed for the current block, and if AF4_flag=0, the AF6 mode is not performed for the current block. Here, performing the AF6 mode means that motion vector prediction is performed using an affine motion model expressed by four parameters.

[0159] The affine flag and the affine parameter flag may be defined at at least one of the following levels: slice, maximum coding unit, coding unit, or prediction unit.

[0160] For example, at least one of AF_flag, AF4_flag, and AF6_flag may be defined at the slice level, and may also be defined at the block level or prediction unit level.

[0161] FIG. 13 shows a syntax structure for performing decoding based on the AF4 mode or AF6 mode as an embodiment (1-3) to which the present invention is applied.

[0162] The decoder obtains merge_flag to check whether the merge mode is applied to the current block (S1310).

[0163] If merge mode is not applied to the current block, the decoder may obtain affine_flag (S1320), where affine_flag indicates whether AF mode is performed.

[0164] If the affine_flag=1, i.e., if the AF mode is performed for the current block, the decoder can obtain affine_param_flag (S1330), where affine_param_flag indicates whether the AF4 mode is performed (or whether affine motion prediction is performed using four parameters).

[0165] When the affine_param_flag=0, that is, when motion vector prediction is performed in AF4 mode, the decoder can obtain two motion vector differential values, mvd_CP0 and mvd_CP1 (S1340), where mvd_CP0 indicates the motion vector differential value for control point 0, and mvd_CP1 indicates the motion vector differential value for control point 1.

[0166] Then, if the affine_param_flag=1, that is, if the motion vector prediction is performed in the AF6 mode, the decoder can obtain three motion vector differential values, mvd_CP0, mvd_CP1, and mvd_CP2 (S1350).

[0167] FIG. 14 shows a flowchart for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode based on condition A as an embodiment (2-1) to which the present invention is applied.

[0168] The encoder performs prediction based on at least one of a skip mode, a merge mode, or an inter mode (S1410).

[0169] The encoder checks whether condition A is satisfied for the current block to determine the optimal coding mode for motion vector prediction (S1420).

[0170] Here, the condition A refers to a condition for the block size. For example, the embodiment shown in Table 1 below can be applied.

[0171] [Table 1]

[0172] In Example 1 of Table 1, the condition A indicates whether the number of pixels (pixNum) of the current block is greater than a threshold (TH1), where the threshold has values ​​such as 64, 128, 256, 512, 1024, etc. For example, TH1=64 means that the block size is 4×16, 8×8, or 16×4, and TH1=128 means that the block size is 32×4, 16×8, 8×16, or 4×32.

[0173] In the case of Example 2, it indicates whether the width and height of the current block are both greater than a threshold value (TH1).

[0174] In the case of Example 3, it indicates whether the width of the current block is greater than a threshold value (TH1) or whether the height of the current block is greater than a threshold value (TH1).

[0175] If the condition A is satisfied, the encoder performs motion vector prediction based on at least one of the AF4 mode or the AF6 mode (S1430).

[0176] The encoder compares the results of steps S1410 and S1430 to determine an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode (S1440).

[0177] On the other hand, if the condition A is not satisfied, the encoder determines the optimal coding mode among the modes other than the AF mode (S1440).

[0178] Thereafter, the encoder may generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector differential value.

[0179] Hereafter, the encoding / decoding processes described in FIG. 1 and FIG. 2 can be applied in the same manner.

[0180] FIG. 15 shows a flowchart of adaptively performing decoding in AF4 mode or AF6 mode based on condition A as an embodiment (2-2) to which the present invention is applied.

[0181] The decoder receives a bitstream (S1510), which includes a video signal and information about a coding mode of a current block.

[0182] The decoder checks whether condition A is satisfied for the current block in order to determine the optimal coding mode for motion vector prediction (S1520). Here, condition A refers to a condition for the block size. For example, the embodiment of Table 1 can be applied.

[0183] If the condition A is satisfied, the decoder checks whether the coding mode of the current block is an AF mode (S1530), where the AF mode refers to an affine motion prediction mode using an affine motion model, and the embodiments described herein may be applied.

[0184] Here, step S1530 may be checked based on an affine flag indicating whether or not the AF mode is performed. For example, the affine flag may be expressed as affine_flag. If affine_flag=1, it indicates that the AF mode is performed for the current block, and if affine_flag=0, it indicates that the AF mode is not performed for the current block.

[0185] If the condition A is not satisfied or if the AF mode is not performed for the current block, the decoder may perform decoding (i.e., motion vector prediction) using a coding mode other than the AF mode (S1540), for example, a skip mode, a merge mode, or an inter mode.

[0186] If the AF mode is to be performed on the current block, the decoder checks whether the AF4 mode is to be applied to the current block (S1550).

[0187] Here, the S1550 step can be confirmed by an affine parameter flag indicating whether or not AF4 mode is performed (or whether or not affine motion prediction is performed using four parameters). For example, the affine parameter flag can be expressed as affine_param_flag. If affine_param_flag=0, it means that motion vector prediction is performed using AF4 mode (S1560), and if affine_param_flag=1, it means that motion vector prediction is performed using AF6 mode (S1570), but the present invention is not limited thereto.

[0188] FIG. 16 shows a syntax structure for performing decoding in AF4 mode or AF6 mode based on condition A as an embodiment (2-3) to which the present invention is applied.

[0189] The decoder obtains merge_flag to check whether the merge mode is applied to the current block (S1610).

[0190] If the merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S1620). Here, condition A refers to a condition on the block size. For example, the embodiment of Table 1 may be applied.

[0191] If the condition A is satisfied, the decoder may obtain affine_flag (S1620), where affine_flag indicates whether or not AF mode is performed.

[0192] If the affine_flag=1, i.e., if the AF mode is performed for the current block, the decoder can obtain affine_param_flag (S1630), where affine_param_flag indicates whether the AF4 mode is performed (or whether affine motion prediction is performed using four parameters).

[0193] When the affine_param_flag=0, that is, when motion vector prediction is performed in AF4 mode, the decoder can obtain two motion vector differential values, mvd_CP0 and mvd_CP1 (S1640), where mvd_CP0 indicates the motion vector differential value for control point 0, and mvd_CP1 indicates the motion vector differential value for control point 1.

[0194] If the affine_param_flag=1, that is, if the motion vector prediction is performed in the AF6 mode, the decoder can obtain three motion vector differential values, mvd_CP0, mvd_CP1, and mvd_CP2 (S1650).

[0195] FIG. 17 shows a flowchart for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode based on at least one of condition B or condition C, as an embodiment (3-1) to which the present invention is applied.

[0196] The present invention provides a method for adaptively selecting between AF4 and AF6 modes based on the size of the current block.

[0197] For example, the AF6 mode transmits one additional motion vector difference value compared to the AF4 mode, and is therefore effective for relatively large blocks. Therefore, if the size of the current block is smaller than (or smaller or equal to) the preset size, encoding is performed taking into consideration only the AF4 mode, and if the size of the current block is larger than (or equal to) the preset size, encoding is performed taking into consideration only the AF6 mode.

[0198] On the other hand, in the case of an area where it is not clearly determined that only one of the AF4 mode and the AF6 mode is advantageous, both the AF4 mode and the AF6 mode can be considered and only the most optimal mode can be signaled.

[0199] As shown in FIG. 17, the encoder performs prediction based on at least one of a skip mode, a merge mode, or an inter mode (S1710).

[0200] The encoder checks whether condition B is satisfied for the current block (S1720), where condition B refers to a condition for the block size. For example, the embodiment shown in Table 2 below may be applied.

[0201] [Table 2]

[0202] In Example 1 of Table 2, the condition B indicates whether the number of pixels (pixNum) of the current block is smaller than a threshold (TH2), where the threshold has values ​​such as 64, 128, 256, 512, 1024, etc. For example, TH2=64 means that the block size is 4×16, 8×8, or 16×4, and TH2=128 means that the block size is 32×4, 16×8, 8×16, or 4×32.

[0203] In the case of Example 2, the condition B indicates whether the width and height of the current block are both smaller than a threshold value (TH2).

[0204] In the case of Example 3, the condition B indicates whether the width of the current block is smaller than a threshold value (TH2) or whether the height of the current block is smaller than a threshold value (TH2).

[0205] If the condition B is satisfied, the encoder performs motion vector prediction based on the AF4 mode (S1730).

[0206] If the condition B is not satisfied, the encoder checks whether the condition C is satisfied for the current block (S1740), where the condition C is a condition for a block size. For example, the following embodiment in Table 3 may be applied.

[0207] [Table 3]

[0208] In Example 1 of Table 3, the condition A indicates whether the number of pixels (pixNum) of the current block is greater than or equal to a threshold (TH3), where the threshold has values ​​such as 64, 128, 256, 512, 1024, etc. For example, TH3=64 means that the block size is 4×16, 8×8, or 16×4, and TH3=128 means that the block size is 32×4, 16×8, 8×16, or 4×32.

[0209] In the case of Example 2, it indicates whether the width and height of the current block are both greater than or equal to a threshold value (TH3).

[0210] In the case of Example 3, it indicates whether the width of the current block is greater than or equal to a threshold value (TH1), or whether the height of the current block is greater than or equal to a threshold value (TH1).

[0211] If the condition C is satisfied, the encoder performs motion vector prediction based on the AF6 mode (S1760).

[0212] If the condition C is not satisfied, the encoder may perform motion vector prediction based on the AF4 mode and the AF6 mode (S1750).

[0213] Meanwhile, in the condition B and the condition C, the threshold value (TH2) and the threshold value (TH3) may be determined to satisfy the following Equation 5.

[0214] [Number 5] TH_2 ≦ TH_3

[0215] The encoder compares the results of steps S1710, S1730, S1750, and S1760 to determine an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode (S1770).

[0216] Thereafter, the encoder may generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector differential value.

[0217] Hereafter, the encoding / decoding processes described in FIG. 1 and FIG. 2 can be applied in the same manner.

[0218] FIG. 18 shows a flowchart of adaptively performing decoding in AF4 mode or AF6 mode based on at least one of condition B or condition C as an embodiment (3-2) to which the present invention is applied.

[0219] The decoder checks whether the coding mode of the current block is the AF mode (S1810). Here, the AF mode refers to an affine motion prediction mode using an affine motion model, and the embodiments described in this specification can be applied, so duplicated descriptions will be omitted.

[0220] If the AF mode is performed for the current block, the decoder checks whether condition B is satisfied for the current block (S1820). Here, condition B refers to a condition for a block size. For example, the embodiment of Table 2 may be applied, and redundant description will be omitted.

[0221] If the condition B is satisfied, the decoder performs motion vector prediction based on the AF4 mode (S1830).

[0222] If the condition B is not satisfied, the decoder checks whether the condition C is satisfied for the current block (S1840). Here, the condition C refers to a condition for a block size. For example, the embodiment of Table 3 can be applied, and the duplicated description will be omitted.

[0223] Meanwhile, in the condition B and the condition C, the thresholds (TH2) and (TH3) may be determined to satisfy Equation 5.

[0224] If the condition C is satisfied, the decoder performs motion vector prediction based on the AF6 mode (S1860).

[0225] If the condition C is not satisfied, the decoder checks whether the AF4 mode is applied to the current block (S1850).

[0226] Here, the step S1850 can be confirmed by an affine parameter flag indicating whether or not the AF4 mode is performed (or whether or not affine motion prediction is performed using four parameters).

[0227] For example, the affine parameter flag may be expressed as affine_param_flag. When affine_param_flag=0, motion vector prediction is performed in AF4 mode (S1830), and when affine_param_flag=1, motion vector prediction is performed in AF6 mode (S1860), but the present invention is not limited thereto.

[0228] On the other hand, if the AF mode is not performed for the current block, the decoder performs decoding (i.e., motion vector prediction) according to a coding mode other than the AF mode (S1870). For example, a skip mode, a merge mode, or an inter mode may be used.

[0229] FIG. 19 shows a syntax structure for performing decoding in AF4 mode or AF6 mode based on at least one of condition B or condition C as an embodiment (3-3) to which the present invention is applied.

[0230] The decoder obtains merge_flag to check whether the merge mode is applied to the current block (S1910).

[0231] If merge mode is not applied to the current block, the decoder obtains affine_flag (S1920), where affine_flag indicates whether AF mode is performed.

[0232] If affine_flag=1, that is, if the AF mode is performed for the current block, the decoder checks whether condition B is satisfied (S1620). Here, condition B refers to a condition for the block size. For example, the embodiment of Table 2 above can be applied.

[0233] If the condition B is satisfied, the decoder sets affine_param_flag to 0 (S1930). Here, affine_param_flag indicates whether or not AF4 mode is performed (or whether or not affine motion prediction is performed using four parameters). affine_param_flag=0 means that motion vector prediction is performed using AF4 mode.

[0234] If condition B is not satisfied but condition C is satisfied, the decoder sets affine_param_flag to 1 (S1940), where affine_param_flag=1 means that motion vector prediction is performed in AF6 mode.

[0235] On the other hand, if the condition B is not satisfied and the condition C is not satisfied, the decoder can obtain affine_param_flag (S1950).

[0236] When the affine_param_flag=0, the decoder can obtain two motion vector differential values, mvd_CP0 and mvd_CP1 (S1960).

[0237] When the affine_param_flag=1, the decoder can obtain three motion vector differential values, mvd_CP0, mvd_CP1, and mvd_CP2 (S1970).

[0238] FIG. 20 shows a flowchart for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode based on the coding mode of a neighboring block, as an embodiment (4-1) to which the present invention is applied.

[0239] The encoder performs prediction based on at least one of a skip mode, a merge mode, or an inter mode (S2010).

[0240] The encoder checks whether the neighboring block is coded in AF mode (S2020). Here, whether the neighboring block is coded in AF mode can be expressed by isNeighborAffine(). For example, isNeighborAffine()=0 means that the neighboring block is not coded in AF mode, and isNeighborAffine()=1 means that the neighboring block is coded in AF mode.

[0241] If the neighboring block is not coded in AF mode, the encoder performs motion vector prediction based on AF4 mode (S2030).

[0242] If the neighboring block is coded in AF mode, the encoder performs motion vector prediction based on AF4 mode and also performs motion vector prediction based on AF6 mode (S2040).

[0243] The encoder compares the results of steps S2030 and S2040 to determine an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode (S2050).

[0244] Thereafter, the encoder may generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector differential value.

[0245] Hereafter, the encoding / decoding processes described in FIG. 1 and FIG. 2 can be applied in the same manner.

[0246] FIG. 21 shows a flowchart of adaptive decoding in AF4 mode or AF6 mode based on the coding mode of an adjacent block, as an embodiment (4-2) to which the present invention is applied.

[0247] The decoder may receive a bitstream (S2110), the bitstream including information regarding a coding mode of a current block in a video signal.

[0248] The decoder checks whether the coding mode of the current block is the AF mode (S2120).

[0249] If the AF mode is not performed for the current block, the decoder performs decoding (i.e., motion vector prediction) according to a coding mode other than the AF mode (S2170). For example, a skip mode, a merge mode, or an inter mode may be used.

[0250] When the AF mode is performed on the current block, the decoder checks whether the neighboring block is coded in the AF mode (S2130). Here, whether the neighboring block is coded in the AF mode can be expressed by isNeighborAffine(). For example, if isNeighborAffine()=0, it means that the neighboring block is not coded in the AF mode, and if isNeighborAffine()=1, it means that the neighboring block is coded in the AF mode.

[0251] If the neighboring block is coded in AF mode, the decoder performs motion vector prediction based on AF4 mode (S2140).

[0252] If the neighboring block is not coded in the AF mode, the decoder checks whether the AF4 mode is applied to the current block (S2150).

[0253] Here, the S2150 step can be confirmed by an affine parameter flag indicating whether or not the AF4 mode is performed (or whether or not affine motion prediction is performed using four parameters). For example, the affine parameter flag can be expressed as affine_param_flag. If the affine_param_flag=0, motion vector prediction is performed using the AF4 mode (S2140), and if the affine_param_flag=1, motion vector prediction is performed using the AF6 mode (S2160).

[0254] FIG. 22 shows a syntax structure for performing decoding in AF4 mode or AF6 mode based on the coding mode of a neighboring block as an embodiment (4-3) to which the present invention is applied.

[0255] The decoder obtains merge_flag to check whether the merge mode is applied to the current block (S2210).

[0256] If merge mode is not applied to the current block, the decoder may obtain affine_flag (S2220), where affine_flag indicates whether AF mode is performed.

[0257] If the affine_flag=1, that is, if the AF mode is performed for the current block, the decoder checks whether the neighboring block is coded in the AF mode (S2230).

[0258] If the neighboring block is coded in AF mode, the decoder may obtain affine_param_flag (S2230), where affine_param_flag indicates whether AF4 mode is performed (or whether affine motion prediction is performed with four parameters).

[0259] If the neighboring block is not coded in AF mode, the decoder sets affine_param_flag to 0 (S2240).

[0260] If the affine_param_flag=0, that is, if the motion vector prediction is performed in the AF4 mode, the decoder can obtain two motion vector differential values, mvd_CP0 and mvd_CP1 (S2250).

[0261] If the affine_param_flag=1, that is, if the motion vector prediction is performed in the AF6 mode, the decoder can obtain three motion vector differential values, mvd_CP0, mvd_CP1, and mvd_CP2 (S2260).

[0262] FIG. 23 shows a flowchart for adaptively determining an optimal coding mode from among motion vector prediction modes including AF4 mode or AF6 mode based on at least one of condition A, condition B, or condition C, as an embodiment (5-1) to which the present invention is applied.

[0263] The present invention shows an embodiment that combines embodiment 2 and embodiment 3. In Fig. 23, an example in which all of conditions A, B, and C are considered is described, and the order of the conditions can be different.

[0264] As shown in FIG. 23, the encoder performs prediction based on at least one of a skip mode, a merge mode, or an inter mode (S2310).

[0265] The encoder checks whether condition A is satisfied for the current block (S2320), where condition A refers to a condition for the block size, and the embodiment of Table 1 can be applied.

[0266] If the condition A is satisfied, the encoder determines the optimal coding mode among the modes excluding the AF mode (S2380).

[0267] On the other hand, if the condition A is not satisfied, the encoder checks whether the condition B is satisfied for the current block (S2330), where the condition B refers to a condition for a block size, and the embodiment of Table 2 can be applied.

[0268] If the condition B is satisfied, the encoder performs motion vector prediction based on the AF4 mode (S2340).

[0269] If condition B is not satisfied, the encoder checks whether condition C is satisfied for the current block (S2350), where condition C refers to a condition for a block size, and the embodiment of Table 3 can be applied.

[0270] If condition C is satisfied, the encoder performs motion vector prediction based on AF6 mode (S2370).

[0271] If the condition C is not satisfied, the encoder performs motion vector prediction based on the AF4 mode and also performs motion vector prediction based on the AF6 mode (S2360).

[0272] Meanwhile, in the condition B and the condition C, the thresholds (TH2) and (TH3) may be determined to satisfy Equation 5.

[0273] The encoder compares the results of steps S2310, S2340, S2360, and S2370 to determine the optimal coding mode (S2380).

[0274] Thereafter, the encoder may generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector differential value.

[0275] Hereafter, the encoding / decoding processes described in FIG. 1 and FIG. 2 can be applied in the same manner.

[0276] FIG. 24 shows a flowchart of an embodiment (5-2) to which the present invention is applied, in which adaptive decoding is performed in AF4 mode or AF6 mode based on at least one of condition A, condition B, and condition C.

[0277] The decoder checks whether condition A is satisfied for the current block (S2410), where condition A refers to a condition for the block size. For example, the embodiment of Table 1 above can be applied.

[0278] If the condition A is satisfied, the decoder checks whether the coding mode of the current block is the AF mode (S2420). Here, the AF mode refers to an affine motion prediction mode using an affine motion model, and the embodiments described in this specification are applicable, so duplicated descriptions will be omitted.

[0279] If the condition A is not satisfied or if the AF mode is not performed for the current block, the decoder performs decoding (i.e., motion vector prediction) according to a coding mode other than the AF mode (S2480). For example, a skip mode, a merge mode, or an inter mode may be used.

[0280] If the AF mode is performed for the current block, the decoder checks whether condition B is satisfied for the current block (S2430). Here, condition B refers to a condition for a block size. For example, the embodiment of Table 2 may be applied, and redundant description will be omitted.

[0281] If the condition B is satisfied, the decoder performs motion vector prediction based on the AF4 mode (S2440).

[0282] If the condition B is not satisfied, the decoder checks whether the condition C is satisfied for the current block (S2450). Here, the condition C refers to a condition for a block size. For example, the embodiment of Table 3 can be applied, and the duplicated description will be omitted.

[0283] Meanwhile, in the condition B and the condition C, the thresholds (TH2) and (TH3) may be determined to satisfy Equation 5.

[0284] If the condition C is satisfied, the decoder performs motion vector prediction based on the AF6 mode (S2470).

[0285] If the condition C is not satisfied, the decoder checks whether the AF4 mode is applied to the current block (S2460).

[0286] Here, step S2460 can be confirmed by an affine parameter flag indicating whether AF4 mode is performed (or whether affine motion prediction is performed using four parameters).

[0287] For example, the affine parameter flag can be expressed as affine_param_flag. When affine_param_flag=0, it means that motion vector prediction is performed in AF4 mode (S2440), and when affine_param_flag=1, it means that motion vector prediction is performed in AF6 mode (S2470), but the present invention is not limited thereto.

[0288] FIG. 25 shows a syntax structure for performing decoding in AF4 mode or AF6 mode based on at least one of condition A, condition B, or condition C, as an embodiment (5-3) to which the present invention is applied.

[0289] The decoder obtains merge_flag to check whether the merge mode is applied to the current block (S2510).

[0290] If the merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S2520). Here, condition A refers to a condition on the block size. For example, the embodiment of Table 1 can be applied.

[0291] If the condition A is satisfied, the decoder may obtain affine_flag (S2520), where affine_flag indicates whether AF mode is performed.

[0292] If affine_flag=1, that is, if the AF mode is performed for the current block, the decoder checks whether condition B is satisfied (S2530). Here, condition B refers to a condition for the block size. For example, the embodiment of Table 2 above can be applied.

[0293] If the condition B is satisfied, the decoder sets affine_param_flag to 0 (S2540). Here, affine_param_flag indicates whether or not AF4 mode is performed (or whether or not affine motion prediction is performed using four parameters). affine_param_flag=0 means that motion vector prediction is performed using AF4 mode.

[0294] If condition B is not satisfied but condition C is satisfied, the decoder sets affine_param_flag to 1 (S2550), where affine_param_flag=1 means that motion vector prediction is performed in AF6 mode.

[0295] On the other hand, if the condition B is not satisfied and the condition C is not satisfied, the decoder can obtain affine_param_flag (S2560).

[0296] When the affine_param_flag=0, the decoder can obtain two motion vector differential values, mvd_CP0 and mvd_CP1 (S2570).

[0297] When the affine_param_flag=1, the decoder can obtain three motion vector differential values, mvd_CP0, mvd_CP1, and mvd_CP2 (S2580).

[0298] FIG. 26 shows a flowchart for adaptively determining an optimal coding mode among motion vector prediction modes including AF4 mode or AF6 mode based on at least one of condition A or the coding mode of an adjacent block as an embodiment (6-1) to which the present invention is applied.

[0299] The encoder may perform prediction based on at least one of a skip mode, a merge mode, or an inter mode (S2610).

[0300] The encoder checks whether condition A is satisfied for the current block (S2620), where condition A refers to a condition for the block size, and the embodiment of Table 1 can be applied.

[0301] If the condition A is satisfied, the encoder determines the optimal coding mode from among the modes excluding the AF mode (S2660).

[0302] On the other hand, if the condition A is not satisfied, the encoder checks whether the neighboring block is coded in AF mode (S2630). Here, whether the neighboring block is coded in AF mode can be expressed by isNeighborAffine(). For example, isNeighborAffine()=0 means that the neighboring block is not coded in AF mode, and isNeighborAffine()=1 means that the neighboring block is coded in AF mode.

[0303] If the neighboring block is not coded in AF mode, the encoder performs motion vector prediction based on AF4 mode (S2640).

[0304] If the neighboring block is coded in AF mode, the encoder performs motion vector prediction based on AF4 mode and also performs motion vector prediction based on AF6 mode (S2650).

[0305] The encoder compares the results of steps S2610, S2640 and S2650 to determine the optimal coding mode (S2660).

[0306] Thereafter, the encoder may generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector differential value.

[0307] Hereafter, the encoding / decoding processes described in FIG. 1 and FIG. 2 can be applied in the same manner.

[0308] FIG. 27 shows a flowchart of adaptive decoding in AF4 mode or AF6 mode based on at least one of condition A or the coding mode of an adjacent block as an embodiment (6-2) to which the present invention is applied.

[0309] The decoder may receive a bitstream (S2710), the bitstream including information regarding the coding mode of a current block in a video signal.

[0310] The decoder checks whether condition A is satisfied for the current block in order to determine the optimal coding mode for motion vector prediction (S2720). Here, condition A refers to a condition for the block size. For example, the embodiment of Table 1 can be applied.

[0311] If the condition A is satisfied, the decoder checks whether the coding mode of the current block is the AF mode (S2730).

[0312] Hereinafter, steps S2730 to S2780 may be the same as those described in steps S2120 to S2170 of FIG. 21, and duplicated descriptions will be omitted.

[0313] FIG. 28 shows a syntax structure for performing decoding in AF4 mode or AF6 mode based on at least one of condition A or the coding mode of a neighboring block as an embodiment (6-3) to which the present invention is applied.

[0314] The decoder obtains merge_flag to check whether the merge mode is applied to the current block (S2810).

[0315] If the merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S2820). Here, condition A refers to a condition on the block size. For example, the embodiment of Table 1 above may be applied.

[0316] If the condition A is satisfied, the decoder may obtain affine_flag (S2820), where affine_flag indicates whether or not AF mode is performed.

[0317] Hereinafter, steps S2830 to S2860 may be the same as those described in steps S2230 to S2260 of FIG. 22, and duplicated descriptions will be omitted.

[0318] FIG. 29 shows a flowchart for generating a motion vector predictor based on at least one of the AF4 mode and the AF6 mode as an embodiment to which the present invention is applied.

[0319] The decoder checks whether the AF mode is applied to the current block (S2910), where the AF mode indicates a motion prediction mode using an affine motion model.

[0320] For example, the decoder may obtain an affine flag from a video signal, and may determine whether the AF mode is applied to the current block based on the affine flag.

[0321] When the AF mode is applied to the current block, the decoder checks whether the AF4 mode is used (S2920), where the AF4 mode indicates a mode of predicting a motion vector using four parameters constituting the affine motion model.

[0322] For example, when the AF mode is applied to the current block by the affine flag, the decoder can obtain an affine parameter flag from the video signal, the affine parameter flag indicating whether the motion vector predictor is generated using the four parameters or the six parameters.

[0323] Here, the affine flag and the affine parameter flag may be defined at least at one level of a slice, a maximum coding unit, a coding unit or a prediction unit.

[0324] If AF4 mode is used, the decoder generates a motion vector predictor using the four parameters, and if AF4 mode is not used, the decoder generates a motion vector predictor using six parameters that constitute the affine motion model (S2930).

[0325] The decoder may obtain a motion vector for the current block based on the motion vector predictor (S2940).

[0326] In one embodiment, the decoder may check whether the size of the current block satisfies a pre-defined condition, where the pre-defined condition indicates whether at least one of the number of pixels in the current block, the width and / or the height of the current block is greater than a pre-defined threshold.

[0327] For example, if the size of the current block satisfies a pre-set condition, the decoder may check whether the AF mode is applied to the current block.

[0328] On the other hand, if the size of the current block does not satisfy the pre-set condition, the current block may be decoded based on a coding mode other than the AF mode.

[0329] In one embodiment, when the AF mode is applied to the current block, the decoder may check whether the AF mode is applied to a neighboring block.

[0330] If AF mode is applied to the neighboring block, a motion vector predictor is generated using the four parameters, and if AF mode is not applied to the neighboring block, the decoder performs a step of checking whether the AF4 mode is used.

[0331] FIG. 30 shows a flowchart for generating a motion vector predictor based on AF4_flag and AF6_flag as an embodiment to which the present invention is applied.

[0332] The decoder obtains at least one of AF4_flag and AF6_flag from the video signal (S3010), where AF4_flag indicates whether or not the AF4 mode is performed for the current block, and AF6_flag indicates whether or not the AF6 mode is performed for the current block.

[0333] Here, at least one of the AF4_flag and the AF6_flag may be defined at a slice level, or may be defined at a block level or a prediction unit level, but the present invention is not limited thereto, and at least one of the AF4_flag and the AF6_flag may be defined at at least one level of a slice, a maximum coding unit, a coding unit, or a prediction unit.

[0334] The decoder checks the values ​​of AF4_flag and AF6_flag (S3020).

[0335] When AF4_flag=1, the AF4 mode is performed for the current block, and when AF4_flag=0, the AF4 mode is not performed for the current block. Here, performing the AF4 mode means performing motion vector prediction using an affine motion model expressed by four parameters.

[0336] When AF6_flag=1, the AF6 mode is performed for the current block, and when AF4_flag=0, the AF6 mode is not performed for the current block. Here, performing the AF6 mode means performing motion vector prediction using an affine motion model expressed by four parameters.

[0337] If AF4_flag=0 and AF6_flag=0, the decoder performs motion vector prediction in a mode other than the AF4 mode and the AF6 mode (S3030).

[0338] If AF4_flag=1 and AF6_flag=0, the decoder performs motion vector prediction in the AF4 mode (S3040).

[0339] If AF4_flag=0 and AF6_flag=1, the decoder performs motion vector prediction in the AF6 mode (S3050).

[0340] If AF4_flag=1 and AF6_flag=1, the decoder performs motion vector prediction in the AF4 mode or the AF6 mode (S3060).

[0341] FIG. 31 shows a flowchart of an embodiment of the present invention in which adaptive decoding is performed in AF4 mode or AF6 mode depending on whether a neighboring block is coded in AF mode.

[0342] The decoder checks whether the AF mode is applied to the current block (S3110).

[0343] If the AF mode is applied to the current block, the decoder checks whether the neighboring block is coded in the AF mode (S3120).

[0344] The decoder may obtain at least one of AF4_flag or AF6_flag if the neighboring block is coded in AF mode (S3130).

[0345] The decoder generates a motion vector predictor using four or six parameters based on at least one of AF4_flag or AF6_flag (S3140). For example, if AF4_flag=1, the decoder can perform motion vector prediction in AF4 mode, and if AF6_flag=1, the decoder can perform motion vector prediction in AF6 mode.

[0346] The decoder may obtain a motion vector for the current block based on the motion vector predictor (S3150).

[0347] FIG. 32 shows a syntax for adaptively decoding based on AF4_flag and AF6_flag as an embodiment to which the present invention is applied.

[0348] The decoder can obtain AF4_flag and AF6_flag at the slice level (S3010). Here, AF4_flag indicates whether the AF4 mode is performed for the current block, and AF6_flag indicates whether the AF6 mode is performed for the current block. The AF4_flag can be expressed by affine_4_flag, and the AF6_flag can be expressed by affine_6_flag.

[0349] The decoder can perform adaptive decoding based on AF4_flag and AF6_flag at the block level or prediction unit level.

[0350] If affine_4_flag is not 0 or affine_6_flag is not 0 (ie, other than the case where affine_4_flag=0 && affine_6_flag=0), the decoder can obtain the affine flag (S3220). The affine flag indicates whether or not the AF mode is performed.

[0351] When the AF mode is performed, the decoder can perform adaptive decoding according to the AF4_flag and AF6_flag values.

[0352] If affine_4_flag=1&& affine_6_flag=0, the decoder can set affine_param_flag to 0. That is, affine_param_flag=0 means that the AF4 mode is performed.

[0353] If affine_4_flag=0&& affine_6_flag=1, the decoder can set affine_param_flag to 1. That is, affine_param_flag=1 means that the AF6 mode is performed.

[0354] If affine_4_flag=1&&affine_6_flag=1, the decoder may parse or obtain affine_param_flag, where the decoder may perform decoding in AF4 mode or AF6 mode depending on the affine_param_flag value at the block level or prediction unit level.

[0355] In addition, the syntax structure may be the same as that of the above-mentioned embodiment, and a duplicated description will be omitted.

[0356] FIG. 33 shows a syntax for adaptively decoding in AF4 mode or AF6 mode depending on whether a neighboring block is coded in AF mode, as an embodiment to which the present invention is applied.

[0357] In the case of this embodiment, the same contents as those in FIG. 32 can be applied to the above description, and only the different parts will be described.

[0358] If affine_4_flag=1&& affine_6_flag=1, the decoder can check whether the neighboring block is coded in AF mode.

[0359] If the neighboring block is coded in AF mode, the decoder parses or obtains affine_param_flag (S3310), where the decoder can perform decoding in AF4 mode or AF6 mode depending on the affine_param_flag value at the block level or prediction unit level.

[0360] On the other hand, if the neighboring block is not coded in AF mode, the decoder may set affine_param_flag to 0. That is, affine_param_flag=0 means that AF4 mode is performed.

[0361] FIG. 34 shows a video coding system to which the present invention is applied.

[0362] A video coding system includes a source device and a receiving device. The source device transmits encoded video / image information or data in file or streaming format to the receiving device via a digital storage medium or a network.

[0363] The source device includes a video source, an encoding apparatus, and a transmitter. The receiving device includes a receiver, a decoding apparatus, and a renderer. The encoding apparatus may be called a video / image encoding apparatus, and the decoding apparatus may be called a video / image decoding apparatus. The transmitter may be included in the encoding apparatus. The receiver may be included in the decoding apparatus. The renderer may include a display unit, which may be configured as a separate device or an external component.

[0364] A video source can obtain video / images by a video / image capture, synthesis or generation process. A video source may include a video / image capture device and / or a video / image generation device. A video / image capture device includes, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device includes, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated by a computer, etc., in which case a process where related data is generated can replace the video / image capture process.

[0365] An encoding device encodes the input video / image. The encoding involves a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) is output in the form of a bitstream.

[0366] The transmitting unit transmits the encoded video / image information or data output in a bitstream format to a receiving unit of a receiving device via a digital storage medium or a network in a file or streaming format. The digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit includes elements for generating a media file according to a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit extracts the bitstream and transmits it to a decoding device.

[0367] The decoding device decodes the video / image by carrying out a series of steps such as inverse quantization, inverse transformation, prediction, etc., which correspond to the operations of the encoding device.

[0368] The renderer renders the decoded video / images, which are then displayed via a display unit.

[0369] FIG. 35 shows a content streaming system to which the present invention is applied.

[0370] As shown in FIG. 35, a content streaming system to which the present invention is applied mainly includes an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0371] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0372] The bitstream is generated by an encoding method or a bitstream generating method to which the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0373] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transfers the request to the streaming server, and the streaming server transmits multimedia data to the user. Here, the content streaming system may include a separate control server, and in this case, the control server controls commands / responses between devices in the content streaming system.

[0374] The streaming server receives content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.

[0375] Examples of the user devices include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, and the like.

[0376] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.

[0377] As mentioned above, the embodiments described in the present invention may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units shown in Figures 1, 2, 34, and 35 may be implemented and performed on a computer, processor, microprocessor, controller, or chip.

[0378] In addition, the decoders and encoders to which the present invention is applied can be included in multimedia broadcast transmitting / receiving devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, custom video (VoD) service providing devices, Internet streaming service providing devices, three-dimensional (3D) video devices, image telephone video devices, and medical video devices, and can be used to process video signals and data signals.

[0379] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present invention can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices in which computer-readable data is stored. The computer-readable storage medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium realized in the form of a carrier wave (e.g., transmission via the Internet). Also, a bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. [Industrial Applicability]

[0380] The above-described preferred embodiments of the present invention have been disclosed for illustrative purposes, and those skilled in the art may improve, modify, substitute or add various other embodiments within the technical spirit and scope of the present invention disclosed in the appended claims.

Claims

1. 1. A method for decoding a video signal including a current block based on an affine motion prediction mode (AF mode), comprising: obtaining a merge flag from the video signal, the merge flag indicating whether a merge mode is applied to the current block, and a set of available candidates in the merge mode consists of spatially neighboring candidates, temporal candidates, and generated candidates; obtaining an affine flag from the video signal based on the merge mode not being applied to the current block and based on the width and height of the current block being equal to or greater than 16, the affine flag indicating whether the AF mode is applied to the current block or not, the AF mode indicating a motion prediction mode using an affine motion model; obtaining an affine parameter flag indicating whether four parameters or six parameters are used for the affine motion model based on the AF mode applied to the current block; obtaining a motion vector predictor based on whether the four parameters or the six parameters are used for the affine motion model; obtaining a prediction sample for the current block based on the motion vector predictor; obtaining a residual sample for the current block; reconstructing the current block based on the predicted sample and the residual sample; filtering the reconstructed current block; The method of claim 1, wherein the affine flag and the affine parameter flag are obtained based on the merge flag indicating that the merge mode does not apply to the current block.

2. The method of claim 1 , wherein the affine flag and the affine parameter flag are defined at a coding unit level.

3. The method of claim 1 , wherein the current block is decoded based on a coding mode other than the AF mode based on the width and the height of the current block being less than 16.

4. 1. A method for encoding a video signal including a current block based on an affine motion prediction mode (AF mode), comprising: generating a merge flag indicating whether a merge mode is applied to the current block, wherein a set of available candidates in the merge mode consists of spatially neighboring candidates, temporal candidates, and generated candidates; generating an affine flag based on the width and height of the current block being equal to or greater than 16 based on the merge mode not being applied to the current block, the affine flag indicating whether the AF mode is applied to the current block, the AF mode indicating a motion prediction mode using an affine motion model; generating an affine parameter flag indicating whether four parameters or six parameters are used for the affine motion model based on which AF mode is applied to the current block; obtaining a motion vector predictor based on whether the four parameters or the six parameters are used for the affine motion model; generating a prediction sample for the current block based on the motion vector predictor; generating a residual sample for the current block based on the predicted sample; transforming, quantizing, and entropy encoding the residual samples; The current block is reconstructed based on the predicted samples and the residual samples; The method of claim 1, wherein the affine flag and the affine parameter flag are generated based on the merge flag indicating that the merge mode does not apply to the current block.

5. 1. A method for transmitting data relative to a video signal, comprising: generating a bitstream for the video signal, the bitstream being generated by an encoding method; transmitting the data including the bitstream; The encoding method includes: generating a merge flag indicating whether a merge mode is applied to the current block, in which a set of available candidates in the merge mode consists of spatially neighboring candidates, temporal candidates, and the generated candidates; generating an affine flag based on the width and height of the current block being equal to or greater than 16 based on the merge mode not being applied to the current block, the affine flag indicating whether an affine motion prediction mode (AF mode) is applied to the current block, the AF mode indicating a motion prediction mode using an affine motion model; generating an affine parameter flag indicating whether four parameters or six parameters are used for the affine motion model based on which AF mode is applied to the current block; obtaining a motion vector predictor based on whether the four parameters or the six parameters are used for the affine motion model; generating a prediction sample for the current block based on the motion vector predictor; generating a residual sample for the current block based on the predicted sample; transforming, quantizing, and entropy encoding the residual samples; The current block is reconstructed based on the predicted samples and the residual samples; The method of claim 1, wherein the affine flag and the affine parameter flag are generated based on the merge flag indicating that the merge mode does not apply to the current block.

Citation Information

Patent Citations

  • Image coding for supporting block division and block integration

    JP2016026454A

  • Motion Vector Prediction for Affine Motion Models in Video Coding

    JP2019535192A

  • JPP7393326B

  • Motion vector prediction for affine motion models in video coding

    US20180098063A1

  • Affine motion prediction in video coding

    US9438910B1

Cited By

  • Method and apparatus for processing video signals using affine prediction

    JP2025100808A