Method and apparatus for processing video signals using affine prediction
Adaptive affine prediction methods in video encoding and decoding optimize coding efficiency by selecting between AF4 and AF6 modes based on block size and adjacent block coding, addressing the challenges of high-resolution video content processing.
Patent Information
- Application Number
- JP2025069754
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-08-03
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2038-08-03
AI Technical Summary
Existing video encoding and decoding technologies face challenges in efficiently processing next-generation video content with high spatial resolution, frame rate, and scene representation, leading to increased memory storage and processing demands.
Adaptive affine prediction methods are employed, selecting between AF4 and AF6 modes based on block size, adjacent block coding, and predefined conditions to optimize coding efficiency and reduce complexity.
The adaptive affine prediction methods improve coding performance by reducing complexity and enhancing efficiency in encoding and decoding high-resolution video content.
Smart Images

Figure 2025100808000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for encoding / decoding video signals, and more particularly, to a method and apparatus for adaptively performing affine prediction.
Background Art
[0002] Compression encoding refers to a series of signal processing techniques for transmitting digitized information via a communication line or storing it in a form suitable for a storage medium. Media such as images, images, and voices are targets of compression encoding, and in particular, a technique for performing compression encoding on an image is called video image compression.
[0003] Next-generation video content has characteristics such as high spatial resolution, high frame rate, and high dimensionality of scene representation. In order to process such content, it will bring a great increase in terms of memory storage, memory access rate, and processing power.
[0004] Therefore, it is necessary to design coding tools for more efficiently processing next-generation video content.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The present invention proposes a method for more efficiently encoding and decoding video signals.
[0006] The present invention also proposes a method for encoding or decoding while considering both the AF4 mode, which is a four-parameter affine prediction mode, and the AF6 mode, which is a six-parameter affine prediction mode.
[0007] The present invention also proposes a method for adaptively determining (or selecting) an optimal coding mode based on at least one of the AF4 mode or the AF6 mode based on the block size.
[0008] The present invention also proposes a method for adaptively determining (or selecting) an optimal coding mode based on at least one of the AF4 mode or the AF6 mode based on whether adjacent blocks are coded by affine prediction.
Means for Solving the Problems
[0009] To solve the above-described technical problems,
[0010] the present invention provides a method for adaptively performing affine prediction based on the block size.
[0011] The present invention also provides a method for adaptively performing affine prediction based on whether adjacent blocks are coded by affine prediction.
[0012] The present invention also provides a method for adaptively determining (or selecting) an optimal coding mode based on at least one of the AF4 mode or the AF6 mode.
[0013] The present invention also provides a method for adaptively performing affine prediction based on whether at least one already set condition is satisfied. In this case, the already set condition may include at least one of block size, number of pixels in the block, width of the block, height of the block, and whether an adjacent block is coded by affine prediction.
Advantages of the Invention
[0014] By providing a method for adaptively performing affine prediction, the present invention can improve the performance of affine prediction and perform more efficient coding by reducing the complexity of affine prediction.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Best Mode for Carrying Out the Invention
[0016] In a method for decoding a video signal including a current block based on an affine motion prediction mode (AF mode, Affine mode), the method includes: checking whether the AF mode is applied to the current block, where the AF mode indicates a motion prediction mode using an affine motion model; when the AF mode is applied to the current block, checking whether the AF4 mode is used, where the AF4 mode indicates a mode for predicting a motion vector using four parameters constituting the affine motion model; when the AF4 mode is used, generating a motion vector predictor using the four parameters, and when the AF4 mode is not used, generating a motion vector predictor using six parameters constituting the affine motion model; and obtaining a motion vector of the current block based on the motion vector predictor.
[0017] In the present invention, the method further includes the step of obtaining an affine flag from the video signal, where the affine flag indicates whether the AF mode is applicable to the current block, and whether the AF mode is applicable to the current block is confirmed based on the affine flag.
[0018] In the present invention, when the AF mode is applicable to the current block according to the affine flag, the method further includes the step of obtaining an affine parameter flag from the video signal, where the affine parameter flag indicates whether the motion vector predictor is generated using the four parameters or the six parameters.
[0019] In the present invention, the affine flag and the affine parameter flag are defined at at least one level of a slice, a maximum coding unit, a coding unit, or a prediction unit.
[0020] In the present invention, the method further includes the step of confirming whether the condition that the size of the current block has already been set is satisfied, where the condition that has already been set indicates whether at least one of the number of pixels in the current block, the width and / or height of the current block is greater than a preset threshold, and when the condition that the size of the current block has already been set is satisfied, the step of confirming whether the AF mode is applicable to the current block is performed.
[0021] In the present invention, when the condition that the size of the current block has already been set is not satisfied, the current block is decoded based on another coding mode that is not the AF mode.
[0022] In the present invention, when the AF mode is applied to the current block, the method further includes a step of checking whether the AF mode has been applied to an adjacent block. When the AF mode has been applied to the adjacent block, a motion vector predictor is generated using the four parameters. When the AF mode has not been applied to the adjacent block, a step of checking whether the AF4 mode is used is performed.
[0023] The present invention provides an apparatus for decoding a video signal including a current block based on an affine motion prediction mode (AF mode), the apparatus including: an inter prediction unit configured to check whether the AF mode is applied to the current block, and when the AF mode is applied to the current block, check whether the AF4 mode is used; when the AF4 mode is used, generate a motion vector predictor using four parameters; when the AF4 mode is not used, generate a motion vector predictor using six parameters that constitute an affine motion model; and obtain a motion vector of the current block based on the motion vector predictor, where the AF mode indicates a motion prediction mode using the affine motion model, and the AF4 mode indicates a mode of predicting a motion vector using four parameters that constitute the affine motion model.
[0024] In the present invention, the apparatus further includes a parsing unit configured to parse an affine flag from the video signal, where the affine flag indicates whether the AF mode is applied to the current block, and whether the AF mode is applied to the current block is checked based on the affine flag.
[0025] In the present invention, when the AF mode is applied to the current block by the affinity flag, the apparatus includes the parsing unit that acquires an affinity parameter flag from the video signal, and the affinity parameter flag indicates whether the motion vector predictor is generated using the four parameters or the six parameters.
[0026] In the present invention, the apparatus includes the inter prediction unit that checks whether a condition that the size of the current block has already been set is satisfied, and the already set condition indicates whether at least one of the number of pixels in the current block, the width and / or height of the current block is greater than a preset threshold. When the condition that the size of the current block has already been set is satisfied, a step of checking whether the AF mode is applied to the current block is performed.
[0027] In the present invention, when the AF mode is applied to the current block, the apparatus includes the inter prediction unit that checks whether the AF mode has been applied to an adjacent block. When the AF mode has been applied to the adjacent block, the motion vector predictor is generated using the four parameters. When the AF mode has not been applied to the adjacent block, a step of checking whether the AF4 mode is used is performed.
Embodiments for Carrying Out the Invention
[0028] Hereinafter, the configuration and operation of embodiments of the present invention will be described with reference to the accompanying drawings. The configuration and operation of the present invention described with reference to the drawings are described as one embodiment, and thereby the technical idea, core configuration, and operation of the present invention are not limited.
[0029] In addition, the terms used in the present invention are generally selected as widely used terms as much as possible, but in specific cases, the terms arbitrarily selected by the applicant are used for explanation. In such cases, the meaning will be clearly described in the detailed description of the corresponding part, so it should be clarified that the terms used in the description of the present invention should not be simply analyzed based on the names of the terms, but should be analyzed by grasping the meaning of the corresponding terms.
[0030] Also, the terms used in the present invention are general terms selected to explain the invention, but when there are other terms having similar meanings, they can be substituted for more appropriate analysis. For example, in the case of signals, data, samples, pictures, frames, blocks, etc., they can be appropriately substituted and analyzed in each coding process. Also, in the case of partitioning, decomposition, splitting, and division, etc., they can be appropriately substituted and analyzed in each coding process.
[0031] FIG. 1 shows a schematic block diagram of an encoder in which video signal encoding is performed as an embodiment to which the present invention is applied.
[0032] As shown in FIG. 1, the encoder 100 includes an image segmentation unit 110, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, a filtering unit 160, a decoded picture buffer (DPB) 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190.
[0033] The image segmentation unit 110 divides the input image (or picture, frame) input to the encoder 100 into one or more processing units. For example, the processing unit may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU). At this time, the division is performed by at least one of the QT (QuadTree), BT (Binary Tree), TT (Ternary Tree), and AT (Asymmetric Tree) methods.
[0034] However, the above terms are only used for the convenience of explaining the present invention, and the present invention is not limited to the definitions of these terms. Also, for the convenience of explanation in this specification, the term "coding unit" is used as the unit used in the process of encoding or decoding a video signal, but the present invention is not limited thereto and can be appropriately analyzed according to the content of the invention.
[0035] The encoder 100 subtracts the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185 from the input image signal to generate a residual signal, and the generated residual signal is transmitted to the conversion unit 120.
[0036] The conversion unit 120 applies a conversion technique to the residual signal to generate transform coefficients. The conversion process may be applied to a pixel block having the same size of a square, or may be applied to a block having a variable size that is not a square.
[0037] The quantization unit 130 quantizes the transform coefficients and transmits them to the entropy encoding unit 190, and the entropy encoding unit 190 entropy-encodes the quantized signal and outputs it as a bitstream.
[0038] The quantized signal output from the quantization unit 130 may be used to generate a prediction signal. For example, the quantized signal can restore the residual signal by applying inverse quantization and inverse transformation by the inverse quantization unit 140 and the inverse transformation unit 150 within the loop. By adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185, a reconstructed signal is generated.
[0039] On the other hand, in the compression process as described above, when adjacent blocks are quantized with different quantization parameters, deterioration where the block boundary can be seen may occur. Such a phenomenon is called blocking artifacts, which is one of the important factors for evaluating image quality. A filtering process may be performed to reduce such deterioration. By such a filtering process, it is possible to remove blocking artifacts and improve the image quality by reducing the error with respect to the current picture.
[0040] The filtering unit 160 applies filtering to the reconstructed signal and outputs it to the playback device or transmits it to the decoded picture buffer 170. The filtered signal transmitted to the decoded picture buffer 170 can be used as a reference picture in the inter prediction unit 180. Thus, by using the filtered picture as a reference picture in the inter-picture prediction mode, not only the image quality but also the coding efficiency is improved.
[0041] The decoded picture buffer 170 stores the filtered picture for use as a reference picture in the inter prediction unit 180.
[0042] The inter prediction unit 180 performs temporal prediction and / or spatial prediction in order to remove temporal redundancy and / or spatial redundancy with reference to a reconstructed picture. Here, since the reference picture used for prediction is a signal that has been transformed through quantization and inverse quantization in block units during encoding / decoding at a previous time, blocking artifacts and ringing artifacts may exist.
[0043] Therefore, in order to solve such signal discontinuities and performance degradation due to quantization, the inter prediction unit 180 can interpolate the signal between pixels in sub-pixel units by applying a lowpass filter. Here, a sub-pixel means a virtual pixel generated by applying an interpolation filter, and an integer pixel means an actual pixel existing in the reconstructed picture. As the interpolation method, linear interpolation, bi-linear interpolation, a wiener filter, etc. may be applied.
[0044] The interpolation filter is applied to the reconstructed picture to improve the precision of prediction. For example, the inter prediction unit 180 can apply an interpolation filter to integer pixels to generate interpolated pixels, and use an interpolated block composed of the interpolated pixels as a prediction block to perform prediction.
[0045] The intra prediction unit 185 can predict the current block by referring to samples around the block to be currently encoded. The intra prediction unit 185 performs the following process to perform intra prediction. First, reference samples necessary for generating a prediction signal are prepared. Then, a prediction signal is generated using the prepared reference samples. Thereafter, the prediction mode is encoded. Here, the reference samples are prepared by reference sample padding and / or reference sample filtering. Since the reference samples have gone through the prediction and restoration processes, there may be quantization errors. Therefore, in order to reduce such errors, a reference sample filtering process is performed for each prediction mode used in intra prediction.
[0046] The prediction signal generated by the inter prediction unit 180 or the intra prediction unit 185 is used to generate a restored signal or to generate a residual signal.
[0047] FIG. 2 shows a schematic block diagram of a decoder in which video signal decoding is performed as an embodiment to which the present invention is applied.
[0048] As shown in FIG. 2, the decoder 200 includes a parsing unit (not shown), an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, a filtering unit 240, a decoded picture buffer (DPB) 250, an inter prediction unit 260, an intra prediction unit 265, and a restoration unit (not shown).
[0049] The decoder 200 receives the signal output from the encoder 100 of FIG. 1 and parses or acquires syntax elements via a parsing unit (not shown). The parsed or acquired signal is entropy decoded by the entropy decoding unit 210.
[0050] In the inverse quantization unit 220, transform coefficients are obtained from the entropy-decoded signal by using quantization step size information.
[0051] In the inverse transform unit 230, the transform coefficients are inversely transformed to obtain a residual signal.
[0052] A restoration unit (not shown) generates a reconstructed signal by adding the obtained residual signal to the prediction signal output from the inter prediction unit 260 or the intra prediction unit 265.
[0053] The filtering unit 240 applies filtering to the reconstructed signal and outputs it to a playback device or transmits it to the decoded picture buffer unit 250. The filtered signal transmitted to the decoded picture buffer unit 250 can be used as a reference picture in the inter prediction unit 260.
[0054] In this specification, the embodiments described in the filtering unit 160, the inter prediction unit 180, and the intra prediction unit 185 of the encoder 100 can be similarly applied to the filtering unit 240, the inter prediction unit 260, and the intra prediction unit 265 of the decoder, respectively.
[0055] The reconstructed video signal output via the decoder 200 can be played back by a playback device.
[0056] FIG. 3 is a diagram for explaining a QT (QuadTree, hereinafter referred to as "QT") block division structure as an embodiment to which the present invention can be applied.
[0057] In video coding, a block can be divided based on a QT (QuadTree). Also, one sub-block divided by QT can be recursively further divided using QT. A leaf block that cannot be further divided by QT can be divided by at least one of the methods of BT (Binary Tree), TT (Ternary Tree), or AT (Asymmetric Tree). BT has two forms of division: horizontal BT (2N×N, 2N×N) and vertical BT (N×2N, N×2N). TT has two forms of division: horizontal TT (2N×1 / 2N, 2N×N, 2N×1 / 2N) and vertical TT (1 / 2N×2N, N×2N, 1 / 2N×2N). AT has four forms of division: horizontal-up AT (2N×1 / 2N, 2N×3 / 2N), horizontal-down AT (2N×3 / 2N, 2N×1 / 2N), vertical-left AT (1 / 2N×2N, 3 / 2N×2N), and vertical-right AT (3 / 2N×2N, 1 / 2N×2N). Each of BT, TT, and AT can be recursively further divided using BT, TT, and AT.
[0058] Figure 3 shows an example of QT division. Block A can be divided into four sub-blocks (A0, A1, A2, A3) by QT. Sub-block A1 can be further divided into four sub-blocks (B0, B1, B2, B3) by QT.
[0059] Figure 4 is a diagram for explaining a BT (Binary Tree, hereinafter referred to as "BT") block division structure as an embodiment to which the present invention can be applied.
[0060] Figure 4 shows an example of BT splitting. Block B3 that cannot be further split by QT can be split by vertical BT (C0, C1) or horizontal BT (D0, D1). Each sub-block, like block C0, can be recursively further split in the form of horizontal BT (E0, E1) or vertical BT (F0, F1).
[0061] Figure 5 is a diagram for explaining a TT (Ternary Tree, hereinafter referred to as "TT") block splitting structure as an embodiment to which the present invention can be applied.
[0062] Figure 5 shows an example of TT splitting. Block B3 that cannot be further split by QT can be split by vertical TT (C0, C1, C2) or horizontal TT (D0, D1, D2). Each sub-block, like block C1, can be recursively further split in the form of horizontal TT (E0, E1, E2) or vertical TT (F0, F1, F2).
[0063] Figure 6 is a diagram for explaining an AT (Asymmetric Tree, hereinafter referred to as "AT") block splitting structure as an embodiment to which the present invention can be applied.
[0064] Figure 6 shows an example of AT splitting. Block B3 that cannot be further split by QT can be split by vertical AT (C0, C1) or horizontal AT (D0, D1). Each sub-block, like block C1, can be recursively further split in the form of horizontal AT (E0, E1) or vertical TT (F0, F1).
[0065] On the one hand, the BT, TT, and AT partitions can all be used for partitioning. For example, the sub-blocks partitioned by BT can be further partitioned by TT or AT. Also, the sub-blocks partitioned by TT can be further partitioned by BT or AT. The sub-blocks partitioned by AT can be further partitioned by BT or TT. For example, after horizontal BT partitioning, each sub-block can be further partitioned by vertical BT, or after vertical BT partitioning, each sub-block can be further partitioned by horizontal BT. Although the two partitioning methods have different partitioning orders, the finally partitioned shapes are the same.
[0066] In addition, when a block is partitioned, the order of searching the block can be defined in various ways. Generally, searching from left to right and from top to bottom means the order of determining whether to perform additional block partitioning on each partitioned sub-block, or when the block is no longer partitioned, it means the encoding order of each sub-block, or the searching order when referring to the information of other adjacent blocks in the sub-block.
[0067] FIG. 7 is a diagram for explaining the inter prediction mode as an embodiment to which the present invention is applied.
[0068] Inter-prediction mode
[0069] In the inter prediction mode to which the invention is applied, in order to reduce the amount of motion information, the Merge mode, the AMVP (Advanced Motion Vector Prediction) mode, or the Affine prediction mode (hereinafter referred to as the "AF mode") is used.
[0070] 1) Merge mode
[0071] The merge mode means a method of deriving motion parameters (or information) from spatially or temporally adjacent blocks.
[0072] The set of candidates available in the merge mode consists of spatial neighbor candidates, temporal candidates, and generated candidates.
[0073] As shown in Fig. 7(a), it is determined whether each spatial candidate block is available in the order of {A1, B1, B0, A0, B2}. At that time, if the candidate block is encoded in the intra prediction mode and there is no motion information, or if the candidate block is located outside the current picture (or slice), the candidate block cannot be used.
[0074] After determining the validity of the spatial candidates, spatial merge candidates can be constructed by removing unnecessary candidate blocks from the candidate blocks of the current processing block. For example, if the candidate block of the current prediction block is the first prediction block within the same coding block, the candidate block can be removed, and candidate blocks having the same motion information can also be removed.
[0075] When the construction of the spatial merge candidates is completed, the process of constructing the temporal merge candidates is performed in the order of {T0, T1}.
[0076] In the temporal candidate composition, if the right bottom block (T0) of the collocated block at the same position of the reference picture is available, configure the block as a temporal merge candidate. The collocated block means the block existing at the position corresponding to the current processing block in the selected reference picture. On the contrary, if not, configure the block (T1) located at the center of the collocated block as a temporal merge candidate.
[0077] The maximum number of merge candidates can be specified in the slice header. If the number of merge candidates is greater than the maximum number, the spatial candidates and temporal candidates of a number smaller than the maximum number are maintained. Otherwise, the number of merge candidates is combined with the candidates added so far until the number of candidates reaches the maximum number, and additional merge candidates (i.e., combined bi-predictive merging candidates) are generated.
[0078] In the encoder, a merge candidate list is configured in the method as described above, and the candidate block information selected from the merge candidate list is signaled to the decoder as a merge index (e.g., merge_idx[x0][y0]') by performing motion estimation. In (b) of FIG. 7, the case where the B1 block is selected from the merge candidate list is illustrated. In this case, "Index 1" is signaled to the decoder as the merge index.
[0079] In the decoder, a merge candidate list is configured in the same way as the encoder, and the motion information regarding the current block is derived from the motion information of the candidate block corresponding to the merge index received from the encoder in the merge candidate list. Then, the decoder generates a prediction block for the current processing block based on the derived motion information.
[0080] 2) AMVP (Advanced Motion Vector Prediction) mode
[0081] The AMVP mode means a method of deriving a motion vector prediction value from surrounding blocks. Therefore, the horizontal and vertical motion vector difference values (MVD: motion vector difference), reference index, and inter prediction mode are signaled to the decoder. The horizontal and vertical motion vector values are calculated using the derived motion vector prediction value and the motion vector difference value (MVD: motion vector difference) provided from the encoder.
[0082] That is, in the encoder, a motion vector prediction value candidate list is configured, and a motion reference flag (i.e., candidate block information) (e.g., mvp_lX_flag[x0][y0]’) selected from the motion vector prediction value candidate list is signaled to the decoder by performing motion estimation. In the decoder, a motion vector prediction value candidate list is configured in the same way as the encoder, and the motion information of the candidate block indicated by the motion reference flag received from the encoder in the motion vector prediction value candidate list is used to derive the motion vector prediction value of the current processing block. Then, the decoder uses the derived motion vector prediction value and the motion vector difference value transmitted from the encoder to obtain the motion vector value for the current processing block. Then, the decoder generates a prediction block for the current processing block based on the derived motion information (i.e., motion compensation).
[0083] In the case of the AMVP mode, two spatial motion candidates out of the five available candidates described above with reference to FIG. 7 are selected. The first spatial motion candidate is selected from the set {A0, A1} located on the left side, and the second spatial motion candidate is selected from the set {B0, B1, B2} located above. Here, if the reference index of adjacent candidate blocks is not the same as the current prediction block, the motion vector is scaled.
[0084] If the number of candidates selected from the search results of spatial motion candidates is two, the candidate configuration is terminated, but if it is less than two, temporal motion candidates are added.
[0085] The decoder (e.g., the inter prediction unit) decodes the motion parameters for the processing block (e.g., the prediction unit).
[0086] For example, when the processing block uses the merge mode, the decoder can decode the merge index signaled from the encoder. Then, the motion parameters of the current processing block can be derived from the motion parameters of the candidate block indicated by the merge index.
[0087] Also, when the AMVP mode is applied to the processing block, the decoder can decode the horizontal and vertical motion vector difference values (MVD), reference index, and inter prediction mode signaled from the encoder. Then, a motion vector prediction value is derived from the motion parameters of the candidate block indicated by the motion reference flag, and the motion vector value of the current processing block can be derived using the motion vector prediction value and the received motion vector difference value.
[0088] The decoder performs motion compensation for the prediction unit using the decoded motion parameters (or information).
[0089] That is, in the encoder / decoder, motion compensation is performed to predict the image of the current unit from the previously decoded picture using the decoded motion parameters.
[0090] 3) AF Mode (Affine Mode)
[0091] The AF mode means a motion prediction mode using an affine motion model, and may include at least one of an affine merge mode or an affine inter mode. The affine inter mode may include at least one of an AF4 mode or an AF6 mode. Here, the AF4 mode indicates an affine prediction (Four parameter affine prediction) mode using four parameters, and the AF6 mode indicates an affine prediction (six parameter affine prediction) mode using six parameters.
[0092] However, in the present invention, for convenience of explanation, it is expressed as the AF4 mode or the AF6 mode, but this does not necessarily need to be defined as a separate prediction mode, and the AF4 mode or the AF6 mode can be understood as being distinguished by whether four parameters or six parameters are used.
[0093] The AF mode will be described in more detail with reference to FIGS. 8 to 10.
[0094] FIG. 8 is a diagram for explaining an affine motion model as an embodiment to which the present invention is applied.
[0095] General image coding techniques use a translation motion model to represent the motion of coding blocks. Here, the translation motion model indicates a prediction method based on a translated block. That is, the motion information of a coding block is represented using one motion vector. However, the optimal motion vector for each pixel within an actual coding block may be different. If the optimal motion vector can be determined for each pixel or for each sub-block with only a small amount of information, the coding efficiency can be improved.
[0096] Therefore, the present invention proposes an inter-prediction-based image processing method that reflects various motions of an image, in addition to a prediction method based on a translated block, in order to enhance the performance of inter-prediction.
[0097] In addition, the present invention proposes an affine motion prediction method that performs encoding / decoding using an affine motion model. The affine motion model indicates a prediction method that derives a motion vector for each pixel or for each sub-block using the motion vectors of control points. In this specification, the affine motion prediction mode using the affine motion model (affine motion model) is referred to as the AF mode (Affine Mode).
[0098] In addition, the present invention provides a method for adaptively performing affine prediction based on the block size.
[0099] In addition, the present invention provides a method for adaptively performing affine prediction based on whether adjacent blocks are coded by affine prediction.
[0100] Furthermore, the present invention provides a method for adaptively determining (or selecting) an optimal coding mode based on at least one of the AF4 mode or the AF6 mode. Here, the AF4 mode represents a four-parameter affine prediction mode that utilizes four parameters, and the AF6 mode represents a six-parameter affine prediction mode that utilizes six parameters.
[0101] As shown in FIG. 8, various methods are used to represent the distortion of an image as motion information. In particular, the affine motion model can represent the four motions shown in FIG. 8.
[0102] For example, the affine motion model can model not only the translation, scale, rotation, and shear of an image, but also any induced image distortion.
[0103] The affine motion model can be represented in various ways. Among them, the present invention proposes a method of utilizing the motion information at a specific reference point (or reference pixel / sample) of a block to display (or identify) the distortion and performing inter prediction using this. Here, the reference point is referred to as a control point (CP), and the motion vector at such a reference point is referred to as a control point motion vector (CPMV). The degree of distortion that can be represented varies according to the number of such control points.
[0104] The affine motion model is represented using six parameters (a, b, c, d, e, f) as shown in the following Equation 1.
[0105]
Equation
[0106] Here, (x, y) indicates the position of the upper left pixel of the coding block. And, v x and v y respectively indicate the motion vectors at (x, y).
[0107] FIG. 9 is a diagram for explaining an affine motion prediction method using a control point motion vector as an embodiment to which the present invention is applied.
[0108] As shown in FIG. 9(a), the upper left control point (CP0) 902 (hereinafter referred to as the first control point), the upper right control point (CP1) 903 (hereinafter referred to as the second control point), and the lower left control point (CP2) 904 (hereinafter referred to as the third control point) of the current block 901 can each have independent motion information. These are respectively expressed as CP0, CP1, and CP2. However, this corresponds to one embodiment of the present invention, and the present invention is not limited thereto. For example, various control points such as the lower right control point, the center control point, and other control points according to the position of the sub-block can be defined.
[0109] As one embodiment of the present invention, at least one of the first to third control points may be a pixel included in the current block. Or, as another example, at least one of the first to third control points may be a pixel adjacent to the current block that is not included in the current block.
[0110] Using the motion information of one or more of the control points, the motion information for each pixel or sub-block of the current block 901 can be induced.
[0111] For example, an affine motion model using the motion vectors of the upper left control point 902, the upper right control point 903, and the lower left control point 904 of the current block 901 is defined as the following Equation 2.
[0112]
Equation
[0113] Here, Let the motion vector of JPEG2025100808000004.jpg19150 be the motion vector of the upper left control point 902, Let the motion vector of JPEG2025100808000005.jpg24150 be the motion vector of the upper right control point 903, When the motion vector of JPEG2025100808000006.jpg25150 is the motion vector of the lower left control point 904, JPEG2025100808000007.jpg21150, JPEG2025100808000008.jpg20150, JPEG2025100808000009.jpg20150 can be defined. And in Equation 2, w represents the width of the current block 901, and h represents the height of the current block 901. And JPEG2025100808000010.jpg19150 represents the motion vector at the {x, y} position.
[0114] In the present invention, among the motions that can be represented by the affine motion model, an affine motion model that represents three motions of translation, scale, and rotation can be defined. In this specification, this will be referred to as a simplified affine motion model or similarity affine motion model.
[0115] The simplified affine motion model is expressed using four parameters (a, b, c, d) as shown in the following Equation 3.
[0116]
Equation
[0117] Here, {v x , v y} indicate the motion vectors at the {x, y} positions respectively. An affine motion model that uses four such parameters is called AF4. The present invention is not limited to this. When six parameters are used, it is called AF6, and the above-described embodiments can be applied in the same manner.
[0118] As shown in FIG. 9(b), JPEG2025100808000012.jpg23150 is the motion vector of the upper left control point 1001 of the current block, assuming that JPEG2025100808000013.jpg20150 is the motion vector of the upper right control point 1002, JPEG2025100808000014.jpg23150, JPEG2025100808000015.jpg22150 can be defined. Here, the affine motion model of AF4 can be defined as the following Equation 4.
[0119]
Equation
[0120] In Equation 4, w represents the width of the current block, and h represents the height of the current block. And, JPEG2025100808000017.jpg19150 indicates the motion vectors at the {x, y} positions respectively.
[0121] The encoder or decoder can determine (or derive) the motion vectors at each pixel position using the control point motion vectors (for example, the motion vectors of the upper left control point 1001 and the upper right control point 1002).
[0122] In the present invention, a set of motion vectors determined by affine motion prediction can be defined as an affine motion vector field. The affine motion vector field is determined using at least one of the above Equations 1 to 4.
[0123] In the symbolization / decryption process, the motion vector obtained by affine motion prediction can be determined in units of pixels or in units of predefined (or preset) blocks (or sub-blocks). For example, when determined in units of pixels, the motion vector is derived based on each pixel in the block, and when determined in units of sub-blocks, the motion vector is derived based on each sub-block unit in the current block. As another example, when determined in units of sub-blocks, the motion vector of the sub-block is derived based on the upper left pixel or the central pixel.
[0124] Hereinafter, for the sake of convenience of explanation in the description of the present invention, the case where the motion vector is determined in units of 4×4 blocks by affine motion prediction will be mainly described. However, the present invention is not limited thereto, and the present invention can be applied in units of pixels or blocks of other sizes.
[0125] On the other hand, as shown in FIG. 9(b), assume that the size of the current block is 16×16. The encoder or decoder determines the motion vector in units of 4×4 sub-blocks using the motion vectors of the upper left control point 1001 and the upper right control point 1002 of the current block. Then, the motion vector of the sub-block is determined based on the central pixel value of each sub-block.
[0126] In FIG. 9(b), the arrows shown in the center of each sub-block indicate the motion vectors obtained by the Affin motion model.
[0127] Affine motion prediction can be used as an affine merge mode (hereinafter referred to as "AF merge mode") and an affine inter mode (hereinafter referred to as "AF inter mode"). The AF merge mode is a method that does not encode the motion vector difference similar to the skip mode or the merge mode, but induces and encodes or decodes two control point motion vectors. The AF inter mode is a method of determining a control point motion vector predictor and a control point motion vector, and then encoding or decoding a control point motion vector difference value (CPMVD) corresponding to the difference. In this case, two control point motion vector difference values are transmitted in the AF4 mode, and three control point motion vector difference values are transmitted in the AF6 mode.
[0128] Here, since the AF4 mode transmits a smaller number of motion vector difference values compared to the AF6 mode, it has the advantage that the control point motion vector (CPMV) can be represented with fewer bits. Since the AF6 mode transmits three CPMVDs, excellent predictor generation is possible, so there is an advantage that the bits for residual coding can be reduced.
[0129] Therefore, the present invention proposes a method of considering both (or simultaneously) the AF4 mode and the AF6 mode in the AF inter mode.
[0130] FIG. 10 is a flowchart for explaining a process of processing a video signal including a current block using an affine prediction mode (hereinafter referred to as "AF mode") as an embodiment to which the present invention is applied.
[0131] The present invention provides a method of processing a video signal including a current block using the AF mode.
[0132] First, the video signal processing device generates a candidate list of a motion vector pair by using motion vectors of pixels or blocks adjacent to at least two control points of the current block (S1010). Here, the control points mean the corner pixels of the current block, and the motion vector pair indicates motion vectors of the upper left corner pixel and the upper right corner pixel of the current block.
[0133] As one embodiment, the control points include at least two of the upper left corner pixel, the upper right corner pixel, the lower left corner pixel, or the lower right corner pixel of the current block, and the candidate list is composed of pixels or blocks adjacent to the upper left corner pixel, the upper right corner pixel, and the lower left corner pixel.
[0134] As one embodiment, the candidate list can be generated based on motion vectors of the diagonal adjacent pixel (A), the upper adjacent pixel (B), and the left adjacent pixel (C) of the upper left corner pixel, motion vectors of the upper adjacent pixel (D) and the diagonal adjacent pixel (E) of the upper right corner pixel, and motion vectors of the left adjacent pixel (F) and the diagonal adjacent pixel (G) of the lower left corner pixel.
[0135] As one embodiment, the method may further include a step of adding an AMVP candidate list to the candidate list when the number of motion vector pairs in the candidate list is less than two.
[0136] As one embodiment, when the current block has a size of N×4, the control point motion vector of the current block is determined as a motion vector derived based on the central positions of the left sub-block and the right sub-block within the current block, and when the current block has a size of 4×N, the control point motion vector of the current block is determined as a motion vector derived based on the central positions of the upper sub-block and the lower sub-block within the current block.
[0137] As one embodiment, when the current block has a size of N×4, the control point motion vector of the left sub-block within the current block is determined by the average value of the first control point motion vector and the third control point motion vector, and the control point motion vector of the right sub-block is determined by the average value of the second control point motion vector and the fourth control point motion vector. When the current block has a size of 4×N, the control point motion vector of the upper sub-block within the current block is determined by the average value of the first control point motion vector and the second control point motion vector, and the control point motion vector of the lower sub-block is determined by the average value of the third control point motion vector and the fourth control point motion vector.
[0138] As another embodiment, the method can signal prediction mode or flag information indicating whether the AF mode is performed.
[0139] In this case, the video signal processing apparatus receives the prediction mode or the flag information, performs the AF mode according to the prediction mode or the flag information, and can derive a motion vector by the AF mode. Here, the AF mode is characterized by indicating a mode of deriving a motion vector in units of pixels or sub-blocks using the control point motion vector of the current block.
[0140] On the other hand, the video signal processing apparatus determines a final candidate list of a preset number of motion vector pairs based on the divergence value of the motion vector pair (S1020). Here, the final candidate list is determined in ascending order of the divergence value, and the divergence value means a value indicating the similarity of the direction of the motion vector.
[0141] The video signal processing apparatus determines the control point motion vector of the current block based on the Rate-Distortion Cost from the final candidate list (S1030).
[0142] The video signal processing device generates a motion vector predictor for the current block based on the control point motion vector (S1040).
[0143] FIG. 11 shows a flowchart for adaptively determining an optimal coding mode based on at least one of an AF4 mode or an AF6 mode as an embodiment (1-1) to which the present invention is applied.
[0144] The video signal processing device performs prediction based on at least one of a skip mode, a merge mode, or an inter mode (S1110). Here, the merge mode may include not only a general merge mode but also the aforementioned AF merge mode, and the inter mode may include not only a general inter mode but also the aforementioned AF inter mode.
[0145] The video signal processing device performs motion vector prediction based on at least one of an AF4 mode or an AF6 mode (S1120). Here, the order of the S1110 step and the S1120 step is not limited.
[0146] The video signal processing device compares the results of the S1120 step to determine the optimal coding mode among the modes (S1130). Here, the results of the S1120 step are compared based on a rate-distortion cost.
[0147] Thereafter, the video signal processing device generates a motion vector predictor for the current block based on the optimal coding mode, subtracts the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.
[0148] Thereafter, the encoding / decoding processes described in FIGS. 1 and 2 above are applied in the same manner.
[0149] FIG. 12 shows a flowchart for adaptively decoding based on the AF4 mode or the AF6 mode as an embodiment (1-2) to which the present invention is applied.
[0150] The decoder receives a bitstream (S1210). The bitstream includes information regarding the coding mode of the current block in the video signal.
[0151] The decoder checks whether the coding mode of the current block is the AF mode (S1220). Here, the AF mode means an affine motion prediction mode that uses an affine motion model, and may include, for example, at least one of an affine merge mode or an affine inter mode. The affine inter mode may include at least one of the AF4 mode or the AF6 mode.
[0152] Here, the step S1220 is checked by an affine flag indicating whether the AF mode is performed. For example, the affine flag can be expressed as affine_flag. When the affine_flag = 1, it indicates that the AF mode is performed on the current block. When the affine_flag = 0, it indicates that the AF mode is not performed on the current block.
[0153] When the AF mode is not performed on the current block, the decoder performs decoding (i.e., motion vector prediction) using a coding mode other than the AF mode (S1230). For example, a skip mode, a merge mode, or an inter mode may be used.
[0154] When the AF mode is performed on the current block, the decoder checks whether the AF4 mode is applied to the current block (S1240).
[0155] Here, the S1240 step can be confirmed by an affine parameter flag indicating whether the AF4 mode is performed (or whether affine motion prediction is performed by four parameters). For example, the affine parameter flag is represented as affine_param_flag. When the affine_param_flag = 0, motion vector prediction is performed by the AF4 mode (S1250), and when the affine_param_flag = 1, it means that motion vector prediction is performed by the AF6 mode (S1260), but the present invention is not limited thereto.
[0156] For example, the affine parameter flag includes at least one of AF4_flag and AF6_flag.
[0157] AF4_flag indicates whether the AF4 mode is performed for the current block. When AF4_flag = 1, the AF4 mode is performed for the current block, and when AF4_flag = 0, the AF4 mode is not performed for the current block. Here, performing the AF4 mode means performing motion vector prediction using an affine motion model represented by four parameters.
[0158] AF6_flag indicates whether the AF6 mode is performed for the current block. When AF6_flag = 1, the AF6 mode is performed for the current block, and when AF4_flag = 0, the AF6 mode is not performed for the current block. Here, performing the AF6 mode means performing motion vector prediction using an affine motion model represented by four parameters.
[0159] The affine flag and the affine parameter flag can be defined at at least one level of a slice, a maximum coding unit, a coding unit, or a prediction unit.
[0160] For example, at least one of AF_flag, AF4_flag, and AF6_flag is defined at the slice level and can also be defined at the block level or the prediction unit level.
[0161] FIG. 13 shows a syntax structure for performing decoding based on the AF4 mode or the AF6 mode as an embodiment (1-3) to which the present invention is applied.
[0162] The decoder obtains merge_flag and checks whether the merge mode is applied to the current block (S1310).
[0163] When the merge mode is not applied to the current block, the decoder can obtain affine_flag (S1320). Here, affine_flag indicates whether the AF mode is performed.
[0164] When the affine_flag = 1, that is, when the AF mode is performed on the current block, the decoder can obtain affine_param_flag (S1330). Here, affine_param_flag indicates whether the AF4 mode is performed (or whether affine motion prediction is performed using four parameters).
[0165] When the affine_param_flag = 0, that is, when motion vector prediction is performed by the AF4 mode, the decoder can obtain mvd_CP0 and mvd_CP1, which are two motion vector difference values (S1340). Here, mvd_CP0 indicates the motion vector difference value for control point 0, and mvd_CP1 indicates the motion vector difference value for control point 1.
[0166] When the affine_param_flag = 1, that is, when motion vector prediction is performed in the AF6 mode, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S1350).
[0167] FIG. 14 shows a flowchart for adaptively determining the optimal coding mode among motion vector prediction modes including the AF4 mode or the AF6 mode based on condition A (condition A) as an embodiment (2-1) to which the present invention is applied.
[0168] The encoder performs prediction based on at least one of the skip mode, the merge mode, or the inter mode (S1410).
[0169] The encoder checks whether condition A is satisfied for the current block in order to determine the optimal coding mode for motion vector prediction (S1420).
[0170] Here, the condition A means a condition for the block size. For example, the following embodiment of Table 1 can be applied.
[0171] [Table 1]
[0172] In Example 1 of Table 1, the condition A indicates whether the number of pixels (pixNum) of the current block is greater than a threshold value (TH1). Here, the threshold value has values such as 64, 128, 256, 512, 1024, ···. For example, TH1 = 64 means that the block size is 4×16, 8×8, or 16×4, and TH1 = 128 means that the block size is 32×4, 16×8, 8×16, or 4×32.
[0173] In the case of Example 2, it indicates whether both the width and height of the current block are greater than the threshold value (TH1).
[0174] In the case of Example 3, it indicates whether the width of the current block is greater than the threshold value (TH1) or the height of the current block is greater than the threshold value (TH1).
[0175] When the condition A is satisfied, the encoder performs motion vector prediction based on at least one of the AF4 mode or the AF6 mode (S1430).
[0176] The encoder compares the results of the S1410 and S1430 steps to determine the optimal coding mode among the motion vector prediction modes including the AF4 mode or the AF6 mode (S1440).
[0177] On the other hand, when the condition A is not satisfied, the encoder determines the optimal coding mode among the modes other than the AF mode (S1440).
[0178] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.
[0179] Thereafter, the encoding / decoding process described in FIGS. 1 and 2 can be applied identically.
[0180] FIG. 15 shows a flowchart for adaptively performing decoding in the AF4 mode or the AF6 mode based on the condition A (condition A) as an embodiment (2-2) to which the present invention is applied.
[0181] The decoder receives a bitstream (S1510). The bitstream includes information regarding a video signal and the coding mode of the current block.
[0182] The decoder checks whether condition A is satisfied for the current block (S1520) in order to determine the optimal coding mode for motion vector prediction. Here, the condition A means a condition for the block size. For example, the embodiment of Table 1 can be applied.
[0183] When the condition A is satisfied, the decoder checks whether the coding mode of the current block is the AF mode (S1530). Here, the AF mode means an affine motion prediction mode using an affine motion model, and the embodiments described in this specification can be applied.
[0184] Here, the step S1530 can be checked by an affine flag indicating whether the AF mode is performed. For example, the affine flag is expressed as affine_flag. When the affine_flag = 1, it indicates that the AF mode is performed for the current block, and when the affine_flag = 0, it indicates that the AF mode is not performed for the current block.
[0185] When the condition A is not satisfied or when the AF mode is not performed for the current block, the decoder can perform decoding (i.e., motion vector prediction) using a coding mode other than the AF mode (S1540). For example, the skip mode, the merge mode, or the inter mode is used.
[0186] When the AF mode is performed for the current block, the decoder checks whether the AF4 mode is applied to the current block (S1550).
[0187] Here, the S1550 step can be confirmed by an affine parameter flag indicating whether the AF4 mode is performed (or whether affine motion prediction is performed based on four parameters). For example, the affine parameter flag can be expressed as affine_param_flag. When the affine_param_flag = 0, motion vector prediction is performed by the AF4 mode (S1560), and when the affine_param_flag = 1, it means that motion vector prediction is performed by the AF6 mode (S1570), but the present invention is not limited thereto.
[0188] FIG. 16 shows a syntax structure for performing decoding by the AF4 mode or the AF6 mode based on condition A (condition A) as an embodiment (2-3) to which the present invention is applied.
[0189] The decoder obtains merge_flag and checks whether the merge mode is applicable to the current block (S1610).
[0190] When the merge mode is not applicable to the current block, the decoder checks whether condition A is satisfied (S1620). Here, condition A means a condition for the block size. For example, the embodiment in Table 1 can be applied.
[0191] When condition A is satisfied, the decoder can obtain affine_flag (S1620). Here, affine_flag indicates whether the AF mode is performed.
[0192] When the affine_flag = 1, that is, when the AF mode is performed on the current block, the decoder can obtain affine_param_flag (S1630). Here, affine_param_flag indicates whether the AF4 mode is performed (or whether affine motion prediction is performed based on four parameters).
[0193] When the said affine_param_flag = 0, that is, when motion vector prediction is performed in the AF4 mode, the decoder can obtain mvd_CP0 and mvd_CP1, which are two motion vector difference values (S1640). Here, mvd_CP0 indicates the motion vector difference value for control point 0, and mvd_CP1 indicates the motion vector difference value for control point 1.
[0194] And when the said affine_param_flag = 1, that is, when motion vector prediction is performed in the AF6 mode, the decoder can obtain mvd_CP0, mvd_CP1, and mvd_CP2, which are three motion vector difference values (S1650).
[0195] FIG. 17 shows a flowchart for adaptively determining the optimal coding mode among motion vector prediction modes including the AF4 mode or the AF6 mode based on at least one of condition B (condition B) or condition C (condition C) as an embodiment (3-1) to which the present invention is applied.
[0196] The present invention provides a method for adaptively selecting the AF4 mode and the AF6 mode based on the size of the current block.
[0197] For example, since the AF6 mode transmits one additional motion vector difference value compared to the AF4 mode, it is effective for relatively large blocks. Therefore, when the size of the current block is smaller than (or equal to) the already set size, encoding is performed considering only the AF4 mode, and when the current block is larger than (or equal to) the already set size, encoding is performed considering only the AF6 mode.
[0198] On the other hand, in the case of an area where it is not determined that only one of the AF4 mode or the AF6 mode is clearly advantageous, both the AF4 mode and the AF6 mode can be considered, and only the optimal mode among them can be signaled.
[0199] As shown in FIG. 17, the encoder makes a prediction based on at least one of a skip mode, a merge mode, or an inter mode (S1710).
[0200] The encoder checks whether condition B is satisfied for the current block (S1720). Here, the condition B means a condition for the block size. For example, the embodiments in Table 2 below can be applied.
[0201]
Table 2
[0202] In Example 1 of Table 2, the condition B indicates whether the number of pixels (pixNum) of the current block is smaller than a threshold value (TH2). Here, the threshold value has values such as 64, 128, 256, 512, 1024, ···. For example, TH2 = 64 means that the block size is 4×16, 8×8, or 16×4, and TH2 = 128 means that the block size is 32×4, 16×8, 8×16, or 4×32.
[0203] In the case of Example 2, the condition B indicates whether both the width and height of the current block are smaller than a threshold value (TH2).
[0204] In the case of Example 3, the condition B indicates whether the width of the current block is smaller than a threshold value (TH2), or whether the height of the current block is smaller than a threshold value (TH2).
[0205] When the condition B is satisfied, the encoder performs motion vector prediction based on the AF4 mode (S1730).
[0206] When the condition B is not satisfied, the encoder checks whether condition C is satisfied for the current block (S1740). Here, the condition C means a condition for the block size. For example, the embodiments in Table 3 below can be applied.
[0207]
Table 3
[0208] In Example 1 of Table 3, the condition A indicates whether the number of pixels (pixNum) of the current block is greater than or equal to a threshold value (TH3). Here, the threshold value has values such as 64, 128, 256, 512, 1024, ···. For example, TH3 = 64 means the block size is 4×16, 8×8, or 16×4, and TH3 = 128 means the block size is 32×4, 16×8, 8×16, or 4×32.
[0209] In the case of Example 2, it indicates whether both the width and height of the current block are greater than or equal to a threshold value (TH3).
[0210] In the case of Example 3, it indicates whether the width of the current block is greater than or equal to a threshold value (TH1), or whether the height of the current block is greater than or equal to a threshold value (TH1).
[0211] When the condition C is satisfied, the encoder performs motion vector prediction based on the AF6 mode (S1760).
[0212] When the condition C is not satisfied, the encoder can perform motion vector prediction based on the AF4 mode and the AF6 mode (S1750).
[0213] On the other hand, in the above-mentioned condition B (condition B) and the above-mentioned condition C (condition C), the threshold value (TH2) and the threshold value (TH3) can be determined so as to satisfy the following mathematical formula 5.
[0214] [Equation 5] TH_2 ≦ TH_3
[0215] The encoder compares the results of the steps S1710, S1730, S1750, and S1760, and determines the optimal coding mode among the motion vector prediction modes including the AF4 mode or the AF6 mode (S1770).
[0216] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.
[0217] Thereafter, the encoding / decoding processes described with reference to FIGS. 1 and 2 can be applied in the same manner.
[0218] FIG. 18 shows a flowchart for adaptively performing decoding in the AF4 mode or the AF6 mode based on at least one of condition B (condition B) or condition C (condition C) as an embodiment (3-2) to which the present invention is applied.
[0219] The decoder checks whether the coding mode of the current block is the AF mode (S1810). Here, the AF mode means an affine motion prediction mode using an affine motion model, and the embodiments described in this specification can be applied, and duplicate descriptions are omitted.
[0220] When the AF mode is performed on the current block, the decoder checks whether condition B is satisfied for the current block (S1820). Here, the condition B means a condition for the block size. For example, the embodiment of Table 2 can be applied, and duplicate explanations are omitted.
[0221] When the condition B is satisfied, the decoder performs motion vector prediction based on the AF4 mode (S1830).
[0222] When the condition B is not satisfied, the decoder checks whether condition C is satisfied for the current block (S1840). Here, the condition C means a condition for the block size. For example, the embodiment of Table 3 can be applied, and duplicate explanations are omitted.
[0223] On the other hand, in the condition B (condition B) and the condition C (condition C), the threshold value (TH2) and the threshold value (TH3) can be determined so as to satisfy the mathematical formula 5.
[0224] When the condition C is satisfied, the decoder performs motion vector prediction based on the AF6 mode (S1860).
[0225] When the condition C is not satisfied, the decoder checks whether the AF4 mode is applicable to the current block (S1850).
[0226] Here, the step S1850 can be checked by an affine parameter flag indicating whether the AF4 mode is performed (or whether affine motion prediction is performed by four parameters).
[0227] For example, the affine parameter flag can be represented by affine_param_flag. When the affine_param_flag = 0, motion vector prediction is performed in the AF4 mode (S1830), and when the affine_param_flag = 1, it means that motion vector prediction is performed in the AF6 mode (S1860), but the present invention is not limited thereto.
[0228] On the other hand, when the AF mode is not performed on the current block, the decoder performs decoding (i.e., motion vector prediction) in a coding mode other than the AF mode (S1870). For example, the skip mode, merge mode, or inter mode can be used.
[0229] FIG. 19 shows a syntax structure for performing decoding in the AF4 mode or the AF6 mode based on at least one of condition B (condition B) or condition C (condition C) as an embodiment (3-3) to which the present invention is applied.
[0230] The decoder obtains merge_flag and checks whether the merge mode is applicable to the current block (S1910).
[0231] When the merge mode is not applicable to the current block, the decoder obtains affine_flag (S1920). Here, affine_flag indicates whether the AF mode is performed.
[0232] When the affine_flag = 1, that is, when the AF mode is performed on the current block, the decoder checks whether condition B is satisfied (S1620). Here, the condition B means a condition for the block size. For example, the embodiment in Table 2 can be applied.
[0233] When the condition B is satisfied, the decoder sets the affine_param_flag to 0 (S1930). Here, the affine_param_flag indicates whether the AF4 mode is performed (or whether affine motion prediction is performed using four parameters). affine_param_flag = 0 means that motion vector prediction is performed in the AF4 mode.
[0234] When the condition B is not satisfied and the condition C is satisfied, the decoder sets the affine_param_flag to 1 (S1940). Here, affine_param_flag = 1 means that motion vector prediction is performed in the AF6 mode.
[0235] On the other hand, when the condition B is not satisfied and the condition C is not satisfied either, the decoder can obtain the affine_param_flag (S1950).
[0236] When the affine_param_flag = 0, the decoder can obtain mvd_CP0 and mvd_CP1, which are two motion vector difference values (S1960).
[0237] When the affine_param_flag = 1, the decoder can obtain mvd_CP0, mvd_CP1, and mvd_CP2, which are three motion vector difference values (S1970).
[0238] FIG. 20 shows a flowchart for adaptively determining the optimal coding mode among motion vector prediction modes including the AF4 mode or the AF6 mode based on the coding mode of adjacent blocks as an embodiment (4-1) to which the present invention is applied.
[0239] The encoder performs prediction based on at least one of the skip mode, the merge mode, or the inter mode (S2010).
[0240] The encoder checks whether an adjacent block is coded in the AF mode (S2020). Here, whether an adjacent block is coded in the AF mode can be expressed by isNeighborAffine(). For example, isNeighborAffine() = 0 means that the adjacent block is not coded in the AF mode, and isNeighborAffine() = 1 means that the adjacent block is coded in the AF mode.
[0241] If the adjacent block is not coded in the AF mode, the encoder performs motion vector prediction based on the AF4 mode (S2030).
[0242] If the adjacent block is coded in the AF mode, the encoder performs motion vector prediction based on the AF4 mode and also performs motion vector prediction based on the AF6 mode (S2040).
[0243] The encoder compares the results of steps S2030 and S2040 to determine the optimal coding mode among the motion vector prediction modes including the AF4 mode or the AF6 mode (S2050).
[0244] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.
[0245] Thereafter, the encoding / decoding process described in FIGS. 1 and 2 can be applied in the same way.
[0246] FIG. 21 shows a flowchart for adaptively performing decoding by the AF4 mode or the AF6 mode based on the coding mode of an adjacent block as an embodiment (4-2) to which the present invention is applied.
[0247] The decoder can receive a bitstream (S2110). The bitstream includes information regarding the coding mode of the current block in the video signal.
[0248] The decoder checks whether the coding mode of the current block is the AF mode (S2120).
[0249] If the AF mode is not performed on the current block, the decoder performs decoding (i.e., motion vector prediction) according to the coding mode that is not the AF mode (S2170). For example, the skip mode, merge mode, or inter mode can be used.
[0250] If the AF mode is performed on the current block, the decoder checks whether an adjacent block is coded in the AF mode (S2130). Here, whether an adjacent block is coded in the AF mode can be represented by isNeighborAffine(). For example, isNeighborAffine() = 0 means that the adjacent block is not coded in the AF mode, and isNeighborAffine() = 1 means that the adjacent block is coded in the AF mode.
[0251] If the adjacent block is coded in the AF mode, the decoder performs motion vector prediction based on the AF4 mode (S2140).
[0252] If the adjacent block is not coded in the AF mode, the decoder checks whether the AF4 mode is applicable to the current block (S2150).
[0253] Here, the step S2150 can be confirmed by an affine parameter flag indicating whether the AF4 mode is performed (or whether affine motion prediction is performed based on four parameters). For example, the affine parameter flag can be represented by affine_param_flag. When the affine_param_flag = 0, motion vector prediction is performed by the AF4 mode (S2140), and when the affine_param_flag = 1, motion vector prediction is performed by the AF6 mode (S2160).
[0254] FIG. 22 shows a syntax structure for performing decoding by the AF4 mode or the AF6 mode based on the coding mode of adjacent blocks as an embodiment (4-3) to which the present invention is applied.
[0255] The decoder obtains merge_flag and checks whether the merge mode is applicable to the current block (S2210).
[0256] When the merge mode is not applicable to the current block, the decoder can obtain affine_flag (S2220). Here, affine_flag indicates whether the AF mode is performed.
[0257] When the affine_flag = 1, that is, when the AF mode is performed on the current block, the decoder checks whether the adjacent block is coded in the AF mode (S2230).
[0258] When the adjacent block is coded in the AF mode, the decoder can obtain affine_param_flag (S2230). Here, affine_param_flag indicates whether the AF4 mode is performed (or whether affine motion prediction is performed based on four parameters).
[0259] If the adjacent block is not coded in the AF mode, the decoder sets the affine_param_flag to 0 (S2240).
[0260] When the affine_param_flag = 0, that is, when motion vector prediction is performed in the AF4 mode, the decoder can obtain two motion vector difference values, mvd_CP0 and mvd_CP1 (S2250).
[0261] And when the affine_param_flag = 1, that is, when motion vector prediction is performed in the AF6 mode, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S2260).
[0262] FIG. 23 shows a flowchart for adaptively determining the optimal coding mode among motion vector prediction modes including the AF4 mode or the AF6 mode based on at least one of condition A (condition A), condition B (condition B), or condition C (condition C) as an embodiment (5-1) to which the present invention is applied.
[0263] The present invention shows an embodiment that combines Embodiment 2 and Embodiment 3. In FIG. 23, an example is described when all of conditions A, B, and C are considered, and the order of the conditions can be applied differently.
[0264] As shown in FIG. 23, the encoder performs prediction based on at least one of the skip mode, the merge mode, or the inter mode (S2310).
[0265] The encoder checks whether condition A is satisfied for the current block (S2320). Here, the condition A means a condition for the block size, and the embodiment of Table 1 can be applied.
[0266] When the condition A is satisfied, the encoder determines the optimal coding mode among the modes excluding the AF mode (S2380).
[0267] On the other hand, when the condition A is not satisfied, the encoder checks whether the condition B is satisfied for the current block (S2330). Here, the condition B means a condition for the block size, and the embodiment of Table 2 can be applied.
[0268] When the condition B is satisfied, the encoder performs motion vector prediction based on the AF4 mode (S2340).
[0269] When the condition B is not satisfied, the encoder checks whether the condition C is satisfied for the current block (S2350). Here, the condition C means a condition for the block size, and the embodiment of Table 3 can be applied.
[0270] When the condition C is satisfied, the encoder performs motion vector prediction based on the AF6 mode (S2370).
[0271] When the condition C is not satisfied, the encoder performs motion vector prediction based on the AF4 mode and also performs motion vector prediction based on the AF6 mode (S2360).
[0272] On the other hand, in the condition B and the condition C, the threshold value (TH2) and the threshold value (TH3) can be determined so as to satisfy the mathematical formula 5.
[0273] The encoder compares the results of the steps S2310, S2340, S2360, and S2370 to determine the optimal coding mode (S2380).
[0274] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.
[0275] Thereafter, the encoding / decoding processes described in FIGS. 1 and 2 can be applied identically.
[0276] FIG. 24 shows a flowchart for adaptively performing decoding in an AF4 mode or an AF6 mode based on at least one of condition A (condition A), condition B (condition B), or condition C (condition C) as an embodiment (5-2) to which the present invention is applied.
[0277] The decoder checks whether condition A is satisfied for the current block (S2410). Here, the condition A means a condition for the block size. For example, the embodiment of Table 1 can be applied.
[0278] When the condition A is satisfied, the decoder checks whether the coding mode of the current block is the AF mode (S2420). Here, the AF mode means an affine motion prediction mode using an affine motion model, and the embodiments described in this specification are applied, and redundant descriptions are omitted.
[0279] When the condition A is not satisfied or when the AF mode is not performed for the current block, the decoder performs decoding (i.e., motion vector prediction) using a coding mode other than the AF mode (S2480). For example, a skip mode, a merge mode, or an inter mode can be used.
[0280] When the AF mode is performed on the current block, the decoder checks whether condition B is satisfied for the current block (S2430). Here, the condition B means a condition for the block size. For example, the embodiment of Table 2 can be applied, and duplicate explanations are omitted.
[0281] When the condition B is satisfied, the decoder performs motion vector prediction based on the AF4 mode (S2440).
[0282] When the condition B is not satisfied, the decoder checks whether condition C is satisfied for the current block (S2450). Here, the condition C means a condition for the block size. For example, the embodiment of Table 3 can be applied, and duplicate explanations are omitted.
[0283] On the other hand, in the condition B (condition B) and the condition C (condition C), the threshold value (TH2) and the threshold value (TH3) can be determined so as to satisfy the mathematical formula 5.
[0284] When the condition C is satisfied, the decoder performs motion vector prediction based on the AF6 mode (S2470).
[0285] When the condition C is not satisfied, the decoder checks whether the AF4 mode is applicable to the current block (S2460).
[0286] Here, the step S2460 can be confirmed by an affine parameter flag indicating whether the AF4 mode is performed (or whether affine motion prediction is performed by four parameters).
[0287] For example, the affine parameter flag can be represented by affine_param_flag. When the affine_param_flag = 0, motion vector prediction is performed in the AF4 mode (S2440), and when the affine_param_flag = 1, it means that motion vector prediction is performed in the AF6 mode (S2470), but the present invention is not limited to this.
[0288] FIG. 25 shows a syntax structure for decoding in the AF4 mode or the AF6 mode based on at least one of condition A (condition A), condition B (condition B), or condition C (condition C) as an embodiment (5-3) to which the present invention is applied.
[0289] The decoder obtains merge_flag and checks whether the merge mode is applied to the current block (S2510).
[0290] When the merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S2520). Here, the condition A means a condition for the block size. For example, the embodiment of Table 1 can be applied.
[0291] When condition A is satisfied, the decoder can obtain affine_flag (S2520). Here, affine_flag indicates whether the AF mode is performed.
[0292] When the affine_flag = 1, that is, when the AF mode is performed on the current block, the decoder checks whether condition B is satisfied (S2530). Here, the condition B means a condition for the block size. For example, the embodiment of Table 2 can be applied.
[0293] When the condition B is satisfied, the decoder sets the affine_param_flag to 0 (S2540). Here, the affine_param_flag indicates whether the AF4 mode is performed (or whether affine motion prediction is performed using four parameters). affine_param_flag = 0 means that motion vector prediction is performed in the AF4 mode.
[0294] When the condition B is not satisfied and the condition C is satisfied, the decoder sets the affine_param_flag to 1 (S2550). Here, affine_param_flag = 1 means that motion vector prediction is performed in the AF6 mode.
[0295] On the other hand, when the condition B is not satisfied and the condition C is not satisfied either, the decoder can obtain the affine_param_flag (S2560).
[0296] If the affine_param_flag = 0, the decoder can obtain two motion vector difference values, mvd_CP0 and mvd_CP1 (S2570).
[0297] If the affine_param_flag = 1, the decoder can obtain three motion vector difference values, mvd_CP0, mvd_CP1, and mvd_CP2 (S2580).
[0298] FIG. 26 shows a flowchart for adaptively determining the optimal coding mode among motion vector prediction modes including the AF4 mode or the AF6 mode based on at least one of the condition A (condition A) or the coding mode of an adjacent block as an embodiment (6-1) to which the present invention is applied.
[0299] The encoder can perform prediction based on at least one of the skip mode, the merge mode, or the inter mode (S2610).
[0300] The encoder checks whether condition A is satisfied for the current block (S2620). Here, the condition A means a condition for the block size, and the embodiment of Table 1 can be applied.
[0301] If the condition A is satisfied, the encoder determines the optimal coding mode among the modes excluding the AF mode (S2660).
[0302] On the other hand, if the condition A is not satisfied, the encoder checks whether the adjacent block is coded in the AF mode (S2630). Here, whether the adjacent block is coded in the AF mode can be represented by isNeighborAffine(). For example, isNeighborAffine() = 0 means that the adjacent block is not coded in the AF mode, and isNeighborAffine() = 1 means that the adjacent block is coded in the AF mode.
[0303] If the adjacent block is not coded in the AF mode, the encoder performs motion vector prediction based on the AF4 mode (S2640).
[0304] If the adjacent block is coded in the AF mode, the encoder performs motion vector prediction based on the AF4 mode and also performs motion vector prediction based on the AF6 mode (S2650).
[0305] The encoder compares the results of steps S2610, S2640, and S2650 to determine the optimal coding mode (S2660).
[0306] Thereafter, the encoder can generate a motion vector predictor for the current block based on the optimal coding mode, and subtract the motion vector predictor from the motion vector of the current block to obtain a motion vector difference value.
[0307] Hereafter, the encoding / decoding processes described with reference to FIGS. 1 and 2 can be applied identically.
[0308] FIG. 27 shows a flowchart for adaptively performing decoding in an AF4 mode or an AF6 mode based on at least one of condition A (condition A) or the coding mode of an adjacent block as an embodiment (6-2) to which the present invention is applied.
[0309] The decoder can receive a bitstream (S2710). The bitstream includes information regarding the coding mode of a current block in a video signal.
[0310] The decoder checks whether condition A is satisfied for the current block in order to determine an optimal coding mode for motion vector prediction (S2720). Here, the condition A means a condition for a block size. For example, the embodiment of Table 1 can be applied.
[0311] When the condition A is satisfied, the decoder checks whether the coding mode of the current block is an AF mode (S2730).
[0312] Hereinafter, the content described in steps S2730 to S2780 can be applied to steps S2120 to S2170 of FIG. 21, and repeated descriptions are omitted.
[0313] FIG. 28 shows a syntax structure for performing decoding in an AF4 mode or an AF6 mode based on at least one of condition A (condition A) or the coding mode of an adjacent block as an embodiment (6-3) to which the present invention is applied.
[0314] The decoder acquires a merge_flag and checks whether a merge mode is applied to the current block (S2810).
[0315] When the merge mode is not applied to the current block, the decoder checks whether condition A is satisfied (S2820). Here, the condition A means a condition for the block size. For example, the embodiment of Table 1 can be applied.
[0316] When the condition A is satisfied, the decoder can obtain the affine_flag (S2820). Here, the affine_flag indicates whether the AF mode is performed.
[0317] Hereinafter, steps S2830 to S2860 can apply the content described in steps S2230 to S2260 of FIG. 22, and duplicate explanations are omitted.
[0318] FIG. 29 shows a flowchart for generating a motion vector predictor based on at least one of AF4 mode or AF6 mode as an embodiment to which the present invention is applied.
[0319] The decoder checks whether the AF mode is applied to the current block (S2910). Here, the AF mode indicates a motion prediction mode using an affine motion model.
[0320] For example, the decoder can obtain an affine flag from the video signal, and whether the AF mode is applied to the current block can be confirmed based on the affine flag.
[0321] When the AF mode is applied to the current block, the decoder checks whether the AF4 mode is used (S2920). Here, the AF4 mode indicates a mode for predicting a motion vector using four parameters constituting the affine motion model.
[0322] For example, when the AF mode is applied to the current block by the affine flag, the decoder can obtain an affine parameter flag from the video signal, and the affine parameter flag indicates whether the motion vector predictor is generated using the four parameters or the six parameters.
[0323] Here, the affine flag and the affine parameter flag can be defined at at least one level of a slice, a maximum coding unit, a coding unit, or a prediction unit.
[0324] When the AF4 mode is used, the decoder generates a motion vector predictor using the four parameters, and when the AF4 mode is not used, the decoder generates a motion vector predictor using the six parameters that constitute the affine motion model (S2930).
[0325] The decoder can obtain the motion vector of the current block based on the motion vector predictor (S2940).
[0326] As an embodiment, the decoder can check whether the condition that the size of the current block has already been set is satisfied. At this time, the already-set condition indicates whether at least one of the number of pixels in the current block, the width and / or height of the current block is greater than a preset threshold.
[0327] For example, when the condition that the size of the current block has already been set is satisfied, the decoder can check whether the AF mode is applied to the current block.
[0328] On the contrary, when the condition that the size of the current block has already been set is not satisfied, the current block can be decoded based on another coding mode that is not the AF mode.
[0329] As one embodiment, when the AF mode is applied to the current block, the decoder can check whether the AF mode has been applied to an adjacent block.
[0330] When the AF mode is applied to the adjacent block, the motion vector predictor is generated using the four parameters. When the AF mode is not applied to the adjacent block, the decoder performs a step of checking whether the AF4 mode is used.
[0331] FIG. 30 shows a flowchart for generating a motion vector predictor based on AF4_flag and AF6_flag as an embodiment to which the present invention is applied.
[0332] The decoder obtains at least one of AF4_flag and AF6_flag from the video signal (S3010). Here, AF4_flag indicates whether the AF4 mode is performed on the current block, and AF6_flag indicates whether the AF6 mode is performed on the current block.
[0333] Here, at least one of the AF4_flag and the AF6_flag can be defined at the slice level and can also be defined at the block level or the prediction unit level. However, the present invention is not limited thereto, and at least one of the AF4_flag and the AF6_flag can be defined at at least one level of a slice, a maximum coding unit, a coding unit, or a prediction unit.
[0334] The decoder checks the values of AF4_flag and AF6_flag (S3020).
[0335] When AF4_flag = 1, the AF4 mode is performed on the current block. When AF4_flag = 0, the AF4 mode is not performed on the current block. Here, performing the AF4 mode means performing motion vector prediction using an affine motion model represented by four parameters.
[0336] When AF6_flag = 1, the AF6 mode is performed on the current block. When AF4_flag = 0, the AF6 mode is not performed on the current block. Here, performing the AF6 mode means performing motion vector prediction using an affine motion model represented by four parameters.
[0337] When AF4_flag = 0 and AF6_flag = 0, the decoder performs motion vector prediction by a mode other than the AF4 mode and the AF6 mode (S3030).
[0338] When AF4_flag = 1 and AF6_flag = 0, the decoder performs motion vector prediction by the AF4 mode (S3040).
[0339] When AF4_flag = 0 and AF6_flag = 1, the decoder performs motion vector prediction by the AF6 mode (S3050).
[0340] When AF4_flag = 1 and AF6_flag = 1, the decoder performs motion vector prediction by the AF4 mode or the AF6 mode (S3060).
[0341] FIG. 31 shows a flowchart for adaptively decoding by the AF4 mode or the AF6 mode based on whether adjacent blocks are coded in the AF mode as an embodiment to which the present invention is applied.
[0342] The decoder checks whether the AF mode is applied to the current block (S3110).
[0343] When the AF mode is applied to the current block, the decoder checks whether an adjacent block is coded in the AF mode (S3120).
[0344] When an adjacent block is coded in the AF mode, the decoder can obtain at least one of AF4_flag or AF6_flag (S3130).
[0345] The decoder generates a motion vector predictor using four or six parameters based on at least one of AF4_flag or AF6_flag (S3140). For example, if AF4_flag = 1, the decoder can perform motion vector prediction in the AF4 mode, and if AF6_flag = 1, the decoder can perform motion vector prediction in the AF6 mode.
[0346] The decoder can obtain the motion vector of the current block based on the motion vector predictor (S3150).
[0347] FIG. 32 shows a syntax for adaptively performing decoding based on AF4_flag and AF6_flag as an embodiment to which the present invention is applied.
[0348] The decoder can obtain AF4_flag and AF6_flag at the slice level (S3010). Here, AF4_flag indicates whether the AF4 mode is performed for the current block, and AF6_flag indicates whether the AF6 mode is performed for the current block. The AF4_flag can be represented by affine_4_flag, and the AF6_flag can be represented by affine_6_flag.
[0349] The decoder can adaptively perform decoding based on AF4_flag and AF6_flag at the block level or prediction unit level.
[0350] When affine_4_flag is not 0 or affine_6_flag is not 0 (i.e., other than the case of affine_4_flag = 0 && affine_6_flag = 0), the decoder can obtain the affine flag (S3220). The affine flag indicates whether the AF mode is performed.
[0351] When the AF mode is performed, the decoder can perform decoding adaptively according to the AF4_flag and AF6_flag values.
[0352] When affine_4_flag = 1 && affine_6_flag = 0, the decoder can set affine_param_flag to 0. That is, affine_param_flag = 0 means that the AF4 mode is performed.
[0353] When affine_4_flag = 0 && affine_6_flag = 1, the decoder can set affine_param_flag to 1. That is, affine_param_flag = 1 means that the AF6 mode is performed.
[0354] When affine_4_flag = 1 && affine_6_flag = 1, the decoder can parse or obtain affine_param_flag. Here, the decoder can perform decoding in the AF4 mode or the AF6 mode according to the affine_param_flag value at the block level or the prediction unit level.
[0355] In other cases, the syntax structure can apply the foregoing embodiments, and duplicate descriptions are omitted.
[0356] FIG. 33 shows a syntax for performing adaptive decoding in the AF4 mode or the AF6 mode based on whether adjacent blocks are coded in the AF mode as an embodiment to which the present invention is applied.
[0357] In the case of this embodiment, the description of the content overlapping with FIG. 32 can be applied as described above, and only the different parts will be described.
[0358] When affine_4_flag = 1 && affine_6_flag = 1, the decoder can check whether the adjacent block is coded in the AF mode.
[0359] When the adjacent block is coded in the AF mode, the decoder parses or obtains affine_param_flag (S3310). Here, the decoder can perform decoding in the AF4 mode or the AF6 mode according to the affine_param_flag value at the block level or the prediction unit level.
[0360] On the contrary, when the adjacent block is not coded in the AF mode, the decoder can set affine_param_flag to 0. That is, affine_param_flag = 0 means that the AF4 mode is performed.
[0361] FIG. 34 shows a video coding system to which the present invention is applied.
[0362] The video coding system includes a source device and a receiving device. The source device transmits the encoded video / image information or data to the receiving device in a digital storage medium or via a network in a file or streaming format.
[0363] The source device includes a video source, an encoding apparatus, and a transmitter. The receiving device includes a receiver, a decoding apparatus, and a renderer. The encoding apparatus may be referred to as a video / image encoding apparatus, and the decoding apparatus may be referred to as a video / image decoding apparatus. A transmitter may be included in the encoding apparatus. A receiver may be included in the decoding apparatus. The renderer may include a display unit, and the display unit may be composed of another device or an external component.
[0364] The video source can obtain video / images through the processes of video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer, etc., and in this case, the process of generating related data can replace the video / image capture process.
[0365] The encoding apparatus encodes the input video / image. Encoding performs a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) is output in bitstream format.
[0366] The transmitting unit transmits the encoded video / image information or data output in bitstream format to the receiving unit of the receiving device in file or streaming format via a digital storage medium or a network. The digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit includes elements for generating media files in a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit extracts the bitstream and transmits it to the decoding device.
[0367] The decoding device decodes the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0368] The renderer renders the decoded video / image. The rendered video / image is displayed via the display unit.
[0369] FIG. 35 shows a content streaming system to which the present invention is applied.
[0370] As shown in FIG. 35, the content streaming system to which the present invention is applied mainly includes an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0371] The encoding server compresses the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.
[0372] The bitstream is generated by an encoding method or a bitstream generation method to which the present invention is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0373] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium for informing the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. Here, the content streaming system may include a separate control server. In this case, the control server plays a role of controlling commands / responses between each device in the content streaming system.
[0374] The streaming server receives content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0375] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, smart glass, HMD (head mounted display)), a digital TV, a desktop computer, a digital signage, and the like.
[0376] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.
[0377] As described above, the embodiments described in the present invention can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in FIGS. 1, 2, 34, and 35 can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip.
[0378] In addition, the decoders and encoders to which the present invention is applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an Internet streaming service providing device, a three-dimensional (3D) video device, an image phone video device, and a medical video device, etc., and can be used to process video signals and data signals.
[0379] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to the present invention can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices in which computer-readable data is stored. The computer-readable storage medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy (registered trademark) disk, and an optical data storage device. Further, the computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission via the Internet). Also, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
Industrial Applicability
[0380] As described above, the preferred embodiments of the present invention described above are disclosed for illustrative purposes, and those skilled in the art can make various other improvements, changes, substitutions, additions, etc. within the technical idea and technical scope of the present invention disclosed in the claims attached below.
Claims
1. In a method of decoding a video signal including a current block based on an affine motion prediction mode (AF mode), obtaining a merge flag from the video signal, wherein the merge flag indicates whether a merge mode is applied to the current block, and a set of candidates available in the merge mode is composed of spatially adjacent candidates, temporal candidates, and generated candidates; obtaining an affine flag from the video signal based on that the merge mode is not applied to the current block or based on that the width and height of the current block are equal to or greater than 16, wherein the affine flag indicates whether the AF mode is applied to the current block, the AF mode represents a motion prediction mode using an affine motion model, and the affine motion model induces a motion vector in pixel units or sub-block units using a control point motion vector; obtaining an affine parameter flag indicating whether four parameters or six parameters are used for the affine motion model based on that the AF mode is applied to the current block; obtaining a motion vector predictor based on that the four parameters or the six parameters are used for the affine motion model; obtaining a prediction sample for the current block based on the motion vector predictor; obtaining a residual sample for the current block; restoring the current block based on the prediction sample and the residual sample; filtering the restored current block, wherein the affine flag and the affine parameter flag are obtained based on that the merge flag indicates that the merge mode is not applied to the current block.
2. The method according to claim 1, wherein the affine flag and the affine parameter flag are defined at the level of a coding unit.
3. The method according to claim 1, wherein based on that the width and height of the current block are less than 16, the current block is decoded based on a coding mode other than the AF mode.
4. A method for encoding a video signal including a current block based on an affine motion prediction mode (AF mode), generating a merge flag indicating whether a merge mode is applied to the current block, wherein a set of candidates available in the merge mode is composed of spatially adjacent candidates, temporal candidates, and generated candidates; generating an affine flag based on the fact that the merge mode is not applied to the current block or based on the width and height of the current block being equal to or greater than 16, wherein the affine flag indicates whether the AF mode is applied to the current block, the AF mode represents a motion prediction mode using an affine motion model, and the affine motion model derives a motion vector in units of pixels or sub-blocks using control point motion vectors; generating an affine parameter flag indicating whether four parameters or six parameters are used for the affine motion model based on the AF mode being applied to the current block; obtaining a motion vector predictor based on the four parameters or the six parameters being used for the affine motion model; generating a prediction sample for the current block based on the motion vector predictor; generating a residual sample for the current block based on the prediction sample; performing transformation, quantization, and entropy encoding on the residual sample, and restoring the current block based on the prediction sample and the residual sample, wherein the affine flag and the affine parameter flag are generated based on the merge flag indicating that the merge mode is not applied to the current block. **Claim 5** A method for transmitting data for a video signal, generating a bitstream for the video signal, wherein the bitstream is generated by an encoding method; transmitting the data including the bitstream, and the encoding method is Generate a merge flag indicating whether the merge mode is currently applied to the current block, and a set of candidates available in the merge mode consists of spatially adjacent candidates, temporal candidates, and generated candidates, step; Based on the fact that the merge mode is not applied to the current block, generate an affine flag based on whether the width and height of the current block are equal to 16 or greater than 16, where the affine flag indicates whether the affine motion prediction mode (AF mode) is applied to the current block, and the AF mode represents a motion prediction mode using an affine motion model, and the affine motion model induces a motion vector in pixel units or sub-block units using a control point motion vector, step; Based on the fact that the AF mode is applied to the current block, generate an affine parameter flag indicating whether four parameters or six parameters are used for the affine motion model, step; Based on the fact that the four parameters or the six parameters are used for the affine motion model, obtain a motion vector predictor, step; Based on the motion vector predictor, generate a prediction sample for the current block, step; Based on the prediction sample, generate a residual sample for the current block, step; Perform transformation, quantization, and entropy encoding on the residual sample, step, and include; Based on the prediction sample and the residual sample, the current block is restored; Based on the fact that the merge flag indicates that the merge mode is not applied to the current block, the affine flag and the affine parameter flag are generated, method.
Citation Information
Patent Citations
Method and apparatus for processing a video signal using affine prediction
JP7673160B2
Method and apparatus for global motion compensation in video coding system
WO2017087751A1
Method and apparatus for affine merge mode prediction for video coding system
WO2017118409A1
Prediction image generation device, moving image decoding device, and moving image encoding device
WO2017130696A1
Affine prediction for video coding
WO2017156705A1