Image coding method and apparatus based on affine motion prediction

JP2025156590A5Pending Publication Date: 2025-11-11GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025134172
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-04-01
Filing Date
2025-08-12
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to higher transmission and storage costs due to increased data volume, necessitating more efficient video coding techniques, particularly in video coding systems.

Method used

A method and apparatus for improving video coding efficiency through affine motion prediction, including generating an affine MVP candidate list, deriving Control Point Motion Vector Predictors (CPMVP) and Differences (CPMVD), and using these to enhance prediction accuracy and reduce data burden.

Benefits of technology

Enhances video coding efficiency by improving the accuracy of motion prediction and reducing data transmission and storage costs for high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method and apparatus which increase image coding efficiency based on affine motion prediction.SOLUTION: There is provided a picture decoding method which includes the steps of: acquiring motion prediction information from a bitstream; generating an affine MVP candidate list comprising affine MVP candidates for the current block; deriving CPMVPs for the respective CPs of the current block based on one affine MVP candidate among the affine MVP candidates included in the affine MVP candidate list; deriving CPMVDs for the CPs of the current block based on information on the CPMVDs for the respective CPs included in the acquired motion prediction information; deriving CPMVs for the CPs of the current block based on the CPMVPs and the CPMVDs; deriving prediction samples for the current block based on the CPMVs; and generating reconstructed samples for the current block based on the derived prediction samples.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video coding technology, and more particularly to a video coding method and apparatus based on affine motion prediction in a video coding system. [Background technology]

[0002] Recently, the demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the resolution and quality of image data increases, the amount of information or bits to be transmitted increases relatively compared to existing image data. Therefore, when image data is transmitted using a medium such as an existing wired or wireless broadband line or when image data is stored using an existing storage medium, transmission costs and storage costs increase.

[0003] This requires highly efficient image compression techniques to effectively transmit, store, and reproduce high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] SUMMARY OF THE INVENTION A technical object of the present invention is to provide a method and apparatus for improving the efficiency of video coding.

[0005] Another technical object of the present invention is to provide a method and apparatus for improving the efficiency of video coding based on affine motion prediction.

[0006] It is still another technical object of the present invention to provide a method and apparatus for efficiently determining a combination of neighboring blocks used in affine motion estimation, thereby improving the efficiency of video coding.

[0007] Another technical object of the present invention is to provide a method and apparatus for improving the efficiency of video coding by signaling information regarding an affine MVP candidate list used in affine motion prediction. [Means for solving the problem]

[0008] According to an embodiment of the present invention, there is provided a picture decoding method performed by a decoding apparatus, the method including the steps of: acquiring motion prediction information from a bitstream; generating an affine MVP candidate list including affine motion vector predictor (MVP) candidates for a current block; deriving Control Point Motion Vector Predictors (CPPMVP) for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list; deriving Control Point Motion Vector Differences (CPMVD) for each CP of the current block based on information on Control Point Motion Vector Differences (CPMVD) for each CP included in the acquired motion prediction information; deriving Control Point Motion Vectors (CPMV) for the CP of the current block based on the CPMVP and the CPMVD; deriving prediction samples for the current block based on the CPMV; and generating reconstructed samples for the current block based on the derived prediction samples.

[0009] According to another embodiment of the present invention, there is provided a decoding apparatus for performing picture decoding, the decoding apparatus including: an entropy decoding unit that acquires motion prediction information from a bitstream; a prediction unit that generates an affine MVP candidate list including affine motion vector predictor (MVP) candidates for a current block, derives Control Point Motion Vector Predictors (CPMVP) for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list, derives Control Point Motion Vector Differences (CPMVD) for each CP of the current block based on information on Control Point Motion Vector Differences (CPMVD) for each CP included in the acquired motion prediction information, derives CPMV for the CP of the current block based on the CPMVP and the CPMVD, and derives predicted samples for the current block based on the CPMV; and an adder that generates reconstructed samples for the current block based on the derived predicted samples.

[0010] According to another embodiment of the present invention, there is provided a picture encoding method performed by an encoding apparatus, the method including the steps of: generating an affine MVP candidate list including affine MVP candidates for a current block, deriving a CPMVP for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list, deriving a CPMV for each CP of the current block, deriving a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each CP, deriving prediction samples for the current block based on the CPMV, deriving residual samples for the current block based on the derived prediction samples, and encoding information on the derived CPMVD and residual information on the residual samples.

[0011] According to another embodiment of the present invention, there is provided an encoding apparatus for performing picture encoding, the encoding apparatus including: a prediction unit that generates an affine MVP candidate list including affine MVP candidates for a current block, derives a CPMVP for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list, derives a CPMV for each of the CPs of the current block, derives a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each of the CPs, and derives predicted samples for the current block based on the CPMV, a residual processing unit that derives residual samples for the current block based on the derived predicted samples, and an entropy encoding unit that encodes information on the derived CPMVD and residual information on the residual samples. [Effects of the Invention]

[0012] The present invention can improve the efficiency of image / video compression in general.

[0013] According to the present invention, the efficiency of video coding based on affine motion prediction can be improved.

[0014] According to the present invention, the efficiency of video coding can be improved by signaling information regarding an affine MVP candidate list used in affine motion prediction. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a diagram illustrating a schematic configuration of an encoding device according to an embodiment; [Figure 2] 1 is a diagram illustrating a configuration of a decoding device according to an embodiment of the present invention; [Figure 3] FIG. 2 is a diagram illustrating an example of motion represented via an affine motion model according to an embodiment. [Figure 4] FIG. 10 is a diagram illustrating an example of an affine motion model using CPMVs (Control Point Motion Vectors) of three CPs (Control Points) for a current block. [Figure 5] FIG. 10 illustrates an example of an affine motion model using CPMV of two CPs for a current block. [Figure 6] FIG. 10 is a diagram showing an example of deriving a motion vector for each sub-block based on an affine motion model. [Figure 7] 1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 8] 1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 9]1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 10] 1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 11] 1 is a flowchart illustrating an operation method of an encoding apparatus according to an embodiment. [Figure 12] 1 is a block diagram showing a configuration of an encoding device according to an embodiment; [Figure 13] 10 is a flowchart illustrating an operation method of a decoding device according to an embodiment. [Figure 14] 1 is a block diagram showing the configuration of a decoding device according to an embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0016] The present invention may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiments. The terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical spirit of the present invention. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0017] Meanwhile, each component in the drawings described in the present invention is illustrated independently for the convenience of explaining different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0018] The following description may be applied to technical fields dealing with video, images, or pictures. For example, the methods or embodiments disclosed in the following description may be related to the launch of the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), a next-generation video / image coding standard after VVC, or a standard before VVC (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), etc.).

[0019] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. In the following, the same reference numerals are used to designate the same components in the drawings, and redundant description of the same components will be omitted.

[0020] In this specification, "video" refers to a collection of a series of images over time. "Picture" generally refers to a unit representing one image at a specific time period, and "slice" refers to a unit constituting a part of a picture in coding. One picture may be composed of multiple slices, and pictures and slices may be used interchangeably as needed.

[0021] A pixel or a pel may refer to the smallest unit constituting one picture (or image). A term corresponding to a pixel may also be used: "sample." A sample generally refers to a pixel or a pixel value, and may refer to only the value of a pixel / pixel of a luminance (luma) component, or may refer to only the value of a pixel / pixel of a chroma component.

[0022] A unit refers to a basic unit of video processing. A unit may include at least one of a specific region of a picture and information about the region. The term unit may be used interchangeably with terms such as block or area. In general, an MxN block may refer to a set of samples or transform coefficients consisting of M columns and N rows.

[0023] 1 is a diagram illustrating a schematic configuration of a video encoding apparatus to which the present invention can be applied. Hereinafter, the encoding / decoding apparatus may include a video encoding / decoding apparatus and / or a video encoding / decoding device, and the video encoding / decoding apparatus may be used as a concept including a video encoding / decoding apparatus, or the video encoding / decoding apparatus may be used as a concept including a video encoding / decoding apparatus.

[0024] 1, the video encoding apparatus 100 may include a picture partitioning module 105, a prediction module 110, a residual processing module 120, an entropy encoding module 130, an adder 140, a filtering module 150, and a memory 160. The residual processing module 120 may include a subtractor 121, a transform module 122, a quantization module 123, a rearrangement module 124, a dequantization module 125, and an inverse transform module 126.

[0025] The picture division unit 105 can divide an input picture into at least one processing unit.

[0026] For example, the processing unit is called a coding unit (CU). In this case, the coding units may be recursively divided from the largest coding unit (LCU) according to a quad-tree binary-tree (QTBT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure and the ternary tree structure. Alternatively, the binary tree structure / ternary tree structure may be applied first. The coding procedure according to the present invention may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, conversion, and restoration, which will be described later.

[0027] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding units may be split into coding units of deeper depths using a quadtree structure, starting from the largest coding unit (LCU). In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to video characteristics, or the coding unit may be recursively split into coding units of lower depths as needed, and the coding unit of the optimal size may be used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that serves as the basis for partitioning or division into prediction units or transform units. A prediction unit is a unit that is partitioned from a coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into subblocks. A transform unit can be divided from a coding unit according to a quadtree structure and is a unit that derives transform coefficients and / or a unit that derives a residual signal from the transform coefficients. Hereinafter, a coding unit is also referred to as a coding block (CB), a prediction unit is also referred to as a prediction block (PB), and a transform unit is also referred to as a transform block (TB). A prediction block or a prediction unit refers to a specific region in a block form within a picture and can include an array of prediction samples.Also, a transform block or transform unit refers to a specific region in a picture in the form of a block, and can include an array of transform coefficients or residual samples.

[0028] The prediction unit 110 performs prediction on a current block (hereinafter, may refer to a current block or a residual block) and generates a predicted block including prediction samples for the current block. The prediction unit 110 performs prediction on a coding block, a transform block, or a prediction block.

[0029] The predictor 110 may determine whether intra prediction or inter prediction is applied to the current block. For example, the predictor 110 may determine whether intra prediction or inter prediction is applied to each CU.

[0030] In intra prediction, the predictor 110 may derive a prediction sample for a current block based on a reference sample outside the current block within a picture to which the current block belongs (hereinafter, the current picture). In this case, the predictor 110 may (i) derive a prediction sample based on an average or interpolation of neighboring reference samples of the current block, or (ii) derive a prediction sample based on a reference sample present in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. (i) is referred to as a non-directional mode or a non-angular mode, and (ii) is referred to as a directional mode or an angular mode. Prediction modes in intra prediction may include, for example, 33 directional prediction modes and at least two or more non-directional modes. Non-directional modes may include a DC prediction mode and a planar mode. The predictor 110 may also determine a prediction mode to be applied to the current block using a prediction mode applied to a neighboring block.

[0031] In the case of inter prediction, the predictor 110 may derive a predicted sample for the current block based on a sample identified by a motion vector on a reference picture. The predictor 110 may derive a predicted sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the skip mode and the merge mode, the predictor 110 may use motion information of a neighboring block as motion information of the current block. In the skip mode, unlike the merge mode, a difference (residual) between a predicted sample and an original sample is not transmitted. In the MVP mode, the motion vector of the current block may be derived by using the motion vector of the neighboring block as a motion vector predictor.

[0032] In the case of inter-prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in a reference picture. The reference picture including the temporal neighboring blocks is also called a collocated picture (colPic). Motion information can include a motion vector and a reference picture index. Information such as prediction mode information and motion information can be (entropy) encoded and output in the form of a bitstream.

[0033] When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture. Reference pictures included in the reference picture list can be sorted based on the POC (Picture Order Count) difference between the current picture and the corresponding reference picture. POC corresponds to the display order of pictures and can be distinguished from the coding order.

[0034] The subtractor 121 generates residual samples, which are the differences between the original samples and the predicted samples. When the skip mode is applied, the residual samples are not generated as described above.

[0035] The transform unit 122 transforms residual samples in units of transform blocks to generate transform coefficients. The transform unit 122 may perform the transform according to the size of the corresponding transform block and a prediction mode applied to a coding block or a prediction block spatially overlapping with the corresponding transform block. For example, if intra prediction is applied to the coding block or the prediction block overlapping with the transform block and the transform block is a 4x4 residual array, the residual samples may be transformed using a Discrete Sine Transform (DST) transform kernel; otherwise, the residual samples may be transformed using a Discrete Cosine Transform (DCT) transform kernel.

[0036] The quantization unit 123 can quantize the transform coefficients to generate quantized transform coefficients.

[0037] The rearrangement unit 124 rearranges the quantized transform coefficients. The rearrangement unit 124 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form through a coefficient scanning method. Here, the rearrangement unit 124 has been described as a separate component, but it may also be a part of the quantization unit 123.

[0038] The entropy encoding unit 130 may perform entropy encoding on the quantized transform coefficients. Entropy encoding may include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 130 may also encode information required for video restoration (e.g., syntax element values) in addition to the quantized transform coefficients using entropy encoding or a preset method, either together with or separately from the quantized transform coefficients. The encoded information may be transmitted or stored in network abstraction layer (NAL) units in the form of a bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. The network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0039] The inverse quantization unit 125 inversely quantizes the values ​​(quantized transformation coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inversely transforms the values ​​inversely quantized by the inverse quantization unit 125 to generate residual samples.

[0040] The adder 140 reconstructs a picture by adding residual samples and predicted samples. The residual samples and predicted samples may be added in block units to generate reconstructed blocks. Although the adder 140 has been described as a separate component, it may be part of the prediction unit 110. Meanwhile, the adder 140 may also be referred to as a reconstruction module or a reconstructed block generator.

[0041] The filter unit 150 may apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through the deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process may be corrected. The sample adaptive offset may be applied on a sample-by-sample basis and may be applied after the deblocking filtering process is completed. The filter unit 150 may also apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filtering and / or the sample adaptive offset have been applied.

[0042] The memory 160 may store a reconstructed picture (a decoded picture) or information necessary for encoding / decoding. Here, a reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter unit 150. The stored reconstructed picture may be used as a reference picture for (inter) prediction of another picture. For example, the memory 160 may store (reference) pictures used for inter prediction. In this case, the pictures used for inter prediction may be specified by a reference picture set or a reference picture list.

[0043] 2 is a diagram illustrating a configuration of a video / image decoding apparatus to which the present invention can be applied. Hereinafter, the term "video decoding apparatus" may include a video decoding apparatus.

[0044] 2, the video decoding apparatus 200 may include an entropy decoding module 210, a residual processing module 220, a prediction module 230, an adder 240, a filtering module 250, and a memory 260. Here, the residual processing module 220 may include a rearrangement module 221, a dequantization module 222, and an inverse transform module 223. Although not shown, the video decoding apparatus 200 may also include a receiver that receives a bitstream including video information. The receiver may be configured as a separate module or may be included in the entropy decoding module 210.

[0045] When a bitstream including video / picture information is input, the video decoding apparatus 200 can restore the video / picture / image corresponding to the process in which the video / picture information was processed in the video encoding apparatus.

[0046] For example, the video decoding apparatus 200 may perform video decoding using a processing unit applied in a video encoding apparatus. Accordingly, a processing unit block for video decoding may be a coding unit, for example, or a coding unit, a prediction unit, or a transform unit, for example. The coding unit may be divided from the largest coding unit into a quad tree structure, a binary tree structure, and / or a ternary tree structure.

[0047] A prediction unit and a transform unit may also be used in some cases. In this case, a prediction block is a block derived or partitioned from a coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into sub-blocks. A transform unit may be divided from a coding unit using a quadtree structure and is a unit that derives transform coefficients or a unit that derives a residual signal from the transform coefficients.

[0048] The entropy decoding unit 210 may parse the bitstream and output information necessary for video or picture reconstruction. For example, the entropy decoding unit 210 may decode information in the bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements necessary for video reconstruction and quantized values ​​of transform coefficients for residuals.

[0049] More specifically, the CABAC entropy decoding method receives BINs corresponding to each syntax element from a bitstream, determines a context model using information on the syntax element to be decoded and decoding information on adjacent and target blocks or information on symbols / BINs decoded in a previous step, predicts the occurrence probability of BINs according to the determined context model, and performs arithmetic decoding of the BINs to generate symbols corresponding to the values ​​of each syntax element. In this case, the CABAC entropy decoding method can update the context model using information on the decoded symbols / BINs for the context model of the next symbol / BIN after determining the context model.

[0050] Among the information decoded by the entropy decoding unit 210, information regarding prediction is provided to the prediction unit 230, and the residual values, i.e., the quantized transform coefficients, on which entropy decoding is performed by the entropy decoding unit 210 can be input to the reordering unit 221.

[0051] The rearrangement unit 221 may rearrange the quantized transform coefficients in a two-dimensional block format. The rearrangement unit 221 may perform rearrangement in response to coefficient scanning performed in the encoding apparatus. Here, although the rearrangement unit 221 has been described as a separate component, it may also be a part of the inverse quantization unit 222.

[0052] The inverse quantization unit 222 may inversely quantize the quantized transform coefficients based on the (inverse) quantization parameter and output the transform coefficients. In this case, information for deriving the quantization parameter may be signaled from the encoding apparatus.

[0053] The inverse transform unit 223 can inversely transform the transform coefficients to derive residual samples.

[0054] The prediction unit 230 may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit 230 performs prediction on a coding block, a transform block, or a prediction block.

[0055] The prediction unit 230 may determine whether to apply intra prediction or inter prediction based on the information regarding the prediction. In this case, the unit for determining whether to apply intra prediction or inter prediction differs from the unit for generating prediction samples. In addition, the unit for generating prediction samples differs between inter prediction and intra prediction. For example, whether to apply inter prediction or intra prediction may be determined on a CU basis. Furthermore, for example, in inter prediction, a prediction mode may be determined on a PU basis to generate prediction samples, and in intra prediction, a prediction mode may be determined on a PU basis to generate prediction samples on a TU basis.

[0056] In the case of intra prediction, the prediction unit 230 may derive a prediction sample for the current block based on neighboring reference samples in the current picture. The prediction unit 230 may derive a prediction sample for the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined using the intra prediction mode of the neighboring block.

[0057] In the case of inter prediction, the predictor 230 may derive a prediction sample for the current block based on a sample identified on the reference picture by a motion vector on the reference picture. The predictor 230 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and an MVP mode. In this case, motion information required for inter prediction of the current block provided by the video encoding apparatus, such as information on a motion vector and a reference picture index, may be obtained or induced based on the information on the prediction.

[0058] In the skip mode and merge mode, motion information of neighboring blocks can be used as motion information of the current block, where the neighboring blocks can include spatial neighboring blocks and temporal neighboring blocks.

[0059] The predictor 230 constructs a merge candidate list using motion information of available neighboring blocks and can use information indicated by a merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding device. The motion information can include a motion vector and a reference picture. When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture.

[0060] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.

[0061] In the MVP mode, the motion vector of the current block can be derived using the motion vector of a neighboring block as a motion vector predictor, where the neighboring block can include a spatial neighboring block and a temporal neighboring block.

[0062] For example, when a merge mode is applied, a merge candidate list may be generated using the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vector corresponding to the Col block, which is a temporally neighboring block. In the merge mode, the motion vector of a candidate block selected from the merge candidate list is used as the motion vector of the current block. The prediction information may include a merge index indicating a candidate block having an optimal motion vector selected from the candidate blocks included in the merge candidate list. In this case, the prediction unit 230 may derive the motion vector of the current block using the merge index.

[0063] As another example, when a Motion Vector Prediction (MVP) mode is applied, a motion vector predictor candidate list may be generated using the motion vector of a reconstructed spatially neighboring block and / or the motion vector corresponding to a Col block, which is a temporally neighboring block. That is, the motion vector of a reconstructed spatially neighboring block and / or the motion vector corresponding to a Col block, which is a temporally neighboring block, may be used as a motion vector candidate. The prediction information may include a predicted motion vector index indicating an optimal motion vector selected from the motion vector candidates included in the list. In this case, the predictor 230 may select a predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. A predictor of the encoding apparatus may obtain a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode it, and output it in the form of a bitstream. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, the prediction unit 230 may obtain a motion vector differential included in the information for the prediction and derive the motion vector of the current block by adding the motion vector differential and the motion vector predictor. In addition, the prediction unit may obtain or induce a reference picture index indicating a reference picture from the information for the prediction.

[0064] The adder 240 may reconstruct a current block or a current picture by adding residual samples and predicted samples. The adder 240 may also reconstruct a current picture by adding residual samples and predicted samples in block units. When a skip mode is applied, residuals are not transmitted, and therefore predicted samples may become reconstructed samples. Although the adder 240 has been described as a separate component, it may also be part of the prediction unit 230. Meanwhile, the adder 240 may also be referred to as a reconstruction module or a reconstructed block generator.

[0065] The filter unit 250 may apply deblocking filtering, sample adaptive offset, and / or ALF to the reconstructed picture. In this case, the sample adaptive offset may be applied in sample units or may be applied after deblocking filtering. The ALF may be applied after deblocking filtering and / or sample adaptive offset.

[0066] The memory 260 may store a reconstructed picture (a decoded picture) or information necessary for decoding. Here, a reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter unit 250. For example, the memory 260 may store a picture used for inter prediction. In this case, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for another picture. The memory 260 may also output the reconstructed picture in an output order.

[0067] Meanwhile, as described above, prediction is performed to improve compression efficiency during video coding. Accordingly, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device, and the encoding device signals information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values ​​of the original block, to the decoding device, thereby improving video coding efficiency. The decoding device derives a residual block including residual samples based on the residual information, and adds the residual block and the predicted block to generate a reconstructed block including reconstructed samples, thereby generating a reconstructed picture including the reconstructed block.

[0068] The residual information may be generated through a transform and quantization procedure. For example, an encoding apparatus may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, thereby signaling the related residual information (via a bitstream) to a decoding apparatus. Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding apparatus may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding apparatus may generate a reconstructed picture based on the predicted block and the residual block. The encoding apparatus may also derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.

[0069] FIG. 3 is a diagram illustrating an example of a motion represented via an affine motion model according to an embodiment.

[0070] In this specification, "CP" is an abbreviation for control point and may refer to a sample or reference point that serves as a reference in the process of applying an affine motion model to a current block. The motion vector of a CP may be referred to as a "CPMV (Control Point Motion Vector)," and the CPMV may be derived based on a CPMV predictor, a "CPMVP (Control Point Motion Vector Predictor)."

[0071] 3, motions that can be represented using an affine motion model according to an embodiment may include translation motion, scale motion, rotation motion, and shear motion. That is, the affine motion model can efficiently represent translation motion in which (a part of) an image moves in a plane over time, scale motion in which (a part of) an image is scaled over time, rotation motion in which (a part of) an image is rotated over time, and shear motion in which (a part of) an image is deformed into the shape of a balanced quadrilateral over time.

[0072] According to an embodiment, affine inter prediction may be performed using an affine motion model. An encoding / decoding apparatus may predict the distortion type of an image based on a motion vector at a CP of a current block through affine inter prediction, thereby improving prediction accuracy and improving image compression performance. In addition, a motion vector for at least one CP of a current block may be derived using motion vectors of neighboring blocks of the current block, thereby reducing the data burden for added side information and improving inter prediction efficiency.

[0073] In one example, affine inter-prediction may be performed based on motion information at three CPs, i.e., three reference points, for the current block. The motion information at the three CPs for the current block may include a CPMV for each CP.

[0074] FIG. 4 exemplarily shows an affine motion model in which motion vectors for three CPs are used.

[0075] If the position of the top-left sample in the current block is (0,0), the width of the current block is w, and the height of the current block is h, then the samples located at (0,0), (w,0), and (0,h) can be determined as CPs for the current block, as shown in Figure 4. Hereinafter, the CP at the sample position (0,0) can be denoted as CP0, the CP at the sample position (w,0) as CP1, and the CP at the sample position (0,h) as CP2.

[0076] An affine motion model according to an embodiment can be applied using each CP and the motion vector for the corresponding CP. The affine motion model can be expressed as Equation 1 below.

[0077]

number

[0078] Here, w represents the width of the current block, h represents the height of the current block, and v represents the width of the current block. 0x , v 0y are the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y are the x and y components of the motion vector of CP1, respectively, and v 2x , v 2y indicate the x and y components of the motion vector of CP2, respectively. Also, x indicates the x component of the position of the target sample within the current block, y indicates the y component of the position of the target sample within the current block, and v x is the x-component of the motion vector of the target sample in the current block, v y denotes the y component of the motion vector of the target sample in the current block.

[0079] Meanwhile, Equation 1 showing an affine motion model is merely an example, and an equation for showing an affine motion model is not limited to Equation 1. For example, the sign of each coefficient disclosed in Equation 1 may differ from Equation 1 depending on the case, and the magnitude of the absolute value of each coefficient may also differ from Equation 1 depending on the case.

[0080] Since the motion vectors of CP0, CP1, and CP2 are known, a motion vector according to the sample position in the current block can be derived based on Equation 1. That is, according to the affine motion model, the motion vector v0 (v 0x ,v 0y ), v1(v 1x ,v 1y ), v2(v 2x ,v 2y ) is scaled, and a motion vector of the target sample according to the position of the target sample can be derived. That is, according to the affine motion model, a motion vector of each sample in the current block can be derived based on the motion vector of the CP. Meanwhile, a set of motion vectors of samples in the current block derived according to the affine motion model can be referred to as an affine motion vector field.

[0081] Meanwhile, the six parameters for Equation 1 can be expressed as a, b, c, d, e, and f as in the following equation, and the equation for the affine motion model expressed by the six parameters is as follows:

[0082]

number

[0083] Here, w represents the width of the current block, h represents the height of the current block, and v represents the width of the current block. 0x , v 0yindicate the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y are the x and y components of the motion vector of CP1, respectively, and v 2x , v 2y indicate the x and y components of the motion vector of CP2, respectively. Also, x indicates the x component of the position of the target sample within the current block, y indicates the y component of the position of the target sample within the current block, and v x is the x-component of the motion vector of the target sample in the current block, v y denotes the y component of the motion vector of the target sample in the current block.

[0084] Meanwhile, Equation 2 representing an affine motion model based on six parameters is merely an example, and an equation for representing an affine motion model based on six parameters is not limited to Equation 2. For example, the sign of each coefficient disclosed in Equation 2 may differ from Equation 2 depending on the case, and the magnitude of the absolute value of each coefficient may also differ from Equation 2 depending on the case.

[0085] The affine motion model or the affine inter prediction using the six parameters may be denoted as a six-parameter affine motion model or AF6.

[0086] In one example, affine inter-prediction may be performed based on motion information at three CPs, i.e., three reference points, for the current block. The motion information at the three CPs for the current block may include a CPMV for each CP.

[0087] In one example, affine inter-prediction may be performed based on motion information at two CPs, i.e., two reference points, for the current block. The motion information at the two CPs for the current block may include a CPMV for each CP.

[0088] FIG. 5 exemplarily shows an affine motion model in which motion vectors for two CPs are used.

[0089] An affine motion model using two CPs can represent three types of motion, including translational motion, scale motion, and rotational motion. An affine motion model that represents three types of motion is sometimes called a similarity affine motion model or a simplified affine motion model.

[0090] If the position of the top-left sample in the current block is (0,0), and the width and height of the current block are w and h, respectively, the samples located at (0,0) and (w,0) may be determined as CPs for the current block, as shown in Figure 5. Hereinafter, the CP at the sample position (0,0) may be referred to as CP0, and the CP at the sample position (w,0) may be referred to as CP1.

[0091] Using each CP and the motion vector for the corresponding CP, an affine motion model based on four parameters can be applied. The affine motion model can be expressed as Equation 3 below.

[0092]

number

[0093] where w represents the width of the current block, and v 0x , v 0y are the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively. Also, x indicates the x component of the position of the target sample within the current block, y indicates the y component of the position of the target sample within the current block, and v x is the x-component of the motion vector of the target sample in the current block, v y denotes the y component of the motion vector of the target sample in the current block.

[0094] Meanwhile, Equation 3 showing an affine motion model based on four parameters is merely an example, and an equation for showing an affine motion model based on four parameters is not limited to Equation 3. For example, the sign of each coefficient disclosed in Equation 3 may differ from Equation 3 depending on the case, and the magnitude of the absolute value of each coefficient may also differ from Equation 3 depending on the case.

[0095] Meanwhile, the four parameters for Equation 3 can be expressed as a, b, c, and d as in Equation 4 below, and Equation 4 for the affine motion model expressed by the four parameters is as follows.

[0096]

number

[0097] where w represents the width of the current block, and v 0x , v 0y are the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively. Also, x indicates the x component of the position of the target sample within the current block, y indicates the y component of the position of the target sample within the current block, and v x is the x-component of the motion vector of the target sample in the current block, v y indicates the y-component of the motion vector of the target sample in the current block. Since the affine motion model using the two CPs can be expressed by four parameters a, b, c, and d as in Equation 4, the affine motion model or the affine inter-prediction using these four parameters can be referred to as a four-parameter affine motion model or AF4. That is, according to the affine motion model, a motion vector for each sample in the current block can be derived based on the motion vectors of the control points. Meanwhile, a set of motion vectors of samples in the current block derived according to the affine motion model can be referred to as an affine motion vector field.

[0098] Meanwhile, Equation 4 showing an affine motion model based on four parameters is merely an example, and an equation for showing an affine motion model based on four parameters is not limited to Equation 4. For example, the sign of each coefficient disclosed in Equation 4 may differ from Equation 4 depending on the case, and the magnitude of the absolute value of each coefficient may also differ from Equation 4 depending on the case.

[0099] Meanwhile, as described above, the affine motion model can derive a motion vector for each sample, thereby significantly improving the accuracy of inter-prediction, although this may result in a significant increase in the complexity of the motion compensation process.

[0100] In another embodiment, instead of deriving a motion vector in a sample unit, a motion vector in a sub-block unit within the current block may be derived.

[0101] FIG. 6 is a diagram showing an example of deriving a motion vector for each sub-block based on an affine motion model.

[0102] 6 illustrates an example in which the size of the current block is 16x16 and motion vectors are derived in units of 4x4 sub-blocks. The sub-blocks can be set to various sizes. For example, if the sub-blocks are set to an nxn size (n is a positive integer, e.g., 4), motion vectors can be derived in units of nxn sub-blocks within the current block based on the affine motion model, and various methods can be applied to derive motion vectors representing each sub-block.

[0103] For example, referring to FIG. 6, a motion vector for each sub-block may be derived using the center or lower right sample position of the center of each sub-block as a representative coordinate. Here, the lower right sample position of the center may refer to the lower right sample position of four samples located at the center of the sub-block. For example, if n is an odd number, one sample may be located at the center of the sub-block, and in this case, the center sample position may be used to derive the motion vector for the sub-block. However, if n is an even number, four samples may be located adjacent to the center of the sub-block, and in this case, the lower right sample position may be used to derive the motion vector. For example, referring to FIG. 6, the representative coordinates for each sub-block may be derived as (2,2), (6,2), (10,2), ..., (14,14), and the encoding / decoding apparatus may derive the motion vector for each sub-block by substituting the representative coordinates of the sub-blocks into Equation 1 or 3. The motion vectors of the sub-blocks within the current block derived through the affine motion model may be referred to as affine MVFs.

[0104] In one embodiment, the affine motion model described above can be organized into two steps: a step of deriving CPMV and a step of performing affine motion compensation.

[0105] Meanwhile, inter prediction using the above-described affine motion model, i.e., affine motion prediction, may include an affine merge mode (AF_MERGE or AAM) and an affine inter mode (AF_INTER or AAMVP).

[0106] The affine merge mode (AAM) according to one embodiment may refer to an encoding / decoding method in which prediction is performed by deriving CPMVs for each of two or three CPs from neighboring blocks of the current block without coding for motion vector difference (MVD), similar to the existing skip / merge mode. The affine inter mode (AAMVP) may refer to a method in which difference information between CPMV and CPMVP is explicitly encoded / decoded, similar to AMVP.

[0107] Meanwhile, the explanation of the affine motion model described above in Figures 3 to 6 is intended to aid in understanding the principles of the encoding / decoding method according to one embodiment of the present invention described later in this specification, and therefore, it should be easily understood by those skilled in the art that the scope of the present invention is not limited by the contents described above in Figures 3 to 6.

[0108] In one embodiment, a method for constructing an affine MVP candidate list for affine inter-prediction is described. In this specification, the affine MVP candidate list is composed of affine MVP candidates, and each affine MVP candidate may refer to a combination of CPMVPs of CP0 and CP1 in a four-parameter (affine) motion model, or a combination of CPMVPs of CP0, CP1, and CP2 in a six-parameter (affine) motion model. The affine MVP candidates described in this specification may be referred to by various names, such as CPMVP candidate, affine CPMVP candidate, CPMVP pair candidate, CPMVP pair, etc. The affine MVP candidate list may include n affine MVP candidates, and if n is an integer greater than 1, encoding and decoding of information indicating the optimal affine MVP candidate may be necessary. If n is 1, encoding and decoding of information indicating the optimal affine MVP candidate may not be necessary. An example of the syntax when n is an integer greater than 1 is shown in Table 1 below, and an example of the syntax when n is 1 is shown in Table 2 below.

[0109] [Table 1]

[0110] [Table 2]

[0111] In Tables 1 and 2, merge_flag is a flag indicating whether or not a merge mode is in effect. When the value of merge_flag is 1, the merge mode is performed, and when the value of merge_flag is 0, the merge mode may not be performed. affine_flag is a flag indicating whether or not affine motion prediction is used. When the value of affine_flag is 1, affine motion prediction is used, and when the value of affine_flag is 0, affine motion prediction may not be used. aamvp_idx is index information indicating the optimal affine MVP candidate among n affine MVP candidates. In Table 1, where n is an integer greater than 1, the optimal affine MVP candidate is indicated based on the aamvp_idx, whereas in Table 2, where n is 1, there is only one affine MVP candidate, so it can be seen that aamvp_idx is not parsed.

[0112] In one embodiment, when determining affine MVP candidates, affine motion models of neighboring blocks coded based on affine motion prediction (hereinafter, also referred to as "affine coding blocks") may be used. In one example, when determining affine MVP candidates, a first step and a second step may be performed. In the first step, neighboring blocks may be scanned in a predefined order, and it may be determined whether each neighboring block is coded based on affine motion prediction. In the second step, affine MVP candidates for the current block may be determined using neighboring blocks coded based on affine motion prediction.

[0113] Up to m blocks coded based on affine motion prediction in the first step may be considered. For example, if m is 1, an affine MVP candidate may be determined using the first affine coding block in the scanning order. For example, if m is 2, at least one affine MVP candidate may be determined using the first and second affine coding blocks in the scanning order. In this case, a pruning check is performed. If the first and second affine MVP candidates are the same, a further scanning process may be performed to determine a further affine MVP candidate. Meanwhile, in one example, m described in this embodiment may not exceed the value of n described above in the description of Tables 1 and 2.

[0114] Meanwhile, there are various embodiments for the process of scanning the neighboring blocks in the first step and determining whether each neighboring block has been coded based on affine motion prediction. Hereinafter, with reference to Figures 7 to 10, an embodiment for the process of scanning the neighboring blocks and determining whether each neighboring block has been coded based on affine motion prediction will be described.

[0115] 7 to 10 show examples of methods for detecting neighboring blocks coded based on affine motion prediction.

[0116] 7, 4x4 blocks A, B, C, D, and E are shown around the current block. Block E, which is the upper left corner peripheral block, is located around CP0, block C, which is the upper right corner peripheral block, and block B, which is the upper peripheral block, are located around CP1, and block D, which is the lower left corner peripheral block, and block A, which is the left peripheral block, are located around CP2. The arrangement shown in FIG. 7 can share a structure with the method according to AMVP or merge mode, which can contribute to reducing design costs.

[0117] 8, 4x4 blocks A, B, C, D, E, F, and G are shown surrounding a current block. Block E, which is the upper left corner neighboring block, block G, which is the first left neighboring block, and block F, which is the first upper neighboring block, are located around CP0. Block C, which is the upper right neighboring block, and block B, which is the second upper neighboring block, are located around CP1. Block D, which is the lower left corner neighboring block, and block A, which is the second left neighboring block, are located around CP2. The arrangement shown in FIG. 8 minimizes an increase in scanning complexity and may be effective in terms of coding performance, since it determines whether a block is coded based on affine motion prediction based only on the 4x4 blocks adjacent to three CPs.

[0118] 9 has the same arrangement of neighboring blocks as FIG. 8 to be scanned when detecting neighboring blocks coded based on affine motion prediction. However, in the embodiment of FIG. 9, affine MVP candidates may be determined based on up to p of the 4x4 neighboring blocks included within the dotted line located to the left of the current block and up to q of the 4x4 neighboring blocks included within the dotted line located above the current block. For example, if p and q are both 1, affine MVP candidates may be determined based on the first affine-coded block in the scanning order among the 4x4 neighboring blocks included within the dotted line located to the left of the current block and the first affine-coded block in the scanning order among the 4x4 neighboring blocks included within the dotted line located above the current block.

[0119] Referring to Figure 10, an affine MVP candidate can be determined based on the first affine coding block in the scanning order among block E, which is the peripheral block in the upper left corner located around CP0, block G, which is the first left peripheral block, and block F, which is the first upper peripheral block, the first affine coding block in the scanning order among block C, which is the upper right peripheral block located around CP1, and block B, which is the second upper peripheral block, and the first affine coding block in the scanning order among block D, which is the peripheral block in the lower left corner located around CP2, and block A, which is the second left peripheral block.

[0120] Meanwhile, the scanning order of the above-described scanning method may be determined based on the probability and performance analysis of a specific encoding or decoding device. Thus, according to one embodiment, the scanning order is not specified, but may be determined based on the statistical characteristics or performance of the encoding or decoding device to which this embodiment is applied.

[0121] FIG. 11 is a flowchart illustrating an operation method of an encoding apparatus according to an embodiment, and FIG. 12 is a block diagram illustrating a configuration of an encoding apparatus according to an embodiment.

[0122] The encoding apparatus according to Figures 11 and 12 may perform operations corresponding to those of the decoding apparatus according to Figures 13 and 14, which will be described later. Therefore, the contents described below with reference to Figures 13 and 14 may also be applied to the encoding apparatus according to Figures 11 and 12.

[0123] The steps disclosed in Fig. 11 may be performed by the encoding apparatus 100 disclosed in Fig. 1. More specifically, steps S1100 to S1140 may be performed by the prediction unit 110 disclosed in Fig. 1, step S1150 may be performed by the residual processing unit 120 disclosed in Fig. 1, and step S1160 may be performed by the entropy encoding unit 130 disclosed in Fig. 1. In addition, the operations of steps S1100 to S1160 are based on some of the content described above with reference to Figs. 3 to 10. Therefore, detailed descriptions that overlap with the content described above with reference to Figs. 1 and 3 to 10 will be omitted or simplified.

[0124] As shown in Fig. 12, an encoding apparatus according to an embodiment may include a prediction unit 110 and an entropy encoding unit 130. However, depending on the case, none of the components shown in Fig. 12 may be essential components of the encoding apparatus, and the encoding apparatus may be implemented with more or fewer components than those shown in Fig. 12.

[0125] In the encoding device according to one embodiment, the prediction unit 110 and the entropy encoding unit 130 may be implemented on separate chips, or at least two or more components may be implemented on a single chip.

[0126] An encoding apparatus according to an embodiment may generate an affine MVP candidate list including affine MVP candidates for a current block (S1100). More specifically, a prediction unit 110 of the encoding apparatus may generate an affine MVP candidate list including affine MVP candidates for a current block.

[0127] An encoding apparatus according to an embodiment may derive a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1110). More specifically, a prediction unit 110 of the encoding apparatus may derive a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list.

[0128] The encoding apparatus according to an embodiment may derive a CPMV for each of the CPs of the current block (S1120). More specifically, the prediction unit 110 of the encoding apparatus may derive a CPMV for each of the CPs of the current block.

[0129] An encoding apparatus according to an embodiment may derive a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each of the CPs (S1130). More specifically, a prediction unit 110 of the encoding apparatus may derive a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each of the CPs.

[0130] The encoding apparatus according to an embodiment may derive predicted samples for the current block based on the CPMV (S1140). More specifically, the prediction unit 110 of the encoding apparatus may derive predicted samples for the current block based on the CPMV.

[0131] The encoding apparatus according to an embodiment may derive residual samples for the current block based on the derived predicted samples (S1150). More specifically, the residual processing unit 120 of the encoding apparatus may derive residual samples for the current block based on the derived predicted samples.

[0132] The encoding apparatus according to an embodiment may encode information on the derived CPMVD and residual information on the residual samples (S1160). More specifically, the entropy encoding unit 130 of the encoding apparatus may encode information on the derived CPMVD and residual information on the residual samples.

[0133] According to the encoding device and the operating method of the encoding device disclosed in Figures 11 and 12, the encoding device generates an affine MVP candidate list including affine MVP candidates for a current block (S1100), derives a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1110), derives a CPMV for each CP of the current block (S1120), derives a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each CP (S1130), derives a predicted sample for the current block based on the CPMV (S1140), derives a residual sample for the current block based on the derived predicted sample (S1150), and encodes information on the derived CPMVD and residual information related to the residual sample (S1160). That is, by signaling information on the affine MVP candidate list used in affine motion prediction, video coding efficiency can be improved.

[0134] FIG. 13 is a flowchart illustrating an operation method of a decoding device according to an embodiment, and FIG. 14 is a block diagram illustrating a configuration of a decoding device according to an embodiment.

[0135] The steps disclosed in Figure 13 may be performed by the decoding apparatus 200 disclosed in Figure 2. More specifically, S1300 may be performed by the entropy decoding unit 210 disclosed in Figure 2, S1310 to S1350 may be performed by the prediction unit 230 disclosed in Figure 2, and S1360 may be performed by the addition unit 240 disclosed in Figure 2. In addition, the operations of S1300 to S1360 are based on some of the content described above with reference to Figures 3 to 10. Therefore, detailed descriptions that overlap with the content described above with reference to Figures 2 to 10 will be omitted or simplified.

[0136] 14, a decoding device according to one embodiment may include an entropy decoding unit 210, a prediction unit 230, and an addition unit 240. However, depending on the case, none of the components shown in FIG. 14 may be essential components of the decoding device, and the decoding device may be implemented with more or fewer components than those shown in FIG.

[0137] In one embodiment of the decoding device, the entropy decoding unit 210, the prediction unit 230, and the addition unit 240 may each be implemented on a separate chip, or at least two or more components may be implemented on a single chip.

[0138] A decoding apparatus according to an embodiment may acquire motion prediction information from a bitstream (S1300). More specifically, an entropy decoding unit 210 of the decoding apparatus may acquire motion prediction information from the bitstream.

[0139] A decoding apparatus according to an embodiment may generate an affine MVP candidate list including affine motion vector predictor (MVP) candidates for a current block (S1310). More specifically, a prediction unit 230 of the decoding apparatus may generate an affine MVP candidate list including affine MVP candidates for the current block.

[0140] In one embodiment, the affine MVP candidates include a first affine MVP candidate and a second affine MVP candidate, where the first affine MVP candidate is derived from a left block group including a bottom-left corner neighboring block and a left neighboring block of the current block, and the second affine MVP candidate is derived from a top block group including a top-right corner neighboring block, a top neighboring block, and a top-left corner neighboring block of the current block. In this case, the first affine MVP candidate may be derived based on a first block included in the left block group, which is coded based on affine motion prediction, and the second affine MVP candidate may be derived based on a second block included in the top block group, which is coded based on affine motion prediction.

[0141] In another embodiment, the affine MVP candidates include a first affine MVP candidate and a second affine MVP candidate, where the first affine MVP candidate is derived from a left block group including a neighboring block at the lower left corner of the current block, a first left neighboring block, and a second left neighboring block, and the second affine MVP candidate is derived from an upper block group including a neighboring block at the upper right corner of the current block, a first upper neighboring block, a second upper neighboring block, and a neighboring block at the upper left corner of the current block. In this case, the first affine MVP candidate may be derived based on a first block included in the left block group, the first block being coded based on affine motion prediction, and the second affine MVP candidate may be derived based on a second block included in the upper block group, the second block being coded based on affine motion prediction.

[0142] In yet another embodiment, the affine MVP candidates include a first affine MVP candidate, a second affine MVP candidate, and a third affine MVP candidate, wherein the first affine MVP candidate is derived from a lower left block group including a peripheral block at the lower left corner of the current block and a first left peripheral block, the second affine MVP candidate is derived from a upper right block group including a peripheral block at the upper right corner of the current block and a first upper peripheral block, and the third affine MVP candidate is derived from an upper left block group including a peripheral block at the upper left corner of the current block, a second upper peripheral block, and a second left peripheral block. In this case, the first affine MVP candidate may be derived based on a first block included in the bottom left block group, and the first block may be coded based on affine motion prediction; the second affine MVP candidate may be derived based on a second block included in the top right block group, and the second block may be coded based on affine motion prediction; and the third affine MVP candidate may be derived based on a third block included in the top left block group, and the third block may be coded based on affine motion prediction.

[0143] A decoding apparatus according to one embodiment may derive control point motion vector predictors (CP MVPs) for each control point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1320). More specifically, a prediction unit 230 of the decoding apparatus may derive a CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list.

[0144] In one embodiment, the one affine MVP candidate may be selected from the affine MVP candidates based on an affine MVP candidate index included in the motion prediction information.

[0145] A decoding apparatus according to an embodiment may derive the CPMVD for the CP of the current block based on information on the CPMVD for each of the CPs included in the acquired motion prediction information (S1330). More specifically, a prediction unit 230 of the decoding apparatus may derive the CPMVD for the CP of the current block based on information on the CPMVD for each of the CPs included in the acquired motion prediction information.

[0146] A decoding apparatus according to an embodiment may derive a CPMV for the CP of the current block based on the CPMVP and the CPMVD (S1340). More specifically, a prediction unit 230 of the decoding apparatus may derive a CPMV for the CP of the current block based on the CPMVP and the CPMVD.

[0147] A decoding apparatus according to an embodiment may derive predicted samples for the current block based on the CPMV (S1350). More specifically, a prediction unit 230 of the decoding apparatus may derive predicted samples for the current block based on the CPMV.

[0148] The decoding apparatus according to an embodiment may generate reconstructed samples for the current block based on the derived predicted samples (S1360). More specifically, the adder 240 of the decoding apparatus may generate reconstructed samples for the current block based on the derived predicted samples.

[0149] In one embodiment, the motion prediction information may include information on a context index indicating whether or not there is a neighboring block for the current block coded based on affine motion prediction.

[0150] In one embodiment, when the value of m in the description of the first step is 1 and the value of n in the description of Tables 1 and 2 is 2, a CABAC context model for encoding and decoding index information for indicating an optimal affine MVP candidate can be configured. When an affine coding block exists around the current block, the affine MVP candidate for the current block can be determined based on the affine motion model as described with reference to FIGS. 7 to 10. However, when an affine coding block does not exist around the current block, this embodiment can be applied. When an affine MVP candidate is determined based on an affine coding block, the reliability of the affine MVP candidate is high. Therefore, a context model can be designed by classifying the case where the affine MVP candidate is determined based on the affine coding block and the case where it is not. In this case, an index of 0 can be assigned to the affine MVP candidate determined based on the affine coding block. The CABAC context index according to this embodiment is expressed as Equation 5 below.

[0151]

number

[0152] The initial value according to the CABAC context index may be determined as shown in Table 3 below, and the CABAC context index and the initial value must satisfy the condition of Equation 6 below.

[0153] [Table 3]

[0154]

number

[0155] According to the decoding apparatus and the operating method of the decoding apparatus of FIGS. 13 and 14, the decoding apparatus acquires motion prediction information from a bitstream (S1300), generates an affine MVP candidate list including affine MVP candidates for a current block (S1310), derives Control Point Motion Vector Predictors (CPMVP) for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1320), derives Control Point Motion Vector Differences (CPMVD) for each CP of the current block based on information on Control Point Motion Vector Differences (CPMVD) for each CP included in the acquired motion prediction information (S1330), and calculates Control Point Motion Vector Differences (CPMVD) for the CP of the current block based on the CPMVP and the CPMVD. The affine motion prediction signal (CPMV) is used to derive a vector (S1340), derive predicted samples for the current block based on the CPMV (S1350), and generate reconstructed samples for the current block based on the derived predicted samples (S1360). That is, by signaling information on an affine MVP candidate list used in affine motion prediction, video coding efficiency can be improved.

[0156] Meanwhile, the methods according to the above-described embodiments of the present specification relate to image and video compression and can be applied to both encoding and decoding devices, and to devices that generate or receive bitstreams, regardless of whether the data is output through a display device in a terminal. For example, an image may be generated as compressed data by a terminal having an encoding device, and the compressed data may be in the form of a bitstream. The bitstream may be stored in various types of storage devices, streamed over a network, and transmitted to a terminal having a decoding device. If the terminal is equipped with a display device, the decoded image may be displayed on the display device, or the bitstream data may simply be stored.

[0157] The above-described method according to the present invention may be implemented in the form of software, and the encoding device and / or decoding device according to the present invention may be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.

[0158] Each of the above-described parts, modules, or units may be a processor or hardware part that executes a sequential execution process stored in a memory (or storage unit). Each step described in the above-described embodiments may be performed by a processor or hardware part. Each of the modules / blocks / units described in the above-described embodiments may operate as hardware / processors. Also, the method presented by the present invention may be implemented as code. This code may be stored in a processor-readable storage medium and thus read by a processor provided by an apparatus.

[0159] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and that different steps may be included, or one or more steps in the flowcharts may be deleted without affecting the scope of the present invention.

[0160] When an embodiment of the present invention is implemented as software, the methods described above may be implemented as modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices.

Claims

1. A video decoding method performed by a decoding device, generating a list of affine MVP candidates for the current block; deriving Control Point Motion Vector Predictors (CPMVPs) for each Control Point (CP) of the current block based on the affine MVP candidate list; deriving Control Point Motion Vector Differences (CPMVDs) based on information about the CPMVDs for each CP; deriving a Control Point Motion Vector (CPMV) for each CP of the current block based on the CPMVP and the CPMVD; deriving a predicted sample for the current block based on the CPMV; the affine MVP candidates in the affine MVP candidate list include a first affine MVP candidate and a second affine MVP candidate; The CPs include CP0, CP1, and CP2, and the CPMVPs for each of the CPs include a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2; The step of generating the affine MVP candidate list comprises: deriving a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2, which constitute the first affine MVP candidate, based on blocks coded based on an affine motion model in a left block group including a neighboring block at a lower left corner of the current block and a left neighboring block; Video decoding method.

2. A video encoding method performed by an encoding device, generating a list of affine MVP candidates for the current block; deriving a CPMVP for each CP (Control Point) of the current block based on the affine MVP candidate list; deriving a CPMV for each CP of the current block; deriving a CPMVD based on the CPMV and the CPMVP; deriving a predicted sample for the current block based on the CPMV; the affine MVP candidates in the affine MVP candidate list include a first affine MVP candidate and a second affine MVP candidate; The CPs include CP0, CP1, and CP2, and the CPMVPs for each of the CPs include a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2; The step of generating the affine MVP candidate list comprises: A video encoding method comprising the step of deriving a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2, which constitute the first affine MVP candidate, based on blocks coded based on an affine motion model in a left block group including a peripheral block at the lower left corner of the current block and a left peripheral block.

3. 1. A non-transitory computer-readable storage medium, comprising: The non-transitory computer-readable storage medium has stored thereon a computer program, which, when executed by a processor, performs steps of a video encoding method for generating a bitstream; The video encoding method includes: generating a list of affine MVP candidates for the current block; deriving a CPMVP for each CP (Control Point) of the current block based on the affine MVP candidate list; deriving a CPMV for each CP of the current block; deriving a CPMVD based on the CPMV and the CPMVP; deriving a predicted sample for the current block based on the CPMV; the affine MVP candidates in the affine MVP candidate list include a first affine MVP candidate and a second affine MVP candidate; The CPs include CP0, CP1, and CP2, and the CPMVPs for each of the CPs include a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2; The step of generating the affine MVP candidate list comprises: A non-transitory computer-readable storage medium comprising: a step of deriving a first MVP for CP0, a second MVP for CP1, and a third MVP for CP2, which constitute the first affine MVP candidate, based on blocks coded based on an affine motion model in a left block group including a peripheral block at the lower left corner of the current block and a left peripheral block.

4. In a method of transmitting data for video, obtaining a bitstream for the video, the bitstream being generated based on generating an affine MVP candidate list for a current block, deriving a CPMVP for each CP (Control Point) of the current block based on the affine MVP candidate list, deriving a CPMV for each CP of the current block, deriving a CPMVD based on the CPMV and the CPMVP, and deriving a predicted sample for the current block based on the CPMV; transmitting the data including the bitstream; the affine MVP candidates in the affine MVP candidate list include a first affine MVP candidate and a second affine MVP candidate; The CPs include CP0, CP1, and CP2, and the CPMVPs for each of the CPs include a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2; generating the affine MVP candidate list, A data transmission method comprising: deriving a first MVP for CP0, a second MVP for CP2 for CP1, and a third MVP for CP2, which constitute the first affine MVP candidate, based on blocks coded based on an affine motion model in a left block group including a peripheral block at the lower left corner of the current block and a left peripheral block.