Method and apparatus for video coding based on affine motion prediction
The method and apparatus enhance video coding efficiency by employing affine motion prediction techniques, specifically through generating affine MVP candidate lists and encoding residual information, addressing the increased costs associated with high-resolution image transmission and storage.
Patent Information
- Application Number
- JP2021502682
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-01
- Filing Date
- 2019-04-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-04-01
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to higher transmission and storage costs due to increased data volume, necessitating more efficient video coding techniques.
A method and apparatus for video coding that improves efficiency by utilizing affine motion prediction, including generating an affine MVP candidate list, deriving Control Point Motion Vector Predictors (CPMVP) and Motion Vector Differences (CPMVD), and encoding residual information.
Enhances the efficiency of video coding by optimizing the use of affine motion prediction, reducing data volume and improving compression performance.
Smart Images

Figure 0007675646000010 
Figure 0007675646000011 
Figure 0007675646000012
Abstract
Description
[Technical field]
[0001] The present invention relates to a video coding technology, and more particularly, to a video coding method and apparatus based on affine motion prediction in a video coding system. [Background technology]
[0002] Recently, the demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images is increasing in various fields. As the image data has higher resolution and quality, the amount of information or bits to be transmitted increases relatively compared to existing image data, so that when the image data is transmitted using a medium such as an existing wired or wireless broadband line or when the image data is stored using an existing storage medium, the transmission cost and storage cost increase.
[0003] This calls for highly efficient image compression techniques to effectively transmit, store and reproduce high resolution, high quality image information. Summary of the Invention [Problem to be solved by the invention]
[0004] SUMMARY OF THE PRESENT EMBODIMENT It is an object of the present invention to provide a method and apparatus for improving the efficiency of video coding.
[0005] Another technical object of the present invention is to provide a method and apparatus for improving the efficiency of video coding based on affine motion prediction.
[0006] It is still another technical object of the present invention to provide a method and apparatus for efficiently determining a combination of neighboring blocks used in affine motion prediction, thereby improving the efficiency of video coding.
[0007] It is yet another technical object of the present invention to provide a method and apparatus for improving the efficiency of video coding by signaling information on an affine MVP candidate list used in affine motion prediction. [Means for solving the problem]
[0008] According to an embodiment of the present invention, there is provided a picture decoding method performed by a decoding apparatus, the method including the steps of: acquiring motion prediction information from a bitstream; generating an affine MVP candidate list including affine motion vector predictor (MVP) candidates for a current block; deriving Control Point Motion Vector Predictors (CPMVP) for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list; deriving the CPMVD for the CP of the current block based on information on Control Point Motion Vector Differences (CPMVD) for each CP included in the acquired motion prediction information; deriving a Control Point Motion Vector (CPMV) for the CP of the current block based on the CPMVP and the CPMVD; deriving a prediction sample for the current block based on the CPMV; and generating a reconstructed sample for the current block based on the derived prediction sample.
[0009] According to another embodiment of the present invention, there is provided a decoding device for performing picture decoding, the decoding device including an entropy decoding unit for acquiring motion prediction information from a bitstream, a prediction unit for generating an affine MVP candidate list including affine motion vector predictor (MVP) candidates for a current block, deriving Control Point Motion Vector Predictors (CPMVP) for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list, deriving the CPMVD for the CP of the current block based on information on Control Point Motion Vector Differences (CPMVD) for each CP included in the acquired motion prediction information, deriving a CPMV for the CP of the current block based on the CPMVP and the CPMVD, and deriving a prediction sample for the current block based on the CPMV, and an adder for generating a reconstructed sample for the current block based on the derived prediction sample.
[0010] According to another embodiment of the present invention, there is provided a picture encoding method performed by an encoding apparatus, the method including the steps of: generating an affine MVP candidate list including affine MVP candidates for a current block, deriving a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list, deriving a CPMV for each CP of the current block, deriving a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each CP, deriving a prediction sample for the current block based on the CPMV, deriving a residual sample for the current block based on the derived prediction sample, and encoding information on the derived CPMVD and residual information on the residual sample.
[0011] According to another embodiment of the present invention, there is provided an encoding apparatus for performing picture encoding, comprising: a prediction unit that generates an affine MVP candidate list including affine MVP candidates for a current block, derives a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list, derives a CPMV for each CP of the current block, derives a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each CP, and derives a prediction sample for the current block based on the CPMV, a residual processing unit that derives a residual sample for the current block based on the derived prediction sample, and an entropy encoding unit that encodes information on the derived CPMVD and residual information on the residual sample. Effect of the Invention
[0012] The present invention can improve the efficiency of image / video compression in general.
[0013] According to the present invention, the efficiency of video coding based on affine motion prediction can be improved.
[0014] According to the present invention, the efficiency of video coding can be improved by signaling information regarding an affine MVP candidate list used in affine motion prediction. [Brief description of the drawings]
[0015] [Figure 1] 1 is a diagram illustrating a schematic configuration of an encoding device according to an embodiment; [Diagram 2] 1 is a diagram illustrating the configuration of a decoding device according to an embodiment of the present invention; [Diagram 3] FIG. 2 is a diagram illustrating an example of motion represented via an affine motion model according to an embodiment. [Figure 4] FIG. 13 is a diagram showing an example of an affine motion model using CPMVs (Control Point Motion Vectors) of three CPs (Control Points) for a current block. [Diagram 5] FIG. 2 shows an example of an affine motion model using CPMV of two CPs for a current block. [Figure 6] FIG. 13 is a diagram showing an example of deriving a motion vector on a sub-block basis based on an affine motion model. [Figure 7] 1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 8] 1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 9]1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 10] 1 illustrates an example of a method for detecting neighboring blocks coded based on affine motion prediction. [Figure 11] 4 is a flowchart illustrating an operation method of an encoding apparatus according to an embodiment. [Figure 12] 1 is a block diagram showing a configuration of an encoding device according to an embodiment; [Figure 13] 4 is a flowchart illustrating an operation method of a decoding device according to an embodiment. [Figure 14] 2 is a block diagram showing the configuration of a decoding device according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] The present invention can be modified in various ways and can have various embodiments, and a specific embodiment will be illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiment. The terms used in this specification are used only to describe a specific embodiment and are not intended to limit the technical idea of the present invention. A singular expression includes a plural expression unless the context clearly indicates otherwise. In this specification, the terms "include" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0017] Meanwhile, each component in the drawings described in the present invention is illustrated independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components among each component may be combined to form one component, or one component may be divided into multiple components. An embodiment in which each component is integrated and / or separated is also included in the scope of the present invention as long as it does not deviate from the essence of the present invention.
[0018] The following description may be applied in technical fields dealing with video, image, or video. For example, the method or embodiment disclosed in the following description may be related to the initiation of the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), a next-generation video / image coding standard after VVC, or a standard before VVC (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), etc.).
[0019] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the following, the same components in the drawings will be designated by the same reference numerals, and duplicated explanations of the same components will be omitted.
[0020] In this specification, video refers to a collection of a series of images over time. A picture generally refers to a unit that indicates one image at a specific time, and a slice is a unit that constitutes a part of a picture in coding. One picture can be composed of multiple slices, and pictures and slices can be used interchangeably as necessary.
[0021] A pixel or pel may refer to the smallest unit constituting one picture (or image). Also, a term corresponding to a pixel may be "sample." A sample generally indicates a pixel or a value of a pixel, and may indicate only a value of a pixel / pixel of a luminance (luma) component, or may indicate only a value of a pixel / pixel of a chroma component.
[0022] A unit refers to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information about the corresponding region. A unit may be used in combination with terms such as a block or an area depending on the case. In a general case, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows.
[0023] 1 is a diagram for explaining a schematic configuration of a video encoding apparatus to which the present invention can be applied. Hereinafter, the encoding / decoding apparatus may include a video encoding / decoding apparatus and / or a video encoding / decoding device, and the video encoding / decoding apparatus may be used as a concept including a video encoding / decoding apparatus, or the video encoding / decoding apparatus may be used as a concept including a video encoding / decoding apparatus.
[0024] 1, the video encoding apparatus 100 may include a picture partitioning module 105, a prediction module 110, a residual processing module 120, an entropy encoding module 130, an adder 140, a filtering module 150, and a memory 160. The residual processing module 120 may include a subtractor 121, a transform module 122, a quantization module 123, a rearrangement module 124, a dequantization module 125, and an inverse transform module 126.
[0025] The picture division unit 105 can divide an input picture into at least one processing unit.
[0026] As an example, the processing unit is called a coding unit (CU). In this case, the coding unit may be recursively divided from the largest coding unit (LCU) according to a quad-tree binary-tree (QTBT) structure. For example, one coding unit may be divided into a plurality of coding units of a deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and the ternary tree structure may be applied. Alternatively, the binary tree structure / ternary tree structure may be applied first. The coding procedure according to the present invention may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of a lower depth as necessary, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, conversion, and restoration, which will be described later.
[0027] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit may be split into coding units of deeper depths from a largest coding unit (LCU) according to a quad tree structure. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively split into coding units of lower depths as necessary, and a coding unit of an optimal size may be used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit may not be split into coding units smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that is a basis for partitioning or splitting into prediction units or transform units. The prediction unit is a unit that is partitioned from the coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into subblocks. A transform unit may be divided from a coding unit according to a quad tree structure, and is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, a coding unit is also called a coding block (CB), a prediction unit is also called a prediction block (PB), and a transform unit is also called a transform block (TB). A prediction block or a prediction unit refers to a specific region in a block form within a picture, and may include an array of prediction samples.Also, a transform block or transform unit refers to a specific region in a picture in the form of a block, and may include an array of transform coefficients or residual samples.
[0028] The prediction unit 110 may perform prediction on a current block (hereinafter, may refer to a current block or a residual block) to generate a predicted block including a prediction sample for the current block. The unit of prediction performed by the prediction unit 110 is a coding block, a transform block, or a prediction block.
[0029] The prediction unit 110 may determine whether intra prediction or inter prediction is applied to the current block. For example, the prediction unit 110 may determine whether intra prediction or inter prediction is applied to the current block on a CU basis.
[0030] In the case of intra prediction, the prediction unit 110 may derive a prediction sample for the current block based on a reference sample outside the current block in a picture to which the current block belongs (hereinafter, the current picture). In this case, the prediction unit 110 may (i) derive a prediction sample based on an average or an interpolation of neighboring reference samples of the current block, and (ii) derive the prediction sample based on a reference sample that exists in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. The case (i) is called a non-directional mode or a non-angular mode, and the case (ii) is called a directional mode or an angular mode. Prediction modes in intra prediction may include, for example, 33 directional prediction modes and at least two or more non-directional modes. The non-directional modes may include a DC prediction mode and a planar mode. The prediction unit 110 may also determine a prediction mode to be applied to the current block using a prediction mode applied to a neighboring block.
[0031] In the case of inter prediction, the prediction unit 110 may derive a prediction sample for the current block based on a sample specified by a motion vector on a reference picture. The prediction unit 110 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the case of the skip mode and the merge mode, the prediction unit 110 may use motion information of a neighboring block as motion information of the current block. In the case of the skip mode, unlike the merge mode, a difference (residual) between a prediction sample and an original sample is not transmitted. In the case of the MVP mode, a motion vector of the current block may be derived by using a motion vector of a neighboring block as a motion vector predictor.
[0032] In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in a reference picture. The reference picture including the temporal neighboring blocks is also called a collocated picture (colPic). The motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and output in the form of a bitstream.
[0033] When motion information of a temporally neighboring block is used in skip mode and merge mode, the top picture on a reference picture list may be used as a reference picture. Reference pictures included in a reference picture list may be sorted based on a picture order count (POC) difference between the current picture and the corresponding reference picture. The POC corresponds to the display order of pictures and may be distinguished from the coding order.
[0034] The subtractor 121 generates a residual sample, which is the difference between the original sample and the predicted sample, and does not generate a residual sample when the skip mode is applied, as described above.
[0035] The transform unit 122 transforms the residual samples in units of transform blocks to generate transform coefficients. The transform unit 122 may perform the transform according to the size of the corresponding transform block and a prediction mode applied to a coding block or a prediction block that spatially overlaps with the corresponding transform block. For example, if intra prediction is applied to the coding block or the prediction block that overlaps with the transform block and the transform block is a 4x4 residual array, the residual samples may be transformed using a Discrete Sine Transform (DST) transform kernel, and otherwise, the residual samples may be transformed using a Discrete Cosine Transform (DCT) transform kernel.
[0036] The quantization unit 123 may quantize the transform coefficients to generate quantized transform coefficients.
[0037] The rearrangement unit 124 rearranges the quantized transform coefficients. The rearrangement unit 124 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form through a coefficient scanning method. Here, the rearrangement unit 124 has been described as a separate configuration, but may be a part of the quantization unit 123.
[0038] The entropy encoding unit 130 may perform entropy encoding on the quantized transform coefficients. The entropy encoding may include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 130 may also encode information required for video restoration (e.g., values of syntax elements, etc.) together with or separately from the quantized transform coefficients by entropy encoding or a preset method. The encoded information may be transmitted or stored in a network abstraction layer (NAL) unit in the form of a bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0039] The inverse quantization unit 125 inversely quantizes the values (quantized transformation coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inversely transforms the values inversely quantized by the inverse quantization unit 125 to generate residual samples.
[0040] The adder 140 reconstructs a picture by adding the residual sample and the prediction sample. The residual sample and the prediction sample may be added in block units to generate a reconstructed block. Although the adder 140 has been described as a separate configuration, it may be a part of the prediction unit 110. Meanwhile, the adder 140 may also be called a reconstruction module or a reconstructed block generator.
[0041] The filter unit 150 may apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through the deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortion in the quantization process may be corrected. The sample adaptive offset may be applied on a sample basis and may be applied after the deblocking filtering process is completed. The filter unit 150 may also apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset are applied.
[0042] The memory 160 may store a reconstructed picture (a decoded picture) or information required for encoding / decoding. Here, a reconstructed picture is a reconstructed picture that has been filtered by the filter unit 150. The stored reconstructed picture may be used as a reference picture for (inter) prediction of another picture. For example, the memory 160 may store (reference) pictures used for inter prediction. In this case, the pictures used for inter prediction may be specified by a reference picture set or a reference picture list.
[0043] 2 is a diagram for explaining a schematic configuration of a video / image decoding apparatus to which the present invention can be applied. Hereinafter, the video decoding apparatus may include an image decoding apparatus.
[0044] Referring to FIG. 2, the video decoding apparatus 200 may include an entropy decoding module 210, a residual processing module 220, a prediction module 230, an adder 240, a filtering module 250, and a memory 260. Here, the residual processing module 220 may include a rearrangement module 221, a dequantization module 222, and an inverse transform module 223. Although not shown, the video decoding apparatus 200 may also include a receiver for receiving a bitstream including video information. The receiver may be configured as a separate module or may be included in the entropy decoding module 210.
[0045] When a bitstream including video / image information is input, the video decoding apparatus 200 can restore a video / image / picture corresponding to the process in which the video / image information was processed in the video encoding apparatus.
[0046] For example, the video decoding apparatus 200 may perform video decoding using a processing unit applied in a video encoding apparatus. Thus, a processing unit block of the video decoding may be a coding unit as an example, and may be a coding unit, a prediction unit, or a transform unit as another example. The coding unit may be divided from a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure.
[0047] A prediction unit and a transform unit may also be used in some cases, where the prediction block is a block derived or partitioned from the coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into sub-blocks. The transform unit may be divided from the coding unit by a quadtree structure and is a unit that derives transform coefficients or a unit that derives a residual signal from the transform coefficients.
[0048] The entropy decoding unit 210 may parse the bitstream and output information required for video reconstruction or picture reconstruction. For example, the entropy decoding unit 210 may decode information in the bitstream based on a coding method such as Exponential Golomb Coding, CAVLC, or CABAC, and output values of syntax elements required for video reconstruction and quantized values of transform coefficients for residuals.
[0049] More specifically, the CABAC entropy decoding method receives BINs corresponding to each syntax element in a bitstream, determines a context model using information on the syntax element to be decoded and decoding information on adjacent and blocks to be decoded or information on symbols / BINs decoded in a previous step, predicts the occurrence probability of BINs according to the determined context model, and performs arithmetic decoding of BINs to generate symbols corresponding to the values of each syntax element. In this case, the CABAC entropy decoding method can update the context model using information on the decoded symbols / BINs for the context model of the next symbol / BIN after determining the context model.
[0050] Among the information decoded by the entropy decoding unit 210, information regarding prediction is provided to the prediction unit 230, and the residual values on which entropy decoding is performed by the entropy decoding unit 210, i.e., the quantized transform coefficients, can be input to the reordering unit 221.
[0051] The rearrangement unit 221 may rearrange the quantized transform coefficients in a two-dimensional block form. The rearrangement unit 221 may perform rearrangement in response to coefficient scanning performed in the encoding apparatus. Here, the rearrangement unit 221 may be a part of the inverse quantization unit 222, although it has been described as a separate configuration.
[0052] The inverse quantization unit 222 may inverse quantize the quantized transform coefficients based on the (inverse) quantization parameter to output the transform coefficients. At this time, information for deriving the quantization parameter may be signaled from an encoding device.
[0053] The inverse transform unit 223 can inversely transform the transform coefficients to derive residual samples.
[0054] The prediction unit 230 may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The unit of prediction performed by the prediction unit 230 may be a coding block, a transform block, or a prediction block.
[0055] The prediction unit 230 may determine whether to apply intra prediction or inter prediction based on the information on the prediction. At this time, the unit for determining whether to apply intra prediction or inter prediction is different from the unit for generating a prediction sample. In addition, the unit for generating a prediction sample is also different in inter prediction and intra prediction. For example, whether to apply inter prediction or intra prediction may be determined in units of CU. In addition, for example, in inter prediction, a prediction mode may be determined in units of PU to generate a prediction sample, and in intra prediction, a prediction mode may be determined in units of PU to generate a prediction sample in units of TU.
[0056] In the case of intra prediction, the prediction unit 230 may derive a prediction sample for the current block based on a neighboring reference sample in the current picture. The prediction unit 230 may derive a prediction sample for the current block by applying a directional mode or a non-directional mode based on the neighboring reference sample of the current block. In this case, a prediction mode to be applied to the current block may be determined using the intra prediction mode of the neighboring block.
[0057] In the case of inter prediction, the prediction unit 230 may derive a prediction sample for the current block based on a sample specified on the reference picture by a motion vector on the reference picture. The prediction unit 230 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and an MVP mode. In this case, motion information required for inter prediction of the current block provided by the video encoding apparatus, for example, information on a motion vector, a reference picture index, etc., may be obtained or induced based on information on the prediction.
[0058] In the case of skip mode and merge mode, motion information of neighboring blocks may be used as motion information of the current block, where the neighboring blocks may include spatial neighboring blocks and temporal neighboring blocks.
[0059] The prediction unit 230 may construct a merge candidate list using motion information of available neighboring blocks, and may use information indicated by a merge index on the merge candidate list as a motion vector of the current block. The merge index may be signaled from an encoding device. The motion information may include a motion vector and a reference picture. When motion information of a temporally neighboring block is used in skip mode and merge mode, the top picture on the reference picture list may be used as a reference picture.
[0060] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.
[0061] In the MVP mode, the motion vector of the current block can be derived by using the motion vector of the neighboring block as a motion vector predictor, where the neighboring block can include a spatial neighboring block and a temporal neighboring block.
[0062] For example, when a merge mode is applied, a merge candidate list may be generated using the motion vectors of the reconstructed spatial neighboring blocks and / or the motion vector corresponding to the Col block, which is a temporal neighboring block. In the merge mode, a motion vector of a candidate block selected from the merge candidate list is used as the motion vector of the current block. The prediction information may include a merge index indicating a candidate block having an optimal motion vector selected from among the candidate blocks included in the merge candidate list. In this case, the prediction unit 230 may derive a motion vector of the current block using the merge index.
[0063] As another example, when a Motion Vector Prediction (MVP) mode is applied, a motion vector predictor candidate list may be generated using the motion vector of the restored spatial neighboring block and / or the motion vector corresponding to the Col block, which is a temporal neighboring block. That is, the motion vector of the restored spatial neighboring block and / or the motion vector corresponding to the Col block, which is a temporal neighboring block, may be used as a motion vector candidate. The prediction information may include a predicted motion vector index indicating an optimal motion vector selected from the motion vector candidates included in the list. In this case, the prediction unit 230 may select a predicted motion vector of the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. A prediction unit of the encoding device may obtain a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode the MVD, and output it in the form of a bitstream. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, the prediction unit 230 may obtain a motion vector differential included in the information on the prediction, and derive the motion vector of the current block by adding the motion vector differential and the motion vector predictor. In addition, the prediction unit may obtain or induce a reference picture index indicating a reference picture from the information on the prediction.
[0064] The adder 240 may reconstruct a current block or a current picture by adding a residual sample and a prediction sample. The adder 240 may also reconstruct a current picture by adding a residual sample and a prediction sample in block units. When a skip mode is applied, the residual is not transmitted, so that the prediction sample may become a reconstructed sample. Here, the adder 240 has been described as a separate configuration, but may be a part of the prediction unit 230. Meanwhile, the adder 240 may also be called a reconstruction module or a reconstructed block generator.
[0065] The filter unit 250 may apply deblocking filtering, sample adaptive offset, and / or ALF to the reconstructed picture. In this case, the sample adaptive offset may be applied on a sample basis or may be applied after deblocking filtering. The ALF may be applied after deblocking filtering and / or sample adaptive offset.
[0066] The memory 260 may store a reconstructed picture (a decoded picture) or information required for decoding. Here, the reconstructed picture is a reconstructed picture for which a filtering procedure by the filter unit 250 has been completed. For example, the memory 260 may store a picture used in inter prediction. At this time, the picture used in inter prediction may be specified by a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for another picture. In addition, the memory 260 may output the reconstructed picture according to an output order.
[0067] Meanwhile, as described above, prediction is performed to improve compression efficiency when performing video coding. As a result, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in a spatial domain (or a pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device, and the encoding device signals information (residual information) on the residual between the original block and the predicted block, which is not an original sample value of the original block, to a decoding device, thereby improving video coding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, and add the residual block and the predicted block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.
[0068] The residual information may be generated through a transform and quantization procedure. For example, the encoding apparatus may derive a residual block between the original block and the predicted block, perform a transform procedure on the residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, thereby signaling the related residual information (through a bitstream) to the decoding apparatus. Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding apparatus may derive a residual sample (or a residual block) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding apparatus may generate a reconstructed picture based on the predicted block and the residual block. In addition, the encoding apparatus may derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.
[0069] FIG. 3 is a diagram illustrating an example of a motion represented via an affine motion model according to an embodiment.
[0070] In this specification, "CP" is an abbreviation of control point and may refer to a sample or a reference point that is a reference in the process of applying an affine motion model to a current block. The motion vector of the CP may be referred to as "CPMV (Control Point Motion Vector)", and the CPMV may be derived based on a CPMV predictor, "CPMVP (Control Point Motion Vector Predictor)".
[0071] 3, the motions that can be expressed by the affine motion model according to an embodiment may include translation motion, scale motion, rotation motion, and shear motion. That is, the affine motion model can efficiently express a translation motion in which an image (part of it) moves in a plane over time, a scale motion in which an image (part of it) is scaled over time, a rotation motion in which an image (part of it) rotates over time, and a shear motion in which an image (part of it) is transformed into a balanced quadrilateral over time.
[0072] Affine inter prediction may be performed using an affine motion model according to an embodiment. The encoding device / decoding device may predict the distortion type of an image based on a motion vector at a CP of a current block through affine inter prediction, thereby improving the accuracy of prediction and improving image compression performance. In addition, a motion vector for at least one CP of the current block may be derived using the motion vectors of neighboring blocks of the current block, thereby reducing the burden of data volume for added additional information and improving inter prediction efficiency.
[0073] In one example, affine inter prediction may be performed based on motion information at three CPs, i.e., three reference points, for the current block. The motion information at three CPs for the current block may include a CPMV for each CP.
[0074] FIG. 4 exemplarily shows an affine motion model in which motion vectors for three CPs are used.
[0075] If the position of the top-left sample in the current block is (0,0), the width of the current block is w, and the height of the current block is h, then the samples located at (0,0), (w,0), and (0,h) may be determined as CPs for the current block, as shown in Figure 4. Hereinafter, the CP at the sample position (0,0) may be denoted as CP0, the CP at the sample position (w,0) as CP1, and the CP at the sample position (0,h) as CP2.
[0076] An affine motion model according to an embodiment can be applied using each CP and a motion vector for the corresponding CP. The affine motion model can be expressed as Equation 1 below.
[0077]
number
[0078] Here, w represents the width of the current block, h represents the height of the current block, and v 0x , v 0y are the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively, and v 2x , v 2y indicate the x and y components of the motion vector of CP2, respectively. Also, x indicates the x component of the position of the target sample in the current block, y indicates the y component of the position of the target sample in the current block, and v x is the x component of the motion vector of the target sample in the current block, v y denotes the y-component of the motion vector of the target sample in the current block.
[0079] Meanwhile, Equation 1 showing the affine motion model is merely an example, and an equation for showing the affine motion model is not limited to Equation 1. For example, the sign of each coefficient disclosed in Equation 1 may be different from that of Equation 1 depending on the case, and the magnitude of the absolute value of each coefficient may also be different from that of Equation 1 depending on the case.
[0080] Since the motion vectors of CP0, CP1, and CP2 are known, a motion vector according to the sample position in the current block can be derived based on Equation 1. That is, according to the affine motion model, the motion vector v0 (v 0x ,v 0y ), v1(v 1x ,v 1y ), v2(v 2x ,v 2y ) is scaled, and a motion vector of the target sample according to the position of the target sample may be derived. That is, according to the affine motion model, a motion vector of each sample in the current block may be derived based on the motion vector of the CP. Meanwhile, a set of motion vectors of samples in the current block derived according to the affine motion model may be referred to as an affine motion vector field.
[0081] Meanwhile, the six parameters for Equation 1 can be expressed as a, b, c, d, e, and f as in the following equation, and the equation for the affine motion model expressed by the six parameters is as follows.
[0082]
number
[0083] Here, w represents the width of the current block, h represents the height of the current block, and v 0x , v 0yare the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively, and v 2x , v 2y indicate the x and y components of the motion vector of CP2, respectively. Also, x indicates the x component of the position of the target sample in the current block, y indicates the y component of the position of the target sample in the current block, and v x is the x component of the motion vector of the target sample in the current block, v y denotes the y-component of the motion vector of the target sample in the current block.
[0084] Meanwhile, Equation 2 showing an affine motion model based on six parameters corresponds to only one example, and an equation for showing an affine motion model based on six parameters is not limited to Equation 2. For example, the sign of each coefficient disclosed in Equation 2 may be different from Equation 2 depending on the case, and the magnitude of the absolute value of each coefficient may also be different from Equation 2 depending on the case.
[0085] The affine motion model or the affine inter prediction using the six parameters may be denoted as a six-parameter affine motion model or AF6.
[0086] In one example, affine inter prediction may be performed based on motion information at three CPs, i.e., three reference points, for the current block. The motion information at three CPs for the current block may include a CPMV for each CP.
[0087] In one example, affine inter prediction may be performed based on motion information at two CPs, i.e., two reference points, for the current block. The motion information at the two CPs for the current block may include a CPMV for each CP.
[0088] FIG. 5 exemplarily shows an affine motion model in which motion vectors for two CPs are used.
[0089] An affine motion model using two CPs can express three types of motion including translational motion, scale motion, and rotational motion. An affine motion model expressing three types of motion is sometimes called a similarity affine motion model or a simplified affine motion model.
[0090] If the position of the top-left sample in the current block is (0,0), and the width and height of the current block are w and h, the samples located at (0,0) and (w,0) may be determined as CPs for the current block, as shown in Figure 5. Hereinafter, the CP at the sample position (0,0) may be denoted as CP0, and the CP at the sample position (w,0) may be denoted as CP1.
[0091] Using each CP and the motion vector for the corresponding CP, an affine motion model based on four parameters can be applied. The affine motion model can be expressed as Equation 3 below.
[0092]
number
[0093] Here, w denotes the width of the current block, and v 0x , v 0y are the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively. Also, x indicates the x component of the position of the target sample in the current block, y indicates the y component of the position of the target sample in the current block, and v x is the x component of the motion vector of the target sample in the current block, v y denotes the y-component of the motion vector of the target sample in the current block.
[0094] Meanwhile, Equation 3 showing an affine motion model based on four parameters corresponds to only one example, and an equation for showing an affine motion model based on four parameters is not limited to Equation 3. For example, the sign of each coefficient disclosed in Equation 3 may be different from Equation 3 depending on the case, and the magnitude of the absolute value of each coefficient may also be different from Equation 3 depending on the case.
[0095] Meanwhile, the four parameters for Equation 3 can be expressed as a, b, c, and d as in Equation 4 below, and Equation 4 for the affine motion model represented by the four parameters is as follows.
[0096]
number
[0097] Here, w denotes the width of the current block, and v 0x , v 0y are the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively. Also, x indicates the x component of the position of the target sample in the current block, y indicates the y component of the position of the target sample in the current block, and v x is the x component of the motion vector of the target sample in the current block, v y indicates the y component of the motion vector of the target sample in the current block. Since the affine motion model using the two CPs may be expressed by four parameters a, b, c, and d as in Equation 4, the affine motion model using the four parameters or the affine inter prediction may be referred to as a four-parameter affine motion model or AF4. That is, according to the affine motion model, a motion vector of each sample in the current block may be derived based on the motion vector of the control point. Meanwhile, a set of motion vectors of samples in the current block derived according to the affine motion model may be referred to as an affine motion vector field.
[0098] Meanwhile, Equation 4 showing an affine motion model based on four parameters corresponds to only one example, and an equation for showing an affine motion model based on four parameters is not limited to Equation 4. For example, the sign of each coefficient disclosed in Equation 4 may be different from Equation 4 depending on the case, and the magnitude of the absolute value of each coefficient may also be different from Equation 4 depending on the case.
[0099] Meanwhile, as described above, a motion vector for each sample can be derived through the affine motion model, and thus the accuracy of inter prediction can be significantly improved, but in this case, the complexity of the motion compensation process can be significantly increased.
[0100] In another embodiment, instead of deriving a sample-wise motion vector, a sub-block-wise motion vector within the current block may be restricted to be derived.
[0101] FIG. 6 is a diagram showing an example of deriving a motion vector for each sub-block based on an affine motion model.
[0102] 6 exemplarily illustrates a case where the size of the current block is 16x16 and a motion vector is derived in units of 4x4 sub-blocks. The sub-blocks can be set to various sizes, and for example, when the sub-blocks are set to a size of nxn (n is a positive integer, e.g., n is 4), a motion vector can be derived in units of nxn sub-blocks in the current block based on the affine motion model, and various methods can be applied to derive a motion vector representing each sub-block.
[0103] For example, referring to FIG. 6, a motion vector of each sub-block may be derived using a sample position at the center or the lower right side of the center as a representative coordinate. Here, the lower right position of the center may refer to a sample position at the lower right side of four samples located at the center of the sub-block. For example, when n is an odd number, one sample may be located at the center of the sub-block, and in this case, the center sample position may be used to derive the motion vector of the sub-block. However, when n is an even number, four samples may be located adjacent to the center of the sub-block, and in this case, the lower right sample position may be used to derive the motion vector. For example, referring to FIG. 6, the representative coordinates of each sub-block may be derived as (2,2), (6,2), (10,2), ..., (14,14), and the encoding device / decoding device may derive the motion vector of each sub-block by substituting each of the representative coordinates of the sub-blocks into the above-mentioned Equation 1 or 3. The motion vectors of the sub-blocks within the current block derived through the affine motion model may be referred to as affine MVFs.
[0104] In one embodiment, the above-described affine motion model can be organized into two steps: a step of deriving CPMV and a step of performing affine motion compensation.
[0105] Meanwhile, inter prediction using the above-mentioned affine motion model, that is, affine motion prediction, may include an affine merge mode (AF_MERGE or AAM) and an affine inter mode (AF_INTER or AAMVP).
[0106] The affine merge mode (AAM) according to one embodiment may indicate an encoding / decoding method in which CPMVs for two or three CPs are derived from neighboring blocks of the current block and prediction is performed without coding for MVD (motion vector difference) like the existing skip / merge mode. The affine inter mode (AAMVP) may indicate a method of explicitly encoding / decoding difference information between CPMV and CPMVP like AMVP.
[0107] Meanwhile, the explanation of the affine motion model described above in Figures 3 to 6 is intended to aid in understanding the principles of the encoding / decoding method according to one embodiment of the present invention described later in this specification, and therefore, it should be easily understood by those of ordinary skill in the art that the scope of the present invention is not limited by the contents described above in Figures 3 to 6.
[0108] In one embodiment, a method of constructing an affine MVP candidate list for affine inter prediction is described. In this specification, the affine MVP candidate list is composed of affine MVP candidates, and each affine MVP candidate may mean a combination of CPMVP of CP0 and CP1 in a four-parameter (affine) motion model, and may mean a combination of CPMVP of CP0, CP1, and CP2 in a six-parameter (affine) motion model. The affine MVP candidates described in this specification may be differently referred to by various names such as CPMVP candidate, affine CPMVP candidate, CPMVP pair candidate, CPMVP pair, etc. The affine MVP candidate list may include n affine MVP candidates, and when n is an integer greater than 1, it may be necessary to encode and decode information indicating an optimal affine MVP candidate. When n is 1, it may not be necessary to encode and decode information indicating an optimal affine MVP candidate. An example of the syntax when n is an integer greater than 1 is shown in Table 1 below, and an example of the syntax when n is 1 is shown in Table 2 below.
[0109] [Table 1]
[0110] [Table 2]
[0111] In Tables 1 and 2, merge_flag is a flag for indicating whether or not a merge mode is in effect. When the value of merge_flag is 1, the merge mode may be performed, and when the value of merge_flag is 0, the merge mode may not be performed. affine_flag is a flag for indicating whether or not affine motion prediction is used. When the value of affine_flag is 1, affine motion prediction may be used, and when the value of affine_flag is 0, affine motion prediction may not be used. aamvp_idx is index information for indicating an optimal affine MVP candidate among n affine MVP candidates. In Table 1, which shows the case where n is an integer greater than 1, an optimal affine MVP candidate is indicated based on the aamvp_idx, whereas in Table 2, which shows the case where n is 1, since there is only one affine MVP candidate, it can be confirmed that aamvp_idx is not parsed.
[0112] In one embodiment, when determining an affine MVP candidate, an affine motion model of a neighboring block (hereinafter, also referred to as an "affine coding block") coded based on affine motion prediction may be used. In one example, when determining an affine MVP candidate, a first step and a second step may be performed. In the first step, while scanning the neighboring blocks according to a predefined order, it may be confirmed whether each neighboring block is coded based on affine motion prediction. In the second step, an affine MVP candidate of the current block may be determined using the neighboring blocks coded based on affine motion prediction.
[0113] In the first step, up to m blocks coded based on affine motion prediction may be considered. For example, if m is 1, an affine MVP candidate may be determined using the first affine coding block in the scanning order. For example, if m is 2, at least one affine MVP candidate may be determined using the first and second affine coding blocks in the scanning order. In this case, a pruning check is performed, and if the first affine MVP candidate and the second affine MVP candidate are the same, a further scanning process may be performed to further determine an affine MVP candidate. Meanwhile, in one example, m described in this embodiment may not exceed the value of n described above in the description of Tables 1 and 2.
[0114] Meanwhile, there are various embodiments for the process of scanning the neighboring blocks in the first step and checking whether each neighboring block is coded based on affine motion prediction. Hereinafter, with reference to Figures 7 to 10, an embodiment of the process of scanning the neighboring blocks and checking whether each neighboring block is coded based on affine motion prediction will be described.
[0115] 7 to 10 show examples of methods for detecting neighboring blocks coded based on affine motion prediction.
[0116] Referring to Figure 7, 4x4 blocks A, B, C, D, and E are shown around the current block. Block E, which is the upper left corner peripheral block, is located around CP0, block C, which is the upper right corner peripheral block, and block B, which is the upper peripheral block, are located around CP1, and block D, which is the lower left corner peripheral block, and block A, which is the left peripheral block, are located around CP2. The arrangement according to Figure 7 can share the method and structure according to AMVP or merge mode, which can contribute to reducing design costs.
[0117] Referring to Figure 8, 4x4 blocks A, B, C, D, E, F, and G are shown around the current block. Block E, which is the upper left corner peripheral block, block G, which is the first left peripheral block, and block F, which is the first upper peripheral block, are located around CP0, block C, which is the upper right peripheral block, and block B, which is the second upper peripheral block, are located around CP1, and block D, which is the lower left corner peripheral block, and block A, which is the second left peripheral block, are located around CP2. The arrangement according to Figure 8 can be effective in terms of coding performance while minimizing the increase in scanning complexity, since it is determined whether or not coding is performed based on affine motion prediction based only on 4x4 blocks adjacent to three CPs.
[0118] In FIG. 9, the arrangement of neighboring blocks scanned when detecting neighboring blocks coded based on affine motion prediction is the same as that in FIG. 8. However, in the embodiment according to FIG. 9, affine MVP candidates may be determined based on a maximum of p of 4x4 neighboring blocks included within a dotted line located to the left of the current block and a maximum of q of 4x4 neighboring blocks included within a dotted line located above the current block. For example, when p and q are each 1, affine MVP candidates may be determined based on the first affine coding block in the scan order among the 4x4 neighboring blocks included within a dotted line located to the left of the current block and the first affine coding block in the scan order among the 4x4 neighboring blocks included within a dotted line located above the current block.
[0119] Referring to FIG. 10, an affine MVP candidate can be determined based on the first affine coding block in the scanning order among block E, which is the upper left corner peripheral block located around CP0, block G, which is the first left peripheral block, and block F, which is the first upper peripheral block, the first affine coding block in the scanning order among block C, which is the upper right peripheral block located around CP1, and block B, which is the second upper peripheral block, and the first affine coding block in the scanning order among block D, which is the lower left corner peripheral block located around CP2, and block A, which is the second left peripheral block.
[0120] Meanwhile, the scanning order of the above-mentioned scanning method may be determined based on the probability and performance analysis of a specific encoding device or decoding device. Thus, according to one embodiment, the scanning order is not specified, and the scanning order may be determined based on the statistical characteristics or performance of the encoding device or decoding device to which the embodiment is applied.
[0121] FIG. 11 is a flowchart showing an operation method of an encoding apparatus according to an embodiment, and FIG. 12 is a block diagram showing a configuration of an encoding apparatus according to an embodiment.
[0122] The encoding apparatus according to Figures 11 and 12 may perform operations corresponding to those of the decoding apparatus according to Figures 13 and 14 described below. Therefore, the contents described below with reference to Figures 13 and 14 may be similarly applied to the encoding apparatus according to Figures 11 and 12.
[0123] Each step disclosed in Fig. 11 may be performed by the encoding apparatus 100 disclosed in Fig. 1. More specifically, S1100 to S1140 may be performed by the prediction unit 110 disclosed in Fig. 1, S1150 may be performed by the residual processing unit 120 disclosed in Fig. 1, and S1160 may be performed by the entropy encoding unit 130 disclosed in Fig. 1. Also, the operations of S1100 to S1160 are based on some of the contents described above in Figs. 3 to 10. Therefore, the detailed description that overlaps with the contents described above in Figs. 1 and 3 to 10 will be omitted or simplified.
[0124] As shown in Fig. 12, the encoding apparatus according to an embodiment may include a prediction unit 110 and an entropy encoding unit 130. However, depending on the case, none of the components shown in Fig. 12 may be essential components of the encoding apparatus, and the encoding apparatus may be realized with more or less components than the components shown in Fig. 12.
[0125] In the encoding device according to an embodiment, the prediction unit 110 and the entropy encoding unit 130 may be implemented as separate chips, or at least two of the components may be implemented as one chip.
[0126] An encoding apparatus according to an embodiment may generate an affine MVP candidate list including affine MVP candidates for a current block (S1100). More specifically, a prediction unit 110 of the encoding apparatus may generate an affine MVP candidate list including affine MVP candidates for a current block.
[0127] The encoding apparatus according to one embodiment may derive a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1110). More specifically, the prediction unit 110 of the encoding apparatus may derive a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list.
[0128] The encoding apparatus according to an embodiment may derive a CPMV for each of the CPs of the current block (S1120). More specifically, a prediction unit 110 of the encoding apparatus may derive a CPMV for each of the CPs of the current block.
[0129] The encoding apparatus according to an embodiment may derive a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each of the CPs (S1130). More specifically, the prediction unit 110 of the encoding apparatus may derive a CPMVD for the CP of the current block based on the CPMVP and the CPMV for each of the CPs.
[0130] The encoding apparatus according to an embodiment may derive a prediction sample for the current block based on the CPMV (S1140). More specifically, a prediction unit 110 of the encoding apparatus may derive a prediction sample for the current block based on the CPMV.
[0131] The encoding apparatus according to an embodiment may derive a residual sample for the current block based on the derived predicted sample (S1150). More specifically, the residual processor 120 of the encoding apparatus may derive a residual sample for the current block based on the derived predicted sample.
[0132] The encoding apparatus according to an embodiment may encode information on the derived CPMVD and residual information on the residual samples (S1160). More specifically, the entropy encoding unit 130 of the encoding apparatus may encode information on the derived CPMVD and residual information on the residual samples.
[0133] According to the encoding device and the operating method of the encoding device disclosed in Figures 11 and 12, the encoding device may generate an affine MVP candidate list including affine MVP candidates for a current block (S1100), derive a CPMVP for each CP (Control Point) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1110), derive a CPMV for each of the CPs of the current block based on the CPMVP and the CPMV for each of the CPs (S1130), derive a predicted sample for the current block based on the CPMV (S1140), derive a residual sample for the current block based on the derived predicted sample (S1150), and encode information on the derived CPMVD and residual information on the residual sample (S1160). That is, by signaling information on an affine MVP candidate list used in affine motion prediction, the efficiency of video coding can be improved.
[0134] FIG. 13 is a flowchart showing an operation method of a decoding device according to an embodiment, and FIG. 14 is a block diagram showing a configuration of a decoding device according to an embodiment.
[0135] Each step disclosed in Fig. 13 may be performed by the decoding apparatus 200 disclosed in Fig. 2. More specifically, S1300 may be performed by the entropy decoding unit 210 disclosed in Fig. 2, S1310 to S1350 may be performed by the prediction unit 230 disclosed in Fig. 2, and S1360 may be performed by the addition unit 240 disclosed in Fig. 2. Also, the operations of S1300 to S1360 are based on some of the contents described above in Figs. 3 to 10. Therefore, detailed descriptions that overlap with the contents described above in Figs. 2 to 10 will be omitted or simplified.
[0136] As shown in Fig. 14, the decoding device according to one embodiment may include an entropy decoding unit 210, a prediction unit 230, and an addition unit 240. However, depending on the case, none of the components shown in Fig. 14 may be essential components of the decoding device, and the decoding device may be realized with more or less components than the components shown in Fig. 14.
[0137] In one embodiment of the decoding device, the entropy decoding unit 210, the prediction unit 230, and the addition unit 240 may each be implemented on a separate chip, or at least two or more components may be implemented on a single chip.
[0138] A decoding apparatus according to an embodiment may acquire motion prediction information from a bitstream (S1300). More specifically, an entropy decoding unit 210 of the decoding apparatus may acquire motion prediction information from the bitstream.
[0139] A decoding apparatus according to an embodiment may generate an affine MVP candidate list including affine motion vector predictor (MVP) candidates for a current block (S1310). More specifically, a prediction unit 230 of the decoding apparatus may generate an affine MVP candidate list including affine MVP candidates for a current block.
[0140] In one embodiment, the affine MVP candidates include a first affine MVP candidate and a second affine MVP candidate, the first affine MVP candidate being derived from a left block group including a bottom-left corner neighboring block and a left neighboring block of the current block, and the second affine MVP candidate being derived from a top block group including a top-right corner neighboring block, a top neighboring block, and a top-left corner neighboring block of the current block. In this case, the first affine MVP candidate is derived based on a first block included in the left block group, and the first block is coded based on affine motion prediction, and the second affine MVP candidate is derived based on a second block included in the top block group, and the second block is coded based on affine motion prediction.
[0141] In another embodiment, the affine MVP candidates include a first affine MVP candidate and a second affine MVP candidate, the first affine MVP candidate being derived from a left block group including a neighboring block of a lower left corner of the current block, a first left neighboring block, and a second left neighboring block, and the second affine MVP candidate being derived from an upper block group including a neighboring block of a upper right corner of the current block, a first upper neighboring block, a second upper neighboring block, and a neighboring block of an upper left corner of the current block. In this case, the first affine MVP candidate may be derived based on a first block included in the left block group, the first block being coded based on affine motion prediction, and the second affine MVP candidate may be derived based on a second block included in the upper block group, the second block being coded based on affine motion prediction.
[0142] In yet another embodiment, the affine MVP candidates include a first affine MVP candidate, a second affine MVP candidate, and a third affine MVP candidate, wherein the first affine MVP candidate is derived from a lower left block group including a peripheral block of the lower left corner of the current block and a first left peripheral block, the second affine MVP candidate is derived from a upper right block group including a peripheral block of the upper right corner of the current block and a first upper peripheral block, and the third affine MVP candidate is derived from an upper left block group including a peripheral block of the upper left corner of the current block, a second upper peripheral block, and a second left peripheral block. In this case, the first affine MVP candidate may be derived based on a first block included in the bottom left block group, which is coded based on affine motion prediction, the second affine MVP candidate may be derived based on a second block included in the top right block group, which is coded based on affine motion prediction, and the third affine MVP candidate may be derived based on a third block included in the top left block group, which is coded based on affine motion prediction.
[0143] A decoding apparatus according to an embodiment may derive Control Point Motion Vector Predictors (CPMVP) for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1320). More specifically, a prediction unit 230 of the decoding apparatus may derive a CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list.
[0144] In one embodiment, the one affine MVP candidate may be selected from the affine MVP candidates based on an affine MVP candidate index included in the motion prediction information.
[0145] The decoding apparatus according to an embodiment may derive the CPMVD for the CP of the current block based on information on the CPMVD for each of the CPs included in the obtained motion prediction information (S1330). More specifically, the prediction unit 230 of the decoding apparatus may derive the CPMVD for the CP of the current block based on information on the CPMVD for each of the CPs included in the obtained motion prediction information.
[0146] The decoding apparatus according to an embodiment may derive a CPMV for the CP of the current block based on the CPMVP and the CPMVD (S1340). More specifically, the prediction unit 230 of the decoding apparatus may derive a CPMV for the CP of the current block based on the CPMVP and the CPMVD.
[0147] The decoding apparatus according to an embodiment may derive a prediction sample for the current block based on the CPMV (S1350). More specifically, the prediction unit 230 of the decoding apparatus may derive a prediction sample for the current block based on the CPMV.
[0148] The decoding apparatus according to an embodiment may generate reconstructed samples for the current block based on the derived predicted samples (S1360). More specifically, the adder 240 of the decoding apparatus may generate reconstructed samples for the current block based on the derived predicted samples.
[0149] In one embodiment, the motion prediction information may include information on a context index indicating whether or not there is a neighboring block for the current block coded based on affine motion prediction.
[0150] In one embodiment, a CABAC context model for encoding and decoding index information for indicating an optimal affine MVP candidate may be configured for the case where the value of m in the description of the first step is 1 and the value of n in the description of Tables 1 and 2 is 2. When an affine coding block is present around the current block, an affine MVP candidate for the current block may be determined based on an affine motion model as described above with reference to FIG. 7 to FIG. 10. However, when an affine coding block is not present around the current block, this embodiment may be applied. When an affine MVP candidate is determined based on an affine coding block, the reliability of the affine MVP candidate is high, so that a context model may be designed by classifying a case where an affine MVP candidate is determined based on an affine coding block and a case where it is not determined. In this case, an index 0 may be assigned to an affine MVP candidate determined based on an affine coding block. The CABAC context index according to this embodiment is expressed by Equation 5 below.
[0151]
number
[0152] The initial value according to the CABAC context index may be determined as shown in Table 3 below, and the CABAC context index and the initial value must satisfy the condition of Equation 6 below.
[0153] [Table 3]
[0154]
number
[0155] According to the decoding apparatus and the operating method of the decoding apparatus of FIGS. 13 and 14, the decoding apparatus acquires motion prediction information from a bitstream (S1300), generates an affine MVP candidate list including affine MVP candidates for a current block (S1310), derives Control Point Motion Vector Predictors (CPMVP) for each Control Point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1320), derives the CPMVD for the CP of the current block based on information on Control Point Motion Vector Differences (CPMVD) for each of the CPs included in the acquired motion prediction information (S1330), and calculates a Control Point Motion Vector Difference (CPMVD) for the CP of the current block based on the CPMVP and the CPMVD. The signal processing unit derives a CPMV vector (S1340), derives prediction samples for the current block based on the CPMV (S1350), and generates reconstructed samples for the current block based on the derived prediction samples (S1360). That is, by signaling information on an affine MVP candidate list used in affine motion prediction, video coding efficiency can be improved.
[0156] Meanwhile, the method according to the above-mentioned embodiment of the present specification relates to image and video compression, and can be applied to both encoding and decoding devices, and to devices that generate bitstreams or devices that receive bitstreams, regardless of whether the data is output through a display device in a terminal. For example, an image is generated as compressed data by a terminal having an encoding device, and the compressed data has the form of a bitstream, and the bitstream can be stored in various types of storage devices, streamed through a network, and transmitted to a terminal having a decoding device. If a terminal is equipped with a display device, the decoded image can be displayed on the display device, or the bitstream data can simply be stored.
[0157] The above-described method according to the present invention may be implemented in the form of software, and the encoding device and / or decoding device according to the present invention may be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0158] Each of the above-mentioned parts, modules or units may be a processor or hardware part that executes a sequence of execution steps stored in a memory (or storage unit). Each step described in the above-mentioned embodiments may be performed by a processor or hardware part. Each of the modules / blocks / units described in the above-mentioned embodiments may operate as hardware / processor. Also, the method presented by the present invention may be executed as code. This code may be embodied in a processor-readable storage medium and thus read by a processor provided by an apparatus.
[0159] In the above embodiments, the method is described based on a flowchart as a series of steps or blocks, but the present invention is not limited to the order of steps, and certain steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and different steps may be included, or one or more steps of the flowcharts may be deleted without affecting the scope of the present invention.
[0160] When embodiments of the present invention are implemented as software, the methods described above may be implemented as modules (steps, functions, etc.) performing the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various ways as is well known. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices.
Claims
1. A picture decoding method performed by a decoding device, comprising: obtaining motion prediction information from the bitstream; generating an affine motion vector predictor (MVP) candidate list including affine MVP candidates for the current block; selecting one of the affine MVP candidates in the affine MVP candidate list based on an affine MVP candidate index in the motion prediction information; deriving Control Point Motion Vector Predictors (CPMVPs) for each Control Point (CP) of the current block based on the selected affine MVP candidates; deriving control point motion vector differences (CPMVDs) for each of the CPs of the current block based on information on control point motion vector differences (CPMVDs) for each of the CPs included in the acquired motion prediction information; deriving a control point motion vector (CPMV) for each of the CPs of the current block based on the CPMVP and the CPMVD; deriving a predicted sample for the current block based on the CPMV; generating a reconstructed sample for the current block based on the derived predicted sample; the affine MVP candidates in the affine MVP candidate list include a first affine MVP candidate and a second affine MVP candidate; The CPs include CP0, CP1, and CP2, and the CPMVPs for each of the CPs include a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2; the first MVP, the second MVP, and the third MVP constituting the first affine MVP candidate are derived based on a first block coded based on an affine motion model in a left block group; the left block group includes a neighboring block at a lower left corner of the current block and a left neighboring block adjacent to and above the neighboring block at the lower left corner; the first MVP, the second MVP, and the third MVP constituting the second affine MVP candidate are derived based on a second block in an upper block group coded based on the affine motion model; The upper block group includes a neighboring block at a top right corner of the current block, an upper neighboring block adjacent to a left side of the neighboring block at the top right corner, and a neighboring block at a top left corner.
2. The method of claim 1 , wherein the motion prediction information includes information on a context index indicating whether a neighboring block for the current block coded based on the affine motion model exists.
3. A picture encoding method performed by an encoding device, comprising: generating an affine MVP candidate list including affine MVP candidates for the current block; selecting one of the affine MVP candidates in the affine MVP candidate list; deriving an affine MVP candidate index associated with the selected affine MVP candidate; deriving a CPMVP for each CP (Control Point) of the current block based on the selected affine MVP candidates; deriving a CPMV for each of the CPs of the current block; deriving a CPMVD for each of the CPs of the current block based on the CPMVP and the CPMV for each of the CPs; deriving a predicted sample for the current block based on the CPMV; deriving a residual sample for the current block based on the derived prediction sample; encoding information related to the affine MVP candidate index, information for the derived CPMVD, and residual information related to the residual samples; the affine MVP candidates in the affine MVP candidate list include a first affine MVP candidate and a second affine MVP candidate; The CPs include CP0, CP1, and CP2, and the CPMVPs for each of the CPs include a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2; the first MVP, the second MVP, and the third MVP constituting the first affine MVP candidate are derived based on a first block coded based on an affine motion model in a left block group; the left block group includes a neighboring block at a lower left corner of the current block and a left neighboring block adjacent to and above the neighboring block at the lower left corner; the first MVP, the second MVP, and the third MVP constituting the second affine MVP candidate are derived based on a second block in an upper block group coded based on the affine motion model; The upper block group includes a neighboring block at a top right corner of the current block, an upper neighboring block adjacent to the left of the neighboring block at the top right corner, and a neighboring block at a top left corner.
4. encoding information regarding a context index for the affine MVP candidate index; The picture encoding method of claim 3 , wherein the information on the context index is related to whether or not there are neighboring blocks for the current block coded based on the affine motion model.
5. A method for transmitting data for a picture, comprising: obtaining a bitstream for the picture; transmitting the data including the bitstream; the bitstream is generated based on: generating an affine MVP candidate list including affine MVP candidates for a current block; selecting one of the affine MVP candidates in the affine MVP candidate list; deriving an affine MVP candidate index associated with the selected affine MVP candidate; deriving a CPMVP for each of the control points (CPs) of the current block based on the selected affine MVP candidate; deriving a CPMV for each of the CPs of the current block based on the CPMVP and the CPMV for each of the CPs; deriving a CPMVD for each of the CPs of the current block based on the CPMVP and the CPMV; deriving a prediction sample for the current block based on the CPMV; deriving a residual sample for the current block based on the derived prediction sample; and encoding information associated with the affine MVP candidate index, information for the derived CPMVD, and residual information for the residual sample; the affine MVP candidates in the affine MVP candidate list include a first affine MVP candidate and a second affine MVP candidate; The CPs include CP0, CP1, and CP2, and the CPMVPs for each of the CPs include a first MVP for the CP0, a second MVP for the CP1, and a third MVP for the CP2; the first MVP, the second MVP, and the third MVP constituting the first affine MVP candidate are derived based on a first block coded based on an affine motion model in a left block group; the left block group includes a neighboring block at a lower left corner of the current block and a left neighboring block adjacent to and above the neighboring block at the lower left corner; the first MVP, the second MVP, and the third MVP constituting the second affine MVP candidate are derived based on a second block in an upper block group coded based on the affine motion model; A data transmission method, wherein the upper block group includes a peripheral block at an upper right corner of the current block, an upper peripheral block adjacent to the left of the peripheral block at the upper right corner, and a peripheral block at an upper left corner.
Citation Information
Patent Citations
Affine motion prediction for video coding
US20170332095A1
Image encoding / decoding method and device
WO2017133243A1