Image encoding / decoding apparatus and image data transmitting apparatus

By using an image coding method based on affine motion prediction, an affine MVP candidate list is generated and the control point motion vector is derived, which solves the problem of high-cost transmission and storage of high-resolution image data and achieves more efficient coding and compression.

CN116781929BActive Publication Date: 2026-02-10GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310910010.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-01
Filing Date
2019-04-01
Publication Date
2026-02-10
Estimated Expiration
2039-04-01

AI Technical Summary

Technical Problem

High-resolution and high-quality image data increases the amount of information during transmission and storage, leading to increased transmission and storage costs, and existing technologies struggle to effectively compress and encode them.

Method used

An image coding method based on affine motion prediction is adopted. By generating an affine MVP candidate list, the predicted values ​​and differences of the control point motion vectors are derived, and then the predicted samples are generated and reconstructed, thereby improving coding efficiency.

Benefits of technology

It improves the efficiency of image encoding, reduces the amount of data transmitted and stored, and lowers costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116781929B_ABST
    Figure CN116781929B_ABST
Patent Text Reader

Abstract

The present application relates to an image encoding / decoding apparatus and an image data transmitting apparatus. According to the present application, a picture decoding method implemented by a decoding device includes the steps of: acquiring motion prediction information from a bitstream; generating an affine MVP candidate list including affine MVP candidates for a current block; deriving a CP MVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list; deriving a CP MVD for the CP of the current block based on information about the CP MVD of each CP included in the acquired motion prediction information; and deriving a CPMV for the CP of the current block based on the CP MVP and the CP MVD.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original invention patent application No. 201980029205.3 (International Application No.: PCT / KR2019 / 003816, Application Date: April 1, 2019, Invention Title: Image Coding Method and Apparatus Based on Affine Motion Prediction). Technical Field

[0002] This disclosure generally relates to image coding techniques, and more specifically, to image coding methods and apparatus based on affine motion prediction in image coding systems. Background Technology

[0003] The demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, is increasing across various sectors. Because image data is high-resolution and high-quality, the amount of information or bits that needs to be transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when storing image data using existing storage media.

[0004] Therefore, there is a need for an efficient image compression technology for effectively sending, storing, and reproducing information from high-resolution and high-quality images. Summary of the Invention

[0005] Technical tasks

[0006] One technical objective of this disclosure is to provide a method and apparatus for improving image coding efficiency.

[0007] Another technical objective of this disclosure is to provide a method and apparatus for improving image coding efficiency based on affine motion prediction.

[0008] Another technical objective of this disclosure is to provide a method and apparatus for improving image coding efficiency by effectively determining combinations of adjacent blocks used in affine motion prediction.

[0009] Another technical objective of this disclosure is to provide a method and apparatus for increasing image coding efficiency by signaling information about a list of affine MVP candidates used in affine motion prediction.

[0010] Solution

[0011] According to examples of this disclosure, an image decoding method performed by a decoding device is provided. The method includes obtaining motion prediction information from a bitstream; generating an affine MVP candidate list including affine motion vector prediction value (MVP) candidates for the current block; deriving control point motion vector prediction values ​​(CPMVP) for each control point (CP) of the current block based on one of the affine MVP candidates included in the affine MVP candidate list; deriving the CPMMVD for the current block of the CP based on information about the control point motion vector difference (CPMVD) for each CP included in the obtained motion prediction information; deriving the control point motion vector (CPMV) for the current block of the CP based on the CPMVP and CPMMVD; deriving a prediction sample for the current block based on the CPMV; and generating a reconstructed sample for the current block based on the derived prediction sample.

[0012] According to another example of this disclosure, a decoding apparatus for performing image decoding is provided. The decoding apparatus includes: an entropy decoder that obtains motion prediction information from a bitstream; a predictor that generates an affine MVP candidate list including affine motion vector prediction value (MVP) candidates for the current block, derives a CPMVP for each CP for the current block based on one of the affine MVP candidates included in the affine MVP candidate list, derives a CPMVD for the current block based on information about the CPMVD for each CP included in the obtained motion prediction information, derives a CPMV for the current block based on the CPMVP and CPMVD, and derives a prediction sample for the current block based on the CPMV; and an adder that generates a reconstructed sample for the current block based on the derived prediction sample.

[0013] According to another embodiment of this disclosure, an image encoding method performed by an encoding device is provided. The method includes generating an affine MVP candidate list including affine MVP candidates for a current block; deriving a CPMVP for each CP for the current block based on one of the affine MVP candidates included in the affine MVP candidate list; deriving a CPMV for each CP for the current block; deriving a CPMVD for the CP for the current block based on the CPMV and CPMVP for each CP; deriving a prediction sample for the current block based on the CPMV; deriving a residual sample for the current block based on the derived prediction sample; and encoding information about the derived CPMVD and residual information about the residual sample.

[0014] According to another embodiment of this disclosure, an encoding apparatus for performing image encoding is provided. The encoding apparatus includes: a predictor that generates an affine MVP candidate list including affine MVP candidates for a current block; derives a CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list; derives a CPMV for each CP of the current block; derives a CPMVD for the CP of the current block based on the CPMV and CPMVP for each CP; and derives a prediction sample for the current block based on the CPMV; a residual processor that derives residual samples for the current block based on the derived prediction samples; and an entropy encoder that encodes information about the derived CPMVD and residual information about the residual samples.

[0015] Beneficial effects

[0016] According to this disclosure, the overall image / video compression efficiency can be improved.

[0017] According to this disclosure, the efficiency of image coding can be improved based on affine motion prediction.

[0018] According to this disclosure, image coding efficiency can be improved by signaling information about the affine MVP candidate list for affine motion prediction. Attached Figure Description

[0019] Figure 1 This is a diagram schematically illustrating the configuration of an encoding device according to an embodiment.

[0020] Figure 2 This is a diagram schematically illustrating the configuration of a decoding device according to an embodiment.

[0021] Figure 3 This is a diagram illustrating an example of motion expressed by an affine motion model according to an implementation method.

[0022] Figure 4 This is a diagram illustrating an example of an affine motion model using control point motion vectors (CPMV) for the three control points (CP) of the current block.

[0023] Figure 5 This is a diagram illustrating an example of an affine motion model using CPMV for the two CPs of the current block.

[0024] Figure 6 This is a diagram illustrating an example of deriving motion vectors in a sub-block unit based on an affine motion model.

[0025] Figures 7 to 10 An example of a method for detecting neighboring blocks encoded based on affine motion prediction is shown.

[0026] Figure 11 This is a flowchart illustrating the operation method of an encoding device according to an embodiment.

[0027] Figure 12 This is a block diagram illustrating the configuration of an encoding device according to an embodiment.

[0028] Figure 13 This is a flowchart illustrating the operation method of a decoding device according to an embodiment.

[0029] Figure 14 This is a block diagram illustrating the configuration of a decoding device according to an embodiment. Detailed Implementation

[0030] According to embodiments of this disclosure, an image decoding method performed by a decoding device is presented. The method includes: obtaining motion prediction information from a bitstream; generating an affine MVP candidate list including affine motion vector prediction value (MVP) candidates for a current block; deriving control point motion vector prediction values ​​(CPMVP) for each control point (CP) for the current block based on one of the affine MVP candidates included in the affine MVP candidate list; deriving a CPMMVD for the current block's CP based on information regarding the control point motion vector difference (CPMVD) for each CP included in the obtained motion prediction information; deriving control point motion vectors (CPMV) for the current block's CP based on the CPMVP and CPMMVD; deriving a prediction sample for the current block based on the CPMV; and generating a reconstructed sample for the current block based on the derived prediction sample.

[0031] Implementation of this disclosure

[0032] This disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit this disclosure. The terminology used in the following description is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions, provided that they are clearly read differently. Terms such as “comprising” and “having” are intended to indicate the presence of the features, numbers, steps, operations, elements, components or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, numbers, steps, operations, elements, components or combinations thereof is not excluded.

[0033] Furthermore, the elements in the accompanying drawings described in this embodiment are drawn independently for ease of explanation of different specific functions, but this does not mean that these elements are implemented by independent hardware or independent software. For example, two or more elements may be combined to form a single element, or a single element may be divided into multiple elements. Embodiments in which elements are combined and / or divided are part of this disclosure without departing from the concept of this embodiment.

[0034] The following description can be applied to the technical field of processing video, images, or pictures. For example, the methods or exemplary implementations disclosed in the following description can be associated with the disclosures of the Universal Video Coding (VVC) standard (ITU-T H.266 Recommendation), next-generation video / image coding standards after VVC, or standards before VVC (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T H.265 Recommendation), etc.).

[0035] Hereinafter, examples of this embodiment will be described in detail with reference to the accompanying drawings. Furthermore, throughout the drawings, similar reference numerals are used to indicate similar elements, and identical descriptions of similar elements will be omitted.

[0036] In this disclosure, video can refer to a collection of images over time. Generally, picture refers to a unit of image representing a specific time, and slice is a unit that constitutes a part of picture. A picture can be composed of multiple slices, and the terms picture and slice can be mixed together as needed.

[0037] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value; it can represent a pixel containing only the luminance component (pixel value) or a pixel containing only the chrominance component (pixel value).

[0038] A unit refers to a basic unit of image processing. A unit may include at least one of a specific region and information associated with that region. Optionally, a unit may be combined with terms such as block, region, etc. Typically, an M×N block may represent a set of samples or transform coefficients arranged in M ​​columns and N rows.

[0039] Figure 1 The structure of the encoding apparatus to which this disclosure applies is briefly illustrated. In the following, encoding / decoding apparatus may include video encoding / decoding apparatus and / or image encoding / decoding apparatus, and video encoding / decoding apparatus may be used as a concept that includes image encoding / decoding apparatus, or image encoding / decoding apparatus may be used as a concept that includes video encoding / decoding apparatus.

[0040] Reference Figure 1 The video encoding device 100 may include an image segmenter 105, a predictor 110, a residual processor 120, an entropy encoder 130, an adder 140, a filter 150, and a memory 160. The residual processor 120 may include a subtractor 121, a transformer 122, a quantizer 123, a rearranger 124, an inverse quantizer 125, and an inverse transformer 126.

[0041] Image segmenter 105 can separate an input image into at least one processing unit.

[0042] In the example, the processing unit can be referred to as a coding unit (CU). In this case, coding units can be recursively separated from the maximum coding unit (LCU) according to a quadtree-binary tree (QTBT) structure. For example, a coding unit can be separated into multiple coding units of deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and a ternary tree structure can be applied later. Alternatively, a binary tree structure / ternary tree structure can be applied first. The coding process according to this embodiment can be performed based on the final coding unit that is no longer further separated. In this case, the maximum coding unit can be used as the final coding unit based on image characteristics such as coding efficiency, or the coding unit can be recursively separated into lower-depth coding units as needed, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, and reconstruction, which will be described later.

[0043] In another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transformer (TU). The coding unit can be separated from the maximum coding unit (LCU) into deeper coding units according to a quadtree structure. In this case, the maximum coding unit can be directly used as the final coding unit based on image characteristics such as coding efficiency, or the coding unit can be recursively separated into deeper coding units as needed, and the coding unit with the optimal size can be used as the final coding unit. When a minimum coding unit (SCU) is set, the coding unit may not be separated into coding units smaller than the minimum coding unit. Here, the final coding unit refers to the coding unit that has been segmented or separated into prediction units or transformers. The prediction unit is a unit segmented from the coding unit and can be a unit for sample prediction. Here, the prediction unit can be divided into sub-blocks. The transformer can be divided from the coding unit according to a quadtree structure, and the transformer can be a unit for deriving the transform coefficients and / or a unit for deriving the residual signal from the transform coefficients. In the following text, the coding unit may be referred to as a coding block (CB), the prediction unit may be referred to as a prediction block (PB), and the transformer may be referred to as a transform block (TB). A prediction block or prediction unit can refer to a specific region in the form of a block in an image and include an array of prediction samples. Similarly, a transform block or transformer can refer to a specific region in the form of a block in an image and include an array of transform coefficients or residual samples.

[0044] Predictor 110 can perform predictions on a target block (hereinafter, it can represent the current block or a residual block) and can generate a prediction block that includes prediction samples for the current block. The unit of prediction performed in predictor 110 can be a coded block, a transform block, or a prediction block.

[0045] Predictor 110 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block. For example, predictor 110 can determine whether to apply intra-frame prediction or inter-frame prediction on a CU-by-CU basis.

[0046] In the case of intra-frame prediction, predictor 110 can derive the prediction sample for the current block based on reference samples outside the current block in the image to which the current block belongs (hereinafter, the current image). In this case, predictor 110 can derive the prediction sample based on the average or interpolation of the neighboring reference samples of the current block (case (i)), or it can derive the prediction sample based on reference samples among the neighboring reference samples of the current block that exist in a specific (prediction) direction relative to the prediction sample (case (ii)). Case (i) can be referred to as a non-directional mode or a non-angular mode, and case (ii) can be referred to as a directional mode or an angular mode. In intra-frame prediction, the prediction modes can include 33 directional modes and at least two non-directional modes, as an example. Non-directional modes can include DC mode and planar mode. Predictor 110 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.

[0047] In the case of inter-frame prediction, predictor 110 can derive predicted samples for the current block based on samples specified by motion vectors on a reference image. Predictor 110 can derive predicted samples for the current block by applying any of the skip mode, merge mode, and motion vector prediction (MVP) mode. In the skip mode and merge mode, predictor 110 can use motion information from neighboring blocks as motion information for the current block. In the skip mode, unlike the merge mode, the difference (residual) between the predicted sample and the original sample is not sent. In the MVP mode, the motion vectors of neighboring blocks are used as motion vector predictors to derive the motion vector of the current block.

[0048] In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image that includes temporally neighboring blocks can also be called a colpic. Motion information can include motion vectors and reference image indices. Information such as prediction mode information and motion information can be (entropy-encoded) and then output as a bitstream.

[0049] When using motion information from temporally adjacent blocks in skip and merge modes, the highest-ranking image in the reference image list can be used as the reference image. Reference images included in the reference image list can be aligned based on the difference in Picture Order Number (POC) between the current image and the corresponding reference image. The POC corresponds to the display order and can be distinguished from the encoding order.

[0050] Subtractor 121 generates a residual sample, which is the difference between the original sample and the predicted sample. If the skip mode is applied, residual samples may not be generated as described above.

[0051] Transformer 122 transforms residual samples on a block-by-block basis to generate transform coefficients. Transformer 122 can perform the transform based on the size of the corresponding transform block and the prediction mode applied to the prediction block or coding block that spatially overlaps with the transform block. For example, if intra-frame prediction is applied to the prediction block or coding block that overlaps with the transform block and the transform block is a 4×4 residual array, a Discrete Sine Transform (DST) kernel can be used to transform the residual samples, and in other cases, a Discrete Cosine Transform (DCT) kernel is used to transform the residual samples.

[0052] Quantizer 123 can quantize the transform coefficients to generate quantized transform coefficients.

[0053] Rearranger 124 rearranges the quantized transform coefficients. Rearranger 124 can rearrange the quantized transform coefficients in block form into a one-dimensional vector using a coefficient sweep method. Although rearranger 124 is described as a separate component, it can be part of quantizer 123.

[0054] The entropy encoder 130 can perform entropy coding on quantized transform coefficients. Entropy coding can include coding methods such as Exponential Columbus, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. In addition to the quantized transform coefficients, the entropy encoder 130 can also encode information required for video reconstruction (e.g., syntax element values, etc.) either together or separately according to entropy coding or according to a pre-configured method. The entropy-coded information can be transmitted or stored in the form of a bitstream at the Network Abstraction Layer (NAL). The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SDD, etc.

[0055] Dequantizer 125 dequantizes the values ​​(transformation coefficients) quantized by quantizer 123, and inverse transformer 126 performs an inverse transformation on the values ​​dequantized by dequantizer 125 to generate residual samples.

[0056] Adder 140 adds residual samples to the prediction samples to reconstruct the image. Residual samples can be added to the prediction samples in blocks to generate reconstructed blocks. Although adder 140 is described as a separate component, adder 140 can be part of predictor 110. Additionally, adder 140 can be referred to as a reconstructor or a reconstructed block generator.

[0057] Filter 150 can apply deblocking filtering and / or adaptive sample shifting to the reconstructed image. Deblocking filtering and / or adaptive sample shifting can correct artifacts at block boundaries or distortions during quantization in the reconstructed image. After deblocking filtering is complete, adaptive sample shifting can be applied on a sample-by-sample basis. Filter 150 can also apply an adaptive loop filter (ALF) to the reconstructed image. An ALF can be applied to a reconstructed image that has already undergone deblocking filtering and / or adaptive sample shifting.

[0058] Memory 160 can store reconstructed images (decoded images) or information required for encoding / decoding. Here, the reconstructed image can be a reconstructed image filtered by filter 150. The stored reconstructed image can be used as a reference image for (inter-frame) prediction of other images. For example, memory 160 can store (reference) images for inter-frame prediction. Here, the images used for inter-frame prediction can be specified according to a set of reference images or a list of reference images.

[0059] Figure 2 The structure of a video / image decoding apparatus to which this disclosure applies is briefly illustrated. Hereinafter, a video decoding apparatus may include an image decoding apparatus.

[0060] Reference Figure 2 The video decoding device 200 may include an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The residual processor 220 may include a rearranger 221, an inverse quantizer 222, and an inverse transformer 223. Additionally, although not depicted, the video decoding device 200 may include a receiver for receiving a bitstream including video information. The receiver may be configured as a separate module or may be included within the entropy decoder 210.

[0061] When the input includes a bitstream containing video / image information, the video decoding device 200 can reconstruct the video / image / picture in association with the process of processing video information in the video encoding device.

[0062] For example, video decoding device 200 can use processing units applied in video encoding devices to perform video decoding. Therefore, the processing unit block for video decoding can be, for example, an encoding unit, and in another example, an encoding unit, a prediction unit, or a transformer. Encoding units can be separated from the maximum encoding unit according to a quadtree structure and / or a binary tree structure and / or a ternary tree structure.

[0063] In some cases, prediction units and transformers can be further used, and in this case, the prediction block is a block derived or segmented from the coding unit, and can be a unit for sample prediction. Here, the prediction unit can be divided into sub-blocks. The transformer can be separated from the coding unit according to a quadtree structure, and can be a unit for deriving the transform coefficients or a unit for deriving the residual signal from the transform coefficients.

[0064] The entropy decoder 210 can parse a bitstream to output the information needed for video reconstruction or image reconstruction. For example, the entropy decoder 210 can decode information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values ​​of the syntax elements needed for video reconstruction and the quantization values ​​of the transform coefficients with respect to the residuals.

[0065] More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of adjacent blocks and the decoding target block, or information of symbols / bins decoded in previous steps, to determine a context model, predict the bin generation probability based on the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model after determining it using information of symbols / bins decoded for the context model of the next symbol / bin.

[0066] Information about the prediction from the information decoded in the entropy decoder 210 can be provided to the predictor 230, and the residual values ​​(i.e., the quantized transform coefficients) that have been entropy decoded by the entropy decoder 210 can be input to the rearranger 221.

[0067] Rearranger 221 can rearrange the quantized transform coefficients into a two-dimensional block form. Rearranger 221 can perform a rearrangement corresponding to the coefficient scan performed by the encoding device. Although rearranger 221 is described as a separate component, rearranger 221 can be part of dequantizer 222.

[0068] The dequantizer 222 can dequantize the quantized transform coefficients based on the (de)quantization parameters to output the transform coefficients. In this case, the encoding device can signal information used to derive the quantization parameters.

[0069] The inverse transformer 223 can perform an inverse transformation on the transformation coefficients to derive the residual samples.

[0070] Predictor 230 can perform prediction on the current block and generate a prediction block that includes prediction samples for the current block. The unit of prediction performed in predictor 230 can be a coded block, a transform block, or a prediction block.

[0071] Predictor 230 can determine whether to apply intra-frame prediction or inter-frame prediction based on information about the prediction. In this case, the unit used to determine which one to use between intra-frame and inter-frame prediction can be different from the unit used to generate prediction samples. Furthermore, the unit used to generate prediction samples can also be different in inter-frame and intra-frame prediction. For example, it can be determined on a CU (unit of measurement) basis to determine which one to apply between inter-frame and intra-frame prediction. Additionally, for example, in inter-frame prediction, prediction samples can be generated by determining the prediction mode on a PU (unit of measurement), and in intra-frame prediction, prediction samples can be generated on a TU (unit of measurement) basis by determining the prediction mode on a PU basis.

[0072] In the case of intra-frame prediction, predictor 230 can derive prediction samples for the current block based on neighboring reference samples in the current image. Predictor 230 can derive prediction samples for the current block by applying either a directional or non-directional mode based on neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block can be determined by using the intra-frame prediction modes of neighboring blocks.

[0073] In the case of inter-frame prediction, predictor 230 can derive prediction samples for the current block based on samples specified in the reference image according to motion vectors. Predictor 230 can use one of skip mode, merge mode, and MVP mode to derive prediction samples for the current block. Here, the motion information (e.g., motion vectors and information about the reference image index) required for inter-frame prediction of the current block provided by the video encoding device can be obtained or derived based on information about the prediction.

[0074] In skip and merge modes, motion information from adjacent blocks can be used as motion information for the current block. Here, adjacent blocks can include spatially adjacent blocks and temporally adjacent blocks.

[0075] Predictor 230 can construct a merge candidate list using motion information from available neighboring blocks, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled by the encoding device. The motion information can include motion vectors and reference images. In skip mode and merge mode, when using motion information from temporally adjacent blocks, the topmost image in the reference image list can be used as the reference image.

[0076] In the skip mode, unlike the merge mode, the difference (residual) between the predicted sample and the original sample is not sent.

[0077] In MVP mode, the motion vectors of neighboring blocks can be used as motion vector prediction values ​​to derive the motion vector of the current block. Here, neighboring blocks can include spatially adjacent blocks and temporally adjacent blocks.

[0078] When applying a merge pattern, for example, a merge candidate list can be generated using the motion vectors of reconstructed spatially adjacent blocks and / or the motion vectors corresponding to Col blocks that are temporally adjacent blocks. The motion vectors of candidate blocks selected from the merge candidate list are used as the motion vectors of the current block in the merge pattern. The aforementioned information about the prediction may include a merge index that indicates the candidate block with the best motion vector selected from the candidate blocks included in the merge candidate list. Here, predictor 230 can use the merge index to derive the motion vector of the current block.

[0079] When applying the MVP (Motion Vector Prediction) mode as another example, a candidate list of motion vector prediction values ​​can be generated using the motion vectors of reconstructed spatially adjacent blocks and / or the motion vectors corresponding to Col blocks that are temporally adjacent blocks. That is, the motion vectors of reconstructed spatially adjacent blocks and / or the motion vectors corresponding to Col blocks that are temporally adjacent blocks can be used as motion vector candidates. The aforementioned information about prediction can include a predicted motion vector index indicating the best motion vector to be selected from the motion vector candidates included in the list. Here, predictor 230 can use the motion vector index to select the predicted motion vector of the current block from the motion vector candidates included in the motion vector candidate list. The predictor of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector prediction value, encode the MVD, and output the encoded MVD as a bitstream. That is, the MVD can be obtained by subtracting the motion vector prediction value from the motion vector of the current block. Here, predictor 230 can obtain the motion vectors included in the information about prediction and derive the motion vector of the current block by adding the motion vector difference to the motion vector prediction value. Additionally, the predictor can obtain or derive a reference image index indicating a reference image from the aforementioned information about prediction.

[0080] Adder 240 can add residual samples to the prediction samples to reconstruct the current block or the current image. Adder 240 can reconstruct the current image by adding residual samples to the prediction samples on a block-by-block basis. When a skip mode is applied, no residuals are sent, and therefore the prediction samples can become the reconstructed samples. Although adder 240 is described as a separate component, adder 240 can be part of predictor 230. Additionally, adder 240 can be referred to as a reconstructor or a reconstructed block generator.

[0081] Filter 250 can apply deblocking filtering, adaptive sample shifting, and / or ALF to the reconstructed image. Here, adaptive sample shifting can be applied on a sample-by-sample basis after deblocking filtering. ALF can be applied after deblocking filtering and / or after applying adaptive sample shifting.

[0082] Memory 260 can store reconstructed images (decoded images) or information required for decoding. Here, the reconstructed image can be a reconstructed image filtered by filter 250. For example, memory 260 can store images used for inter-frame prediction. Here, the images used for inter-frame prediction can be specified according to a set of reference images or a list of reference images. The reconstructed image can be used as a reference image for other images. Memory 260 can output the reconstructed images in the output order.

[0083] Furthermore, as mentioned above, prediction is performed during video encoding to improve compression efficiency. Therefore, a prediction block can be generated that includes prediction samples for the current block, which is the block to be encoded (i.e., the target block for encoding). Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same manner in both the encoding and decoding devices, and the encoding device can signal information about the residual between the original block and the prediction block (residual information) rather than the original sample values ​​of the original block to the decoding device, thereby improving image encoding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block to the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed image including the reconstructed block.

[0084] Residual information can be generated through transformation and quantization processes. For example, the encoding device can derive a residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and can signal the relevant residual information to the decoding device (via bitstream). Here, the residual information can include the value information, position information, transform technique, transform kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform inverse quantization / inverse transform processes based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. In addition, for inter-frame prediction of the reference image later, the encoding device can also perform inverse quantization / inverse transform on the quantized transform coefficients to derive residual blocks and generate a reconstructed image based on the residual blocks.

[0085] Figure 3 This is a diagram illustrating an example of motion expressed by an affine motion model according to an implementation method.

[0086] In this specification, "CP," an abbreviation for control point, can refer to a reference point or sample used as a reference when applying an affine motion model to the current block. The motion vector of the CP can be called the "Control Point Motion Vector (CPMV)," and the CPMV can be derived based on the "Control Point Motion Vector Predicted Value (CPMVP)" which serves as the CPMV predictor.

[0087] Reference Figure 3 The motion that can be expressed by the affine motion model according to the implementation method can include translational motion, scaling motion, rotational motion, and shearing motion. That is, the affine motion model can effectively express the translational motion of an image (or a part thereof) moving in a plane over time, the scaling motion of an image (or a part thereof) scaling over time, the rotational motion of an image (or a part thereof) rotating over time, and the shearing motion of an image (or a part thereof) deforming into a parallelogram over time.

[0088] Affine inter-frame prediction can be performed using an affine motion model according to the implementation. The encoding / decoding device can predict the distortion shape of the image based on the motion vector at the CP of the current block through affine inter-frame prediction, which can lead to increased prediction accuracy and thus improve image compression performance. Furthermore, the motion vector of at least one CP of the current block can be derived using the motion vectors of neighboring blocks, thus reducing the amount of additional information required and improving inter-frame prediction efficiency.

[0089] In one example, affine inter-frame prediction can be performed based on motion information at three CPs (i.e., three reference points) for the current block. The motion information at the three CPs for the current block can include the CPMV for each CP.

[0090] Figure 4 An affine motion model is schematically shown in which three CP motion vectors are used.

[0091] When the position of the top-left sample within the current block is (0, 0), the width of the current block is 'w', and its height is 'h', as follows: Figure 4 As shown, the samples located at (0, 0), (w, 0), and (0, h) can be identified as CPs for the current block. In the following text, the CP at sample position (0, 0) can be denoted as CP0, the CP at sample position (w, 0) can be denoted as CP1, and the CP at sample position (0, h) can be denoted as CP2.

[0092] The affine motion model according to the implementation method can be applied using the motion vectors of the aforementioned CPs and their respective CPs. The affine motion model can be expressed as Equation 1 below.

[0093] [Formula 1]

[0094]

[0095] Here, w represents the width of the current block, h represents the height of the current block, and v 0x and v 0y Let v represent the x-component and y-component of the motion vector of CP0, respectively. 1x and v 1y Let v represent the x-component and y-component of the motion vector of CP1, respectively. 2x and v 2y These represent the x and y components of the motion vector of CP2, respectively. Additionally, x represents the x-component of the target sample's position within the current block, y represents the y-component of the target sample's position within the current block, and v... x Let v represent the x-component of the motion vector of the target sample within the current block, and v y This represents the y-component of the motion vector of the target sample within the current block.

[0096] Furthermore, Equation 1, which represents the affine motion model, is merely an example, and the formulas used to represent affine motion models are not limited to Equation 1. For example, in some cases, the sign of each coefficient disclosed in Equation 1 may be changed from the sign of Equation 1, and in some cases, the magnitude of the absolute value of each coefficient may also be changed from the magnitude of the absolute value of Equation 1.

[0097] Since the motion vectors of CP0, CP1, and CP2 are known, the motion vector based on the sample position within the current block can be derived using Equation 1 above. That is, according to the affine motion model, the motion vector at CP is v0(v 0x v 0y v1(v) 1x v 1y v2(v) 2x v 2y The model can be scaled based on the distance ratio between the target sample's coordinates (x, y) and the three CPs, allowing the derivation of the target sample's motion vector based on its position. In other words, based on the affine motion model, the motion vector of each sample within the current block can be derived from the motion vectors of the CPs. Furthermore, the set of motion vectors derived from the affine motion model for the samples within the current block can be called the affine motion vector field.

[0098] Furthermore, the six parameters in Equation 1 above can be expressed as a, b, c, d, e, and f in the following equations, and the equations for the affine motion model expressed using these six parameters can be as follows:

[0099] [Equation 2]

[0100]

[0101]

[0102]

[0103] Where w represents the width of the current block, h represents the height of the current block, and v 0x and v 0y Let v represent the x-component and y-component of the motion vector of CP0, respectively. 1x and v 1y Let v represent the x-component and y-component of the motion vector of CP1, respectively. 2x and v 2y These represent the x and y components of the motion vector of CP2, respectively. Additionally, x represents the x-component of the target sample's position within the current block, y represents the y-component of the target sample's position within the current block, and v... x The x-component of the motion vector of the target sample within the current block, v y This represents the y-component of the motion vector of the target sample within the current block.

[0104] Furthermore, Equation 2, which represents an affine motion model based on six parameters, is merely an example, and the formulas used to represent affine motion models based on six parameters are not limited to Equation 2. For example, in some cases, the sign of each coefficient disclosed in Equation 2 may be changed from the sign of Equation 2, and in some cases, the magnitude of the absolute value of each coefficient may also be changed from the magnitude of the absolute value of Equation 2.

[0105] Affine inter-frame prediction or affine motion model using six parameters can be called a six-parameter affine motion model or AF6.

[0106] In one example, affine inter-frame prediction can be performed based on motion information at three CPs (i.e., three reference points) for the current block. The motion information at the three CPs of the current block can include the CPMV for each CP.

[0107] In one example, affine inter-frame prediction can be performed based on motion information at two CPs (i.e., two reference points) for the current block. The motion information at the two CPs of the current block can include the CPMV of each CP.

[0108] Figure 5 An affine motion model is schematically shown in which two CP motion vectors are used.

[0109] An affine motion model using two CPs can represent three motions: translation, scaling, and rotation. An affine motion model representing these three motions can be called a similar affine motion model or a simplified affine motion model.

[0110] When the position of the top-left sample within the current block is (0, 0), the width of the current block is 'w', and its height is 'h', as follows: Figure 5 As shown, the samples located at (0, 0) and (w, 0) can be identified as the CP of the current block. In the following text, the CP at sample position (0, 0) can be denoted as CP0, and the CP at sample position (w, 0) can be denoted as CP1.

[0111] The affine motion model based on four parameters can be applied using the aforementioned CPs and their corresponding motion vectors. The affine motion model can be represented by Equation 3 below.

[0112] [Formula 3]

[0113]

[0114] Here, w represents the width of the current block, and v 0x and v 0y Let x and y represent the motion vector of CP0, respectively, and v 1x and v 1yLet x and y represent the x and y components of the motion vector of CP1, respectively. Additionally, x represents the x-component of the target sample's position within the current block, and y represents the y-component of the target sample's position within the current block. x The x-component of the motion vector of the target sample within the current block, v y This represents the y-component of the motion vector of the target sample within the current block.

[0115] Furthermore, Equation 3, which represents an affine motion model based on four parameters, is merely an example, and the formulas used to represent affine motion models based on four parameters are not limited to Equation 3. For example, in some cases, the sign of each coefficient disclosed in Equation 3 may be changed from the sign of Equation 3, and in some cases, the magnitude of the absolute value of each coefficient may also be changed from the magnitude of the absolute value of Equation 3.

[0116] Furthermore, the four parameters of Equation 3 above can be expressed as a, b, c, and d in Equation 4 below, and Equation 4 of the affine motion model using the four parameters can be expressed as follows:

[0117] [Formula 4]

[0118] c = v 0x d = v oy

[0119]

[0120] Here, w represents the width of the current block, and v 0x and v 0y Let x and y represent the motion vector of CP0, respectively, and v 1x and v 1y Let x and y represent the x and y components of the motion vector of CP1, respectively. Additionally, x represents the x-component of the target sample's position within the current block, and y represents the y-component of the target sample's position within the current block. x The x-component of the motion vector of the target sample within the current block, v y This represents the y-component of the motion vector of the target sample within the current block. Since the affine motion model using two CPs can be represented by four parameters a, b, c, and d as shown in Equation 4, the affine inter-frame prediction or affine motion model using four parameters can be called a four-parameter affine motion model or AF4. That is, based on the affine motion model, the motion vector of each sample within the current block can be derived from the motion vectors of the control points. Furthermore, the set of motion vectors of the samples within the current block derived from the affine motion model can be called the affine motion vector field.

[0121] Furthermore, Equation 4, which represents an affine motion model based on four parameters, is merely an example, and the formulas used to represent affine motion models based on four parameters are not limited to Equation 4. For example, in some cases, the sign of each coefficient disclosed in Equation 4 may be changed from the sign of Equation 4, and in some cases, the magnitude of the absolute value of each coefficient may also be changed from the magnitude of the absolute value of Equation 4.

[0122] Furthermore, as mentioned above, the motion vector of the sample unit can be derived through an affine motion model, which can significantly improve the accuracy of inter-frame prediction. However, in this case, the complexity of motion compensation processing may increase considerably.

[0123] In another implementation, the focus can be on deriving the motion vectors of sub-block units within the current block rather than the motion vectors of sample units.

[0124] Figure 6 This is a diagram illustrating an example of deriving motion vectors in a sub-block unit based on an affine motion model.

[0125] Figure 6 The example illustrates a case where the current block size is 16×16 and the motion vectors are derived in 4×4 sub-block cells. Sub-blocks can be set to various sizes, and for example, if the sub-blocks are set to an n×n size (n is a positive integer, and for example, n is 4), motion vectors can be derived in n×n sub-block cells within the current block based on an affine motion model, and various methods for deriving the motion vectors representing each sub-block can be applied.

[0126] For example, refer to Figure 6 The motion vector for each sub-block can be derived by setting the center or lower-right sample position of each sub-block as the representative coordinate. Here, the lower-right position can represent the lower-right sample position among the four samples located at the center of the sub-block. For example, if n is odd, a sample can be located at the center of the sub-block, and in this case, the center sample position can be used to derive the motion vector of the sub-block. However, if n is even, four samples can be located near the center of the sub-block, and in this case, the lower-right sample position can be used to derive the motion vector. For example, refer to... Figure 6 The representative coordinates of each sub-block can be derived as (2, 2), (6, 2), (10, 2), ..., (14, 14), and the encoding / decoding device can derive the motion vector of each sub-block by inputting each of the representative coordinates of the sub-block into Equations 1 to 3 above. The motion vector of the sub-blocks within the current block derived through the affine motion model can be called the affine MVF.

[0127] In one embodiment, when the above-described affine motion model is summarized into two steps, it may include the step of deriving CPMV and the step of performing affine motion compensation.

[0128] In addition, in the inter-frame prediction (i.e., affine motion prediction) using the above-mentioned affine motion model, there can be an affine merging mode (AF_MERGE or AAM) and an affine inter-frame mode (AF_INTER or AAMVP).

[0129] Similar to traditional skip / merge patterns, the affine merge pattern, according to its implementation, can represent an encoding / decoding method that performs prediction by deriving the CPMV of each of two or three CPs from the neighboring blocks of the current block without encoding the motion vector difference (MVD). Similar to AMVP, the affine inter-frame pattern (AAMVP) can explicitly represent a method for encoding / decoding the difference information between the CPMV and CPMVP.

[0130] In addition, the above Figures 3 to 6 The description of the affine motion model is intended to aid in understanding the principles of the encoding / decoding methods according to embodiments of this disclosure, which will be described later in this specification. Therefore, those skilled in the art will readily understand that the scope of this disclosure is not limited to the foregoing references. Figures 3 to 6 Limitations on the content described.

[0131] In one embodiment, a method for constructing an affine MVP candidate list for affine inter-frame prediction will be described. In this specification, the affine MVP candidate list includes affine MVP candidates, and each affine MVP candidate can represent a combination of CPMVPs of CP0 and CP1 in a four-parameter (affine) motion model, and can represent a combination of CPMVPs of CP0, CP1, and CP2 in a six-parameter (affine) motion model. The affine MVP candidates described in this specification can be referred to by various names such as CPMVP candidate, affine CPMVP candidate, CPMVP pair candidate, and CPMVP pair. The affine MVP candidate list can include n affine MVP candidates, and when n is an integer greater than 1, it may be necessary to encode and decode information indicating the best affine MVP candidate. When n is 1, it may not be necessary to encode and decode information indicating the best affine MVP candidate. Examples of the syntax when n is an integer greater than 1 are shown in Table 1 below, and examples of the syntax when n is 1 are shown in Table 2 below.

[0132] [Table 1]

[0133]

[0134] [Table 2]

[0135]

[0136] In Tables 1 and 2, `merge_flag` is a flag indicating whether it is in merge mode. A value of 1 indicates that merge mode is enabled, and a value of 0 indicates that merge mode is disabled. `affine_flag` is a flag indicating whether affine motion prediction is used. A value of 1 indicates that affine motion prediction is used, and a value of 0 indicates that affine motion prediction is disabled. `aamvp_idx` is an index indicating the best affine MVP candidate among n affine MVP candidates. It can be understood that in Table 1, where n is an integer greater than 1, `aamvp_idx` represents the best affine MVP candidate, while in Table 2, where n is 1, there is only one affine MVP candidate, therefore `aamvp_idx` is not resolved.

[0137] In one embodiment, when determining an affine MVP candidate, an affine motion model based on neighboring blocks (hereinafter also referred to as "affine-coded blocks") encoded based on affine motion prediction can be used. In one embodiment, when determining an affine MVP candidate, a first step and a second step can be performed. In the first step, while scanning neighboring blocks in a predetermined order, it can be checked whether each neighboring block has been encoded based on affine motion prediction. In the second step, neighboring blocks encoded based on affine motion prediction can be used to determine the affine MVP candidate for the current block.

[0138] In the first step, up to m blocks encoded based on affine motion prediction can be considered. For example, when m is 1, an affine MVP candidate can be determined using the affine coded block that is first in the scan order. For example, when m is 2, an affine MVP candidate can be determined using the affine coded blocks that are first and second in the scan order. In this case, when a pruning check is performed and the first and second affine MVP candidates are the same, an additional scan process can be performed to determine additional affine MVP candidates. Furthermore, in one embodiment, m described in this embodiment may not exceed the value of n described above in Tables 1 and 2.

[0139] Furthermore, in the first step, the process of checking whether each neighboring block is encoded based on affine motion prediction while scanning neighboring blocks can be implemented in various ways. These will be discussed below. Figures 7 to 10 The document describes an implementation of a process that checks whether each neighboring block is encoded based on affine motion prediction while scanning neighboring blocks.

[0140] Figures 7 to 10An example of a method for detecting neighboring blocks encoded based on affine motion prediction is shown.

[0141] Reference Figure 7 4×4 blocks A, B, C, D, and E are displayed in their adjacent positions to the current block. Block E, as the top-left adjacent block, is adjacent to CP0; block C, as the top-right adjacent block, and block B, as the top adjacent block, are adjacent to CP1; and block D, as the bottom-left adjacent block, and block A, as the left adjacent block, are adjacent to CP2. Figure 7 The layout can help reduce design costs because it can share the structure with methods based on AMVP or merge patterns.

[0142] Reference Figure 8 4×4 blocks A, B, C, D, E, F, and G are displayed adjacent to the current block. Block E, the top-left adjacent block, block G, the first left adjacent block, and block F, the top adjacent block, are adjacent to CP0. Block C, the top-right adjacent block, and block B, the second top adjacent block, are adjacent to CP1. Block D, the bottom-left adjacent block, and block A, the second left adjacent block, are adjacent to CP2. Figure 8 The arrangement is determined based solely on the 4×4 blocks adjacent to the three CPs to determine whether to encode them based on affine motion prediction, thus minimizing the increase in scan complexity and being efficient in terms of coding performance.

[0143] Figure 9 This shows the arrangement of neighboring blocks scanned when detecting neighboring blocks encoded based on affine motion prediction, which is consistent with... Figure 8 The arrangement shown is the same. However, according to Figure 9 In this implementation, the affine MVP candidate can be determined based on at most p 4×4 neighboring blocks contained within the closed dashed line to the left of the current block and at most q 4×4 neighboring blocks contained within the closed dashed line above the current block. For example, if both p and q are 1, the affine MVP candidate can be determined based on the first affine-coded block in the scan order among the 4×4 neighboring blocks contained within the closed dashed line to the left of the current block and the first affine-coded block in the scan order among the 4×4 neighboring blocks contained within the closed dashed line above the current block.

[0144] Reference Figure 10The affine MVP candidate can be determined based on the first affine coding block in the scanning order among the blocks E (top-left neighbor), G (first left neighbor), and F (first top neighbor) located adjacent to CP0; the first affine coding block in the scanning order among the blocks C (top-right neighbor) and B (second top neighbor) located adjacent to CP1; and the first affine coding block in the scanning order among the blocks D (bottom-left neighbor) and A (second left neighbor) located adjacent to CP2.

[0145] Furthermore, the scanning order of the above-described scanning method can be determined based on probability and performance analysis of a specific encoding or decoding device. Therefore, according to one embodiment, the scanning order can be determined based on the statistical characteristics or performance of the encoding or decoding device to which this embodiment is applied, rather than specifying the scanning order.

[0146] Figure 11 This is a flowchart illustrating an operation method of an encoding device according to an embodiment, and Figure 12 This is a block diagram illustrating the configuration of an encoding device according to an embodiment.

[0147] according to Figure 11 and Figure 12 Encoding devices can perform operations as described later. Figure 13 and Figure 14 The corresponding operations for the decoding device. Therefore, later in Figure 13 and Figure 14 The content described herein can be similarly applied to Figure 11 and Figure 12 Encoding devices.

[0148] Figure 11 Each step disclosed in the document can be made by Figure 1 The encoding device 100 disclosed herein performs the operation. More specifically, S1100 to S1140 can be performed by... Figure 1 The predictor 1150 disclosed in the document executes the function, and S1150 can be performed by... Figure 1 The residual processor 120 disclosed herein executes the S1160, and S1160 can be executed by... Figure 1 The entropy encoder 130 disclosed herein is executed. Furthermore, the operations according to S1100 to S1160 are based on the above. Figures 3 to 10 Some of the content described above. Therefore, the above will be omitted or briefly explained. Figure 1 and Figures 3 to 10 The specific content that is repeated within the content.

[0149] like Figure 12 As shown, the encoding device according to the embodiment may include a predictor 110 and an entropy encoder 130. However, in some cases, Figure 12 All components shown may not be essential components of the encoding device, and the encoding device may be composed of components that are not necessary for the encoding device. Figure 12 The components shown can be implemented with more or fewer components.

[0150] In the encoding device according to the embodiments, the predictor 110 and the entropy encoder 130 may be implemented by separate chips, or at least two or more components may be implemented by a single chip.

[0151] The encoding device according to the embodiment can generate an affine MVP candidate list including affine MVP candidates for the current block (S1100). More specifically, the predictor 110 of the encoding device can generate an affine MVP candidate list including affine MVP candidates for the current block.

[0152] According to the implementation, the encoding device can derive the CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1110). More specifically, the predictor 110 of the encoding device can derive the CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list.

[0153] The encoding device according to the implementation can derive CPMV for each CP of the current block (S1120). More specifically, the predictor 110 of the encoding device can derive CPMV for each CP of the current block.

[0154] According to the implementation, the encoding device can derive the CPMVD of the current block's CP based on the CPMV and CPMVP of each CP (S1130). More specifically, the predictor 110 of the encoding device can derive the CPMVD of the current block's CP based on the CPMV and CPMVP of each CP.

[0155] The encoding device according to the implementation can derive a prediction sample for the current block based on CPMV (S1140). More specifically, the predictor 110 of the encoding device can derive a prediction sample for the current block based on CPMV.

[0156] The encoding device according to the embodiment can derive residual samples for the current block based on the derived prediction samples (S1150). More specifically, the residual processor 120 of the encoding device can derive residual samples for the current block based on the derived prediction samples.

[0157] The encoding device according to the embodiment can encode information about the derived CPMVD and residual information about the residual samples (S1160). More specifically, the entropy encoder 130 of the encoding device can encode information about the derived CPMVD and residual information about the residual samples.

[0158] According to Figure 11 and Figure 12 The disclosed encoding device and its operation method allow the encoding device to generate an affine MVP candidate list including affine MVP candidates for the current block (S1100), derive a CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1110), derive a CPMV for each CP of the current block (S1120), derive a CPMVD for the current block based on the CPMV and CPMVP of each CP (S1130), derive a prediction sample for the current block based on the CPMV (S1140), derive a residual sample for the current block based on the derived prediction sample (S1150), and encode information about the derived CPMVD and residual information about the residual sample (S1160). In other words, image encoding efficiency can be increased by signaling information about the affine MVP candidate list used for affine motion prediction.

[0159] Figure 13 This is a flowchart illustrating an operation method of the decoding device according to an embodiment. Figure 14 This is a block diagram illustrating the configuration of a decoding device according to an embodiment.

[0160] Figure 13 Each step disclosed in the document can be made by Figure 2 The video decoding device 200 disclosed in the document performs this operation. More specifically, S1300 can be performed by... Figure 2 The entropy decoder 210 disclosed in the document executes S1310 to S1350, which can be performed by... Figure 2 The predictor 230 disclosed in the document is executed, and S1360 can be performed by... Figure 2 The adder 240 disclosed herein is executed. Furthermore, the operations according to S1300 to S1360 are based on the above. Figures 3 to 10 Some of the content described above. Therefore, the above will be omitted or briefly explained. Figures 2 to 10 The specific content that is repeated within the content.

[0161] The decoding device according to the embodiment may include an entropy decoder 210, a predictor 230, and an adder 240. However, in some cases, Figure 14 All the components shown may not be necessary components of the decoding device, and the decoding device may be composed of components that are not necessary for the decoding device. Figure 14The components shown can be implemented with more or fewer components.

[0162] In the decoding device according to the embodiment, the entropy decoder 210, the predictor 230 and the adder 240 can be implemented by a single chip, or at least two or more components can be implemented by a single chip.

[0163] The decoding device according to the embodiment can obtain motion prediction information from the bitstream (S1300). More specifically, the entropy decoder 210 of the decoding device can obtain motion prediction information from the bitstream.

[0164] According to the implementation, the decoding device can generate an affine MVP candidate list including affine motion vector prediction (MVP) candidates for the current block (S1310). More specifically, the predictor 230 of the decoding device can generate an affine MVP candidate list including affine MVP candidates for the current block.

[0165] In one embodiment, the affine MVP candidate may include a first affine MVP candidate and a second affine MVP candidate. The first affine MVP candidate can be derived from the left block group, which includes the lower-left neighbor block and the left neighbor block of the current block. The second affine MVP candidate can be derived from the upper block group, which includes the upper-right neighbor block, the upper neighbor block, and the upper-left neighbor block of the current block. In this regard, the first affine MVP candidate may be derived based on a first block included in the left block group, which may be encoded based on affine motion prediction. The second affine MVP candidate may be derived based on a second block included in the upper block group, which may be encoded based on affine motion prediction.

[0166] In another embodiment, the affine MVP candidate may include a first affine MVP candidate and a second affine MVP candidate. The first affine MVP candidate can be derived from the left block group, which includes the lower-left neighbor block of the current block, the first left neighbor block, and the second left neighbor block. The second affine MVP candidate can be derived from the upper-right neighbor block of the current block, the upper block group, the second upper neighbor block, and the upper-left neighbor block. In this regard, the first affine MVP candidate may be derived based on a first block included in the left block group, which may be encoded based on affine motion prediction. The second affine MVP candidate may be derived based on a second block included in the upper block group, which may be encoded based on affine motion prediction.

[0167] In another embodiment, the affine MVP candidate may include a first affine MVP candidate, a second affine MVP candidate, and a third affine MVP candidate. The first affine MVP candidate can be derived from the lower-left block group including the lower-left adjacent block of the current block and the first left adjacent block. The second affine MVP candidate can be derived from the upper-right block group including the upper-right adjacent block of the current block and the first upper adjacent block. The third affine MVP candidate can be derived from the upper-left block group including the upper-left adjacent block of the current block, the second upper adjacent block, and the second left adjacent block. In this regard, the first affine MVP candidate may be derived based on the first block included in the lower-left block group, and the first block may be encoded based on affine motion prediction. The second affine MVP candidate may be derived based on the second block included in the upper-right block group, and the second block may be encoded based on affine motion prediction. The third affine MVP candidate may be derived based on the third block included in the upper-left block group, and the third block may be encoded based on affine motion prediction.

[0168] According to the implementation, the decoding device can derive the CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1320). More specifically, the predictor 230 of the decoding device can derive the CPMVP for each CP of the current block based on one of the affine MVP candidates included in the affine MVP candidate list.

[0169] In one implementation, an affine MVP candidate can be selected from the affine MVP candidates based on the affine MVP candidate index included in the motion prediction information.

[0170] According to the implementation, the decoding device can deduce the CPMVD for the current block's CP based on information about the CPMVD for each CP included in the acquired motion prediction information (S1330). More specifically, the predictor 230 of the decoding device can deduce the CPMVD for the current block's CP based on information about the CPMVD for each CP included in the acquired motion prediction information.

[0171] According to the implementation, the decoding device can derive the CPMV for the CP of the current block based on CPMVP and CPMVD (S1340). More specifically, the predictor 230 of the decoding device can derive the CPMV for the CP of the current block based on CPMVP and CPMVD.

[0172] The decoding device according to the embodiment can derive a prediction sample for the current block based on CPMV (S1350). More specifically, the predictor 230 of the decoding device can derive a prediction sample for the current block based on CPMV.

[0173] According to the implementation method, the decoding device can generate a reconstruction sample for the current block based on the derived prediction sample (S1360). More specifically, the adder 240 of the decoding device can generate a reconstruction sample for the current block based on the derived prediction sample.

[0174] In an implementation, motion prediction information may include context indexes indicating whether there are neighboring blocks of the current block encoded based on affine motion prediction.

[0175] In the implementation, for the case where the value of m described in the first step above is 1, and the value of n described in Tables 1 and 2 above is 2, a CABAC context model for encoding and decoding index information indicating the best affine MVP candidate can be constructed. When the affine coded block exists in the vicinity of the current block, it can be based on the above reference... Figures 7 to 10 The described affine motion model determines the affine MVP candidate for the current block, but this implementation can be applied when the affine coded block does not exist in the vicinity of the current block. Since the affine MVP candidate has high reliability when determined based on the affine coded block, a context model can be designed to distinguish between the case where the affine MVP candidate is determined based on the affine coded block and the case where the affine MVP candidate is determined in a different manner. In this case, index 0 can be assigned to the affine MVP candidate determined based on the affine coded block. The CABAC context index according to this implementation is shown in Equation 5 below.

[0176] [Formula 5]

[0177]

[0178] The initial value based on the CABAC context index can be determined as shown in Table 3 below, and the CABAC context index and the initial value need to satisfy the condition in Equation 6 below.

[0179] [Table 3]

[0180] ctx_idx_for_aamvp_idx 0 1 Init_val <![CDATA[N0]]> <![CDATA[N1]]>

[0181] [Formula 6]

[0182] p(aamvp_idx=0|init_val=N0)>p(1|N0)

[0183] p(aamvp_idx=0|init_val=N0)>p(0|N1)

[0184] p(aamvp_idx=0|init_val=N0)>p(1|N1)

[0185] according to Figure 13 and Figure 14 The decoding device and its operation method are described. The decoding device can obtain motion prediction information from the bitstream (S1300), generate an affine MVP candidate list including affine MVP candidates for the current block (S1310), derive the CPMVP for each CP for the current block based on one of the affine MVP candidates included in the affine MVP candidate list (S1320), derive the CPMVD for the current block based on the information about the CPMVD for each CP included in the obtained motion prediction information (S1330), derive the control point motion vector (CPMV) for the current block based on the CPMVP and CPMVD (S1340), derive the prediction sample for the current block based on the CPMV (S1350), and generate a reconstructed sample for the current block based on the derived prediction sample (S1360). In other words, image coding efficiency can be increased by signaling information about the affine MVP candidate list used for affine motion prediction.

[0186] Furthermore, the methods described above for image and video compression according to this specification can be applied to encoding and decoding devices, to devices that generate bitstreams and devices that receive bitstreams, and can be applied regardless of whether the terminal outputs data through a display device. For example, an image can be generated as compressed data by a terminal equipped with an encoding device. This compressed data can be in bitstream form, and the bitstream can be stored in various types of storage devices and streamed over a network to a terminal equipped with a decoding device. When the terminal is equipped with a display device, the decoded image can be displayed on the display device, or the bitstream data can simply be stored in the terminal.

[0187] The methods described above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included in image processing apparatus such as TVs, computers, smartphones, set-top boxes, display devices, etc.

[0188] Each of the aforementioned components, modules, or units can be a processor or hardware component that performs sequential processing stored in a memory (or storage unit). Each step described in the above embodiments can be executed by a processor or hardware component. Each module / block / unit in the above embodiments can operate as hardware / processor. Furthermore, the methods proposed in this disclosure can be implemented using code. This code can be written to a storage medium that can be read by a processor, and therefore can be read by the processor provided by the device.

[0189] In the above embodiments, the method is explained based on a flowchart using a series of steps or block diagrams. However, this disclosure is not limited to the order of the steps, and a step may occur in a different order than described above, or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be removed without affecting the scope of this disclosure.

[0190] When the embodiments of this disclosure are implemented in software, the methods described above can be implemented using modules (processes, functions, etc.) that perform the functions described above. These modules can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor in various well-known ways. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices.

Claims

1. A decoding device for image decoding, the decoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: Obtain motion prediction information from the bitstream; Generate an affine MVP candidate list that includes affine motion vector prediction MVP candidates for the current block; An affine MVP candidate is selected from the affine MVP candidates in the affine MVP candidate list by using the affine MVP candidate index in the motion prediction information; Based on the selected affine MVP candidate, derive the control point motion vector prediction value CPMVP for each control point CP of the current block; The CPMVD for each CP of the current block is derived based on the information about the control point motion vector difference (CPMVD) for each CP included in the obtained motion prediction information. Based on the CPMVP and the CPMVD, the control point motion vector CPMV for each CP of the current block is derived; The sub-block motion vector for the current block is derived based on the CPMV of each CP for the current block; The prediction sample for the current block is derived based on the sub-block motion vectors for the current block; and Based on the predicted samples, a reconstruction sample is generated for the current block. The affine MVP candidate list includes a first affine MVP candidate and a second affine MVP candidate. The at least one processor is further configured to: The first MVP, second MVP, and third MVP constituting the first affine MVP candidate are derived from blocks encoded based on an affine motion model in the left block group, including the lower-left adjacent block and the left adjacent block of the current block; and The fourth, fifth, and sixth MVPs constituting the second affine MVP candidate are derived from the blocks encoded based on the affine motion model in the upper block group, which includes the upper right adjacent block, the upper adjacent block, and the upper left adjacent block of the current block. Wherein, the left block group does not include the upper right adjacent block, the upper adjacent block, and the upper left adjacent block, and The upper block group does not include the lower left adjacent block and the left adjacent block.

2. An encoding device for image encoding, the encoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: Generate a list of affine MVP candidates, including affine MVP candidates for the current block; Select one affine MVP candidate from the affine MVP candidates in the affine MVP candidate list; The derivation represents the affine MVP candidate index of the selected affine MVP candidate; The CPMVP for each CP of the current block is derived based on the selected affine MVP candidate. Derive the CPMV for each CP of the current block; The CPMVD for each CP of the current block is derived based on the CPMVP and CPMV for each CP. The sub-block motion vector for the current block is derived based on the CPMV of each CP for the current block; The prediction sample for the current block is derived based on the sub-block motion vectors for the current block; Based on the predicted samples, residual samples for the current block are derived; and Information related to the affine MVP candidate index, information about the CPMVD, and residual information about the residual samples are encoded. The affine MVP candidate list includes a first affine MVP candidate and a second affine MVP candidate. The at least one processor is further configured to: The first MVP, second MVP, and third MVP constituting the first affine MVP candidate are derived from blocks encoded based on an affine motion model in the left block group, including the lower-left adjacent block and the left adjacent block of the current block; and The fourth, fifth, and sixth MVPs constituting the second affine MVP candidate are derived from the blocks encoded based on the affine motion model in the upper block group, which includes the upper right adjacent block, the upper adjacent block, and the upper left adjacent block of the current block. Wherein, the left block group does not include the upper right adjacent block, the upper adjacent block, and the upper left adjacent block, and The upper block group does not include the lower left adjacent block and the left adjacent block.

3. An apparatus for transmitting data for an image, the apparatus comprising: At least one processor is configured to obtain a bitstream for the image, wherein the bitstream is generated based on the following operations: generating an affine MVP candidate list including affine MVP candidates for a current block; selecting an affine MVP candidate from the affine MVP candidates in the affine MVP candidate list; deriving an affine MVP candidate index representing the selected affine MVP candidate; deriving a CPMVP for each CP of the current block based on the selected affine MVP candidate; deriving a CPMV for each CP of the current block; deriving a CPMVD for each CP of the current block based on the CPMVP and CPMV for each CP; deriving a sub-block motion vector for the current block based on the CPMV for each CP of the current block; deriving a prediction sample for the current block based on the sub-block motion vector for the current block; deriving a residual sample for the current block based on the prediction sample; and encoding information related to the affine MVP candidate index, information about the CPMVD, and residual information about the residual sample; and A transmitter configured to transmit the data comprising the bit stream. The affine MVP candidate list includes a first affine MVP candidate and a second affine MVP candidate. Generating the affine MVP candidate list includes the following operations: The first MVP, second MVP, and third MVP constituting the first affine MVP candidate are derived from blocks encoded based on an affine motion model in the left block group, including the lower-left adjacent block and the left adjacent block of the current block; and The fourth, fifth, and sixth MVPs constituting the second affine MVP candidate are derived from the blocks encoded based on the affine motion model in the upper block group, which includes the upper right adjacent block, the upper adjacent block, and the upper left adjacent block of the current block. Wherein, the left block group does not include the upper right adjacent block, the upper adjacent block, and the upper left adjacent block, and The upper block group does not include the lower left adjacent block and the left adjacent block.

Citation Information

Patent Citations

  • Method And Device For Encoding Three-Dimensional Image, And Decoding Method And Device

    CN104025601A

  • Method and device for image prediction

    CN105163116A