Image decoding method and apparatus based on affine motion prediction using an affine MVP candidate list in an image coding system
The image decoding method constructs an affine MVP candidate list using neighboring blocks to enhance compression efficiency and reduce hardware costs, addressing the high transmission and storage costs of high-resolution images.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-08-14
- Publication Date
- 2026-06-02
AI Technical Summary
The increasing demand for high-resolution and high-quality images has led to a rise in transmission and storage costs due to the increased amount of information, necessitating a more efficient image compression technology.
An image decoding method and apparatus that constructs an affine MVP candidate list based on neighboring blocks, deriving control point motion vector predictors and differences, and using these to perform prediction on the current block, ensuring efficient image decoding even when the number of available candidates is limited.
This approach improves the overall image/video compression efficiency by reducing the complexity of deriving affine MVP candidates and minimizing hardware costs, while maintaining effective image decoding.
Smart Images

Figure 0007869387000012 
Figure 0007869387000013 
Figure 0007869387000014
Abstract
Description
Technical Field
[0001] This document relates to image encoding technology, and more particularly, to an image decoding method and apparatus based on affine motion prediction in an image encoding system.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing image data. Therefore, when transmitting image data using a medium such as an existing wired / wireless broadband line or storing image data using an existing storage medium, the transmission cost and storage cost increase.
[0003] Thus, in order to effectively transmit, store, and reproduce information of high-resolution and high-quality images, a highly efficient image compression technology is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The technical problem of this document is to provide a method and apparatus for increasing image encoding efficiency.
[0005] Another technical problem of this document is to derive a constructed affine MVP candidate based on neighboring blocks and configure an affine MVP candidate list for the current block only when all candidate motion vectors for the CP are available, and to provide an image decoding method and apparatus for performing prediction on the current block based on the configured affine MVP candidate list.
[0006] Another technical problem of this document is to provide an image decoding method and apparatus that derive affine MVP candidates using candidate motion vectors derived in the constructed affine MVP candidate derivation process as additional affine MVP candidates when the number of available inherited affine MVP candidates and constructed affine MVP candidates is smaller than the maximum number of candidates in the MVP candidate list, and perform prediction on the current block based on the constructed affine MVP candidate list.
Means for Solving the Problem
[0007] According to one embodiment of this document, an image decoding method performed by a decoding device is provided. The method includes the steps of: obtaining motion prediction information for the current block from a bitstream; constructing a list of candidate affine motion vector predictors (MVPs) for the current block; deriving CPMVPs (Control Point Motion Vector Predictors) for the control points (CPs) of the current block based on the affine MVP candidate list; deriving CPMVPs (Control Point Motion Vector Differences) for the control points (CPs) of the current block based on the motion prediction information; and deriving CPMVPs (Control Point Motion Vector Differences) for the control points (CPs) of the current block based on the CPMVPs and CPMVPs. The steps include deriving Vectors, deriving predicted samples for the current block based on the CPMV, and generating a reconstructed picture for the current block based on the derived predicted samples, wherein the step of constructing the affine MVP candidate list includes checking whether an inherited affine MVP candidate for the current block is available, and if an inherited affine MVP candidate is available, the step of deriving the inherited affine MVP candidate, the constructed affine MVP for the current block The process checks if a constructed affine MVP candidate is available, and if so, the constructed affine MVP candidate is derived, the constructed affine MVP candidate includes a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block. If the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, a first affine MVP candidate is derived, the first affine MVP candidate is,Steps include: deriving an affine MVP candidate that includes the motion vector for CP0 as a candidate motion vector for CP; if the number of derived affine MVP candidates is less than 2 and the motion vector for CP1 is available, deriving a second affine MVP candidate, the second affine MVP candidate being an affine MVP candidate that includes the motion vector for CP1 as a candidate motion vector for CP; if the number of derived affine MVP candidates is less than 2 and the motion vector for CP2 is available, deriving a third affine MVP candidate, the third affine MVP candidate being an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for CP; if the number of derived affine MVP candidates is less than 2, deriving a fourth affine MVP candidate that includes the temporal MVP derived based on the temporal surrounding blocks of the current block as a candidate motion vector for CP; and if the number of derived affine MVP candidates is less than 2, deriving a zero motion vector (zero motion). The method is characterized by including a step of deriving a fifth affine MVP candidate that includes a vector as a candidate motion vector for the CP.
[0008] According to another embodiment of this document, a decoding device for image decoding is provided. The decoding device includes an entropy decoding unit that acquires motion prediction information for the current block from a bitstream, a list of candidate affine motion vector predictors (MVPs) for the current block, a CPMVP (Control Point Motion Vector Predictors) for the current block's CP (Control Point) based on the affine MVP candidate list, a CPMVP (Control Point Motion Vector Predictors) for the current block's CP based on the motion prediction information, and a CMVD (Control Point Motion Vector Differences) for the current block's CP based on the CPMVP and CMVD. The affine MVP candidate list includes a prediction unit that derives vectors and derives predicted samples for the current block based on the CPMV, and an addition unit that generates a restored picture for the current block based on the derived predicted samples, and the affine MVP candidate list checks whether an inherited affine MVP candidate for the current block is available, and if an inherited affine MVP candidate is available, the inherited affine MVP candidate is derived, and the constructed affine MVP candidate for the current block is The steps include checking if the constructed affine MVP candidate is available, deriving the constructed affine MVP candidate, which includes a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; if the number of derived affine MVP candidates is less than 2 and the motion vector for CP0 is available, deriving a first affine MVP candidate, which is the first affine MVP candidate.Steps include: deriving an affine MVP candidate that includes the motion vector for CP0 as a candidate motion vector for CP; if the number of derived affine MVP candidates is less than 2 and the motion vector for CP1 is available, deriving a second affine MVP candidate, the second affine MVP candidate being an affine MVP candidate that includes the motion vector for CP1 as a candidate motion vector for CP; if the number of derived affine MVP candidates is less than 2 and the motion vector for CP2 is available, deriving a third affine MVP candidate, the third affine MVP candidate being an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for CP; if the number of derived affine MVP candidates is less than 2, deriving a fourth affine MVP candidate that includes the temporal MVP derived based on the temporal surrounding blocks of the current block as a candidate motion vector for CP; and if the number of derived affine MVP candidates is less than 2, deriving a zero motion vector (zero motion). The method is characterized by being based on the step of deriving a fifth affine MVP candidate that includes a vector as a candidate motion vector for the CP.
[0009] Another embodiment of this document provides a video encoding method performed by an encoding device. The method comprises the steps of: configuring a list of candidate affine motion vector predictors (MVPs) for a current block; deriving CPMVPs (Control Point Motion Vector Predictors) for the control points (CPs) of the current block based on the affine MVP candidate list; deriving CPMVs for the control points (CPs) of the current block; deriving CPMVPs (Control Point Motion Vector Differences) for the control points (CPs) of the current block based on the CPMVPs and CPMVs; and providing motion prediction information including information about the CPMVPs.The steps include encoding information, and the steps of constructing the affine MVP candidate list include checking whether an inherited affine MVP candidate for the current block is available, and if an inherited affine MVP candidate is available, the inherited affine MVP candidate is derived, checking whether a constructed affine MVP candidate for the current block is available, and if a constructed affine MVP candidate is available, the constructed affine MVP candidate is derived, and the constructed affine MVP candidate includes a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block, and if the number of derived affine MVP candidates is less than 2 and a motion vector for CP0 is available, a first affine MVP candidate is derived, and the first affine MVP candidate includes a motion for CP0 Steps include: deriving an affine MVP candidate that includes a vector as a candidate motion vector for the CP; if the number of derived affine MVP candidates is less than 2 and a motion vector for CP1 is available, deriving a second affine MVP candidate, the second affine MVP candidate being an affine MVP candidate that includes the motion vector for CP1 as a candidate motion vector for the CP; if the number of derived affine MVP candidates is less than 2 and a motion vector for CP2 is available, deriving a third affine MVP candidate, the third affine MVP candidate being an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for the CP; if the number of derived affine MVP candidates is less than 2, deriving a fourth affine MVP candidate that includes a temporal MVP derived based on the temporal surrounding blocks of the current block as a candidate motion vector for the CP; and if the number of derived affine MVP candidates is less than 2, deriving a zero motion vector (zero motionThe method is characterized by including the step of deriving a fifth affine MVP candidate that includes a vector as a candidate motion vector for the CP.
[0010] A video encoding device is provided according to yet another embodiment of this document. The encoding device comprises a prediction unit which configures a list of candidate affine motion vector predictors (MVPs) for the current block, derives CPMVPs (Control Point Motion Vector Predictors) for the current block's CPs (Control Points) based on the affine MVP candidate list, and derives CPMVs for the current block's CPs; a subtraction unit which derives CPMVPDs (Control Point Motion Vector Differences) for the current block's CPs based on the CPMVPs and CPMVs; and motion prediction information which includes information about the CPMVPD.The affine MVP candidate list includes an entropy encoding unit that encodes information, and the steps include: checking whether an inherited affine MVP candidate of the current block is available, and if an inherited affine MVP candidate is available, the inherited affine MVP candidate is derived; checking whether a constructed affine MVP candidate of the current block is available, and if a constructed affine MVP candidate is available, the constructed affine MVP candidate is derived, and the constructed affine MVP candidate includes a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; if the number of derived affine MVP candidates is less than 2 and a motion vector for CP0 is available, a first affine MVP candidate is derived, and the first affine MVP candidate includes a motion vector for CP0 Steps include: deriving an affine MVP candidate that includes a vector as a candidate motion vector for the CP; if the number of derived affine MVP candidates is less than 2 and a motion vector for CP1 is available, deriving a second affine MVP candidate, the second affine MVP candidate being an affine MVP candidate that includes the motion vector for CP1 as a candidate motion vector for the CP; if the number of derived affine MVP candidates is less than 2 and a motion vector for CP2 is available, deriving a third affine MVP candidate, the third affine MVP candidate being an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for the CP; if the number of derived affine MVP candidates is less than 2, deriving a fourth affine MVP candidate that includes a temporal MVP derived based on the temporal surrounding blocks of the current block as a candidate motion vector for the CP; and if the number of derived affine MVP candidates is less than 2, deriving a zero motion vector (zero motionThe method is characterized by being configured based on the step of deriving a fifth affine MVP candidate that includes a vector as a candidate motion vector for the CP. [Effects of the Invention]
[0011] According to this document, it is possible to improve the overall image / video compression efficiency.
[0012] According to this document, the efficiency of image coding based on affine motion prediction can be improved.
[0013] According to this document, when deriving the affine MVP candidate list, a constructed affine MVP candidate can be added only if all candidate motion vectors for the CP of the constructed affine MVP candidate are available. This reduces the complexity of the process of deriving the constructed affine MVP candidate and the process of constructing the affine MVP candidate list, thereby improving coding efficiency.
[0014] According to this paper, in deriving the affine MVP candidate list, additional affine MVP candidates can be derived based on the candidate motion vectors for the CP derived in the process of deriving the constructed affine MVP candidates. This reduces the complexity of the process of constructing the affine MVP candidate list and improves coding efficiency.
[0015] According to this document, the inherited affine MVP candidate can be derived using the upper peripheral block only if the upper peripheral block is currently included in the CTU during the process of deriving the inherited affine MVP candidate. This reduces the amount of line buffer storage required for affine prediction and minimizes hardware costs. [Brief explanation of the drawing]
[0016] [Figure 1] This document outlines an example of a video / image encoding system to which the embodiments described herein may be applied. [Figure 2] This figure outlines the configuration of a video / image encoding device to which the embodiments described herein may be applied. [Figure 3] This diagram illustrates the schematic configuration of a video / image decoding device to which the embodiments described herein may be applied. [Figure 4] The motion represented through the aforementioned affine motion model is illustrated in the following example. [Figure 5] An affine motion model in which motion vectors for three control points are used is illustrated as an example. [Figure 6] An affine motion model in which motion vectors for two control points are used is illustrated as an example. [Figure 7] An illustrative method for deriving motion vectors at the subblock level based on the aforementioned affine motion model is shown. [Figure 8] A schematic diagram illustrating the sequence of an affine motion prediction method according to one embodiment of this document is shown as an example. [Figure 9] This figure illustrates a method for deriving a motion vector predictor at a control point according to one embodiment of this document. [Figure 10] This figure illustrates a method for deriving a motion vector predictor at a control point according to one embodiment of this document. [Figure 11] This shows an example of affine prediction performed when surrounding block A is selected as a candidate for affine merge. [Figure 12] The surrounding blocks for deriving the inherited affine candidate are shown as an example. [Figure 13] An illustrative example of a spatial candidate for the aforementioned constructed affine candidate is shown below. [Figure 14] An example of how to construct an Affine MVP list is shown below. [Figure 15] An example of deriving the aforementioned constructed candidate is shown below. [Figure 16] An example of deriving the aforementioned constructed candidate is shown below. [Figure 17]The surrounding block locations scanned to derive inherited affine candidates are shown exemplarily. [Figure 18] An example of deriving the constructed candidate when a 4-affine motion model is applied to the aforementioned current block is shown. [Figure 19] An example of deriving the constructed candidate when a 6-affine motion model is applied to the aforementioned current block is shown. [Figure 20a] An exemplary embodiment for deriving the inherited affine candidate is shown below. [Figure 20b] An exemplary embodiment for deriving the inherited affine candidate is shown below. [Figure 21] This document outlines the image encoding method using the encoding device described herein. [Figure 22] This document outlines an encoding device that performs the image encoding method described herein. [Figure 23] This document outlines the image decoding method using the decoding device described herein. [Figure 24] This document outlines a decoding device that performs the image decoding method described herein. [Figure 25] An illustrative diagram of a content streaming system structure to which the embodiments described herein apply is shown. [Modes for carrying out the invention]
[0017] This document can be modified in various ways and may have various embodiments; however, specific embodiments are illustrated in the drawings and described in detail. This does not mean that this document is intended to limit itself to any particular embodiment. Terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as “includes” or “has” herein are intended to specify the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0018] On the other hand, each configuration shown in the diagrams described in this document is illustrated independently for the purpose of explaining its distinct characteristic functions, etc., and does not mean that each configuration is implemented with separate hardware or separate software. For example, two or more of the configurations can be combined to form one configuration, and one configuration can be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included within the scope of the rights of this document, as long as they do not deviate from the essence of this document.
[0019] Desired embodiments of this document will be described in more detail below with reference to the attached drawings. Hereafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components will be omitted.
[0020] Figure 1 shows a schematic example of a video / image coding system to which the embodiments described in this document may be applied.
[0021] As shown in Figure 1, a video / image encoding system may include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.
[0022] The source device may comprise a video source, an encoding device, and a transmitter. The receiving device may comprise a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may comprise a display unit, which may consist of a separate device or external component.
[0023] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. A video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. A video / image generation device may include, for example, a computer, tablet, and smartphone, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.
[0024] An encoding device can encode input video / images. For compression and encoding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.
[0025] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0026] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.
[0027] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0028] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0029] This document presents various embodiments relating to video / image encoding, and unless otherwise noted, these embodiments may be combined with each other.
[0030] In this document, "video" can mean a collection of images or other elements following a flow of time. "Picture" generally refers to a unit representing a single image at a specific time point in time, while "slice" or "tile" is a unit that constitutes part of a picture in encoding. A slice or tile can contain one or more CTUs (coding tree units). A single picture can consist of one or more slices or tiles. A single picture can consist of one or more tile groups. A tile group can contain one or more tiles. A brick can represent a rectangular region of CTU rows within a tile in a picture. A tile can be partitioned into multiple bricks, each of which consists of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan can represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in a CTU raster scan in a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture. A tile scan can represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice may contain an integer number of bricks of a picture, and these integer number of bricks may be contained in a single NAL unit.A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0031] A pixel or pel can refer to the smallest unit that makes up a picture (or image). The term "sample" can also be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value.
[0032] A unit can represent a basic unit of image processing. A unit can contain at least one of the following: a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a sample (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients, etc.
[0033] In this document, the terms " / " and "," should be interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Similarly, "A, B, C" also means "at least one of A, B, and / or C."
[0034] In addition, in this document, "or" should be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) both "A and B." In other words, "or" in this document can mean "additionally or alternatively."
[0035] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments described in this document may be applied. Hereinafter, the term "video encoding device" may include an image encoding device.
[0036] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) can include an inter-prediction unit (221) and an intra-prediction unit (222). The residual processor (230) can include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) can further include a subtractor (231). The addition unit 250 may be called a reconstructor or a reconstructed block generator. The image segmentation unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 described above may be composed of one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be composed of a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0037] The image splitting unit 210 can split an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure and / or the ternary structure. Alternatively, the binary tree structure may be applied first. The coding procedure described in this document may be performed based on the final coding unit that is not further split. In this case, based on the coding efficiency according to the image characteristics, the largest coding unit can be immediately used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of deeper depth so that the optimally sized coding unit can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.
[0038] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.
[0039] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes the predicted sample for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.
[0040] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) of the current block or at a distance, depending on the prediction mode. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and Planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0041] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include information on the interprediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, surrounding blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, colCU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 221 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidates are used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 221 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0042] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content image / video coding, such as in games, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be done similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be seen as an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information about the palette table and palette index.
[0043] The prediction signal generated by starting the pre-prediction unit (including the inter-prediction unit 221 and / or pre-intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can generate transformation coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means a transformation obtained from a graph when the relationship information between pixels is represented by this graph. CNT means a transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process can be applied to pixel blocks of the same size and are square, or to non-square, variable-sized blocks.
[0044] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this document may be included in the video / image information. The video / image information may be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.
[0045] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, and can also be used for inter-prediction of the next picture after filtering, as described later.
[0046] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.
[0047] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream form.
[0048] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.
[0049] Memory 270DPB can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information has been derived (or encoded) in the current picture and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.
[0050] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document may be applied.
[0051] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) can include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) can include an inverse quantizer (321) and an inverse transformer (322). The entropy decoder (310), residual processor (320), predictor (330), adder (340), and filter (350) described above can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory (360) can also include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.
[0052] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image corresponding to the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit may be, for example, an encoding unit, and the encoding unit can be divided from an encoding tree unit or a maximum encoding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units may be derived from the encoding unit. The reconstructed image signal decoded and output via the decoding device 300 can then be reproduced via a playback device.
[0053] The decoding device 300 can receive the signal output from the encoding device shown in Figure 2 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction, quantized values of conversion coefficients related to residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the information of the syntax element to be decoded, the decoded information of the surrounding and decoded blocks, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values entropy-decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, inter-prediction unit 332, and intra-prediction unit 331.
[0054] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.
[0055] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0056] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.
[0057] The prediction unit 320 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on intra-block copy (IBC) prediction mode or palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be done similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be seen as an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0058] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample can be located in the vicinity (neighbor) of the current block or at a distance from it, depending on the prediction mode. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0059] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted from the interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.
[0060] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 332 and / or intra-prediction unit 331). If there is no residual signal for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.
[0061] The summing unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.
[0062] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0063] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0064] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.
[0065] In this specification, the embodiments described for the filtering unit 260, inter-prediction unit 221, and intra-prediction unit 222 of the encoding device 100 can be applied similarly or in a corresponding manner to the filtering unit 350, inter-prediction unit 332, and intra-prediction unit 331 of the decoding device 300, respectively.
[0066] On the other hand, in relation to inter prediction, inter prediction methods that take image distortion into account have been proposed. Specifically, an affine motion model has been proposed that efficiently derives motion vectors for subblocks or sample points of the current block, and improves the accuracy of inter prediction despite deformations such as image rotation, zoom in, or zoom out. That is, an affine motion model has been proposed that derives motion vectors for subblocks or sample points of the current block. Predictions using this affine motion model can be called affine inter prediction or affine motion prediction.
[0067] For example, the affine interpretation using the affine motion model can efficiently represent four types of motion, or four types of deformation, as described later.
[0068] Figure 4 illustrates motion represented by the affine motion model. As shown in Figure 4, motion that can be represented by the affine motion model can include translation, scaling, rotation, and shear motion. That is, not only translation, in which part of an image moves in a plane due to the passage of time as shown in Figure 4, but also scaling, in which part of an image is scaled due to the passage of time, rotation, in which part of an image is rotated due to the passage of time, and shear, in which part of an image is deformed into a parallelogram due to the passage of time can be efficiently represented by the affine interpretation.
[0069] The encoding / decoding device can predict the distortion pattern of the image based on the motion vector at the control point (CP) of the current block via the affine interpretation, thereby improving the accuracy of the prediction and thus improving the image compression performance. Furthermore, since the motion vector for at least one control point of the current block can be derived using the motion vectors of the surrounding blocks of the current block, the burden of data volume for additional information can be reduced, and the interpretation efficiency can be significantly improved.
[0070] As an example of the aforementioned affine interpretation prediction, motion information from three control points, i.e., three reference points, may be required.
[0071] Figure 5 illustrates the affine motion model in which motion vectors for three control points are used.
[0072] Currently, if the top-left sample position within block 500 is (0,0), then the (0,0), (w,0), and (0,h) sample positions can be determined as control points, as shown in Figure 5. Hereinafter, the control point for the (0,0) sample position can be represented as CP0, the control point for the (w,0) sample position as CP1, and the control point for the (0,h) sample position as CP2.
[0073] Using the control points and motion vectors for each of the control points described above, a mathematical formula for the affine motion model can be derived. The mathematical formula for the affine motion model can be expressed as follows:
[0074]
number
[0075] Here, w represents the width of the current block 500, h represents the height of the current block 500, and v 0x , v 0y These represent the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y These represent the x and y components of the motion vector of CP1, respectively, and v 2x , v 2y These represent the x and y components of the motion vector of CP2, respectively. Furthermore, x represents the x component of the position of the target sample within the current block 500, and y represents the y component of the position of the target sample within the current block 500, and v x The x-component of the motion vector of the target sample within the current block 500, v y This shows the y component of the motion vector of the target sample currently within block 500.
[0076] Since the motion vectors of the CP0, the motion vector of the CP1, and the motion vector of the CP2 are known, a motion vector based on the sample position within the current block can be derived according to the formula 1. That is, according to the affine motion model, based on the distance ratios between the coordinates (x, y) of the target sample and the three control points, the motion vectors v0(v 0x , v 0y ) at the control points, v1(v 1x , v 1y ), and v2(v 2x , v 2y ) are scaled, and the motion vector of the target sample according to the target sample position can be derived. That is, according to the affine motion model, the motion vector of each sample within the current block can be derived based on the motion vectors of the control points. On the other hand, a set such as the motion vectors of the samples within the current block derived by the affine motion model can be represented as an affine motion vector field (MVF).
[0077] On the other hand, the six parameters for the formula 1 can be represented by a, b, c, d, e, f as in the following formula, and the formula for the affine motion model represented by the six parameters can be as follows.
[0078]
Equation
[0079] Here, w represents the width of the current block 500, h represents the height of the current block 500, v 0x , v 0y respectively represent the x - component and y - component of the motion vector of CP0, v 1x , v 1y respectively represent the x - component and y - component of the motion vector of CP1, v 2x , v 2yThese represent the x and y components of the motion vector of CP2, respectively. Furthermore, x represents the x component of the position of the target sample within the current block 500, and y represents the y component of the position of the target sample within the current block 500, and v x The x-component of the motion vector of the target sample within the current block 500, v y This shows the y component of the motion vector of the target sample currently within block 500.
[0080] The affine motion model or affine interpretation using the six parameters can be represented as a 6-parameter affine motion model or AF6.
[0081] Furthermore, as an example of the aforementioned affine interpretation, motion information from two control points, i.e., two reference points, may be required.
[0082] Figure 6 illustrates the affine motion model in which motion vectors are used for two control points. The affine motion model using two control points can represent three types of motion, including translational motion, scaling motion, and rotational motion. The affine motion model that represents three types of motion can also be referred to as a similarity affine motion model or a simplified affine motion model.
[0083] If the top-left sample position within block 600 is currently set to (0,0), then, as shown in Figure 6, the (0,0) and (w,0) sample positions can be determined as the control points. Hereafter, the control point for the (0,0) sample position can be denoted as CP0, and the control point for the (w,0) sample position can be denoted as CP1.
[0084] Using the control points and motion vectors for each of the control points described above, a mathematical formula for the affine motion model can be derived. The mathematical formula for the affine motion model can be expressed as follows:
[0085]
number
[0086] Here, w represents the width of the current block 600, and v 0x , v 0y These represent the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y These represent the x and y components of the motion vector of CP1, respectively. Furthermore, x represents the x component of the position of the target sample within the current block 600, and y represents the y component of the position of the target sample within the current block 600, and v x The x-component of the motion vector of the target sample within the current block 600, v y This shows the y component of the motion vector of the target sample currently within block 600.
[0087] On the other hand, the four parameters for equation 3 can be represented by a, b, c, and d, as shown in the following equation, and the equation for the affine motion model represented by the four parameters may be as follows.
[0088]
number
[0089] Here, w represents the width of the current block 600, and v 0x , v 0y These represent the x and y components of the motion vector of CP0, respectively, and v 1x , v 1yThese represent the x and y components of the motion vector of CP1, respectively. Furthermore, x represents the x component of the position of the target sample within the current block 600, and y represents the y component of the position of the target sample within the current block 600, and v x The x-component of the motion vector of the target sample within the current block 600, v y This represents the y-component of the motion vector of the target sample currently in block 600. The affine motion model using the two control points can be expressed by four parameters a, b, c, and d, as shown in equation 4, and the affine motion model or affine interpretation using the four parameters can be expressed as a four-parameter affine motion model or AF4. That is, according to the affine motion model, the motion vector of each sample in the current block can be derived based on the motion vectors of the control points. On the other hand, the set of motion vectors of the samples in the current block derived by the affine motion model can be expressed as an affine motion vector field (MVF).
[0090] On the other hand, as described above, sample-level motion vectors can be derived through the affine motion model, and the accuracy of interpretation can be significantly improved through this. However, in this case, the complexity of the motion compensation process can increase significantly.
[0091] This restricts the derivation of motion vectors to be at the sub-block level within the current block, rather than at the sample level.
[0092] Figure 7 illustrates a method for deriving motion vectors in subblock units based on the affine motion model. Figure 7 illustrates a case where the size of the current block is 16 × 16 and motion vectors are derived in 4 × 4 subblock units. The subblocks can be set to various sizes; for example, if the subblocks are set to n × n size (where n is a positive integer, e.g., n is 4), motion vectors can be derived in n × n subblock units within the current block based on the affine motion model, and various methods can be applied to derive motion vectors representing each subblock.
[0093] For example, as shown in Figure 7, the motion vector of each subblock can be derived using the center or lower right side of the center sample position of each subblock as the representative coordinate. Here, the lower right side of the center position can represent the sample position located on the lower right side among the four samples located at the center of the subblock. For example, when n is odd, one sample can be located in the middle of the subblock, in which case the center sample position can be used to derive the motion vector of the subblock. However, when n is even, four samples can be located adjacent to each other in the center of the subblock, in which case the lower right side sample position can be used to derive the motion vector. For example, as shown in Figure 7, the representative coordinates for each subblock can be derived as (2, 2), (6, 2), (10, 2), ..., (14, 14), and the encoding / decoding device can derive the motion vector of each subblock by substituting each of the representative coordinates of the subblock into equation 1 or 3 described above. The motion vectors of subblocks within the current block, derived via the aforementioned affine motion model, can be represented as affine MVF.
[0094] On the other hand, as an example, the size of the subblock within the current block can also be derived based on the following formula.
[0095]
number
[0096] Here, M represents the width of the subblock, and N represents the height of the subblock. Also, v 0x , v 0y These represent the x and y components of CPMV0 of the current block, respectively, and v 0x , v 0y , where x and , respectively represent the x and y components of the CPMV1 of the current block, w represents the width of the current block, h represents the height of the current block, and MvPre represents the motion vector fraction accuracy. For example, the motion vector fraction accuracy can be set to 1 / 16.
[0097] On the other hand, inter-prediction using the affine motion model described above, i.e., affine motion prediction, can have two modes: affine merge mode (AF_MERGE) and affine inter mode (AF_INTER). Here, the affine inter mode can also be expressed as affine MVP mode (affine motion vector prediction mode, AF_MVP).
[0098] The affine merge mode is similar to existing merge modes in that it does not transmit MVDs for the motion vectors of the control points. That is, the affine merge mode can represent an encoding / decoding method that, like existing skip / merge modes, derives and predicts CPMVs for each of two or three control points from the surrounding blocks of the current block without encoding the MVD (motion vector difference).
[0099] For example, when the AF_MRG mode is applied to the current block, the MVs (i.e., CPMV0 and CPMV1) for CP0 and CP1 can be derived from the surrounding blocks of the current block to which the affine mode is applied. That is, the CPMV0 and CPMV1 of the surrounding blocks to which the affine mode is applied can be derived as merge candidates, and the CPMV0 and CPMV1 for the current block can be derived based on the merge candidates. An affine motion model can be derived based on the CPMV0 and CPMV1 of the surrounding blocks represented by the merge candidates, and the CPMV0 and CPMV1 for the current block can be derived based on the affine motion model.
[0100] The affine intermode can represent an interprediction that derives an MVP (Motion Vector Predictor) for the motion vector of the control point, derives the motion vector of the control point based on the received MVD (motion vector difference) and the MVP, derives an affine MVF of the current block based on the motion vector of the control point, and performs prediction based on the affine MVF. Here, the motion vector of the control point can be represented as CPMV (Control Point Motion Vector), the MVP of the control point as CPMVP (Control Point Motion Vector Predictor), and the MVD of the control point as CPMVD (Control Point Motion Vector Difference). Specifically, for example, an encoding device can derive a CPMVP (Control Point Motion Vector Predictor) and a CPMV (Control Point Motion Vector) for each of CP0 and CP1 (or CP0, CP1, and CP2), and can transmit or store information about the CPMVP and / or the CPMVD, which is the difference between the CPMVP and CPMV.
[0101] Here, when the affine intermode is applied to the current block, the encoding / decoding device can construct an affine MVP candidate list based on the surrounding blocks of the current block, the affine MVP candidates may be referred to as CPMVP pair candidates, and the affine MVP candidate list may also be referred to as a CPMVP candidate list.
[0102] Furthermore, each affine MVP candidate can represent a combination of CPMVPs (CP0 and CP1) in a four-parameter affine motion model, and a combination of CPMVPs (CP0, CP1, and CP2) in a six-parameter affine motion model.
[0103] Figure 8 illustrates a sequence diagram of an affine motion prediction method according to one embodiment of this document.
[0104] As shown in Figure 8, affine motion prediction methods can be broadly described as follows. When the affine motion prediction method starts, a CPMV pair may be obtained first (S800). Here, the CPMV pair may include CPMV0 and CPMV1 when using a 4-parameter affine model.
[0105] Subsequently, affine motion compensation can be performed based on the CPMV pair (S810), and the affine motion prediction can be completed.
[0106] Furthermore, two affine prediction modes may exist to determine CPMV0 and CPMV1. Here, the two affine prediction modes may include an affine intermode and an affine merge mode. The affine intermode can clearly determine CPMV0 and CPMV1 by signaling two motion vector difference (MVD) pieces of information for CPMV0 and CPMV1. In contrast, the affine merge mode can derive a CPMV pair without MVD information signaling.
[0107] In other words, if the affine merge mode allows the CPMV of the current block to be derived using the CPMV of the surrounding blocks encoded in affine mode, and the motion vector is determined on a subblock basis, then the affine merge mode can also be called the subblock merge mode.
[0108] In affine merge mode, the encoder can signal to the decoder an index to an affine-encoded surrounding block for deriving the CPMV of the current block, and can further signal the difference between the CPMV of the surrounding block and the CPMV of the current block. Here, affine merge mode can construct an affine merge candidate list based on the surrounding blocks, and the index to the surrounding block can represent the surrounding block from the affine merge candidate list that is referenced for deriving the CPMV of the current block. The affine merge candidate list can also be called a subblock merge candidate list.
[0109] The affine intermode can also be called the affine MVP mode. In the affine MVP mode, the CPMV of the current block can be derived based on the CPMVP (Control Point Motion Vector Predictor) and CPMVD (Control Point Motion Vector Difference). In other words, the encoding device can determine the CPMVP for the current block's CPMV, derive the CPMVD which is the difference between the current block's CPMV and CPMVP, and signal information about the CPMVP and CPMVD to the decoding device. Here, the affine MVP mode can construct an affine MVP candidate list based on surrounding blocks, and the information about the CPMVP can represent the surrounding blocks from the affine MVP candidate list that are referenced to derive the CPMVP for the current block's CPMV. The affine MVP candidate list can also be called the control point motion vector predictor candidate list.
[0110] For example, when the affine intermode of a 6-parameter affine motion model is applied, the current block can be encoded as described later.
[0111] Figure 9 illustrates a method for deriving a motion vector predictor at a control point according to one embodiment of this document.
[0112] As shown in Figure 9, the motion vector of CP0 in the current block can be represented as v0, the motion vector of CP1 as v1, the motion vector of the control point at the bottom-left sample position as v2, and the motion vector of CP2 as v3. That is, v0 can represent the CPMVP of CP0, v1 as the CPMVP of CP1, and v2 as the CPMVP of CP2.
[0113] The Affine MVP candidate may be a combination of the CPMVP candidate for CP0, the CPMVP candidate for CP1, and the candidate for CP2.
[0114] For example, the aforementioned affine MVP candidate can be derived as follows:
[0115] Specifically, up to 12 combinations of CPMVP candidates can be determined as shown in the following formula.
[0116]
number
[0117] Here, v A The motion vector of the surrounding block A, v B The motion vector of the surrounding block B, v C The motion vector of the surrounding block C, v D The motion vector of the surrounding block D, v E The motion vector of the surrounding block E, v F The motion vector of the surrounding block F, v G This can represent the motion vector of the surrounding block G.
[0118] Furthermore, peripheral block A can represent a peripheral block located at the upper left end of the upper left end sample position of the current block, peripheral block B can represent a peripheral block located at the upper end of the upper left end sample position of the current block, and peripheral block C can represent a peripheral block located to the left of the upper left end sample position of the current block. Furthermore, peripheral block D can represent a peripheral block located at the upper end of the upper right end sample position of the current block, and peripheral block E can represent a peripheral block located at the upper right end of the upper right end sample position of the current block. Furthermore, peripheral block F can represent a peripheral block located to the left of the lower left end sample position of the current block, and peripheral block G can represent a peripheral block located at the lower left end of the lower left end sample position of the current block.
[0119] In other words, referring to the equation 6 above, the CPMVP candidate for CP0 is the motion vector v of the surrounding block A. A , the motion vector v of the surrounding block B B , and / or the motion vector v of the surrounding block C C The CPMVP candidate of CP1 is the motion vector v of the surrounding block D. D , and / or the motion vector v of the surrounding block E E The CPMVP candidate of CP2 is the motion vector v of the surrounding block F. F , and / or the motion vector v of the surrounding block G G It can include...
[0120] In other words, the CPMVP v0 of CP0 can be derived based on the motion vector of at least one of the surrounding blocks A, B, and C of the upper-left sample position. Here, surrounding block A can mean the block located at the upper-left end of the current block's upper-left sample position, surrounding block B can mean the block located at the upper end of the current block's upper-left sample position, and surrounding block C can mean the block located to the left of the current block's upper-left sample position.
[0121] Based on the motion vectors of the surrounding blocks, a combination of up to 12 CPMVP candidates can be derived, including the CPMVP candidate for CP0, the CPMVP candidate for CP1, and the CPMVP candidate for CP2.
[0122] Subsequently, the derived CPMVP candidate combinations are sorted in ascending order of DV, and the top two CPMVP candidate combinations can be derived as the affine MVP candidates.
[0123] The DV of the CPMVP candidate combination can be derived using the following formula.
[0124]
number
[0125] Subsequently, the encoding device can determine the CPMV for each of the affine MVP candidates, compare the Rate Distortion (RD) costs for the CPMVs, and select the affine MVP candidate with the smaller RD cost as the optimal affine MVP candidate for the current block. The encoding device can encode and signal an index and CPMVD pointing to the optimal candidate.
[0126] Furthermore, for example, when affine merge mode is applied, the current block can be encoded as described later.
[0127] Figure 10 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to one embodiment of this document.
[0128] A list of affine merge candidates for the current block can be constructed based on the surrounding blocks of the current block shown in Figure 10. The surrounding blocks may include surrounding block A, surrounding block B, surrounding block C, surrounding block D, and surrounding block E. Surrounding block A may represent the left surrounding block of the current block, surrounding block B may represent the upper surrounding block of the current block, surrounding block C may represent the upper right corner surrounding block of the current block, surrounding block D may represent the lower left corner surrounding block of the current block, and surrounding block E may represent the upper left corner surrounding block of the current block.
[0129] For example, if the size of the current block is W × H, and the x-component and y-component of the top-left sample position of the current block are 0, then the left peripheral block may be a block containing a sample at coordinates (-1, H-1), the upper peripheral block may be a block containing a sample at coordinates (W-1, -1), the upper right corner peripheral block may be a block containing a sample at coordinates (W, -1), the lower left corner peripheral block may be a block containing a sample at coordinates (-1, H), and the upper left corner peripheral block may be a block containing a sample at coordinates (-1, -1).
[0130] Specifically, for example, the encoding device can scan the surrounding blocks A, B, C, D, and E of the current block in a specific scanning order, and can determine the surrounding block that is encoded first in the affine prediction mode in the scanning order as a candidate block for the affine merge mode, i.e., an affine merge candidate. Here, for example, the specific scanning order may be alphabetical. That is, the specific scanning order may be in the order of surrounding block A, surrounding block B, surrounding block C, surrounding block D, and surrounding block E.
[0131] Subsequently, the encoding device can determine the affine motion model of the current block using the CPMV of the determined candidate block, determine the CPMV of the current block based on the affine motion model, and determine the affine MVF of the current block based on the CPMV.
[0132] For example, if surrounding block A is determined to be a candidate block for the current block, it can be encoded as described later.
[0133] Figure 11 shows an example of affine prediction performed when surrounding block A is selected as an affine merge candidate.
[0134] As shown in Figure 11, the encoding device can determine a surrounding block A of the current block as a candidate block and derive an affine motion model of the current block based on the CPMV, v2, and v3 of the surrounding block. Subsequently, the encoding device can determine the CPMV, v0, and v1 of the current block based on the affine motion model. The encoding device can determine the affine MVF based on the CPMV, v0, and v1 of the current block and perform the encoding process for the current block based on the affine MVF.
[0135] On the other hand, in relation to the affine interpretation, the construction of the affine MVP candidate list takes into account inherited affine candidates and constructed affine candidates.
[0136] Here, the inherited affine candidates may be as follows:
[0137] For example, if the surrounding blocks of the current block are affine blocks, and the reference picture of the current block and the reference picture of the surrounding block are the same, the affine MVP pair of the current block can be determined from the affine motion model of the surrounding block. Here, the affine block can represent the block to which the affine interpretation is applied. The inherited affine candidate can represent a CPMVP (e.g., the affine MVP pair) derived based on the affine motion model of the surrounding block.
[0138] Specifically, as an example, the inherited affine candidate can be derived as described later.
[0139] Figure 12 illustrates the peripheral blocks used to derive the inherited affine candidate.
[0140] As shown in Figure 12, the surrounding blocks of the current block may include the left surrounding block A0, the left lower corner surrounding block A1, the upper surrounding block B0, the right upper corner surrounding block B1, and the left upper corner surrounding block B2.
[0141] For example, if the size of the current block is W × H, and the x-component and y-component of the top-left sample position of the current block are 0, then the left peripheral block may be a block containing a sample at coordinates (-1, H-1), the upper peripheral block may be a block containing a sample at coordinates (W-1, -1), the upper right corner peripheral block may be a block containing a sample at coordinates (W, -1), the lower left corner peripheral block may be a block containing a sample at coordinates (-1, H), and the upper left corner peripheral block may be a block containing a sample at coordinates (-1, -1).
[0142] The encoding / decoding device can sequentially check peripheral blocks A0, A1, B0, B1, and B2. If the peripheral blocks are encoded using an affine motion model and the reference picture of the current block is the same as the reference picture of the peripheral blocks, it can derive two or three CPMVs for the current block based on the affine motion model of the peripheral blocks. The CPMVs can be derived as affine MVP candidates for the current block. The affine MVP candidates can represent the inherited affine candidates.
[0143] As an example, up to two inherited affine candidates can be derived based on the surrounding blocks.
[0144] For example, an encoding / decoding device can derive a first affine MVP candidate for the current block based on a first block in the surrounding blocks. Here, the first block can be encoded using an affine motion model, and the reference picture of the first block may be identical to the reference picture of the current block. That is, the first block may be the first block found to satisfy a condition after checking the surrounding blocks in a specific order. The condition can be encoded using an affine motion model, and the reference picture of the block may be identical to the reference picture of the current block.
[0145] Subsequently, the encoding / decoding device can derive a second affine MVP candidate for the current block based on a second block in the surrounding block. Here, the second block can be encoded using an affine motion model, and the reference picture of the second block may be identical to the reference picture of the current block. That is, the second block may be a block that satisfies the second condition found when the surrounding blocks are checked in a specific order. The condition can be encoded using an affine motion model, and the reference picture of the block may be identical to the reference picture of the current block.
[0146] On the other hand, for example, if the number of available inherited affine candidates is less than 2 (i.e., if the number of derived inherited affine candidates is less than 2), a constructed affine candidate may be considered. The constructed affine candidate can be derived as follows.
[0147] Figure 13 illustrates a spatial candidate for the constructed affine candidate.
[0148] As shown in Figure 13, the motion vectors of the surrounding blocks of the current block can be divided into three groups. As shown in Figure 13, the surrounding blocks may include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G.
[0149] Peripheral block A can represent a peripheral block located at the upper left end of the upper left end sample position of the current block, peripheral block B can represent a peripheral block located at the upper end of the upper left end sample position of the current block, and peripheral block C can represent a peripheral block located at the left end of the upper left end sample position of the current block. Furthermore, peripheral block D can represent a peripheral block located at the upper end of the upper right end sample position of the current block, and peripheral block E can represent a peripheral block located at the upper right end of the upper right end sample position of the current block. Furthermore, peripheral block F can represent a peripheral block located at the left end of the lower left end sample position of the current block, and peripheral block G can represent a peripheral block located at the lower left end of the lower left end sample position of the current block.
[0150] For example, the three groups may include S0, S1, and S2, and S0, S1, and S2 can be derived as shown in the following table.
[0151] [Table 1]
[0152] Here, mv A The motion vector of the surrounding block A, mv B The motion vector of the surrounding block B, mv C The motion vector of the surrounding block C, mv D The motion vector of the surrounding block D, mv E The motion vector of the surrounding block E, mvF The motion vector of the surrounding block F, mv G This represents the motion vector of the surrounding block G. S0 can also be represented as the first group, S1 as the second group, and S2 as the third group.
[0153] The encoding / decoding device can derive mv0 from S0, mv1 from S1, mv2 from S2, and derive an affine MVP candidate that includes mv0, mv1, and mv2. The affine MVP candidate can represent the constructed affine candidate. Furthermore, mv0 may be a CPMVP candidate for CP0, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.
[0154] Here, the reference picture for mv0 may be the same as the reference picture of the current block. That is, mv0 may be a motion vector that satisfies the first condition found when checking motion vectors in S0 according to a specific order. The condition is that the reference picture for the motion vector may be the same as the reference picture of the current block. The specific order may be surrounding block A → surrounding block B → surrounding block C in S0. Furthermore, the process may be performed in an order other than that described above, and is not limited to the examples described above.
[0155] Furthermore, the reference picture for mv1 may be the same as the reference picture of the current block. That is, mv1 may be a motion vector that satisfies the first condition found when checking motion vectors in S1 according to a specific order. The condition is that the reference picture for the motion vector may be the same as the reference picture of the current block. The specific order may be the surrounding block D → surrounding block E in S1. Furthermore, the process may be performed in an order other than that described above, and is not limited to the examples described above.
[0156] Furthermore, the reference picture for mv2 may be the same as the reference picture of the current block. That is, mv2 may be a motion vector that satisfies the first condition found when checking motion vectors in S2 according to a specific order. The condition is that the reference picture for the motion vector may be the same as the reference picture of the current block. The specific order may be the surrounding block F → surrounding block G in S2. Furthermore, the process may be performed in an order other than that described above, and is not limited to the examples described above.
[0157] On the other hand, if only mv0 and mv1 are available, that is, if only mv0 and mv1 are derived, then mv2 can be derived as shown in the following formula.
[0158]
number
[0159] Here, mv2 x This represents the x component of mv2, and mv2 y This represents the y component of mv2, and mv0 x This represents the x component of mv0, and mv0 y This represents the y component of mv0, and mv1 x This represents the x component of mv1, and mv1 y represents the y component of mv1. Also, w represents the width of the current block, and h represents the height of the current block.
[0160] On the other hand, if only mv0 and mv2 are derived, mv1 can be derived as shown in the following formula.
[0161]
number
[0162] Here, mv1 x This represents the x component of mv1, and mv1 yThis represents the y component of mv1, and mv0 x This represents the x component of mv0, and mv0 y This represents the y component of mv0, and mv2 x This represents the x component of mv2, and mv2 y represents the y component of mv2. Also, w represents the width of the current block, and h represents the height of the current block.
[0163] Furthermore, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the AMVP process of the existing HEVC standard can be applied to the affine MVP list construction. That is, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the process of constructing MVP candidates in the existing HEVC standard can be performed.
[0164] On the other hand, the sequence diagram of the embodiments and other components that constitute the above-mentioned Affine MVP list is as follows:
[0165] Figure 14 illustrates an example of how an affine MVP list can be constructed.
[0166] As shown in Figure 14, the encoding / decoding device can add an inherited candidate to the affine MVP list of the current block (S1400). The inherited candidate can represent the inherited affine candidate described above.
[0167] Specifically, the encoding / decoding device can derive up to two inherited affine candidates from the surrounding blocks of the current block (S1405). Here, the surrounding blocks may include the left surrounding block A0, the left lower corner surrounding block A1, the upper surrounding block B0, the right upper corner surrounding block B1, and the left upper corner surrounding block B2 of the current block.
[0168] For example, an encoding / decoding device can derive a first affine MVP candidate for the current block based on a first block in the surrounding blocks. Here, the first block can be encoded using an affine motion model, and the reference picture of the first block may be identical to the reference picture of the current block. That is, the first block may be the first block found to satisfy a condition after checking the surrounding blocks in a specific order. The condition can be encoded using an affine motion model, and the reference picture of the block may be identical to the reference picture of the current block.
[0169] Subsequently, the encoding / decoding device can derive a second affine MVP candidate for the current block based on a second block in the surrounding block. Here, the second block can be encoded using an affine motion model, and the reference picture of the second block may be identical to the reference picture of the current block. That is, the second block may be a block that satisfies the second condition found by checking the surrounding blocks in a specific order. The condition can be encoded using an affine motion model, and the reference picture of the block may be identical to the reference picture of the current block.
[0170] On the other hand, the specified order could be left-side perimeter block A0 → left-lower corner perimeter block A1 → upper-side perimeter block B0 → right-upper corner perimeter block B1 → left-upper corner perimeter block B2. Furthermore, the process can be carried out in any order other than those described above, and is not limited to the examples given.
[0171] The encoding / decoding device can add a constructed candidate to the current block's affine MVP list (S1410). The constructed candidate can represent the constructed affine candidate described above. The constructed candidate can also be represented as a constructed affine MVP candidate. If the number of available inherited candidates is less than two, the encoding / decoding device can add a constructed candidate to the current block's affine MVP list. For example, the encoding / decoding device can derive one constructed affine candidate.
[0172] On the other hand, the method for deriving the constructed affine candidate may differ depending on whether the affine motion model applied to the current block is a 6-affine motion model or a 4-affine motion model. The specific details of the method for deriving the constructed candidate will be described later.
[0173] The encoding / decoding device can add HEVC AMVP candidates to the affine MVP list of the current block (S1420). If the number of available inherited and / or constructed candidates is less than two, the encoding / decoding device can add HEVC AMVP candidates to the affine MVP list of the current block. That is, if the number of available inherited and / or constructed candidates is less than two, the encoding / decoding device may perform the process of constructing MVP candidates in the existing HEVC standard.
[0174] On the other hand, the method for deriving the constructed candidate may be as follows:
[0175] For example, if the affine motion model applied to the current block is a 6-affine motion model, the constructed candidate can be derived as shown in the embodiment in Figure 15.
[0176] Figure 15 shows an example of how to derive the constructed candidate.
[0177] As shown in Figure 15, the encoding / decoding device can check mv0, mv1, and mv2 for the current block (S1500). That is, the encoding / decoding device can determine if there are mv0, mv1, and mv2 available in the surrounding blocks of the current block. Here, mv0 may be a candidate CPMVP for CP0 of the current block, mv1 may be a candidate CPMVP for CP1, and mv2 may be a candidate CPMVP for CP2. Furthermore, mv0, mv1, and mv2 can be represented as candidate motion vectors for the CP.
[0178] For example, an encoding / decoding device can check whether the motion vectors of surrounding blocks in a first group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of a surrounding block that satisfies the conditions first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the first group in a specific order. If the motion vector of the surrounding block in the first group does not satisfy the specific conditions, there may be no available mv0. Here, for example, the specific order may be from surrounding block A to surrounding block B and then surrounding block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0179] Furthermore, for example, the encoding / decoding device can check whether the motion vectors of surrounding blocks in the second group satisfy specific conditions according to a specific order. The encoding / decoding device can derive the motion vector of a surrounding block that satisfies the conditions first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the second group according to a specific order. If the motion vector of the surrounding block in the second group does not satisfy the specific conditions, there may be no available mv1. Here, for example, the specific order may be from surrounding block D to surrounding block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0180] Furthermore, for example, the encoding / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy specific conditions according to a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the third group according to a specific order. If the motion vector of the surrounding block in the third group does not satisfy the specific conditions, there may be no available mv2. Here, for example, the specific order may be the order from surrounding block F to surrounding block G in the third group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0181] On the other hand, the first group may include the motion vectors of peripheral block A, peripheral block B, and peripheral block C; the second group may include the motion vectors of peripheral block D and peripheral block E; and the third group may include the motion vectors of peripheral block F and peripheral block G. Peripheral block A may represent a peripheral block located at the upper left end of the upper left end sample position of the current block; peripheral block B may represent a peripheral block located at the upper end of the upper left end sample position of the current block; peripheral block C may represent a peripheral block located at the left end of the upper left end sample position of the current block; peripheral block D may represent a peripheral block located at the upper end of the upper right end sample position of the current block; peripheral block E may represent a peripheral block located at the upper right end of the upper right end sample position of the current block; peripheral block F may represent a peripheral block located at the left end of the lower left end sample position of the current block; and peripheral block G may represent a peripheral block located at the lower left end of the lower left end sample position of the current block.
[0182] If only mv0 and mv1 are available for the current block, that is, if only mv0 and mv1 are derived for the current block, the encoding / decoding device can derive mv2 for the current block based on the above-described formula 8 (S1510). The encoding / decoding device can derive mv2 by substituting the derived mv0 and mv1 into the above-described formula 8.
[0183] If only mv0 and mv2 are available for the current block, that is, if only mv0 and mv2 have been derived for the current block, the encoding / decoding device can derive mv1 for the current block based on the above-described formula 9 (S1520). The encoding / decoding device can derive mv1 by substituting the derived mv0 and mv2 into the above-described formula 9.
[0184] The encoding / decoding device can derive the derived mv0, mv1, and mv2 as constructed candidates for the current block (S1530). If the mv0, mv1, and mv2 are available, that is, if the mv0, mv1, and mv2 are derived based on the surrounding blocks of the current block, the encoding / decoding device can derive the derived mv0, mv1, and mv2 as constructed candidates for the current block.
[0185] Furthermore, if only mv0 and mv1 are available for the current block, that is, if only mv0 and mv1 are derived for the current block, the encoding / decoding device can derive the derived mv0, mv1 and mv2 derived based on the above-described formula 8 as the constructed candidate for the current block.
[0186] Furthermore, if only mv0 and mv2 are available for the current block, that is, if only mv0 and mv2 are derived for the current block, the encoding / decoding device can derive the derived mv0, mv2 and mv1 derived based on the above-described formula 9 as the constructed candidate for the current block.
[0187] Furthermore, for example, if the affine motion model applied to the current block is a 4-affine motion model, the constructed candidate can be derived as shown in the embodiment in Figure 15.
[0188] Figure 16 shows an example of how to derive the constructed candidate.
[0189] As shown in Figure 16, the encoding / decoding device can check mv0, mv1, and mv2 for the current block (S1600). That is, the encoding / decoding device can determine if there are mv0, mv1, and mv2 available in the surrounding blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.
[0190] For example, an encoding / decoding device can check whether the motion vectors of surrounding blocks in a first group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of a surrounding block that satisfies the conditions first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the first group in a specific order. If the motion vector of the surrounding block in the first group does not satisfy the specific conditions, there may be no available mv0. Here, for example, the specific order may be from surrounding block A to surrounding block B and then surrounding block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0191] Furthermore, for example, the encoding / decoding device can check whether the motion vectors of surrounding blocks in the second group satisfy specific conditions according to a specific order. The encoding / decoding device can derive the motion vector of a surrounding block that satisfies the conditions first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the second group according to a specific order. If the motion vector of the surrounding block in the second group does not satisfy the specific conditions, there may be no available mv1. Here, for example, the specific order may be from surrounding block D to surrounding block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0192] Furthermore, for example, the encoding / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy specific conditions according to a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the third group according to a specific order. If the motion vector of the surrounding block in the third group does not satisfy the specific conditions, there may be no available mv2. Here, for example, the specific order may be the order from surrounding block F to surrounding block G in the third group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0193] On the other hand, the first group may include the motion vectors of peripheral block A, peripheral block B, and peripheral block C; the second group may include the motion vectors of peripheral block D and peripheral block E; and the third group may include the motion vectors of peripheral block F and peripheral block G. Peripheral block A may represent a peripheral block located at the upper left end of the upper left end sample position of the current block; peripheral block B may represent a peripheral block located at the upper end of the upper left end sample position of the current block; peripheral block C may represent a peripheral block located at the left end of the upper left end sample position of the current block; peripheral block D may represent a peripheral block located at the upper end of the upper right end sample position of the current block; peripheral block E may represent a peripheral block located at the upper right end of the upper right end sample position of the current block; peripheral block F may represent a peripheral block located at the left end of the lower left end sample position of the current block; and peripheral block G may represent a peripheral block located at the lower left end of the lower left end sample position of the current block.
[0194] If only mv0 and mv1 are available for the current block, or if mv0, mv1, and mv2 are available for the current block, that is, if only mv0 and mv1 are derived for the current block, or if mv0, mv1, and mv2 are derived for the current block, the encoding / decoding device can derive the derived mv0 and mv1 as constructed candidates for the current block (S1610).
[0195] On the other hand, if only mv0 and mv2 are available for the current block, that is, if only mv0 and mv2 are derived for the current block, the encoding / decoding device can derive mv1 for the current block based on the above-described formula 9 (S1620). The encoding / decoding device can derive mv1 by substituting the derived mv0 and mv2 into the above-described formula 9.
[0196] Subsequently, the encoding / decoding device can derive the derived mv0 and mv1 as constructed candidates for the current block (S1610).
[0197] On the other hand, this document proposes other embodiments for deriving the inherited affine candidates. The proposed embodiments can reduce the complexity of the operations in deriving the inherited affine candidates and improve coding performance.
[0198] Figure 17 illustrates the peripheral block locations scanned to derive inherited affine candidates.
[0199] The encoding / decoding device can derive up to two inherited affine candidates from the surrounding blocks of the current block. Figure 17 can represent the surrounding blocks for the inherited affine candidates. For example, the surrounding blocks may include surrounding block A and surrounding block B shown in Figure 17. Surrounding block A may represent the left surrounding block A0 described above, and surrounding block B may represent the upper surrounding block B0 described above.
[0200] For example, an encoding / decoding device can check whether the surrounding blocks are available in a specific order and derive an inherited affine candidate for the current block based on the first available surrounding block identified. That is, an encoding / decoding device can check whether the surrounding blocks satisfy specific conditions in a specific order and derive an inherited affine candidate for the current block based on the first available surrounding block identified. Furthermore, an encoding / decoding device can derive an inherited affine candidate for the current block based on the second surrounding block that satisfies the specific conditions identified. That is, an encoding / decoding device can derive an inherited affine candidate for the current block based on the second surrounding block that satisfies the specific conditions identified. Here, availability is encoded in an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. That is, the specific conditions are encoded in an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. Also, for example, the specific order may be surrounding block A → surrounding block B. On the other hand, a pruning check process between two inherited affine candidates (i.e., derived inherited affine candidates, etc.) may be omitted. This pruning check process can represent a process of checking whether they are identical to each other, and if they are identical candidates, removing the candidate derived in the final order.
[0201] The above-described embodiment proposes a method for deriving the inherited affine candidate by checking only two peripheral blocks (i.e., peripheral block A and peripheral block B) instead of checking all existing peripheral blocks (i.e., peripheral block A, peripheral block B, peripheral block C, peripheral block D, and peripheral block E) to derive the inherited affine candidate. Here, peripheral block C can represent the upper right corner peripheral block B1 described above, peripheral block D can represent the lower left corner peripheral block A1 described above, and peripheral block E can represent the upper left corner peripheral block B2 described above.
[0202] To analyze the spatial correlation between the surrounding blocks and the current block based on affine interpretation, the probability that affine prediction will be applied to the current block if affine prediction is applied to each surrounding block can be referenced. The probability that affine prediction will be applied to the current block if affine prediction is applied to each surrounding block can be derived as shown in the following table.
[0203] [Table 2]
[0204] Referring to Table 2 above, it can be confirmed that peripheral block A and peripheral block B have a high spatial correlation with the current block. Therefore, by using only peripheral blocks A and peripheral block B, which have a high spatial correlation, an embodiment can be obtained that reduces processing time while deriving high decoding performance.
[0205] On the other hand, the pruning check process can be performed to prevent the same candidate from existing in the candidate list. While the pruning check process can eliminate redundancy and thus offer advantages in terms of encoding efficiency, it has the disadvantage of increasing the complexity of the computation. In particular, the pruning check process for affine candidates is computationally very complex because it must be performed on the affine type (e.g., whether the affine motion model is a 4-affine motion model or a 6-affine motion model), the reference picture (or reference picture index), and the MVs of CP0, CP1, and CP2. Therefore, this embodiment proposes a method that does not perform a pruning check process between inherited affine candidates derived based on peripheral block A (e.g., inherited_A) and inherited affine candidates derived based on peripheral block B (e.g., inherited_B). In the case of surrounding blocks A and B, the distance is large, and therefore the spatial correlation is low, making it unlikely that inherited_A and inherited_B are identical. Therefore, it may be appropriate not to perform the pruning check process between the inherited affine candidates.
[0206] Alternatively, a method can be proposed to perform a minimal pruning check process based on the above grounds. For example, an encoding / decoding device can perform a pruning check process by comparing only the MV of the inherited affine candidate's CP0.
[0207] Furthermore, this document proposes a method for deriving constructed candidates that differs from the embodiments described above. The proposed embodiments can reduce complexity and improve coding performance compared to the embodiments for deriving constructed candidates described above. The proposed embodiments are described below. In addition, if the number of available inherited affine candidates is less than 2 (i.e., if the number of derived inherited affine candidates is less than 2), constructed affine candidates may be considered.
[0208] For example, an encoding / decoding device can check mv0, mv1, and mv2 for the current block. That is, the encoding / decoding device can determine if there are mv0, mv1, and mv2 available in the surrounding blocks of the current block. Here, mv0 may be a candidate for the CPMVP of CP0 of the current block, mv1 may be a candidate for the CPMVP of CP1, and mv2 may be a candidate for the CPMVP of CP2.
[0209] Specifically, the surrounding blocks of the current block can be divided into three groups, and these surrounding blocks may include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G. The first group may include the motion vectors of surrounding block A, surrounding block B, and surrounding block C; the second group may include the motion vectors of surrounding block D and surrounding block E; and the third group may include the motion vectors of surrounding block F and surrounding block G. Peripheral block A can represent a peripheral block located at the upper left end of the upper left end sample position of the current block; peripheral block B can represent a peripheral block located at the upper end of the upper left end sample position of the current block; peripheral block C can represent a peripheral block located at the left end of the upper left end sample position of the current block; peripheral block D can represent a peripheral block located at the upper end of the upper right end sample position of the current block; peripheral block E can represent a peripheral block located at the upper right end of the upper right end sample position of the current block; peripheral block F can represent a peripheral block located at the left end of the lower left end sample position of the current block; and peripheral block G can represent a peripheral block located at the lower left end of the lower left end sample position of the current block.
[0210] The encoding / decoding device can determine whether there is an mv0 available in the first group, whether there is an mv1 available in the second group, and whether there is an mv2 available in the third group.
[0211] Specifically, for example, an encoding / decoding device can check whether the motion vectors of peripheral blocks in a first group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of a peripheral block that satisfies the conditions first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the first group in a specific order. If the motion vector of the peripheral block in the first group does not satisfy the specific conditions, there may be no available mv0. Here, for example, the specific order may be from peripheral block A to peripheral block B and then peripheral block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the peripheral block is the same as the reference picture of the current block.
[0212] Furthermore, the encoding / decoding device can check whether the motion vectors of the surrounding blocks in the second group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the second group in a specific order. If the motion vector of the surrounding block in the second group does not satisfy the specific conditions, there may be no available mv1. Here, for example, the specific order may be from surrounding block D to surrounding block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0213] Furthermore, the encoding / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the third group in a specific order. If the motion vector of the surrounding block in the third group does not satisfy the specific conditions, there may be no available mv2. Here, for example, the specific order may be from surrounding block F to surrounding block G in the third group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0214] Subsequently, if the affine motion model applied to the current block is a 4-affine motion model, and mv0 and mv1 are available for the current block, the encoding / decoding device can derive the derived mv0 and mv1 as constructed candidates for the current block. On the other hand, if mv0 and / or mv1 are not available for the current block, i.e., if at least one of mv0 and mv1 is not derived from the surrounding blocks of the current block, the encoding / decoding device may not add the constructed candidates to the affine MVP list of the current block.
[0215] Furthermore, if the affine motion model applied to the current block is a 6-affine motion model, and mv0, mv1, and mv2 are available for the current block, the encoding / decoding device can derive the derived mv0, mv1, and mv2 as constructed candidates for the current block. On the other hand, if mv0, mv1, and / or mv2 are not available for the current block, i.e., if at least one of mv0, mv1, and mv2 is not derived from the surrounding blocks of the current block, the encoding / decoding device may not add constructed candidates to the affine MVP list of the current block.
[0216] The proposed embodiment described above is a method of considering a constructed candidate only if all motion vectors of the CPs for generating the affine motion model of the current block are available. Here, "available" can mean that the reference picture of the surrounding block and the reference picture of the current block are identical. That is, the constructed candidate can be derived only if there are motion vectors among the motion vectors of the surrounding blocks for each of the CPs of the current block that satisfy the above condition. Therefore, if the affine motion model applied to the current block is a 4-affine motion model, the constructed candidate can be considered only if the MVs of CP0 and CP1 of the current block (i.e., mv0 and mv1) are available. Also, if the affine motion model applied to the current block is a 6-affine motion model, the constructed candidate can be considered only if the MVs of CP0, CP1, and CP2 of the current block (i.e., mv0, mv1, and mv2) are available. Therefore, according to the proposed embodiment, additional configuration for deriving the motion vector for CP based on equation 8 or equation 9 described above may not be necessary. This reduces the complexity of the calculations for deriving the constructed candidate. Furthermore, the constructed candidate is determined only if a CPMVP candidate having only the same reference picture is available, thereby improving overall coding performance.
[0217] On the other hand, a pruning check process between the derived inherited affine candidate and the constructed affine candidate may be omitted. This pruning check process may represent a process of checking whether they are identical to each other and, if they are identical candidates, removing the candidate derived in the final order.
[0218] The embodiments described above can be shown in Figures 18 and 19.
[0219] Figure 18 shows an example of deriving the constructed candidate when a 4-affine motion model is applied to the current block.
[0220] As shown in Figure 18, the encoding / decoding device can determine whether mv0 and mv1 are available for the current block (S1800). That is, the encoding / decoding device can determine whether there are available mv0 and mv1 in the surrounding blocks of the current block. Here, mv0 may be a candidate for the CPMVP of CP0 of the current block, and mv1 may be a candidate for the CPMVP of CP1.
[0221] The encoding / decoding device can determine if there is an mv0 available in the first group, and if there is an mv1 available in the second group.
[0222] Specifically, the surrounding blocks of the current block can be divided into three groups, and these surrounding blocks may include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G. The first group may include the motion vectors of surrounding block A, surrounding block B, and surrounding block C; the second group may include the motion vectors of surrounding block D and surrounding block E; and the third group may include the motion vectors of surrounding block F and surrounding block G. Peripheral block A can represent a peripheral block located at the upper left end of the upper left end sample position of the current block; peripheral block B can represent a peripheral block located at the upper end of the upper left end sample position of the current block; peripheral block C can represent a peripheral block located at the left end of the upper left end sample position of the current block; peripheral block D can represent a peripheral block located at the upper end of the upper right end sample position of the current block; peripheral block E can represent a peripheral block located at the upper right end of the upper right end sample position of the current block; peripheral block F can represent a peripheral block located at the left end of the lower left end sample position of the current block; and peripheral block G can represent a peripheral block located at the lower left end of the lower left end sample position of the current block.
[0223] The encoding / decoding device can check whether the motion vectors of the surrounding blocks in the first group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the first group in a specific order. If the motion vector of the surrounding block in the first group does not satisfy the specific conditions, there may be no available mv0. Here, for example, the specific order may be from surrounding block A to surrounding block B and then surrounding block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0224] Furthermore, the encoding / decoding device can check whether the motion vectors of the surrounding blocks in the second group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the second group in a specific order. If the motion vector of the surrounding block in the second group does not satisfy the specific conditions, there may be no available mv1. Here, for example, the specific order may be from surrounding block D to surrounding block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0225] If mv0 and mv1 are available for the current block, that is, if mv0 and mv1 are derived for the current block, the encoding / decoding device can derive the derived mv0 and mv1 as constructed candidates for the current block (S1810). On the other hand, if mv0 and / or mv1 are not available for the current block, that is, if at least one of mv0 and mv1 is not derived from the surrounding blocks of the current block, the encoding / decoding device may not add constructed candidates to the affine MVP list of the current block.
[0226] On the other hand, a pruning check process between the derived inherited affine candidate and the constructed affine candidate may be omitted. This pruning check process may represent a process of checking whether they are identical to each other and, if they are identical candidates, removing the candidate derived in the final order.
[0227] Figure 19 shows an example of deriving the constructed candidate when a 6-affine motion model is applied to the current block.
[0228] As shown in Figure 19, the encoding / decoding device can determine whether mv0, mv1, and mv2 are available for the current block (S1900). That is, the encoding / decoding device can determine whether there are available mv0, mv1, and mv2 in the surrounding blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.
[0229] The encoding / decoding device can determine if there is an mv0 available in the first group, if there is an mv1 available in the second group, and if there is an mv2 available in the third group.
[0230] Specifically, the surrounding blocks of the current block can be divided into three groups, and these surrounding blocks may include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G. The first group may include the motion vectors of surrounding block A, surrounding block B, and surrounding block C; the second group may include the motion vectors of surrounding block D and surrounding block E; and the third group may include the motion vectors of surrounding block F and surrounding block G. Peripheral block A can represent a peripheral block located at the upper left end of the upper left end sample position of the current block; peripheral block B can represent a peripheral block located at the upper end of the upper left end sample position of the current block; peripheral block C can represent a peripheral block located at the left end of the upper left end sample position of the current block; peripheral block D can represent a peripheral block located at the upper end of the upper right end sample position of the current block; peripheral block E can represent a peripheral block located at the upper right end of the upper right end sample position of the current block; peripheral block F can represent a peripheral block located at the left end of the lower left end sample position of the current block; and peripheral block G can represent a peripheral block located at the lower left end of the lower left end sample position of the current block.
[0231] The encoding / decoding device can check whether the motion vectors of the surrounding blocks in the first group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the first group in a specific order. If the motion vector of the surrounding block in the first group does not satisfy the specific conditions, there may be no available mv0. Here, for example, the specific order may be from surrounding block A to surrounding block B and then surrounding block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0232] Furthermore, the encoding / decoding device can check whether the motion vectors of the surrounding blocks in the second group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the second group in a specific order. If the motion vector of the surrounding block in the second group does not satisfy the specific conditions, there may be no available mv1. Here, for example, the specific order may be from surrounding block D to surrounding block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0233] Furthermore, the encoding / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy specific conditions in a specific order. The encoding / decoding device can derive the motion vector of the surrounding block that satisfies the conditions first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific conditions first confirmed after checking the motion vectors in the third group in a specific order. If the motion vector of the surrounding block in the third group does not satisfy the specific conditions, there may be no available mv2. Here, for example, the specific order may be from surrounding block F to surrounding block G in the third group. Also, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.
[0234] If mv0, mv1, and mv2 are available for the current block, that is, if mv0, mv1, and mv2 are derived for the current block, the encoding / decoding device can derive the derived mv0, mv1, and mv2 as constructed candidates for the current block (S1910). On the other hand, if mv0, mv1, and / or mv2 are not available for the current block, that is, if at least one of mv0, mv1, and mv2 is not derived from the surrounding blocks of the current block, the encoding / decoding device may not add constructed candidates to the affine MVP list of the current block.
[0235] On the other hand, the pruning check process between the derived inherited affine candidate and the constructed affine candidate may be omitted.
[0236] On the one hand, when the number of derived affine candidates is less than 2 (i.e., when the number of inherited affine candidates and / or constructed affine candidates is less than 2), HEVC AMVP candidates can be added to the affine MVP list of the current block.
[0237] For example, the HEVC AMVP candidates can be derived in the following order.
[0238] Specifically, when the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is available, the CPMV0 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is available (i.e., the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is derived), a first affine MVP candidate including the CPMV0 of the constructed affine candidate as CPMV0, CPMV1, and CPMV2 can be derived.
[0239] Next, when the number of derived affine candidates is less than 2 and the CPMV1 of the constructed affine candidate is available, the CPMV1 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2 and the CPMV1 of the constructed affine candidate is available (i.e., the number of derived affine candidates is less than 2 and the CPMV1 of the constructed affine candidate is derived), a second affine MVP candidate including the CPMV1 of the constructed affine candidate as CPMV0, CPMV1, and CPMV2 can be derived.
[0240] Next, when the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is available, the CPMV2 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is available (i.e., the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is derived), a third affine MVP candidate including the CPMV2 of the constructed affine candidate as CPMV0, CPMV1, and CPMV2 can be derived.
[0241] Next, when the number of derived affine candidates is less than 2, the HEVC TMVP (Temporal Motion Vector Predictor) can be used as the affine MVP candidate. The HEVC TMVP can be derived based on the motion information of the temporal neighboring blocks of the current block. That is, when the number of derived affine candidates is less than 2, a third affine MVP candidate including the motion vectors of the temporal neighboring blocks of the current block as CPMV0, CPMV1, and CPMV2 can be derived. The temporal neighboring blocks can represent the collocated blocks at the same position within the collocated picture corresponding to the current block.
[0242] Next, when the number of derived affine candidates is less than 2, the zero motion vector (zero MV) can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2, a third affine MVP candidate including the zero motion vector as CPMV0, CPMV1, and CPMV2 can be derived. The zero motion vector can represent a motion vector with a value of 0.
[0243] This is because the step of using the CPMV of the constructed affine candidate can reduce complexity compared to the method of deriving existing HEVC AMVP candidates, since it reuses the MV that has already been considered for the generation of the constructed affine candidate.
[0244] On the other hand, this document proposes other embodiments for deriving the inherited affine candidates.
[0245] In order to derive the inherited affine candidates, affine prediction information of the surrounding blocks is necessary, and specifically, the following affine prediction information is required.
[0246] 1) An affine flag (affine_flag) indicating whether or not affine prediction-based encoding has been applied to the surrounding block.
[0247] 2) Movement information of the surrounding blocks
[0248] When a 4-affine motion model is applied to the peripheral block, the motion information of the peripheral block may include L0 motion information and L1 motion information for CP0, and L0 motion information and L1 motion information for CP1. Furthermore, when a 6-affine motion model is applied to the peripheral block, the motion information of the peripheral block may include L0 motion information and L1 motion information for CP0, and L0 motion information and L1 motion information for CP2. Here, the L0 motion information may represent motion information for L0 (List 0), and the L1 motion information may represent motion information for L1 (List 1). The L0 motion information may include an L0 reference picture index and an L0 motion vector, and the L1 motion information may include an L1 reference picture index and an L1 motion vector.
[0249] As described above, in the case of affine prediction, a large amount of information must be stored, and therefore, this can be a major cause of increased hardware costs in actual implementations in encoding / decoding devices. In particular, when a peripheral block is located above the current block and is a CTU boundary, a line buffer should be used to store the affine prediction-related information of the peripheral block, which can lead to even greater cost problems. This problem can be referred to as the line buffer issue. To address this, this document proposes an embodiment that minimizes hardware costs and derives inherited affine candidates by either not storing or reducing the amount of affine prediction-related information in the line buffer. The proposed embodiment can improve encoding performance by reducing the complexity of the calculations when deriving the inherited affine candidates. On the other hand, for reference, if the line buffer already stores motion information for a 4x4 size block, and the affine prediction-related information is additionally stored, the amount of information to be stored can increase threefold compared to the existing storage amount.
[0250] In this embodiment, no additional information for affine prediction may be stored in the line buffer, and if the information in the line buffer must be referenced for the generation of the inherited affine candidate, the generation of the inherited affine candidate may be restricted.
[0251] Figures 20a and 20b illustrate an exemplary embodiment for deriving the inherited affine candidate.
[0252] As shown in Figure 20a, if the peripheral block B of the current block (i.e., the upper peripheral block of the current block) does not exist in the same CTU as the current block (i.e., the current CTU), the peripheral block B may not be used to generate the inherited affine candidate. On the other hand, peripheral block A also does not exist in the same CTU as the current block, but information about peripheral block A is not stored in the line buffer and can therefore be used to generate the inherited affine candidate. Thus, in this embodiment, the upper peripheral block of the current block can only be used to derive the inherited affine candidate if it is included in the same CTU as the current block. Furthermore, if the upper peripheral block of the current block is not included in the same CTU as the current block, the upper peripheral block may not be used to derive the inherited affine candidate.
[0253] As shown in Figure 20b, the surrounding block B of the current block (i.e., the upper surrounding block of the current block) can reside in the same CTU as the current block. In this case, the encoding / decoding device can generate the inherited affine candidate by referring to the surrounding block B.
[0254] Figure 21 shows an outline of an image encoding method using an encoding device relating to this document. The method disclosed in Figure 21 can be performed by the encoding device disclosed in Figure 2. Specifically, for example, steps S2100 to S2120 in Figure 21 can be performed by the prediction unit of the encoding device, step S2130 can be performed by the subtraction unit of the encoding device, and step S2140 can be performed by the entropy encoding unit of the encoding device. Also, for example, although not shown, the process of deriving a predicted sample for the current block based on the CPMV can be performed by the prediction unit of the encoding device, the process of deriving a residual sample for the current block based on the original sample and the predicted sample for the current block can be performed by the subtraction unit of the encoding device, the process of generating information about the residual for the current block based on the residual sample can be performed by the conversion unit of the encoding device, and the process of encoding the information about the residual can be performed by the entropy encoding unit of the encoding device.
[0255] The encoding device configures a list of candidate affine motion vector predictors (MVPs) for the current block (S2100). The encoding device can configure an affine MVP candidate list that includes the affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list may be 2.
[0256] Furthermore, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The encoding device can check whether inherited affine MVP candidates for the current block are available, and if so, the inherited affine MVP candidates may be derived. For example, the inherited affine MVP candidates may be derived based on the surrounding blocks of the current block, and the maximum number of inherited affine MVP candidates may be two. The surrounding blocks may be checked for availability in a specific order, and the inherited affine MVP candidates may be derived based on the checked available surrounding blocks. That is, the surrounding blocks may be checked for availability in a specific order, and a first inherited affine MVP candidate may be derived based on the first checked available surrounding block, and a second inherited affine MVP candidate may be derived based on the second checked available surrounding block. The aforementioned availability can be encoded in an affine motion model, and the reference picture of the surrounding block can be the same as the reference picture of the current block. That is, an available surrounding block can be encoded in an affine motion model (i.e., an affine prediction is applied), and its reference picture can be the same surrounding block as the reference picture of the current block. Specifically, the encoding device can derive a motion vector for the current block relative to the CP based on the affine motion model of the first checked available surrounding block, and can derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The encoding device can also derive a motion vector for the current block relative to the CP based on the affine motion model of the second checked available surrounding block, and can derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model can be derived as shown in Equation 1 or Equation 3 above.
[0257] In other words, the surrounding blocks can be checked to see if they satisfy specific conditions in a specific order, and the inherited affine MVP candidates can be derived based on the checked surrounding blocks that satisfy the specific conditions. That is, the surrounding blocks can be checked to see if they satisfy the specific conditions in a specific order, and a first inherited affine MVP candidate can be derived based on the first checked surrounding block that satisfies the specific conditions, and a second inherited affine MVP candidate can be derived based on the second checked surrounding block that satisfies the specific conditions. Specifically, the encoding device can derive a motion vector of the current block relative to the CP based on the affine motion model of the first checked surrounding block that satisfies the specific conditions, and can derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The encoding device can also derive a motion vector of the current block relative to the CP based on the affine motion model of the second checked surrounding block that satisfies the specific conditions, and can derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model can be derived as shown in Equation 1 or Equation 3 above. On the other hand, the specific condition can be encoded in the affine motion model and represent that the reference picture of the surrounding block is the same as the reference picture of the current block. That is, a surrounding block that satisfies the specific condition can be encoded in the affine motion model (i.e., affine prediction is applied) and its reference picture may be the same as the reference picture of the current block.
[0258] Here, for example, the surrounding blocks may include the surrounding block to the left of the current block, the surrounding block above, the surrounding block at the upper right corner, the surrounding block at the lower left corner, and the surrounding block at the upper left corner. In this case, the specific order may be from the surrounding block on the left to the surrounding block at the lower left corner, the surrounding block above, the surrounding block at the upper right corner, and the surrounding block at the upper left corner.
[0259] Alternatively, for example, the peripheral block may include only the left peripheral block and the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block.
[0260] Alternatively, for example, the peripheral block may include the left peripheral block, and if the upper peripheral block is included in the current CTU which includes the current block, the peripheral block may further include the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block. Also, if the upper peripheral block is not included in the current CTU, the peripheral block may not include the upper peripheral block. In this case, only the left peripheral block may be checked.
[0261] On the other hand, if the size is W×H and the x-component and y-component of the top-left sample position of the current block are 0, then the block around the lower left corner may be a block containing a sample at coordinates (-1, H), the left peripheral block may be a block containing a sample at coordinates (-1, H-1), the upper right corner peripheral block may be a block containing a sample at coordinates (W, -1), the upper peripheral block may be a block containing a sample at coordinates (W-1, -1), and the upper left corner peripheral block may be a block containing a sample at coordinates (-1, -1). That is, the left peripheral block may be the leftmost peripheral block among the left peripheral blocks of the current block, and the upper peripheral block may be the leftmost peripheral block among the upper peripheral blocks of the current block.
[0262] Also, as an example, when a constructed affine MVP candidate is available, the affine MVP candidate list can include the constructed affine MVP candidate. The encoding device can check whether the constructed affine MVP candidate of the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate can be derived. Also, for example, after the inherited affine MVP candidate is derived, the constructed affine MVP candidate can be derived. If the number of derived affine MPV candidates (i.e., the inherited affine MVP candidates) is less than 2 and the constructed affine MVP candidate is available, the affine MVP candidate list can include the constructed affine MVP candidate. Here, the constructed affine MVP candidate can include candidate motion vectors for the CP. The constructed affine MVP candidate can be available when all the candidate motion vectors are available.
[0263] For example, when a 4 - affine motion model is applied to the current block, the CP of the current block can include CP0 and CP1. If the candidate motion vector for CP0 is available and the candidate motion vector for CP1 is available, the constructed affine MVP candidate may be available, and the affine MVP candidate list can include the constructed affine MVP candidate. Here, CP0 can represent the upper - left position of the current block, and CP1 can represent the upper - right position of the current block.
[0264] The constructed affine MVP candidate may include a candidate motion vector for CP0 and a candidate motion vector for CP1. The candidate motion vector for CP0 may be the motion vector of a first block, and the candidate motion vector for CP1 may be the motion vector of a second block.
[0265] Furthermore, the first block checks the surrounding blocks in the first group according to a first specific order, and the first reference picture identified may be the same block as the reference picture of the current block. That is, the candidate motion vector for CP1 may be the motion vector of the same block as the reference picture of the current block, after checking the surrounding blocks in the first group according to a first order and the first reference picture identified. The availability can indicate that the surrounding block exists and that the surrounding block is encoded by interpretation. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. Also, for example, the first group may include surrounding block A, surrounding block B, and surrounding block C, and the first specific order may be from surrounding block A to surrounding block B and then surrounding block C.
[0266] Furthermore, the second block checks the surrounding blocks in the second group according to a second specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. Also, for example, the second group may include surrounding block D and surrounding block E, and the second specific order may be from surrounding block D to surrounding block E.
[0267] On the other hand, if the size of the current block is W × H, and the x-component and y-component of the top-left sample position of the current block are 0, then the surrounding block A may be a block containing a sample at (-1, -1) coordinates, the surrounding block B may be a block containing a sample at (0, -1) coordinates, the surrounding block C may be a block containing a sample at (-1, 0) coordinates, the surrounding block D may be a block containing a sample at (W-1, -1) coordinates, and the surrounding block E may be a block containing a sample at (W, -1) coordinates. That is, surrounding block A may be the upper left corner surrounding block of the current block, surrounding block B may be the uppermost upper surrounding block of the upper surrounding blocks of the current block, surrounding block C may be the uppermost left surrounding block of the left surrounding blocks of the current block, surrounding block D may be the uppermost upper surrounding block of the upper surrounding blocks of the current block, and surrounding block E may be the upper right corner surrounding block of the current block.
[0268] On the other hand, if at least one of the candidate motion vectors of CP0 and CP1 is unavailable, the constructed affine MVP candidate may be unavailable.
[0269] Alternatively, for example, if a 6-affine motion model is applied to the current block, the CP of the current block may include CP0, CP1, and CP2. If candidate motion vectors for CP0, CP1, and CP2 are available, then the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, CP0 may represent the upper left corner of the current block, CP1 may represent the upper right corner of the current block, and CP2 may represent the lower left corner of the current block.
[0270] The constructed affine MVP candidate may include a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2. The candidate motion vector for CP0 may be the motion vector of a first block, the candidate motion vector for CP1 may be the motion vector of a second block, and the candidate motion vector for CP2 may be the motion vector of a third block.
[0271] Furthermore, the first block checks the surrounding blocks in the first group according to a first specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, a candidate motion vector for CP0 may be available. Also, for example, the first group may include surrounding block A, surrounding block B, and surrounding block C, and the first specific order may be from surrounding block A to surrounding block B and then to surrounding block C.
[0272] Furthermore, the second block checks the surrounding blocks in the second group according to a second specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. Also, for example, the second group may include surrounding block D and surrounding block E, and the second specific order may be from surrounding block D to surrounding block E.
[0273] Furthermore, the third block checks the surrounding blocks in the third group according to a third specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the third block in the third group is the same as the reference picture of the current block, a candidate motion vector for CP2 may be available. Also, for example, the third group may include surrounding block F and surrounding block G, and the third specific order may be from surrounding block F to surrounding block G.
[0274] On the other hand, if the size of the current block is W × H, and the x-component and y-component of the top-left sample position of the current block are 0, then peripheral block A may be a block containing a sample at (-1, -1) coordinates, peripheral block B may be a block containing a sample at (0, -1) coordinates, peripheral block C may be a block containing a sample at (-1, 0) coordinates, peripheral block D may be a block containing a sample at (W-1, -1) coordinates, peripheral block E may be a block containing a sample at (W, -1) coordinates, peripheral block F may be a block containing a sample at (-1, H-1) coordinates, and peripheral block G may be a block containing a sample at (-1, H) coordinates. That is, surrounding block A may be the upper left corner surrounding block of the current block, surrounding block B may be the uppermost upper surrounding block among the upper surrounding blocks of the current block, surrounding block C may be the uppermost left surrounding block among the left surrounding blocks of the current block, surrounding block D may be the uppermost upper surrounding block among the upper surrounding blocks of the current block, surrounding block E may be the upper right corner surrounding block of the current block, surrounding block F may be the lowermost left surrounding block among the left surrounding blocks of the current block, and surrounding block G may be the lower left corner surrounding block of the current block.
[0275] On the other hand, if at least one of the candidate motion vectors for CP0, CP1, and CP2 is unavailable, the constructed affine MVP candidate may be unavailable.
[0276] Subsequently, the affine MVP candidate list can be derived based on the steps in the order described later.
[0277] For example, if the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, the encoding device can derive a first affine MVP candidate. Here, the first affine MVP candidate may be an affine MVP candidate that includes the motion vector for CP0 as a candidate motion vector for CP.
[0278] Furthermore, for example, if the number of derived affine MVP candidates is less than two and the motion vector for CP1 is available, the encoding device can derive a second affine MVP candidate. Here, the second affine MVP candidate may be an affine MVP candidate that includes the motion vector for CP1 as a candidate motion vector for CP.
[0279] Furthermore, for example, if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the encoding device can derive a third affine MVP candidate. Here, the third affine MVP candidate may be an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for CP.
[0280] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the encoding device can derive a fourth affine MVP candidate that includes a temporal MVP derived based on the temporally surrounding blocks of the current block as a candidate motion vector for the CP. The temporally surrounding blocks can represent a collocated block within a collocated picture corresponding to the current block. The temporal MVP can be derived based on the motion vector of the temporally surrounding blocks.
[0281] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the encoding device can derive a fifth affine MVP candidate that includes a zero motion vector as a candidate motion vector for the CP. The zero motion vector can represent a motion vector with a value of 0.
[0282] The encoding device derives CPMVP (Control Point Motion Vector Predictors) for the Control Point (CP) of the current block based on the affine MVP candidate list (S2110). The encoding device can derive a CPMV for the CP of the current block that has the optimal RD cost, and can select the affine MVP candidate that is most similar to the CPMV from among the affine MVP candidates as the affine MVP candidate for the current block. The encoding device can derive CPMVP (Control Point Motion Vector Predictors) for the Control Point (CP) of the current block based on the selected affine MVP candidate from among the affine MVP candidates included in the affine MVP candidate list. Specifically, if an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP for CP0, and the candidate motion vector for CP1 of the affine MVP candidate can be derived as the CPMVP for CP1. Furthermore, if an affine MVP candidate includes candidate motion vectors for CP0, candidate motion vectors for CP1, and candidate motion vectors for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived using the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived using the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived using the CPMVP of CP2.
[0283] The encoding device can encode an affine MVP candidate index that points to the selected affine MVP candidate from among the affine MVP candidates. The affine MVP candidate index can point to one of the affine MVP candidates included in the list of affine motion vector predictor (MVP) candidates for the current block.
[0284] The encoding device derives the CPMV for the CP of the current block (S2120). The encoding device can derive the CPMV for each of the CPs of the current block.
[0285] The encoding device derives the CPMVD (Control Point Motion Vector Differences) for the current block with respect to the CP based on the CPMVP and CPMV (S2130). The encoding device can derive the CPMVD for the current block with respect to the CP based on the CPMVP and CPMV for each of the CPs.
[0286] The encoding device encodes motion prediction information, which includes information about the CPMVD (S2140). The encoding device can output the motion prediction information, which includes information about the CPMVD, in bitstream format. That is, the encoding device can output image information, which includes the motion prediction information, in bitstream format. The encoding device can encode information about the CPMVD for each of the CPs, and the motion prediction information may include information about the CPMVD.
[0287] Furthermore, the motion prediction information may include the affine MVP candidate index. The affine MVP candidate index may refer to the selected affine MVP candidate from among the affine MVP candidates included in the list of affine motion vector predictor (MVP) candidates for the current block.
[0288] On the other hand, as an example, the encoding device can derive predicted samples for the current block based on the CPMV, derive residual samples for the current block based on the original samples and predicted samples for the current block, generate information about the residual for the current block based on the residual samples, and encode the information about the residual. The image information may include the information about the residual.
[0289] On the other hand, the bitstream can be transmitted to a decoding device via a network or a (digital) storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0290] Figure 22 shows an schematic of an encoding device that performs the image encoding method relating to this document. The method disclosed in Figure 21 can be performed by the encoding device disclosed in Figure 22. Specifically, for example, the prediction unit of the encoding device in Figure 22 can perform S2100 to S2130 of Figure 21, and the entropy encoding unit of the encoding device in Figure 22 can perform S2140 of Figure 21. Also, for example, although not shown, the process of deriving a predicted sample for the current block based on the CPMV can be performed by the prediction unit of the encoding device in Figure 22, the process of deriving a residual sample for the current block based on the original sample and the predicted sample for the current block can be performed by the subtraction unit of the encoding device in Figure 22, the process of generating information about the residual for the current block based on the residual sample can be performed by the conversion unit of the encoding device in Figure 22, and the process of encoding the information about the residual can be performed by the entropy encoding unit of the encoding device in Figure 22.
[0291] Figure 23 shows an outline of an image decoding method using a decoding device relating to this document. The method disclosed in Figure 23 can be performed by the decoding device disclosed in Figure 3. Specifically, for example, S2300 in Figure 23 can be performed by the entropy decoding unit of the decoding device, S2310 to S2350 can be performed by the prediction unit of the decoding device, and S2360 can be performed by the addition unit of the decoding device. Also, for example, although not shown, the process of obtaining information about the current block's residual via a bitstream can be performed by the entropy decoding unit of the decoding device, and the process of deriving the residual sample for the current block based on the residual information can be performed by the inverse transform unit of the decoding device.
[0292] The decoding device obtains motion prediction information for the current block from the bitstream (S2300). The decoding device can obtain image information including the motion prediction information from the bitstream.
[0293] Furthermore, for example, the motion prediction information may include information regarding the Control Point Motion Vector Differences (CPMVD) for the Control Points (CP) of the current block. That is, the motion prediction information may include information regarding the CPMVD for each of the Control Points of the current block.
[0294] Furthermore, for example, the motion prediction information may include an affine MVP candidate index for the current block. The affine MVP candidate index may point to one of the affine MVP candidates included in the list of affine motion vector predictor (MVP) candidates for the current block.
[0295] The decoding device constructs a list of candidate affine motion vector predictors (MVPs) for the current block (S2310). The decoding device can construct an affine MVP candidate list that includes affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list may be 2.
[0296] Furthermore, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The decoding device can check whether inherited affine MVP candidates for the current block are available, and if so, the inherited affine MVP candidates may be derived. For example, the inherited affine MVP candidates may be derived based on the surrounding blocks of the current block, and the maximum number of inherited affine MVP candidates may be two. The surrounding blocks may be checked for availability in a specific order, and the inherited affine MVP candidates may be derived based on the checked available surrounding blocks. That is, the surrounding blocks may be checked for availability in a specific order, and a first inherited affine MVP candidate may be derived based on the first checked available surrounding block, and a second inherited affine MVP candidate may be derived based on the second checked available surrounding block. The aforementioned availability can be encoded using an affine motion model, and the reference picture of the surrounding block can be the same as the reference picture of the current block. That is, an available surrounding block can be encoded using an affine motion model (i.e., an affine prediction is applied), and its reference picture can be the same surrounding block as the reference picture of the current block. Specifically, the decoding device can derive a motion vector for the current block relative to the CP based on the affine motion model of the first checked available surrounding block, and can derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The decoding device can also derive a motion vector for the current block relative to the CP based on the affine motion model of the second checked available surrounding block, and can derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model can be derived as shown in Equation 1 or Equation 3 above.
[0297] In other words, the surrounding blocks can be checked to see if they satisfy specific conditions in a specific order, and the inherited affine MVP candidates can be derived based on the checked surrounding blocks that satisfy the specific conditions. That is, the surrounding blocks can be checked to see if they satisfy the specific conditions in a specific order, and a first inherited affine MVP candidate can be derived based on the first checked surrounding block that satisfies the specific conditions, and a second inherited affine MVP candidate can be derived based on the second checked surrounding block that satisfies the specific conditions. Specifically, the decoding device can derive a motion vector of the current block relative to the CP based on the affine motion model of the first checked surrounding block that satisfies the specific conditions, and can derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The decoding device can also derive a motion vector of the current block relative to the CP based on the affine motion model of the second checked surrounding block that satisfies the specific conditions, and can derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model can be derived as shown in Equation 1 or Equation 3 above. On the other hand, the specific condition can be encoded in the affine motion model and represent that the reference picture of the surrounding block is the same as the reference picture of the current block. That is, a surrounding block that satisfies the specific condition can be encoded in the affine motion model (i.e., affine prediction is applied) and its reference picture may be the same as the reference picture of the current block.
[0298] Here, for example, the surrounding blocks may include the surrounding block to the left of the current block, the surrounding block above, the surrounding block at the upper right corner, the surrounding block at the lower left corner, and the surrounding block at the upper left corner. In this case, the specific order may be from the surrounding block on the left to the surrounding block at the lower left corner, the surrounding block above, the surrounding block at the upper right corner, and the surrounding block at the upper left corner.
[0299] Alternatively, for example, the peripheral block may include only the left peripheral block and the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block.
[0300] Alternatively, for example, the peripheral block may include the left peripheral block, and if the upper peripheral block is included in the current CTU which includes the current block, the peripheral block may further include the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block. Also, if the upper peripheral block is not included in the current CTU, the peripheral block may not include the upper peripheral block. In this case, only the left peripheral block may be checked. That is, if the upper peripheral block of the current block is included in the current CTU (Coding tree unit) which includes the current block, the upper peripheral block may be used for the inherited affine MVP candidate derivation, and if the upper peripheral block of the current block is not included in the current CTU, the upper peripheral block may not be used for the inherited affine MVP candidate derivation.
[0301] On the other hand, if the size is W×H and the x-component and y-component of the top-left sample position of the current block are 0, then the block around the lower left corner may be a block containing a sample at coordinates (-1, H), the left peripheral block may be a block containing a sample at coordinates (-1, H-1), the upper right corner peripheral block may be a block containing a sample at coordinates (W, -1), the upper peripheral block may be a block containing a sample at coordinates (W-1, -1), and the upper left corner peripheral block may be a block containing a sample at coordinates (-1, -1). That is, the left peripheral block may be the leftmost peripheral block among the left peripheral blocks of the current block, and the upper peripheral block may be the leftmost peripheral block among the upper peripheral blocks of the current block.
[0302] Furthermore, as an example, if a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. The decoding device can check if a constructed affine MVP candidate is available for the current block, and if so, the constructed affine MVP candidate may be derived. Also, for example, the constructed affine MVP candidate may be derived after the inherited affine MVP candidate has been derived. If the number of derived affine MVP candidates (i.e., the inherited affine MVP candidates) is less than two and a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. Here, the constructed affine MVP candidate may include candidate motion vectors for the CP. The constructed affine MVP candidate can be available if all of the candidate motion vectors are available.
[0303] For example, if a 4-affine motion model is applied to the current block, the CP of the current block may include CP0 and CP1. If candidate motion vectors for CP0 and candidate motion vectors for CP1 are available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, CP0 may represent the upper left corner position of the current block, and CP1 may represent the upper right corner position of the current block.
[0304] The constructed affine MVP candidate may include a candidate motion vector for CP0 and a candidate motion vector for CP1. The candidate motion vector for CP0 may be the motion vector of a first block, and the candidate motion vector for CP1 may be the motion vector of a second block.
[0305] Furthermore, the first block checks the surrounding blocks in the first group according to a first specific order, and the first reference picture identified may be the same block as the reference picture of the current block. That is, the candidate motion vector for CP1 may be the motion vector of the same block as the reference picture of the current block, after checking the surrounding blocks in the first group according to a first order and the first reference picture identified. The availability can indicate that the surrounding block exists and that the surrounding block is encoded by interpretation. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. Also, for example, the first group may include surrounding block A, surrounding block B, and surrounding block C, and the first specific order may be from surrounding block A to surrounding block B and then surrounding block C.
[0306] Furthermore, the second block checks the surrounding blocks in the second group according to a second specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. Also, for example, the second group may include surrounding block D and surrounding block E, and the second specific order may be from surrounding block D to surrounding block E.
[0307] On the other hand, if the size of the current block is W × H, and the x-component and y-component of the top-left sample position of the current block are 0, then the surrounding block A may be a block containing a sample at (-1, -1) coordinates, the surrounding block B may be a block containing a sample at (0, -1) coordinates, the surrounding block C may be a block containing a sample at (-1, 0) coordinates, the surrounding block D may be a block containing a sample at (W-1, -1) coordinates, and the surrounding block E may be a block containing a sample at (W, -1) coordinates. That is, surrounding block A may be the upper left corner surrounding block of the current block, surrounding block B may be the uppermost upper surrounding block of the upper surrounding blocks of the current block, surrounding block C may be the uppermost left surrounding block of the left surrounding blocks of the current block, surrounding block D may be the uppermost upper surrounding block of the upper surrounding blocks of the current block, and surrounding block E may be the upper right corner surrounding block of the current block.
[0308] On the other hand, if at least one of the candidate motion vectors of CP0 and CP1 is unavailable, the constructed affine MVP candidate may be unavailable.
[0309] Alternatively, for example, if a 6-affine motion model is applied to the current block, the CP of the current block may include CP0, CP1, and CP2. If candidate motion vectors for CP0, CP1, and CP2 are available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, CP0 may represent the upper left corner of the current block, CP1 may represent the upper right corner of the current block, and CP2 may represent the lower left corner of the current block.
[0310] The constructed affine MVP candidate may include a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2. The candidate motion vector for CP0 may be the motion vector of a first block, the candidate motion vector for CP1 may be the motion vector of a second block, and the candidate motion vector for CP2 may be the motion vector of a third block.
[0311] Furthermore, the first block checks the surrounding blocks in the first group according to a first specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, a candidate motion vector for CP0 may be available. Also, for example, the first group may include surrounding block A, surrounding block B, and surrounding block C, and the first specific order may be from surrounding block A to surrounding block B and then to surrounding block C.
[0312] Furthermore, the second block checks the surrounding blocks in the second group according to a second specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. Also, for example, the second group may include surrounding block D and surrounding block E, and the second specific order may be from surrounding block D to surrounding block E.
[0313] Furthermore, the third block checks the surrounding blocks in the third group according to a third specific order, and the first reference picture identified may be the same block as the reference picture of the current block. Here, if the reference picture of the third block in the third group is the same as the reference picture of the current block, a candidate motion vector for CP2 may be available. Also, for example, the third group may include surrounding block F and surrounding block G, and the third specific order may be from surrounding block F to surrounding block G.
[0314] On the other hand, if the size of the current block is W × H, and the x-component and y-component of the top-left sample position of the current block are 0, then peripheral block A may be a block containing a sample at (-1, -1) coordinates, peripheral block B may be a block containing a sample at (0, -1) coordinates, peripheral block C may be a block containing a sample at (-1, 0) coordinates, peripheral block D may be a block containing a sample at (W-1, -1) coordinates, peripheral block E may be a block containing a sample at (W, -1) coordinates, peripheral block F may be a block containing a sample at (-1, H-1) coordinates, and peripheral block G may be a block containing a sample at (-1, H) coordinates. That is, surrounding block A may be the upper left corner surrounding block of the current block, surrounding block B may be the uppermost upper surrounding block among the upper surrounding blocks of the current block, surrounding block C may be the uppermost left surrounding block among the left surrounding blocks of the current block, surrounding block D may be the uppermost upper surrounding block among the upper surrounding blocks of the current block, surrounding block E may be the upper right corner surrounding block of the current block, surrounding block F may be the lowermost left surrounding block among the left surrounding blocks of the current block, and surrounding block G may be the lower left corner surrounding block of the current block.
[0315] On the other hand, if at least one of the candidate motion vectors for CP0, CP1, and CP2 is unavailable, the constructed affine MVP candidate may be unavailable.
[0316] On the other hand, a pruning check between the inherited affine MVP candidate and the constructed affine MVP candidate may be omitted. The pruning check can represent a process of checking whether the constructed affine MVP candidate is identical to the inherited affine MVP candidate, and if they are identical, not deriving the constructed affine MVP candidate.
[0317] Subsequently, the affine MVP candidate list can be derived based on the steps in the order described later.
[0318] For example, if the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, the decoding device can derive a first affine MVP candidate. Here, the first affine MVP candidate may be an affine MVP candidate that includes the motion vector for CP0 as a candidate motion vector for CP.
[0319] Furthermore, for example, if the number of derived affine MVP candidates is less than two and the motion vector for CP1 is available, the decoding device can derive a second affine MVP candidate. Here, the second affine MVP candidate may be an affine MVP candidate that includes the motion vector for CP1 as a candidate motion vector for CP.
[0320] Furthermore, for example, if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the decoding device can derive a third affine MVP candidate. Here, the third affine MVP candidate may be an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for CP.
[0321] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the decoding device can derive a fourth affine MVP candidate that includes a temporal MVP derived based on the temporally surrounding blocks of the current block as a candidate motion vector for the CP. The temporally surrounding blocks can represent a collocated block within a collocated picture corresponding to the current block. The temporal MVP can be derived based on the motion vector of the temporally surrounding blocks.
[0322] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the decoding device can derive a fifth affine MVP candidate that includes a zero motion vector as a candidate motion vector for the CP. The zero motion vector can represent a motion vector with a value of 0.
[0323] The decoding device derives the CPMVP (Control Point Motion Vector Predictors) for the current block's CP (Control Point) based on the affine MVP candidate list (S2320).
[0324] The decoding device can select a specific affine MVP candidate from among the affine MVP candidates included in the affine MVP candidate list, and derive the selected affine MVP candidate using CPMVP for the CP of the current block. For example, the decoding device can obtain the affine MVP candidate index for the current block from the bitstream, and derive the affine MVP candidate pointed to by the affine MVP candidate index from among the affine MVP candidates included in the affine MVP candidate list using CPMVP for the CP of the current block. Specifically, if an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate can be derived using CPMVP for CP0, and the candidate motion vector for CP1 of the affine MVP candidate can be derived using CPMVP for CP1. Furthermore, if an affine MVP candidate includes candidate motion vectors for CP0, candidate motion vectors for CP1, and candidate motion vectors for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived using the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived using the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived using the CPMVP of CP2.
[0325] The decoding device derives the CPMVD (Control Point Motion Vector Differences) for the current block with respect to the CP based on the motion prediction information (S2330). The motion prediction information may include information regarding the CPMVD for each of the CPs, and the decoding device can derive the CPMVD for each of the current block with respect to the CP based on the information regarding the CPMVD for each of the CPs.
[0326] The decoding device derives the CPMV (Control Point Motion Vectors) for the current block's CP based on the CPMVP and CPMVD (S2340). The decoding device can derive the CPMV for each CP based on the CPMVP and CPMVD for each of the CPs. For example, the decoding device can derive the CPMV for each CP by adding the CPMVP and CPMVD for each CP.
[0327] The decoding device derives predicted samples for the current block based on the CPMV (S2350). The decoding device can derive motion vectors for each subblock or sample of the current block based on the CPMV. That is, the decoding device can derive motion vectors for each subblock or each sample of the current block based on the CPMV. The motion vectors for each subblock or sample can be derived based on equation 1 or equation 3 described above. The motion vectors can be represented as an affine motion vector field (MVF) or a motion vector array.
[0328] The decoding device can derive predicted samples for the current block based on the motion vectors of the subblocks or samples. The decoding device can derive a reference region in a reference picture based on the motion vectors of the subblocks or samples, and can generate predicted samples for the current block based on the restored samples in the reference region.
[0329] The decoding device generates a restored picture for the current block based on the derived predicted sample (S2360). The decoding device can generate a restored picture for the current block based on the derived predicted sample. The decoding device can use the predicted sample immediately as a restored sample in prediction mode, or it can generate a restored sample by adding a residual sample to the predicted sample. If a residual sample for the current block exists, the decoding device can obtain information about the residual for the current block from the bitstream. The information about the residual may include conversion coefficients for the residual sample. The decoding device can derive the residual sample (or a residual sample array) for the current block based on the residual information. The decoding device can generate a restored sample based on the predicted sample and the residual sample, and can derive a restored block or restored picture based on the restored sample. Subsequently, as described above, the decoding device can apply deblocking filtering and / or in-loop filtering procedures such as the SAO procedure to the restored picture to improve subjective / objective image quality as needed.
[0330] Figure 24 shows a schematic diagram of a decoding device that performs the image decoding method relating to this document. The method disclosed in Figure 23 can be performed by the decoding device disclosed in Figure 24. Specifically, for example, the entropy decoding unit of the decoding device in Figure 24 can perform S2300 in Figure 23, the prediction unit of the decoding device in Figure 24 can perform S2310 to S2350 in Figure 23, and the addition unit of the decoding device in Figure 24 can perform S2360 in Figure 23. In addition, for example, although not shown, the process of acquiring image information including information about the residual of the current block via a bitstream can be performed by the entropy decoding unit of the decoding device in Figure 24, and the process of deriving the residual sample for the current block based on the residual information can be performed by the inverse transform unit of the decoding device in Figure 24.
[0331] According to the document mentioned above, the efficiency of image coding based on affine motion prediction can be improved.
[0332] Furthermore, according to this document, when deriving the affine MVP candidate list, a constructed affine MVP candidate can be added only if all candidate motion vectors for the CP of the constructed affine MVP candidate are available. This reduces the complexity of the process of deriving the constructed affine MVP candidate and the process of constructing the affine MVP candidate list, thereby improving coding efficiency.
[0333] Furthermore, according to this document, when deriving the affine MVP candidate list, additional affine MVP candidates can be derived based on the candidate motion vectors for the CP derived in the process of deriving the constructed affine MVP candidates. This reduces the complexity of the process of constructing the affine MVP candidate list and improves coding efficiency.
[0334] Furthermore, according to this document, the inherited affine MVP candidate can be derived using the upper peripheral block only if the upper peripheral block is currently included in the CTU during the process of deriving the inherited affine MVP candidate. This reduces the amount of line buffer storage required for affine prediction, thereby minimizing hardware costs.
[0335] In the embodiments described above, the methods, etc., are explained based on a sequence diagram as a series of steps or blocks. However, this document is not limited to the order of the steps, etc., and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the sequence diagram are not exclusive, and other steps may be included, or one or more steps in the sequence diagram may be deleted without affecting the scope of this document.
[0336] The embodiments described in this document can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be implemented on a computer, processor, microprocessor, controller, or chip. In this case, implementation information (e.g., information on instructions) or algorithms can be stored on a digital storage medium.
[0337] Furthermore, the decoding and encoding devices to which the embodiments of this document apply can include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, transportation terminals (e.g., vehicle terminals, airplane terminals, ship terminals, etc.), and medical video equipment, and can be used to process video signals or data signals. For example, an over-the-top (OTT) video equipment may include a game console, Blu-ray player, internet-connected TV, home theater system, smartphone, tablet PC, DVR (Digital Video Recorder), etc.
[0338] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having the data structures relating to this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices that store data that can be read by a computer. The computer-readable recording medium can include, for example, Blu-ray discs (BDs), Universal Serial Bus (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media implemented in the form of carrier waves (e.g., transmission over the Internet). In addition, bitstreams generated by encoding methods can be stored on a computer-readable recording medium or transmitted over wired / wireless networks.
[0339] Furthermore, the embodiments described herein can be implemented in a computer program product using program code, and the program code can be performed on a computer according to the embodiments described herein. The program code can be stored on a computer-readable carrier.
[0340] Figure 25 illustrates a content streaming system structure diagram to which the embodiments described in this document apply.
[0341] The content streaming system to which the embodiments described herein apply may broadly include an encoding server, a streaming server, a web server, a media storage facility, user equipment, and multimedia input devices.
[0342] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted.
[0343] The bitstream can be generated by an encoding method or bitstream generation method to which an embodiment of this document applies, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0344] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0345] The streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0346] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage. Each server in the content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
Claims
1. In a video decoding method performed by a decoding device, Steps include obtaining movement prediction information for the current block from the bitstream, The steps include: constructing a list of candidate affine MVPs (Motion Vector Predictors) for the current block; The steps include: deriving a CPMVP (Control Point Motion Vector Predictor) for the Control Point (CP) of the current block based on the aforementioned affine MVP candidate list; The steps include: deriving the Control Point Motion Vector Difference (CPMVD) of the current block with respect to the CP based on the motion prediction information; The steps include: deriving the CPMV (Control Point Motion Vector) for the CP of the current block based on the CPMVP and CPMVD; The steps include: deriving a predicted sample for the current block based on the CPMV; The step of generating a restored picture for the current block based on the derived predicted sample, The steps described above for constructing the Affine MVP candidate list are: A step to check whether the inherited Affine MVP candidate is available, The inherited affine MVP candidate is an affine MVP candidate in which the motion vector derived from the affine model of the surrounding block relative to the inherited affine MVP candidate of the current block is configured as the candidate motion vector for the CP. The inherited affine MVP candidate is derived when the inherited affine MVP candidate is available. The condition for the availability of the inherited affine MVP candidate is whether the reference picture of the surrounding block for the inherited affine MVP candidate is the same as the reference picture of the current block, step, A step to check whether a candidate for Constructed Affine MVP is available, The aforementioned Constructed Affine MVP candidate is, The motion vector derived from the first peripheral block within the first peripheral block group of the current block is used as a candidate motion vector for CP0. The motion vector derived from the second peripheral block within the second peripheral block group of the current block is used as a candidate motion vector for CP1. The motion vector derived from the third peripheral block within the third peripheral block group of the current block is configured as a candidate motion vector for CP2, and is an affine MVP candidate. The constructed affine MVP candidate is derived when the constructed affine MVP candidate is available. The conditions for the availability of the constructed affine MVP candidate are whether all of the motion vectors derived from the first peripheral block, the second peripheral block, and the third peripheral block are available. The condition for the availability of the motion vector derived from the first peripheral block is whether the reference picture of the first peripheral block and the reference picture of the current block are the same. The condition for the availability of the motion vector derived from the second peripheral block is whether the reference picture of the second peripheral block and the reference picture of the current block are the same. The condition for the availability of the motion vector derived from the third peripheral block is whether the reference picture of the third peripheral block and the reference picture of the current block are the same, step, If the number of derived affine MVP candidates, including the inherited affine MVP candidate and the constructed affine MVP candidate, is less than two, the step of deriving a first affine MVP candidate is: The first affine MVP candidate is an affine MVP candidate that includes a specific motion vector as a candidate motion vector for the CP, The specific motion vector is one of the available motion vectors derived from the first peripheral block, the second peripheral block, and the third peripheral block, step, If the number of derived affine MVP candidates is less than two, the step is to derive a second affine MVP candidate that includes the temporal MVP derived based on the temporally surrounding blocks of the current block as a candidate motion vector for the CP, A method comprising the step of deriving a third affine MVP candidate that includes a zero motion vector as a candidate vector for the CP if the number of derived affine MVP candidates is less than two.
2. The aforementioned CP0 represents the upper left corner position of the current block, The aforementioned CP1 represents the upper right corner position of the current block, The method according to claim 1, wherein CP2 represents the lower left end position of the current block.
3. In a video encoding method performed by an encoding device, The current steps involve constructing a list of candidate affine MVPs (Motion Vector Predictors) for a given block, The steps include: deriving a CPMVP (Control Point Motion Vector Predictor) for the Control Point (CP) of the current block based on the aforementioned affine MVP candidate list; The steps include: deriving the CPMV (Control Point Motion Vector) for the current block with respect to the CP, The steps include: deriving the CPMVD (Control Point Motion Vector Difference) of the current block with respect to the CP based on the CPMVP and CPMV; The step includes encoding motion prediction information that includes information about the CPMVD, The steps described above for constructing the Affine MVP candidate list are: A step to check whether the inherited Affine MVP candidate is available, The inherited affine MVP candidate is an affine MVP candidate in which the motion vector derived from the affine model of the surrounding block relative to the inherited affine MVP candidate of the current block is configured as the candidate motion vector for the CP. The inherited affine MVP candidate is derived when the inherited affine MVP candidate is available. The condition for the availability of the inherited affine MVP candidate is whether the reference picture of the surrounding block for the inherited affine MVP candidate is the same as the reference picture of the current block, step, A step to check whether a candidate for Constructed Affine MVP is available, The aforementioned Constructed Affine MVP candidate is, The motion vector derived from the first peripheral block within the first peripheral block group of the current block is used as a candidate motion vector for CP0. The motion vector derived from the second peripheral block within the second peripheral block group of the current block is used as a candidate motion vector for CP1. The motion vector derived from the third peripheral block within the third peripheral block group of the current block is configured as a candidate motion vector for CP2, and is an affine MVP candidate. The constructed affine MVP candidate is derived when the constructed affine MVP candidate is available. The conditions for the availability of the constructed affine MVP candidate are whether all of the motion vectors derived from the first peripheral block, the second peripheral block, and the third peripheral block are available. The condition for the availability of the motion vector derived from the first peripheral block is whether the reference picture of the first peripheral block and the reference picture of the current block are the same. The condition for the availability of the motion vector derived from the second peripheral block is whether the reference picture of the second peripheral block and the reference picture of the current block are the same. The condition for the availability of the motion vector derived from the third peripheral block is whether the reference picture of the third peripheral block and the reference picture of the current block are the same, step, If the number of derived affine MVP candidates, including the inherited affine MVP candidate and the constructed affine MVP candidate, is less than two, the step of deriving a first affine MVP candidate is: The first affine MVP candidate is an affine MVP candidate that includes a specific motion vector as a candidate motion vector for the CP, The specific motion vector is one of the available motion vectors derived from the first peripheral block, the second peripheral block, and the third peripheral block, step, If the number of derived affine MVP candidates is less than two, the step is to derive a second affine MVP candidate that includes the temporal MVP derived based on the temporally surrounding blocks of the current block as a candidate motion vector for the CP, A method comprising the step of deriving a third affine MVP candidate that includes a zero motion vector as a candidate vector for the CP if the number of derived affine MVP candidates is less than two.
4. Regarding methods for transmitting image-related data, A step of obtaining a bitstream of image information, The aforementioned bitstream is The current steps involve constructing a list of candidate affine MVPs (Motion Vector Predictors) for a given block, The steps include: deriving a CPMVP (Control Point Motion Vector Predictor) for the Control Point (CP) of the current block based on the aforementioned affine MVP candidate list; The steps include: deriving the CPMV (Control Point Motion Vector) for the current block with respect to the CP, The steps include: deriving the CPMVD (Control Point Motion Vector Difference) of the current block with respect to the CP based on the CPMVP and CPMV; A step of encoding motion prediction information including information about the CPMVD, and a step of generating based on, The step of transmitting the data, which includes the bitstream, The steps described above for constructing the Affine MVP candidate list are: A step to check whether the inherited Affine MVP candidate is available, The inherited affine MVP candidate is an affine MVP candidate in which the motion vector derived from the affine model of the surrounding block relative to the inherited affine MVP candidate of the current block is configured as the candidate motion vector for the CP. The inherited affine MVP candidate is derived when the inherited affine MVP candidate is available. The condition for the availability of the inherited affine MVP candidate is whether the reference picture of the surrounding block for the inherited affine MVP candidate is the same as the reference picture of the current block, step, A step to check whether a candidate for Constructed Affine MVP is available, The aforementioned Constructed Affine MVP candidate is, The motion vector derived from the first peripheral block within the first peripheral block group of the current block is used as a candidate motion vector for CP0. The motion vector derived from the second peripheral block within the second peripheral block group of the current block is used as a candidate motion vector for CP1. The motion vector derived from the third peripheral block within the third peripheral block group of the current block is configured as a candidate motion vector for CP2, and is an affine MVP candidate. The constructed affine MVP candidate is derived when the constructed affine MVP candidate is available. The conditions for the availability of the constructed affine MVP candidate are whether all of the motion vectors derived from the first peripheral block, the second peripheral block, and the third peripheral block are available. The condition for the availability of the motion vector derived from the first peripheral block is whether the reference picture of the first peripheral block and the reference picture of the current block are the same. The condition for the availability of the motion vector derived from the second peripheral block is whether the reference picture of the second peripheral block and the reference picture of the current block are the same. The condition for the availability of the motion vector derived from the third peripheral block is whether the reference picture of the third peripheral block and the reference picture of the current block are the same, step, If the number of derived affine MVP candidates, including the inherited affine MVP candidate and the constructed affine MVP candidate, is less than two, the step of deriving a first affine MVP candidate is: The first affine MVP candidate is an affine MVP candidate that includes a specific motion vector as a candidate motion vector for the CP, The specific motion vector is one of the available motion vectors derived from the first peripheral block, the second peripheral block, and the third peripheral block, step, If the number of derived affine MVP candidates is less than two, the step is to derive a second affine MVP candidate that includes the temporal MVP derived based on the temporally surrounding blocks of the current block as a candidate motion vector for the CP, A method comprising the step of deriving a third affine MVP candidate that includes a zero motion vector as a candidate vector for the CP if the number of derived affine MVP candidates is less than two.