Image decoding method and apparatus based on subblock-based motion prediction in an image coding system

The image decoding method constructs affine MVP candidates using surrounding blocks to enhance coding efficiency, addressing the need for efficient compression of high-resolution images/videos in immersive media formats.

JP7738145B2Active Publication Date: 2025-09-11LG ELECTRONICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024177084
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-09-12
Filing Date
2024-10-09
Publication Date
2025-09-11
Estimated Expiration
2039-09-11

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, particularly in immersive media formats like VR and AR, necessitates a highly efficient image/video compression technique to reduce transmission and storage costs while maintaining image quality.

Method used

An image decoding method and apparatus that constructs affine MVP candidates based on surrounding blocks and derives motion prediction information using a motion candidate list, adding candidates when necessary to improve coding efficiency.

Benefits of technology

This approach enhances image coding efficiency by reducing complexity and hardware costs, improving overall image/video compression efficiency and minimizing line buffer storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007738145000011
    Figure 0007738145000011
  • Figure 0007738145000012
    Figure 0007738145000012
  • Figure 0007738145000013
    Figure 0007738145000013
Patent Text Reader

Abstract

To provide a method and apparatus for improving the efficiency of image coding.SOLUTION: An image decoding method may comprise the steps of: obtaining motion prediction information for a current block from a bitstream; generating an affine MVP candidate list for the current block; deriving CPMVPs for CPs of the current block based on the affine MVP candidate list; deriving CPMVDs for the CPs of the current block based on the motion prediction information; deriving CPMVs for the CPs of the current block based on the CPMVPs and the CPMVDs; and deriving prediction samples for the current block based on the CPMVs.SELECTED DRAWING: Figure 26
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This document relates to image coding techniques, and more particularly to an image coding method and apparatus based on motion prediction using a motion candidate list for deriving sub-block-wise motion information in an image coding system. [Background technology]

[0002] Recently, the demand for high-resolution, high-quality images / videos such as 4K or 8K or higher UHD (Ultra High Definition) images / videos is increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to existing image / video data, which increases the transmission and storage costs when transmitting image data using existing media such as wired or wireless broadband lines or storing image / video data using existing storage media.

[0003] In addition, interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content and holograms has been increasing recently, and the broadcast of images / videos with different image characteristics from real images, such as game images, is increasing.

[0004] Accordingly, a highly efficient image / video compression technique is required to effectively compress, transmit, store, and play back high-resolution, high-quality image / video information having the above-mentioned various characteristics. Summary of the Invention [Problem to be solved by the invention]

[0005] The technical problem of this document is to provide a method and apparatus for increasing the efficiency of image coding.

[0006] Another technical problem of this document is to provide an image decoding method and apparatus that derives constructed affine MVP candidates based on surrounding blocks only when all candidate motion vectors for a CP are available, constructs an affine MVP candidate list for the current block, and performs prediction for the current block based on the constructed affine MVP candidate list.

[0007] Another technical problem of this document is to provide an image decoding method and apparatus that, when the number of available inherited affine MVP candidates and constructed affine MVP candidates is smaller than the maximum number of candidates in the MVP candidate list, derives an affine MVP candidate using a derived candidate motion vector in the process of deriving the constructed affine MVP candidate as an added affine MVP candidate, and performs prediction for the current block based on the constructed affine MVP candidate list. [Means for solving the problem]

[0008] According to one embodiment of the present document, an image decoding method performed by a decoding device includes the steps of: acquiring motion prediction information for a current block from a bitstream; constructing an affine motion vector predictor (MVP) candidate list for the current block; deriving Control Point Motion Vector Predictors (CPMVP) for a Control Point (CP) of the current block based on the affine MVP candidate list; deriving Control Point Motion Vector Differences (CPMVD) for the CP of the current block based on the motion prediction information; and deriving Control Point Motion Vector Differences (CPMVD) for the CP of the current block based on the CPMVP and the CPMVD. and generating a reconstructed picture for the current block based on the derived predicted samples, wherein constructing the affine MVP candidate list includes checking whether a first affine MVP candidate is available, and the first affine MVP candidate is a candidate for a first block in a left block group that is coded with an affine motion model and whose reference picture index is equal to or greater than that of the current block. If the number of available affine MVP candidates is less than two, checking whether a third affine MVP candidate is available, and the third affine MVP candidate is available if a four-parameter affine model is applied to inter prediction.a first motion vector for CP0 of the current block and a second motion vector for CP1 of the current block can be derived from a block group at an upper left position of the current block and a block group at an upper right position of the current block, respectively, and if a 6-parameter affine model is applied to the inter prediction, a first motion vector for CP0 of the current block, a second motion vector for CP1 of the current block, and a third motion vector for CP2 of the current block can be derived from a block group at an upper left position of the current block, a block group at an upper right position of the current block, and a block group at the left side of the current block, respectively; and if the number of available affine MVP candidates is less than two and the first motion vector is available, deriving a fourth affine MVP candidate, the fourth affine MVP candidate being an affine MVP candidate that includes the motion vector for CP0 in candidate motion vectors for the CPs. and deriving a fifth affine MVP candidate when the number of available affine MVP candidates is less than two and the second motion vector is available, the fifth affine MVP candidate being an affine MVP candidate including the motion vector for the CP1 as a candidate motion vector for the CP. When the number of available affine MVP candidates is less than two and a third motion vector for CP2 of the current block is available, deriving a sixth affine MVP candidate when the number of available affine MVP candidates is less than two and the third motion vector for CP2 of the current block is available, the sixth affine MVP candidate being an affine MVP candidate including the third motion vector as a candidate motion vector for the CP. When the number of available affine MVP candidates is less than two and a temporal MVP candidate derived based on a temporally neighboring block of the current block is available, deriving a seventh affine MVP candidate including the temporal MVP as a candidate motion vector for the CP. When the number of available affine MVP candidates is less than two, a zero motion vector (zero and deriving an eighth affine candidate including the candidate motion vector for the CP.

[0009] According to an embodiment of the present document, an image encoding method performed by an encoding device includes the steps of: constructing an affine motion vector predictor (MVP) candidate list for a current block; deriving control point motion vector predictors (CPMVPs) for a control point (CP) of the current block based on the affine MVP candidate list; deriving CPMVs for the CPs of the current block; deriving control point motion vector differences (CPMVDs) for the CPs of the current block based on the CPMVPs and the CPMVs; and generating motion prediction information (motion prediction information) including information on the CPMVDs. and encoding the affine MVP candidate list, wherein constructing the affine MVP candidate list includes checking whether a first affine MVP candidate is available, the first affine MVP candidate being available if a first block in a left group of blocks is coded with an affine motion model and the reference picture index of the first block is the same as the reference picture index of the current block; and checking whether a second affine MVP candidate is available, the second affine MVP candidate being available if a second block in a top group of blocks is coded with an affine motion model and the reference picture index of the second block is the same as the reference picture index of the current block. and if the number of available affine MVP candidates is less than two, checking whether a third affine MVP candidate is available. The third affine MVP candidate is available if, when a four-parameter affine model is applied to inter prediction, a first motion vector for CP0 of the current block and a second motion vector for CP1 of the current block are derived from a block group at the top left of the current block and a block group at the top right of the current block, respectively; and, when a six-parameter affine model is applied to inter prediction, the first motion vector for CP0 of the current block,a second motion vector for CP1 of the current block and a third motion vector for CP2 of the current block are derived from the upper left block group of the current block, the upper right block group of the current block, and the left block group; a fourth affine MVP candidate is derived if the number of available affine MVP candidates is less than two and the first motion vector is available, the fourth affine MVP candidate being an affine MVP candidate that includes the motion vector for CP0 as a candidate motion vector for the CP; a fifth affine MVP candidate is derived if the number of available affine MVP candidates is less than two and the second motion vector is available, the fifth affine MVP candidate being an affine MVP candidate that includes the motion vector for CP1 deriving a sixth affine MVP candidate that includes the third motion vector as a candidate motion vector for the CP if the number of available affine MVP candidates is less than two and a third motion vector for CP2 of the current block is available, the sixth affine MVP candidate being an affine MVP candidate that includes the third motion vector as a candidate motion vector for the CP if the number of available affine MVP candidates is less than two and a temporal MVP candidate derived based on a temporal neighboring block of the current block is available; deriving a seventh affine MVP candidate that includes the temporal MVP as a candidate motion vector for the CP if the number of available affine MVP candidates is less than two and a temporal MVP candidate derived based on a temporal neighboring block of the current block is available; and deriving an eighth affine MVP candidate that includes a zero motion vector as a candidate motion vector for the CP if the number of available affine MVP candidates is less than two. [Effects of the Invention]

[0010] According to an embodiment of this document, the overall image / video compression efficiency can be improved.

[0011] According to this document, the efficiency of image coding based on affine motion prediction can be increased.

[0012] According to this document, when deriving an affine MVP candidate list, a constructed affine MVP candidate can be added only if all candidate motion vectors for the CP of the constructed affine MVP candidate are available, thereby reducing the complexity of the process of deriving a constructed affine MVP candidate and the process of constructing an affine MVP candidate list and improving coding efficiency.

[0013] According to this document, when deriving an affine MVP candidate list, in the process of deriving a constructed affine MVP candidate, further affine MVP candidates can be derived based on candidate motion vectors for the derived CP, thereby reducing the complexity of the process of constructing an affine MVP candidate list and improving coding efficiency.

[0014] According to this document, in the process of deriving an inherited affine MVP candidate, the upper surrounding block can be used only if the upper surrounding block is included in the current CTU, and the inherited affine MVP candidate can be derived. This can reduce the amount of line buffer storage for affine prediction and minimize hardware costs. [Brief explanation of the drawings]

[0015] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which this document is applicable. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which this document is applicable; [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which this document is applicable; [Figure 4] 1 illustrates an example of an inter-prediction based video / image encoding method. [Figure 5] 1 illustrates an example of an inter-prediction based video / image encoding method. [Figure 6]1 illustrates an exemplary inter-prediction procedure. [Figure 7] 10 illustrates an exemplary motion represented via an affine motion model. [Figure 8] 1 shows an example of the affine motion model in which motion vectors for three control points are used. [Figure 9] 10 exemplarily illustrates the affine unit motion model in which motion vectors for two control points are used. [Figure 10] A method for deriving a motion vector on a sub-block basis based on an affine motion model will be described below as an example. [Figure 11] 10 exemplarily shows surrounding blocks for deriving the inherited affine candidates. [Figure 12] 10 shows an example of spatial candidates for the constructed affine candidates. [Figure 13] An example of constructing an affine MVP list is shown below. [Figure 14] An example of deriving the constructed candidates will be shown below. [Figure 15] An example of deriving the constructed candidates will be shown below. [Figure 16] 10 shows exemplary locations of surrounding blocks scanned to derive inherited affine candidates. [Figure 17] 10 shows exemplary locations of surrounding blocks scanned to derive inherited affine candidates. [Figure 18] 10 shows exemplary positions for deriving inherited affine candidates. [Figure 19] An example of constructing a merge candidate list for the current block is shown below. [Figure 20] 1 illustrates neighboring blocks of the current block for deriving constructed candidates according to one embodiment of the present document; [Figure 21] An example of deriving the constructed candidates when four affine motion models are applied to the current block will be shown below. [Figure 22]An example of deriving the constructed candidates when a 6-affine motion model is applied to the current block will be shown below. [Figure 23a] An example of deriving the inherited affine candidates is shown below. [Figure 23b] An example of deriving the inherited affine candidates is shown below. [Figure 24] 1 illustrates a schematic diagram of an image encoding method using an encoding device according to the present document. [Figure 25] 1 shows a schematic representation of an encoding device for performing an image encoding method according to the present document; [Figure 26] 1 illustrates a schematic diagram of an image decoding method using a decoding device according to the present document; [Figure 27] 1 shows a schematic diagram of a decoding device for performing an image decoding method according to the present document; [Figure 28] 1 illustrates an exemplary structural diagram of a content streaming system to which the embodiments disclosed herein can be applied. DETAILED DESCRIPTION OF THE INVENTION

[0016] Because this document can be modified in various ways and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to the specific embodiments. Common terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. A singular expression includes a plural expression unless the context clearly indicates otherwise. In this specification, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.

[0017] Meanwhile, each component in the drawings described in this document is shown independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Implementations in which each component is integrated and / or separated are also within the scope of this document as long as they do not deviate from the essence of this document.

[0018] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and duplicated descriptions of the same components may be omitted.

[0019] FIG. 1 shows a schematic diagram of an example of a video / image coding system to which this document can be applied.

[0020] 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or streaming data via a digital storage medium or a network.

[0021] The source device may include a video source, an encoding device, and a sending unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device is called a video / image encoding device, and the decoding device is called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be configured as a separate device or an external component.

[0022] A video source can acquire video / images through a video / image capture, synthesis, or generation process, etc. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which the associated data is generated.

[0023] An encoding device can encode input video / images. The encoding device can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0024] The transmitter may transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver may receive / extract the bitstream and transmit it to a decoding device.

[0025] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0026] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.

[0027] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the versatile video coding (VVC) standard, the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2) or next generation video / image coding standards (e.g., H.267 or H.268).

[0028] This document presents various embodiments for video / image coding, which may be combined with one another unless otherwise stated.

[0029] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that shows one image at a specific time, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile can include one or more coding tree units (CTUs). A picture can be composed of one or more slices / tiles. A picture can be composed of one or more tile groups. A tile group can include one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile can be partitioned into multiple bricks, and each brick can consist of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan can indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs can be aligned in a CTU raster scan within a brick, bricks within a tile can be aligned consecutively in a raster scan of the bricks in the tile, and tiles within a picture can be aligned consecutively in a raster scan of the tiles in the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the height of the picture. A tile scan may indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs may be consecutively aligned in a CTU raster scan in a tile, and tiles in a picture may be consecutively aligned in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit. A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile.In this document, tile group and slice may be used interchangeably, for example, tile group / tile group header is referred to as slice / slice header in this document.

[0030] A pixel or a pel can refer to the smallest unit that makes up a picture (or image). A "sample" can also be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, or can indicate only a pixel / pixel value of a luma component, or can indicate only a pixel / pixel value of a chroma component.

[0031] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information about the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.

[0032] In this document, the terms " / " and "," should be interpreted to indicate "and / or." For example, "A / B" means "A and / or B," and "A, B" means "A and / or B." Furthermore, "A / B / C" means "at least one of A, B, and / or C." Similarly, "A, B, C" means "at least one of A, B, and / or C." (In this document, the terms " / " and "," should be interpreted to indicate "and / or." For instance, the expression "A / B" may mean "A and / or B." Further, "A,B" may mean "A and / or B." Further, "A / B / C" may mean "at least one of A, B, and / or C." Also, "A / B / C" may mean "at least one of A, B, and / or C.")

[0033] Further, in this document, "or" should be interpreted to mean "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Expressed differently, "or" in this document can mean "additionally or alternatively." (Further, in the document, the term "or" should be interpreted to indicate "and / or." For instance, the expression "A or B" may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted to indicate "additionally or alternatively.")

[0034] 2 is a diagram for explaining the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the term "video encoding device" may include an image encoding device.

[0035] Referring to FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 is referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Furthermore, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0036] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) using a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to this document is performed based on the final coding unit that is not further divided. In this case, based on coding efficiency according to image characteristics, the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of lower depth, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0037] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally refer to a pixel or a pixel value, and can refer to only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample can be used as a term that corresponds to one pixel or pel of a picture (or image).

[0038] The encoding apparatus 200 may subtract a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown, a unit in the encoder 200 that subtracts a prediction signal (predicted block, prediction sample array) from an input image signal (original block, original sample array) is called a subtraction unit 231. The prediction unit may perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0039] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away, depending on the prediction mode. In intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of precision of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0040] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block is also called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block is also called a collocated picture (colPic). For example, the inter predictor 221 may construct a candidate list of motion information based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction is performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of a neighboring block as a motion vector predictor and signaling the motion vector difference.

[0041] The prediction unit 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This is called combined inter and intra prediction (CIIP). The prediction unit may also use intra block copy (IBC) prediction mode or palette mode for predicting a block. The IBC prediction mode or palette mode may be used for coding content images / videos, such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but is similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. Palette mode can be seen as an example of intra coding or intra prediction. When palette mode is applied, sample values ​​within a picture may be signaled based on information about a palette table and a palette index.

[0042] The prediction signal generated by the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may be a discrete cosine transform (DCT), a discrete sine transform (DST), a sine transform (SST), a sine transform (SCT), a sine transform (SST ... JPEG0007738145000001.jpg7116, can include at least one of GBT (Graph-Based Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transformation obtained based on a predicted signal generated using all previously reconstructed pixels. In addition, the transformation process may be applied to pixel blocks having the same square size, or to non-square blocks of variable size.

[0043] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients is called residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit 240 may also encode information required for video / image restoration other than the quantized transform coefficients (e.g., values ​​of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) unit units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information is encoded through the encoding procedure described above and included in the bitstream.The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding apparatus 200, or the transmitter may be included in the entropy encoding unit 240.

[0044] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter predictor 221 or the intra predictor 222. When there is no residual for the current block, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 is referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.

[0045] Meanwhile, luma mapping with chroma scaling (LMCS) may be applied during picture encoding and / or reconstruction.

[0046] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0047] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. Through this, when inter prediction is applied, the encoding apparatus can avoid prediction mismatch between the encoding apparatus 100 and the decoding apparatus, and can also improve encoding efficiency.

[0048] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.

[0049] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which this document is applicable.

[0050] Referring to FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoding unit 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware components may further include memory 360 in the internal / external components.

[0051] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image by processing the video / image information in the encoding apparatus of FIG. 2. For example, the decoding apparatus 300 can derive units / blocks based on information about block division obtained from the bitstream. The decoding apparatus 300 can perform decoding using the processing units applied in the encoding apparatus. Accordingly, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided into a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Furthermore, the reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced through a reproduction device.

[0052] The decoding apparatus 300 may receive a signal output from the encoding apparatus of FIG. 2 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding apparatus may further decode pictures based on the information on the parameter sets and / or the general constraint information. Signaled / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method may receive bins corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information on neighboring and target blocks, or information on symbols / bins decoded in a previous step, predict the occurrence probability of the bins according to the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element.In this case, after determining a context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoding unit 310 is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values ​​entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). Furthermore, filtering-related information from the information decoded by the entropy decoding unit 310 may be provided to the filtering unit 350. Meanwhile, a receiving unit (not shown) for receiving a signal output from the encoding apparatus may be configured as an internal / external element of the decoding apparatus 300, or the receiving unit may be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document is called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.

[0053] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The inverse quantization unit 321 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0054] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0055] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.

[0056] The prediction unit 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This is called combined inter and intra prediction (CIIP). The prediction unit may also use intra block copy (IBC) prediction mode or palette mode for predicting a block. The IBC prediction mode or palette mode may be used for coding images / videos of content such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but is similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be seen as an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / image information and signaled.

[0057] The intra prediction unit 331 may predict a current block by referring to samples in a current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.

[0058] The inter prediction unit 332 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 332 may construct a candidate list of motion information based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction is performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.

[0059] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the reconstructed block.

[0060] The adder 340 is called a reconstruction unit or a reconstruction block generator. The generated reconstruction signal can be used for intra prediction of the next block to be processed in the current picture, and can be output after filtering, as described below, or can be used for inter prediction of the next picture.

[0061] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during picture decoding.

[0062] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0063] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.

[0064] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.

[0065] As described above, prediction is performed to improve compression efficiency during video coding. Through this, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device. The encoding device signals information (residual information) regarding the residual between the original block and the predicted block to the decoding device, rather than the original sample values ​​of the original block, thereby improving image coding efficiency. The decoding device derives a residual block including residual samples based on the residual information, combines the residual block with the predicted block to generate a reconstructed block including reconstructed samples, and generates a reconstructed picture including the reconstructed block.

[0066] The residual information may be generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and signal the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device may perform an inverse quantization / inverse transform procedure based on the residual information to derive residual samples (or a residual block). The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device may also derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a subsequent picture, and generate a reconstructed picture based on the residual block.

[0067] When inter prediction is applied, a prediction unit of an encoding / decoding apparatus may perform inter prediction on a block-by-block basis to derive a prediction sample. Inter prediction may refer to a prediction derived in a manner that is dependent on data elements (e.g., sample values ​​or motion information) of pictures other than the current picture. When inter prediction is applied to a current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector on a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted on a block-by-block, sub-block-by-sample basis based on correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction type (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks are called collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including the temporal neighboring blocks is also called a collocated picture (colPic).For example, a motion information candidate list may be constructed based on neighboring blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter prediction may be performed based on various prediction modes. For example, in skip mode and (normal) merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. Unlike merge mode, in skip mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0068] A video / image encoding procedure based on inter prediction may generally include, for example:

[0069] FIG. 4 shows an example of an inter-prediction based video / image encoding method.

[0070] An encoding apparatus performs inter prediction on a current block (S400). The encoding apparatus may derive an inter prediction mode and motion information for the current block and generate a predicted sample for the current block. Here, the procedures for determining the inter prediction mode, deriving the motion information, and generating the predicted sample may be performed simultaneously, or one procedure may precede the other. For example, the inter prediction unit of the encoding apparatus may include a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit. The prediction mode determination unit may determine a prediction mode for the current block, the motion information derivation unit may derive motion information for the current block, and the predicted sample derivation unit may derive a predicted sample for the current block. For example, the inter prediction unit of the encoding apparatus may search for a block similar to the current block within a certain region (search region) of a reference picture through motion estimation, and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located can be derived, and a motion vector can be derived based on a difference between the positions of the reference block and the current block. The encoding apparatus can determine a mode to be applied to the current block from various prediction modes. The encoding apparatus can compare RD costs for the various prediction modes to determine an optimal prediction mode for the current block.

[0071] For example, when a skip mode or a merge mode is applied to the current block, the encoding apparatus may construct a merge candidate list (described below) and derive a reference block, among reference blocks indicated by merge candidates included in the merge candidate list, whose difference between the current block and the current block is minimum or equal to or less than a certain criterion. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to a decoding apparatus. Motion information of the current block may be derived using motion information of the selected merge candidate.

[0072] As another example, when the (A)MVP mode is applied to the current block, the encoding apparatus may construct an (A)MVP candidate list (described below) and use the motion vector of a selected MVP (motion vector predictor) candidate from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the motion estimation described above may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block may be the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding apparatus. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding apparatus.

[0073] The encoding apparatus may derive residual samples based on the predicted samples (S410) by comparing the original samples of the current block with the predicted samples.

[0074] The encoding apparatus encodes image information including prediction information and residual information (S420). The encoding apparatus may output the encoded image information in the form of a bitstream. The prediction information may include prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information about motion information as information about the prediction procedure. The information about the motion information may include candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) that is information for deriving a motion vector. The information about the motion information may also include the above-mentioned information about MVD and / or reference picture index information. The information about the motion information may also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples.

[0075] The output bitstream can be stored in a (digital) storage medium and then transmitted to the decoding device, or can be transmitted to the decoding device via a network.

[0076] Meanwhile, as described above, the encoding apparatus can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the encoding apparatus can derive the same prediction result as that performed by the decoding apparatus, thereby improving coding efficiency. Therefore, the encoding apparatus can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in a memory and use it as a reference picture for inter-prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture.

[0077] A video / image decoding procedure based on inter prediction may generally include, for example:

[0078] FIG. 5 illustrates an example of an inter-prediction based video / picture decoding method.

[0079] 5, the decoding apparatus may perform operations corresponding to those performed by the encoding apparatus, such as performing prediction on a current block based on received prediction information and deriving predicted samples.

[0080] Specifically, the decoding device may determine a prediction mode for the current block based on received prediction information (S500). The decoding device may determine which inter-prediction mode is applied to the current block based on prediction mode information in the prediction information.

[0081] For example, it may determine whether the merge mode or (A)MVP mode is applied to the current block based on the merge flag. Alternatively, it may select one of various inter prediction mode candidates based on the mode index. The inter prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or various inter prediction modes described below.

[0082] The decoding apparatus derives motion information of the current block based on the determined inter prediction mode (S510). For example, when a skip mode or a merge mode is applied to the current block, the decoding apparatus may construct a merge candidate list (described below) and select one merge candidate from among the merge candidates included in the merge candidate list. The selection is performed based on the selection information (merge index) described above. Motion information of the selected merge candidate may be used to derive motion information of the current block. The motion information of the selected merge candidate may be used as motion information of the current block.

[0083] As another example, when the (A)MVP mode is applied to the current block, the decoding apparatus may construct an (A)MVP candidate list (described below) and use a motion vector of a motion vector predictor (MVP) candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Furthermore, the decoding apparatus may derive a reference picture index of the current block based on the reference picture index information. Within the list of reference pictures for the current block, a picture indicated by the reference picture index may be derived as a reference picture referenced for inter-prediction of the current block.

[0084] On the other hand, as will be described later, the motion information of the current block can be derived without constructing a candidate list, in which case the motion information of the current block can be derived by a procedure started in a prediction mode, as will be described later. In this case, the construction of the candidate list as described above can be omitted.

[0085] The decoding apparatus may generate predictive samples for the current block based on the motion information of the current block (S520). In this case, the reference picture may be derived based on the reference picture index of the current block, and the predictive samples for the current block may be derived using samples of the reference block pointed to in the reference picture by the motion vector of the current block. In this case, as described below, a predictive sample filtering procedure may further be performed on all or some of the predictive samples of the current block, depending on the circumstances.

[0086] For example, the inter-prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and may determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit, derive motion information (such as a motion vector and / or a reference picture index) of the current block based on information regarding the motion information received by the motion information derivation unit, and derive a prediction sample of the current block by the prediction sample derivation unit.

[0087] The decoding apparatus generates residual samples for the current block based on the received residual information (S530). The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and can generate a reconstructed picture based on the reconstructed samples (S540). As described above, an in-loop filtering procedure can then be applied to the reconstructed picture.

[0088] FIG. 6 exemplarily illustrates an inter-prediction procedure.

[0089] 6, as described above, the inter prediction procedure may include an inter prediction mode determination step, a motion information deriving step according to the determined prediction mode, and a prediction execution step (generation of prediction samples) based on the derived motion information. The inter prediction procedure is performed by an encoding device and a decoding device as described above. In this document, a coding device may include an encoding device and / or a decoding device.

[0090] Referring to FIG. 6, a coding apparatus determines an inter prediction mode for a current block (S600). Various inter prediction modes can be used for predicting a current block in a picture. For example, merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, etc. can be used. Decoder side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weight (BCW), bi-directional optical flow (BDOF), etc. can be used in addition to or instead of the additional mode. Affine mode is also called affine motion prediction mode. MVP mode is also called advanced motion vector prediction mode. In this document, some modes and / or motion information candidates derived by some modes may be included among the motion information candidates of other modes. For example, an HMVP candidate may be added as a merge candidate in the merge / skip mode, or as an MVP candidate in the MVP mode. When the HMVP candidate is used as a motion information candidate in the merge mode or skip mode, the HMVP candidate is called an HMVP merge candidate.

[0091] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the encoding apparatus to the decoding apparatus. The prediction mode information may be included in a bitstream and received by the decoding apparatus. The prediction mode information may include index information indicating one of multiple candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applied, and if the skip mode is not applied, a merge flag may be signaled to indicate whether the merge mode is applied, and if the merge mode is not applied, an MVP mode may be applied, or a flag for further classification may be further signaled. The affine mode may be signaled as an independent mode, or may be signaled as a mode dependent on the merge mode, MVP mode, etc. For example, the affine mode may include affine merge mode and affine MVP mode.

[0092] The coding apparatus derives motion information for the current block (S610). The motion information may be derived based on the inter prediction mode.

[0093] A coding apparatus may perform inter-prediction using motion information of a current block. An encoding apparatus may derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding apparatus may use an original block in an original picture for the current block to search for a similar reference block with high correlation within a predetermined search range in the reference picture in fractional pixel units, thereby deriving motion information. Block similarity may be derived based on a phase-based sample value difference. For example, block similarity may be calculated based on the SAD between the current block (or a template of the current block) and a reference block (or a template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD within the search range. The derived motion information may be signaled to a decoding apparatus in various ways based on the inter-prediction mode.

[0094] The coding apparatus performs inter prediction based on motion information for the current block (S620). The coding apparatus may derive predictive samples for the current block based on the motion information. The current block including the predictive samples is called a predicted block.

[0095] Meanwhile, in the case of inter prediction, an inter prediction method that takes image distortion into consideration has been proposed. Specifically, an affine motion model has been proposed that efficiently derives motion vectors for sub-blocks or sample points of a current block and improves the accuracy of inter prediction despite deformations such as image rotation, zoom-in, or zoom-out. That is, the affine motion model derives motion vectors for sub-blocks or sample points of a current block, and prediction using the affine motion model is called affine motion prediction, sub-block-based motion prediction, or sub-block motion prediction.

[0096] For example, the sub-block motion prediction using the affine motion model can efficiently represent four motions, ie, four transformations, as described below.

[0097] 7 exemplarily illustrates motions represented using an affine motion model. Referring to FIG. 7, motions that can be represented using the affine motion model include translational motion, scaled motion, rotated motion, and sheared motion. That is, not only the translational motion in which an image (or a part thereof) moves planarly over time as shown in FIG. 7, but also the scaled motion in which an image (or a part thereof) scales over time, the rotational motion in which an image (or a part thereof) rotates over time, and the sheared motion in which an image (or a part thereof) is transformed into a parallelogram over time can be efficiently represented through the sub-block-based motion prediction.

[0098] The encoding / decoding apparatus can predict the distortion type of the image based on the motion vector at the control point (CP) of the current block through the affine inter prediction, thereby improving the accuracy of prediction and thereby improving the image compression performance. Also, since the motion vector for at least one control point of the current block can be derived using the motion vectors of the neighboring blocks of the current block, the burden of data volume for added side information can be reduced and the efficiency of inter prediction can be significantly improved.

[0099] As an example of the affine motion prediction, three control points, i.e., motion information at three reference points, may be required.

[0100] FIG. 8 exemplarily shows the affine motion model in which motion vectors for three control points are used.

[0101] If the position of the top-left sample in the current block 800 is (0,0), the sample positions (0,0), (w,0), and (0,h) can be determined as the control points, as shown in Figure 8. Hereinafter, the control point of the (0,0) sample position can be denoted as CP0, the control point of the (w,0) sample position can be denoted as CP0, and the control point of the (0,h) sample position can be denoted as CP1.

[0102] Using the above-described control points and the motion vectors for the control points, the mathematical formula for the affine motion model can be derived. The mathematical formula for the affine motion model can be expressed as follows.

[0103]

number

[0104] Here, w represents the width of the current block 800, h represents the height of the current block 800, and v 0x , v 0y represent the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y represent the x and y components of the motion vector of CP0, respectively, and v 2x , v 2y represent the x and y components of the motion vector of CP1, respectively. Also, x represents the x component of the position of the target sample within the current block 800, y represents the y component of the position of the target sample within the current block 800, and v x is the x-component of the motion vector of the target sample in the current block 800, v y represents the y component of the motion vector of the target sample in the current block 800.

[0105] Since the motion vector of CP0, the motion vector of CP1, and the motion vector of CP2 are known, a motion vector according to the sample position in the current block can be derived based on Equation 1. That is, according to the affine motion model, the motion vector at the control point, v0(v 0x ,v 0y ), v1(v 1x ,v 1y ), v2(v 2x ,v 2y ) is scaled, and a motion vector of the target sample according to the position of the target sample can be derived. That is, according to the affine motion model, a motion vector of each sample in the current block can be derived based on the motion vector of the control point. Meanwhile, a set of motion vectors of samples in the current block derived by the affine motion model can be referred to as an affine motion vector field (MVF).

[0106] Meanwhile, the six parameters for Equation 1 can be expressed as a, b, c, d, e, and f as in the following equation, and the equation for the affine motion model expressed by the six parameters is as follows:

[0107]

number

[0108] Here, w represents the width of the current block 800, h represents the height of the current block 800, and v 0x , v 0y represent the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y represent the x and y components of the motion vector of CP0, respectively, and v 2x , v 2yrepresent the x and y components of the motion vector of CP1, respectively. Also, x represents the x component of the position of the target sample within the current block 800, y represents the y component of the position of the target sample within the current block 800, and v x is the x-component of the motion vector of the target sample in the current block 800, v y represents the y component of the motion vector of the target sample in the current block 800.

[0109] The affine motion model or the affine inter-prediction using the six parameters may be denoted as a six-parameter affine motion model or AF6.

[0110] Also, as an example of the affine motion prediction, two control points, i.e., motion information at two reference points, may be required.

[0111] 9 exemplarily illustrates the affine unit motion model in which motion vectors for two control points are used. The affine motion model using two control points can express three types of motion, including translational motion, scaling motion, and rotational motion. The affine motion model expressing three types of motion can also be referred to as a similarity affine motion model or a simplified affine motion model.

[0112] If the position of the top-left sample in the current block 900 is (0,0), the sample positions (0,0) and (w,0) can be determined as the control points, as shown in Figure 9. Hereinafter, the control point of the sample position (0,0) can be denoted as CP0, and the control point of the sample position (w,0) can be denoted as CP0.

[0113] Using the above-described control points and the motion vectors for the corresponding control points, the mathematical formula for the affine motion model can be derived. The mathematical formula for the affine motion model can be expressed as follows.

[0114]

number

[0115] where w represents the width of the current block 900, and v 0x , v 0y represent the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y represent the x and y components of the motion vector of CP0, respectively. Also, x represents the x component of the position of the target sample within the current block 900, y represents the y component of the position of the target sample within the current block 900, and v x is the x-component of the motion vector of the target sample in the current block 900, v y represents the y component of the motion vector of the target sample in the current block 900.

[0116] Meanwhile, the four parameters for Equation 3 can be expressed as a, b, c, and d as in the following equation, and the equation for the affine motion model expressed by the four parameters is as follows.

[0117]

number

[0118] where w represents the width of the current block 900, and v 0x , v 0y represent the x and y components of the motion vector of CP0, respectively, and v 1x , v 1yrepresent the x and y components of the motion vector of CP0, respectively. Also, x represents the x component of the position of the target sample within the current block 900, y represents the y component of the position of the target sample within the current block 900, and v x is the x-component of the motion vector of the target sample in the current block 900, v y represents the y-component of the motion vector of the target sample in the current block 900. Since the affine motion model using the two control points can be expressed by four parameters a, b, c, and d as in Equation 4, the affine motion model or affine motion prediction using the four parameters can be expressed as a four-parameter affine motion model or AF4. That is, according to the affine motion model, a motion vector for each sample in the current block can be derived based on the motion vectors of the control points. Meanwhile, a set of motion vectors of samples in the current block derived by the affine motion model can be expressed as an affine motion vector field (MVF).

[0119] Meanwhile, as described above, the affine motion model can derive a motion vector for each sample, thereby significantly improving the accuracy of inter-prediction, although this may result in a significant increase in the complexity of the motion compensation process.

[0120] Therefore, instead of deriving a motion vector in units of samples, a motion vector in units of sub-blocks within the current block can be restricted to be derived.

[0121] 10 exemplarily illustrates a method for deriving a motion vector in units of sub-blocks based on an affine motion model. Fig. 10 exemplarily illustrates a case where the size of the current block is 16x16 and a motion vector is derived in units of 4x4 sub-blocks. The sub-blocks can be set to various sizes. For example, if the sub-blocks are set to a size of nxn (n is a positive integer, e.g., 4), a motion vector can be derived in units of nxn sub-blocks within the current block based on the affine motion model, and various methods can be applied to derive a motion vector representing each sub-block.

[0122] For example, referring to FIG. 10, a motion vector for each sub-block can be derived using the sample position at the center or the lower right side of the center as the representative coordinate. Here, the lower right position of the center may refer to the sample position at the lower right of four samples located at the center of the sub-block. For example, if n is an odd number, one sample may be located at the center of the sub-block, and in this case, the center sample position may be used to derive the motion vector for the sub-block. However, if n is an even number, four samples may be located adjacent to the center of the sub-block, and in this case, the lower right sample position may be used to derive the motion vector. For example, referring to FIG. 10, the representative coordinates for each sub-block can be derived as (2,2), (6,2), (10,2), ... (14,14), and the encoding / decoding apparatus can substitute each of the representative coordinates of the sub-block into Equation 1 or 3 to derive the motion vector for each sub-block. Predicting the motion of a sub-block within a current block using the affine motion model is called sub-block-based motion prediction or sub-block motion prediction, and the motion vector of such a sub-block can be denoted as MVF.

[0123] Meanwhile, as an example, the size of the sub-block within the current block may be derived based on the following formula:

[0124]

number

[0125] Here, M represents the width of the sub-block, and N represents the height of the sub-block. 0x , v 0y represent the x and y components of the CPMV0 of the current block, respectively, and v 1x , v 1y where x and y represent the x and y components of CPMV1 of the current block, w represents the width of the current block, h represents the height of the current block, and MvPre represents the motion vector fraction accuracy. For example, the motion vector fraction accuracy can be set to 1 / 16.

[0126] Meanwhile, inter prediction using the affine motion model, i.e., affine motion prediction, may include a merge mode (AF_MERGE) and an affine inter mode (AF_INTER). Here, the affine inter mode may be referred to as an affine motion vector prediction mode (AF_MVP).

[0127] The merge mode using the affine motion model is similar to the existing merge mode in that MVD for the motion vector of the control point is not transmitted. That is, like the existing skip / merge mode, the merge mode using the affine motion model can represent an encoding / decoding method in which CPMV for each of two or three control points from neighboring blocks of the current block is derived and prediction is performed without coding for MVD (motion vector difference).

[0128] For example, when the AF_MRG mode is applied to the current block, CP0 and MVs for CP0 (i.e., CPMV0 and CPMV1) can be derived from neighboring blocks of the current block to which affine mode, i.e., a prediction mode using affine motion prediction, is applied. That is, CPMV0 and CPMV1 of the neighboring blocks to which the affine mode is applied can be derived as merge candidates, and the merge candidates can be derived from CPMV0 and CPMV1 for the current block.

[0129] The affine inter mode may indicate inter prediction, which derives a motion vector predictor (MVP) for the motion vector of the control point, derives the motion vector of the control point based on a received motion vector difference (MVD) and the MVP, derives an affine MVF of the current block based on the motion vector of the control point, and performs prediction based on the affine MVF. Here, the motion vector of the control point may be represented as a control point motion vector (CPMV), the MVP of the control point may be represented as a control point motion vector predictor (CPMVP), and the MVD of the control point may be represented as a control point motion vector difference (CPMVD). Specifically, for example, an encoding apparatus may derive a control point motion vector predictor (CPMVP) and a control point motion vector (CPMVD) for each of CP0 and CP1 (or CP0, CP1, and CP2), and may transmit or store information on the CPMVP and / or CPMVD, which is a difference between the CPMVP and CPMV.

[0130] Here, when the affine inter mode is applied to the current block, the encoding device / decoding device can construct an affine MVP candidate list based on the neighboring blocks of the current block, and the affine MVP candidates are called CPMVP pair candidates, and the affine MVP candidate list is also called a CPMVP candidate list.

[0131] Furthermore, each affine MVP candidate may represent a combination of CPMVPs of CP0 and CP1 in a four parameter affine motion model, or a combination of CPMVPs of CP0, CP1, and CP2 in a six parameter affine motion model.

[0132] Meanwhile, for affine inter-prediction, inherited affine candidates or inherited and constructed affine candidates are considered for constructing an affine MVP candidate list. An inherited candidate refers to a candidate in which the motion information of the neighboring blocks of the current block, i.e., the CPMV of the neighboring blocks, is added to the motion candidate list of the current block without any modification or combination. Here, the neighboring blocks can include a neighboring block A0 at the lower left corner of the current block, a neighboring block A1 on the left side, a neighboring block B0 above the current block, a neighboring block B1 at the upper right corner, and a neighboring block B2 at the upper left corner. A constructed affine candidate refers to an affine candidate that constructs the CPMV of the current block by combining the CPMV of at least two neighboring blocks. The derivation of a constructed affine candidate will be described in detail below.

[0133] Here, the inherited affine candidates are:

[0134] For example, if a neighboring block of the current block is an affine block and the reference picture of the current block is the same as the reference picture of the neighboring block, an affine MVP pair of the current block can be determined from the affine motion model of the neighboring block. Here, the affine block may indicate a block to which the affine inter prediction is applied. The inherited affine candidate may indicate a CPMVP (e.g., the affine MVP pair) derived based on the affine motion model of the neighboring block.

[0135] Specifically, as an example, the inherited affine candidates can be derived as described below.

[0136] FIG. 11 exemplarily shows surrounding blocks for deriving the inherited affine candidates.

[0137] Referring to FIG. 11, the neighboring blocks of the current block may include neighboring block A0 on the left side of the current block, neighboring block A1 in the lower left corner of the current block, neighboring block B0 above the current block, neighboring block B1 in the upper right corner of the current block, and neighboring block B2 in the upper left corner of the current block.

[0138] For example, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the left neighboring block may be a block including a sample at (-1, H-1) coordinates, the upper neighboring block may be a block including a sample at (W-1, -1) coordinates, the upper right neighboring block may be a block including a sample at (W, -1) coordinates, the lower left neighboring block may be a block including a sample at (-1, H) coordinates, and the upper left neighboring block may be a block including a sample at (-1, -1) coordinates.

[0139] The encoding / decoding device may sequentially check neighboring blocks A0, A1, B0, B1, and B2, and if the neighboring blocks are coded using an affine motion model and the reference picture of the current block and the reference picture of the neighboring blocks are the same, derive two or three CPMVs of the current block based on the affine motion models of the neighboring blocks. The CPMVs may be derived from affine MVP candidates of the current block. The affine MVP candidates may indicate the inherited affine candidates.

[0140] As an example, up to two successive affine candidates can be derived based on the surrounding blocks.

[0141] For example, the encoding / decoding apparatus may derive a first affine MVP candidate for the current block based on a first block among neighboring blocks. Here, the first block may be coded using an affine motion model, and the reference picture of the first block may be the same as the reference picture of the current block. That is, the first block may be a block that satisfies a condition that is first confirmed by checking the neighboring blocks in a specific order. The condition may be that the block is coded using an affine motion model, and the reference picture of the block is the same as the reference picture of the current block.

[0142] Thereafter, the encoding / decoding apparatus may derive a second affine MVP candidate for the current block based on a second block among neighboring blocks. Here, the second block may be coded using an affine motion model, and the reference picture of the second block may be the same as the reference picture of the current block. That is, the second block may be a block that satisfies a second confirmed condition when checking the neighboring blocks in a specific order. The condition may be that the block is coded using an affine motion model, and the reference picture of the block is the same as the reference picture of the current block.

[0143] On the other hand, for example, if the available number of inherited affine candidates is less than 2 (i.e., the number of derived inherited affine candidates is less than 2), a constructed affine candidate can be considered. The constructed affine candidate can be derived as follows:

[0144] FIG. 12 exemplarily shows spatial candidates for the constructed affine candidates.

[0145] As shown in Figure 12, the motion vectors of the neighboring blocks of the current block are divided into three groups. Referring to Figure 12, the neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G.

[0146] The peripheral block A may refer to a peripheral block located at the upper left of the upper left sample position of the current block, the peripheral block B may refer to a peripheral block located above the upper left sample position of the current block, and the peripheral block C may refer to a peripheral block located at the left of the upper left sample position of the current block. The peripheral block D may refer to a peripheral block located above the upper right sample position of the current block, and the peripheral block E may refer to a peripheral block located at the upper right of the upper right sample position of the current block. The peripheral block F may refer to a peripheral block located at the left of the lower left sample position of the current block, and the peripheral block G may refer to a peripheral block located at the lower left of the lower left sample position of the current block.

[0147] For example, the three groups may include S0, S1, and S2, and S0, S1, and S2 can be derived as shown in the following table.

[0148] [Table 1]

[0149] Here, mv A is the motion vector of the surrounding block A, mv B is the motion vector of the surrounding block B, mv C is the motion vector of the surrounding block C, mv D is the motion vector of the surrounding block D, mv E is the motion vector of the surrounding block E, mv F is the motion vector of the surrounding block F, mv G indicates the motion vector of the surrounding block G. S0 may be indicated as the first group, S1 may be indicated as the second group, and S2 may be indicated as the third group.

[0150] The encoding device / decoding device can derive mv0 from S0, mv1 from S1, and mv2 from S2, and can derive affine MVP candidates including mv0, mv1, and mv2. The affine MVP candidates can represent the constructed affine candidates. Also, mv0 may be a CPMVP candidate for CP0, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0151] Here, the reference picture for mv0 may be the same as the reference picture of the current block. That is, mv0 may be a motion vector that satisfies a condition that is first confirmed by checking motion vectors in S0 in a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be the neighboring block A, the neighboring block B, and the neighboring block C in S0. Also, the order may be other than the above-mentioned order and is not limited to the above-mentioned example.

[0152] Furthermore, the reference picture for mv1 may be the same as the reference picture of the current block. That is, mv1 may be a motion vector that satisfies a condition that is first confirmed by checking motion vectors in S1 according to a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be neighboring block D → neighboring block E in S1. An order other than the above-mentioned order may also be used, and is not limited to the above-mentioned example.

[0153] Furthermore, the reference picture block for mv2 may be the same as the reference picture of the current block. That is, mv2 may be a motion vector that satisfies a condition that is first confirmed by checking motion vectors in S2 in a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be the neighboring block F→the neighboring block G in S2. An order other than the above-mentioned order may also be used, and is not limited to the above-mentioned example.

[0154] On the other hand, when only the mv0 and the mv1 are available, that is, when only the mv0 and the mv1 are derived, the mv2 can be derived as follows.

[0155]

number

[0156] where mv2 x represents the x component of mv2, and mv2 y represents the y component of mv2, and mv0 x represents the x component of mv0, and mv0 y represents the y component of mv0, and mv1 x represents the x-component of mv1, and mv1 y represents the y component of mv1, w represents the width of the current block, and h represents the height of the current block.

[0157] Meanwhile, when only the mv0 and mv2 are derived, the mv1 can be derived as follows:

[0158]

number

[0159] where mv1 x represents the x-component of mv1, and mv1 yrepresents the y component of mv1, and mv0 x represents the x component of mv0, and mv0 y represents the y component of mv0, and mv2 x represents the x component of mv2, and mv2 y represents the y component of mv2, w represents the width of the current block, and h represents the height of the current block.

[0160] Also, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the AMVP process of the existing HEVC standard can be applied to construct the affine MVP list. That is, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the process of constructing MVP candidates in the existing HEVC standard is performed.

[0161] Meanwhile, a flowchart of an embodiment for constructing the above-mentioned affine MVP list is as follows.

[0162] FIG. 13 exemplarily shows an example of constructing an affine MVP list.

[0163] 13, the encoding / decoding device may add an inherited candidate to the affine MVP list of the current block (S1300). The inherited candidate may indicate the inherited affine candidate described above.

[0164] Specifically, the encoding / decoding apparatus may derive up to two inherited affine candidates from neighboring blocks of the current block (S1305), where the neighboring blocks may include a neighboring block A0 on the left side of the current block, a neighboring block A1 in the lower left corner, a neighboring block B0 above the current block, a neighboring block B1 in the upper right corner, and a neighboring block B2 in the upper left corner.

[0165] For example, the encoding / decoding apparatus may derive a first affine MVP candidate for the current block based on a first block among neighboring blocks. Here, the first block may be coded using an affine motion model, and the reference picture of the first block may be the same as the reference picture of the current block. That is, the first block may be a block that satisfies a condition that is first confirmed by checking the neighboring blocks in a specific order. The condition may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block.

[0166] Thereafter, the encoding / decoding apparatus may derive a second affine MVP candidate for the current block based on a second block among neighboring blocks. Here, the second block may be coded using an affine motion model, and the reference picture of the second block may be the same as the reference picture of the current block. That is, the second block may be a block that satisfies a second confirmed condition when checking the neighboring blocks in a specific order. The condition may be that the block is coded using an affine motion model, and the reference picture of the block is the same as the reference picture of the current block.

[0167] Alternatively, the specific order may be the left peripheral block A0, the lower left peripheral block A1, the upper peripheral block B0, the upper right peripheral block B1, and the upper left peripheral block B2. The order may be other than the above-mentioned order and is not limited to the above-mentioned example.

[0168] The encoding device / decoding device may add a constructed candidate to the affine MVP list of the current block (S1310). The constructed candidate may represent the above-mentioned constructed affine candidate. The constructed candidate may also be referred to as a constructed affine MVP candidate. If the number of available inherited candidates is less than two, the encoding device / decoding device may add a constructed candidate to the affine MVP list of the current block. For example, the encoding device / decoding device may derive one constructed affine candidate.

[0169] Meanwhile, the method for deriving the constructed affine candidate may differ depending on whether the affine motion model applied to the current block is a 6-affine motion model or a 4-affine motion model. The method for deriving the constructed affine candidate will be described in detail later.

[0170] The encoding device / decoding device may add an HEVC AMVP candidate to the affine MVP list of the current block (S1320). If the number of available inherited and / or constructed candidates is less than two, the encoding device / decoding device may add an HEVC AMVP candidate to the affine MVP list of the current block. That is, if the number of available inherited and / or constructed candidates is less than two, the encoding device / decoding device performs a process of configuring MVP candidates in the existing HEVC standard.

[0171] Meanwhile, the proposed method for deriving the constructed candidates is as follows.

[0172] For example, if the affine motion model applied to the current block is a 6-affine motion model, the constructed candidates can be derived as in the embodiment shown in FIG.

[0173] FIG. 14 shows an example of deriving the constructed candidates.

[0174] 14, the encoding / decoding apparatus may check mv0, mv1, and mv2 for the current block (S1400). That is, the encoding / decoding apparatus may determine whether mv0, mv1, and mv2 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2. In addition, mv0, mv1, and mv2 may be indicated as candidate motion vectors for the CP.

[0175] For example, the encoding / decoding apparatus may check whether motion vectors of neighboring blocks in a first group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of a neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no usable mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0176] For example, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no usable mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0177] For example, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no usable mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0178] Meanwhile, the first group may include motion vectors of peripheral blocks A, B, and C, the second group may include motion vectors of peripheral blocks D and E, and the third group may include motion vectors of peripheral blocks F and G. The peripheral block A may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block B may indicate a peripheral block located above the upper left sample position of the current block, the peripheral block C may indicate a peripheral block located left corner of the upper left sample position of the current block, the peripheral block D may indicate a peripheral block located above the upper right sample position of the current block, the peripheral block E may indicate a peripheral block located above the upper right sample position of the current block, the peripheral block F may indicate a peripheral block located left corner of the lower left sample position of the current block, and the peripheral block G may indicate a peripheral block located below the lower left sample position of the current block.

[0179] If only the mv0 and mv1 for the current block are available, i.e., if only the mv0 and mv1 for the current block are derived, the encoding / decoding apparatus may derive mv2 for the current block based on the above-described Equation 6 (S1410). The encoding / decoding apparatus may derive mv2 by substituting the derived mv0 and mv1 into the above-described Equation 6.

[0180] If only mv0 and mv2 for the current block are available, i.e., if only mv0 and mv2 for the current block are derived, the encoding / decoding apparatus may derive mv1 for the current block based on Equation 7 (S1420). The encoding / decoding apparatus may derive mv1 by substituting the derived mv0 and mv2 into Equation 7.

[0181] The encoding device / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block (S1430). If the mv0, mv1, and mv2 are available, i.e., if the mv0, mv1, and mv2 are derived based on neighboring blocks of the current block, the encoding device / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block.

[0182] Also, if only mv0 and mv1 for the current block are available, i.e., if only mv0 and mv1 for the current block are derived, the encoding device / decoding device can derive mv2 derived based on the derived mv0, mv1 and the above-mentioned Equation 6 as the constructed candidate for the current block.

[0183] Also, if only mv0 and mv2 for the current block are available, i.e., if only mv0 and mv2 for the current block are derived, the encoding device / decoding device can derive the derived mv0, mv2 and mv1 derived based on the above-mentioned Equation 7 as the constructed candidate for the current block.

[0184] Also, for example, if the affine motion model applied to the current block is a 4-affine motion model, the constructed candidates can be derived as in the embodiment shown in FIG.

[0185] FIG. 15 shows an example of deriving the constructed candidates.

[0186] 15, the encoding / decoding device may check mv0, mv1, and mv2 for the current block (S1500). That is, the encoding / decoding device may determine whether mv0, mv1, and mv2 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0187] For example, the encoding / decoding apparatus may check whether motion vectors of neighboring blocks in a first group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of a neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no usable mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0188] Furthermore, for example, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no usable mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0189] For example, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no usable mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0190] Meanwhile, the first group may include motion vectors of peripheral blocks A, B, and C, the second group may include motion vectors of peripheral blocks D and E, and the third group may include motion vectors of peripheral blocks F and G. The peripheral block A may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block B may indicate a peripheral block located above the upper left sample position of the current block, the peripheral block C may indicate a peripheral block located left corner of the upper left sample position of the current block, the peripheral block D may indicate a peripheral block located above the upper right sample position of the current block, the peripheral block E may indicate a peripheral block located above the upper right sample position of the current block, the peripheral block F may indicate a peripheral block located left corner of the lower left sample position of the current block, and the peripheral block G may indicate a peripheral block located below the lower left sample position of the current block.

[0191] If only mv0 and mv1 for the current block are available, or if mv0, mv1, and mv2 for the current block are available, i.e., if only mv0 and mv1 for the current block are derived, or if mv0, mv1, and mv2 for the current block are derived, the encoding device / decoding device can derive the derived mv0 and mv1 as constructed candidates for the current block (S1510).

[0192] On the other hand, if only mv0 and mv2 for the current block are available, i.e., if only mv0 and mv2 for the current block are derived, the encoding / decoding apparatus may derive mv1 for the current block based on Equation 7 (S1520). The encoding / decoding apparatus may derive mv1 by substituting the derived mv0 and mv2 into Equation 7.

[0193] Thereafter, the encoding / decoding apparatus can derive the derived mv0 and mv1 as constructed candidates for the current block (S1510).

[0194] Meanwhile, this document proposes another embodiment for deriving the inherited affine candidates, which can reduce the computational complexity and improve coding performance when deriving the inherited affine candidates.

[0195] Meanwhile, this document proposes another embodiment for deriving the inherited affine candidates, which can reduce the computational complexity and improve coding performance when deriving the inherited affine candidates.

[0196] FIG. 16 exemplarily shows the locations of the surrounding blocks scanned to derive the inherited affine candidates.

[0197] The encoding / decoding apparatus may derive up to two inherited affine candidates from neighboring blocks of the current block. Figure 16 may show the neighboring blocks for the inherited affine candidates. For example, the neighboring blocks may include neighboring block A and neighboring block B shown in Figure 16. The neighboring block A may represent the left neighboring block A0 described above, and the neighboring block B may represent the upper neighboring block B0 described above.

[0198] For example, the encoding / decoding device may check whether the neighboring blocks are available in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. That is, the encoding / decoding device may check whether the neighboring blocks satisfy a specific condition in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. The encoding / decoding device may also derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies a specific condition. That is, the encoding / decoding device may derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies a specific condition. Here, the availability may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. That is, the specific condition may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. Also, for example, the specific order may be the neighboring block A → the neighboring block B. Meanwhile, a pruning check process between two inherited affine candidates (i.e., derived inherited affine candidates) may not be performed. The pruning check process may refer to a process of checking whether the candidates are identical to each other, and if they are the same candidate, removing the candidate derived in the later order.

[0199] The above-described embodiment proposes a scheme in which, instead of checking all existing neighboring blocks (i.e., neighboring block A, neighboring block B, neighboring block C, neighboring block D, and neighboring block E) to derive the inherited affine candidate, only two neighboring blocks (i.e., neighboring block A and neighboring block B) are checked to derive the inherited affine candidate. Here, the neighboring block C may indicate the neighboring block B1 in the upper right corner, the neighboring block D may indicate the neighboring block A1 in the lower left corner, and the neighboring block E may indicate the neighboring block B2 in the upper left corner.

[0200] In order to analyze the spatial correlation between the neighboring blocks and the current block through affine inter-prediction, the probability that affine prediction is applied to the current block when affine prediction is applied to each neighboring block may be referenced. The probability that affine prediction is applied to the current block when affine prediction is applied to each neighboring block may be derived as shown in the following table.

[0201] [Table 2]

[0202] Referring to Table 2, it can be seen that among the neighboring blocks, neighboring block A and neighboring block B have high spatial correlation with the current block. Therefore, by using only neighboring block A and neighboring block B with high spatial correlation to derive the inherited affine candidate, it is possible to obtain an effect of deriving high decoding performance while reducing processing time.

[0203] Meanwhile, the pruning check process may be performed to prevent the same candidate from existing in the candidate list. The pruning check process may be advantageous in terms of encoding efficiency since it can eliminate redundancy, but has the disadvantage of increasing computational complexity. In particular, the pruning check process for affine candidates must be performed on the affine type (e.g., whether the affine motion model is a 4-affine motion model or a 6-affine motion model), reference picture (or reference picture index), and MVs CP0, CP1, and CP2, resulting in very high computational complexity. Therefore, this embodiment proposes a scheme in which the pruning check process is not performed between an inherited affine candidate (e.g., inherited_A) derived based on the neighboring block A and an inherited affine candidate (e.g., inherited_B) derived based on the neighboring block B. In the case of neighboring blocks A and B, the distance between them is large, and therefore the spatial correlation is low, so it is unlikely that inherited_A and inherited_B are the same. Therefore, it may be appropriate not to perform a pruning check process between the inherited affine candidates.

[0204] Alternatively, based on the above-mentioned reasons, a scheme of performing a minimal pruning check process may be proposed. For example, the encoding device / decoding device may perform a pruning check process by comparing only the MVs of the CP0s of the inherited affine candidates.

[0205] However, in this document, another embodiment is proposed to derive the inherited affine candidates.

[0206] FIG. 17 exemplarily shows the locations of the surrounding blocks scanned to derive the inherited affine candidates.

[0207] The encoding / decoding apparatus may derive up to two inherited affine candidates from the neighboring blocks of the current block. Figure 17 may show the neighboring blocks for the inherited affine candidates. For example, the neighboring blocks may include neighboring blocks A to D shown in Figure 17. The neighboring block A may represent the left neighboring block A0, the neighboring block B may represent the upper neighboring block B0, the neighboring block C may represent the upper right neighboring block B1, and the neighboring block D may represent the lower left neighboring block A1.

[0208] For example, the encoding / decoding device may check whether the neighboring blocks are available in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. That is, the encoding / decoding device may check whether the neighboring blocks satisfy a specific condition in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. The encoding / decoding device may also derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies a specific condition. That is, the encoding / decoding device may derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies a specific condition. Here, the availability may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. That is, the specific condition may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block.

[0209] In Figure 17, surrounding blocks A and D can be used to derive the left predictor among the inherited affine candidates, and surrounding blocks B and C can be used to derive the upper predictor among the inherited affine candidates.

[0210] The left predictor, i.e., the motion candidate that can be added from the left peripheral block, can be added to the inherited candidate in the order of blocks A → D or D → A, with the "neighboring valid block" that is determined to be available first. The upper predictor, i.e., the motion candidate that can be added from the upper peripheral block, can be added to the inherited candidate in the order of blocks B → C or C → B, with the "neighboring valid block" that is determined to be available first. That is, the maximum number of inherited candidates that can be derived from each of the left predictor and the upper predictor is 1.

[0211] If the "neighboring valid block" is coded with a four-parameter affine motion model, the inherited candidate can be determined using the four-parameter affine motion model, and if the "neighboring valid block" is coded with a six-parameter affine motion model, the inherited candidate can be determined using the six-parameter affine motion model.

[0212] When there are two inherited candidates determined by the left predictor and the upper predictor, a pruning check process may or may not be performed. It is common to perform a pruning check process to prevent duplicate candidates from being added to a candidate list. However, the pruning check process increases complexity because the MVs of each CP must be compared in motion prediction using an affine model. However, when inherited candidates are constructed using the embodiment described with reference to FIG. 17, the distance between the candidates determined by the left predictor and the upper predictor is large, so the candidates are highly likely to be different from each other. Therefore, even if a pruning check process is not performed, there is an advantage that there is almost no degradation in coding performance.

[0213] However, in this document, another embodiment is proposed for deriving the inherited affine candidates.

[0214] FIG. 18 exemplarily shows the positions for deriving inherited affine candidates.

[0215] The encoding / decoding apparatus may derive up to two inherited affine candidates from the neighboring blocks of the current block. FIG. 18 illustrates the neighboring blocks for the inherited affine candidates according to this embodiment. For example, the neighboring blocks may include neighboring blocks A to E shown in FIG. 18. The neighboring block A may represent the left neighboring block A0, the neighboring block B may represent the upper neighboring block B0, the neighboring block C may represent the upper right neighboring block B1, the neighboring block D may represent the lower left neighboring block A1, and the neighboring block E may represent the left neighboring block located adjacent to the lower left neighboring block B2.

[0216] For example, the encoding / decoding device may check whether the neighboring blocks are available in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. That is, the encoding / decoding device may check whether the neighboring blocks satisfy a specific condition in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. The encoding / decoding device may also derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies a specific condition. That is, the encoding / decoding device may derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies a specific condition. Here, the availability may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. That is, the specific condition may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block.

[0217] The surrounding blocks A, D, and E in FIG. 18 can be used to derive the left predictor of the inherited affine candidates, and the surrounding blocks B and C can be used to derive the upper predictor of the inherited affine candidates.

[0218] The left predictor, i.e., the motion candidate that can be added from the left peripheral block, can be added to the inherited candidate by the "neighboring valid block" that is determined to be first available in the block order of A → D → E (or A → E → D, D → A → E). The upper predictor, i.e., the motion candidate that can be added from the upper peripheral block, can be added to the inherited candidate by the "neighboring valid block" that is determined to be first available in the block order of B → C or C → B. That is, the maximum number of inherited candidates that can be derived from each of the left predictor and the upper predictor is one.

[0219] If the "neighboring valid block" is coded with a four-parameter affine motion model, the inherited candidate can be determined using the four-parameter affine motion model, and if the "neighboring valid block" is coded with a six-parameter affine motion model, the inherited candidate can be determined using the six-parameter affine motion model.

[0220] When there are two inherited candidates determined by the left predictor and the upper predictor, a pruning check process may or may not be performed. It is common to perform a pruning check process to prevent duplicate candidates from being added to a candidate list. However, the pruning check process increases complexity because the MVs of each CP must be compared in motion prediction using an affine model. However, when inherited candidates are constructed using the embodiment described with reference to FIG. 18, the distance between the candidates determined by the left predictor and the upper predictor is large, so the candidates are highly likely to be different from each other. Therefore, even if a pruning check process is not performed, there is an advantage that there is almost no degradation in coding performance.

[0221] Alternatively, a low-complexity pruning check method can be used instead of performing the pruning check process, for example, by comparing only the MVs of CP0.

[0222] The reason why E is determined as the position of the neighboring block to be scanned for the inherited candidate is as follows: In the line buffer reduction method described below, if the reference blocks (i.e., neighboring blocks B and C) located above the current block are not present in the same CTU as the current block, the line buffer reduction method cannot be used. Therefore, when applying the line buffer reduction method while generating the inherited candidate, the positions of the neighboring blocks shown in FIG. 18 are used to maintain coding performance.

[0223] In addition, this method can generate up to one inherited candidate and use it as an affine MVP candidate. In this case, the motion vector of the first valid neighboring block in the order A→B→C→D→E can be used as the inherited candidate, regardless of whether it is a left predictor or an upper predictor.

[0224] However, in this document, another embodiment is proposed for deriving the inherited affine candidates.

[0225] In this embodiment, the neighboring blocks shown in FIG. 18 can be used to derive the inherited candidates.

[0226] That is, the encoding / decoding apparatus can derive up to two inherited affine candidates from the neighboring blocks of the current block.

[0227] Furthermore, the encoding / decoding apparatus may check whether the neighboring blocks are available in a specific order and derive inherited affine candidates for the current block based on the first identified available neighboring block. That is, the encoding / decoding apparatus may check whether the neighboring blocks satisfy a specific condition in a specific order and derive inherited affine candidates for the current block based on the first identified available neighboring block. The encoding / decoding apparatus may also derive inherited affine candidates for the current block based on the second identified neighboring block that satisfies a specific condition. That is, the encoding / decoding apparatus may derive inherited affine candidates for the current block based on the second identified neighboring block that satisfies a specific condition. Here, "availability" may mean that a block is coded using an affine motion model and that a reference picture for the block is the same as the reference picture for the current block. That is, the specific condition may mean that a block is coded using an affine motion model and that a reference picture for the block is the same as the reference picture for the current block.

[0228] As mentioned above, surrounding blocks A, D, and E can be used to derive the left predictors of the inherited affine candidates, and surrounding blocks B and C can be used to derive the upper predictors of the inherited affine candidates.

[0229] The left predictor, i.e., the motion candidate that can be added from the left peripheral block, can be added to the inherited candidate by the "neighboring valid block" that is determined to be first available in the block order of A → D → E (or A → E → D, D → A → E). The upper predictor, i.e., the motion candidate that can be added from the upper peripheral block, can be added to the inherited candidate by the "neighboring valid block" that is determined to be first available in the block order of B → C or C → B. That is, the maximum number of inherited candidates that can be derived from each of the left predictor and the upper predictor is one.

[0230] If the "neighboring valid block" is coded using a four-parameter affine motion model, the inherited candidate can be determined using the four-parameter affine motion model, and if the "neighboring valid block" is coded using a six-parameter affine motion model, the inherited candidate can be determined using the six-parameter affine motion model.

[0231] Also, in this embodiment, if there are two inherited candidates determined by the left predictor and the upper predictor, a pruning check process may or may not be performed. While it is common to perform a pruning check process to prevent duplicate candidates from being added to a candidate list, the pruning check process increases complexity because the MVs of each CP must be compared in motion prediction using an affine model. However, when inherited candidates are constructed using the embodiment described with reference to FIG. 18, the distance between the candidates determined by the left predictor and the upper predictor is large, so the candidates are highly likely to be different from each other. Therefore, even if a pruning check process is not performed, there is an advantage that there is almost no degradation in coding performance.

[0232] Meanwhile, instead of performing the pruning check process, a low-complexity pruning check method can be used. For example, only when neighboring block E is a "neighboring valid block," it is possible to determine whether neighboring block E is included in the same coding block as neighboring block A, and then perform the pruning check process. This method has low complexity because only one pruning check is performed. The reason for performing the pruning check only on neighboring block E is that the reference blocks of the predictors above neighboring block E (neighboring block B, neighboring block C) are located sufficiently far from the reference blocks of the predictors on the left side (neighboring block A, neighboring block D), making it very unlikely that they will form the same inherited candidate. On the other hand, in the case of neighboring block E, if it is included in the same block as neighboring block A, there is a possibility that they will form the same inherited candidate.

[0233] The reason why E is determined as the position of the neighboring block to be scanned for the inherited candidate is as follows: In the line buffer reduction method described below, if the reference blocks (i.e., neighboring blocks B and C) located above the current block are not in the same CTU as the current block, the line buffer reduction method cannot be used. Therefore, when applying the line buffer reduction method while generating the inherited candidate, the positions of the neighboring blocks shown in FIG. 18 are used to maintain coding performance.

[0234] In addition, this method can generate up to one inherited candidate and use it as an affine MVP candidate. In this case, the motion vector of the first valid neighboring block in the order A→B→C→D→E can be used as the inherited candidate, regardless of whether it is a left predictor or an upper predictor.

[0235] Meanwhile, according to one embodiment of the present document, the method for generating an affine MVP list described with reference to Figures 16 to 18 can also be applied to deriving inherited candidates for a merge candidate list based on an affine motion model. In this embodiment, the same process can be applied to generating an affine MVP list and a merge candidate list, which is advantageous in terms of design costs. An example of generating a merge candidate list based on an affine motion model is as follows, and this process can be applied to constructing inherited candidates when generating other merge candidates.

[0236] Specifically, the merge candidate list can be constructed as follows.

[0237] FIG. 19 shows an example of constructing a merge candidate list for the current block.

[0238] Referring to FIG. 19, the encoding / decoding device may add inherited merge candidates to a merge candidate list (S1900).

[0239] Specifically, the encoding / decoding apparatus can derive successive candidates based on neighboring blocks of the current block.

[0240] The neighboring blocks of the current block for deriving the inherited candidates are shown in Figure 11. That is, the neighboring blocks of the current block may include a neighboring block A0 at the lower left corner of the current block, a neighboring block A1 on the left side of the current block, a neighboring block B0 at the upper right corner of the current block, a neighboring block B1 above the current block, and a neighboring block B2 at the upper left corner of the current block.

[0241] The inherited candidates may be derived based on valid neighboring reconstructed blocks coded in affine mode. For example, the encoding / decoding device may sequentially check neighboring blocks A0, A1, B0, B1, and B2 or sequentially check A1, B1, B0, A0, and B2. If the neighboring blocks are coded in affine mode (i.e., the neighboring blocks are validly reconstructed using an affine motion model), two or three CPMVs for the current block may be derived based on the affine motion models of the neighboring blocks, and the CPMVs may be derived as inherited candidates for the current block. For example, up to five inherited candidates may be added to the merge candidate list. That is, up to five inherited candidates may be derived based on the neighboring blocks.

[0242] In this embodiment, the peripheral blocks of Figures 16 to 18 can be used instead of the peripheral blocks of Figure 11 to derive inherited candidates, and the embodiment described with reference to Figures 16 to 18 can be applied.

[0243] Thereafter, the encoding / decoding device can add the constructed candidate to the merge candidate list (S1910).

[0244] For example, if the number of merge candidates in the merge candidate list is less than five, the constructed candidate may be added to the merge candidate list. The constructed candidate may indicate a merge candidate generated by combining neighboring motion information (i.e., motion vectors and reference picture indexes of neighboring blocks) for each CP of the current block. The motion information for each CP may be derived based on spatial or temporal neighboring blocks for the corresponding CP. The motion information for each CP may be indicated as a candidate motion vector for the corresponding CP.

[0245] FIG. 20 shows neighboring blocks of the current block for deriving constructed candidates according to one embodiment of the present document.

[0246] 20, the neighboring blocks may include spatial neighboring blocks and temporal neighboring blocks. The spatial neighboring blocks may include neighboring block A0, neighboring block A1, neighboring block A2, neighboring block B0, neighboring block B1, neighboring block B2, and neighboring block B3. The neighboring block T shown in FIG. 20 may represent the temporal neighboring block.

[0247] Here, the peripheral block B2 may refer to a peripheral block located at the upper left of the upper left sample position of the current block, the peripheral block B3 may refer to a peripheral block located above the upper left sample position of the current block, and the peripheral block A2 may refer to a peripheral block located at the left of the upper left sample position of the current block. Also, the peripheral block B1 may refer to a peripheral block located above the upper right sample position of the current block, and the peripheral block B0 may refer to a peripheral block located at the upper right of the upper right sample position of the current block. Also, the peripheral block A1 may refer to a peripheral block located at the left of the lower left sample position of the current block, and the peripheral block A0 may refer to a peripheral block located at the lower left of the lower left sample position of the current block.

[0248] 20, the CPs of the current block may include CP0, CP1, CP2, and / or CP3. CP0 may indicate the position of the upper left corner of the current block, CP1 may indicate the position of the upper right corner of the current block, CP2 may indicate the position of the lower left corner of the current block, and CP3 may indicate the position of the lower right corner of the current block. For example, if the size of the current block is WxH and the x component and y component of the top-left sample position of the current block are 0, CP0 may indicate the position of (0,0) coordinates, CP1 may indicate the position of (W,0) coordinates, CP2 may indicate the position of (0,H) coordinates, and CP3 may indicate the position of (W,H) coordinates.

[0249] Candidate motion vectors for each of the above CPs can be derived as follows.

[0250] For example, the encoding / decoding apparatus may check whether neighboring blocks in a first group are available in a first order, and derive the motion vector of the available neighboring block first identified in the checking process as the candidate motion vector for CP0. That is, the candidate motion vector for CP0 may be the motion vector of the available neighboring block first identified in the checking process by checking neighboring blocks in the first group in a first order. The availability may indicate the existence of a motion vector for the neighboring block. That is, the available neighboring block may be a block coded using inter prediction (i.e., a block to which inter prediction is applied). Here, for example, the first group may include neighboring block B2, neighboring block B3, and neighboring block A2. The first order may be from neighboring block B2 to neighboring block B3 to neighboring block A2 in the first group. As an example, if the surrounding block B2 is available, the motion vector of the surrounding block B2 can be derived as a candidate motion vector for the CP0; if the surrounding block B2 is not available and the surrounding block B3 is available, the motion vector of the surrounding block B3 can be derived as a candidate motion vector for the CP0; if the surrounding blocks B2 and B3 are not available and the surrounding block A2 is available, the motion vector of the surrounding block A2 can be derived as a candidate motion vector for the CP0.

[0251] Furthermore, for example, the encoding / decoding apparatus may check whether neighboring blocks in the second group are available in a second order, and derive the motion vector of the available neighboring block first identified in the checking process as the candidate motion vector for CP1. That is, the candidate motion vector for CP1 may be the motion vector of the available neighboring block first identified by checking neighboring blocks in the second group in the second order. The availability may indicate the existence of a motion vector for the neighboring block. That is, the available neighboring block may be a block coded using inter prediction (i.e., a block to which inter prediction is applied). Here, the second group may include the neighboring block B1 and the neighboring block B0. The second order may be from the neighboring block B1 to the neighboring block B0 in the second group. As an example, if the surrounding block B1 is available, the motion vector of the surrounding block B1 can be derived as a candidate motion vector for the CP1, and if the surrounding block B1 is not available but the surrounding block B0 is available, the motion vector of the surrounding block B0 can be derived as a candidate motion vector for the CP1.

[0252] Furthermore, for example, the encoding / decoding apparatus may check whether neighboring blocks in a third group are available in a third order, and derive the motion vector of the available neighboring block first identified in the checking process as the candidate motion vector for CP2. That is, the candidate motion vector for CP2 may be the motion vector of the available neighboring block first identified by checking neighboring blocks in the third group in a third order. The availability may indicate the existence of a motion vector for the neighboring block. That is, the available neighboring block may be a block coded using inter prediction (i.e., a block to which inter prediction is applied). Here, the third group may include the neighboring block A1 and the neighboring block A0. The third order may be from the neighboring block A1 to the neighboring block A0 in the third group. As an example, if the surrounding block A1 is available, the motion vector of the surrounding block A1 can be derived as a candidate motion vector for the CP2, and if the surrounding block A1 is not available but the surrounding block A0 is available, the motion vector of the surrounding block A0 can be derived as a candidate motion vector for the CP2.

[0253] Also, for example, the encoding device / decoding device can check whether the temporal surrounding block (i.e., the surrounding block T) is available, and if the temporal surrounding block (i.e., the surrounding block T) is available, it can derive the motion vector of the temporal surrounding block (i.e., the surrounding block T) as a candidate motion vector for CP3.

[0254] A combination of the candidate motion vector for CP0, the candidate motion vector for CP1, the candidate motion vector for CP2, and / or the candidate motion vector for CP3 can be derived as a constructed candidate.

[0255] For example, as described above, a 6-affine model requires motion vectors of three CPs. Three CPs can be selected from CP0, CP1, CP2, and CP3 for the 6-affine model. For example, the CPs can be selected as one of {CP0, CP1, CP3}, {CP0, CP1, CP2}, {CP1, CP2, CP3}, and {CP0, CP2, CP3}. As an example, the 6-affine model can be constructed using CP0, CP1, and CP2. In this case, the CPs can be represented as {CP0, CP1, CP2}.

[0256] Also, for example, as described above, a 4-affine model requires motion vectors of two CPs. Two CPs can be selected from CP0, CP1, CP2, and CP3 for the 4-affine model. For example, the CPs can be selected as one of {CP0, CP3}, {CP1, CP2}, {CP0, CP1}, {CP1, CP3}, {CP0, CP2}, and {CP2, CP3}. As an example, the 4-affine model can be constructed using CP0 and CP1. In this case, the CPs can be represented as {CP0, CP1}.

[0257] Constructed candidates, which are combinations of candidate motion vectors, may be added to the merge candidate list in the following order: That is, after the candidate motion vectors for the CP are derived, constructed candidates may be derived in the following order:

[0258] {CP0, CP1, CP2}, {CP0, CP1, CP3}, {CP0, CP2, CP3}, {CP1, CP2, CP3}, {CP0, CP1}, {CP0, CP2}, {CP1, CP2}, {CP0, CP3}, {CP1, CP3}, {CP2, CP3}

[0259] That is, for example, constructed candidates including the candidate motion vector for CP0, the candidate motion vector for CP1, and the candidate motion vector for CP2, constructed candidates including the candidate motion vector for CP0, the candidate motion vector for CP1, and the candidate motion vector for CP3, constructed candidates including the candidate motion vector for CP0, the candidate motion vector for CP2, and the candidate motion vector for CP3, constructed candidates including the candidate motion vector for CP1, the candidate motion vector for CP2, and the candidate motion vector for CP3, constructed candidates including the candidate motion vector for CP0, ... The constructed candidates may be added to the merge candidate list in the following order: constructed candidates including the candidate motion vector, constructed candidates including the candidate motion vector for CP0 and the candidate motion vector for CP2, constructed candidates including the candidate motion vector for CP1 and the candidate motion vector for CP2, constructed candidates including the candidate motion vector for CP0 and the candidate motion vector for CP3, constructed candidates including the candidate motion vector for CP1 and the candidate motion vector for CP3, constructed candidates including the candidate motion vector for CP2 and the candidate motion vector for CP3.

[0260] Thereafter, the encoding / decoding device can add the 0 motion vector as a merging candidate to the merging candidate list (S1920).

[0261] For example, if the number of merge candidates in the merge candidate list is less than 5, merge candidates including zero motion vectors may be added to the merge candidate list until the merge candidate list is configured with the maximum number of merge candidates, which may be 5. In addition, the zero motion vector may indicate a motion vector whose vector value is 0.

[0262] Meanwhile, the scanning method for positioning neighboring blocks and configuring candidates used in the affine MVP list generation method described with reference to Figures 16 to 18 can also be used for normal merge and normal MVP. Here, normal merge refers to a merge mode that can be used in HEVC, etc., rather than an affine merge mode, and normal MVP can also refer to AMVP that can be used in HEVC, etc., rather than affine MVP. For example, applying the method described with reference to Figure 16 to normal merge and / or normal MVP specifically means scanning neighboring blocks at spatial positions in Figure 16, configuring left predictors and upper predictors using neighboring blocks in Figure 16, performing pruning checks, or performing them in a low-complexity manner. Applying such a method to normal merge or MVP can be effective in terms of design costs.

[0263] This document also proposes a method for deriving constructed candidates that is different from the above-described embodiment. The proposed embodiment can reduce complexity and improve coding performance compared to the above-described embodiment for deriving constructed candidates. The proposed embodiment is described below. Furthermore, if the available number of inherited affine candidates is less than two (i.e., if the number of derived inherited affine candidates is less than two), a constructed affine candidate can be considered.

[0264] For example, the encoding / decoding device may check mv0, mv1, and mv2 for the current block. That is, the encoding / decoding device may determine whether mv0, mv1, and mv2 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0265] Specifically, the neighboring blocks of the current block are divided into three groups, and the neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G. The first group may include the motion vectors of neighboring block A, neighboring block B, and neighboring block C, the second group may include the motion vectors of neighboring block D and neighboring block E, and the third group may include the motion vectors of neighboring block F and neighboring block G. The peripheral block A may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block B may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block C may indicate a peripheral block located in the left corner of the upper left sample position of the current block, the peripheral block D may indicate a peripheral block located in the upper right corner of the current block, the peripheral block E may indicate a peripheral block located in the upper right corner of the upper right sample position of the current block, the peripheral block F may indicate a peripheral block located in the left corner of the lower left sample position of the current block, and the peripheral block G may indicate a peripheral block located in the lower left corner of the lower left sample position of the current block.

[0266] The encoding device / decoding device can determine whether there is an available mv0 in the first group, can determine whether there is an available mv1 in the second group, and can determine whether there is an available mv2 in the third group.

[0267] Specifically, for example, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the first group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed by checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no usable mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0268] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no usable mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0269] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no usable mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0270] Hereinafter, when the affine motion model applied to the current block is a 4-affine motion model, if mv0 and mv1 for the current block are available, the encoding / decoding device may derive the derived mv0 and mv1 as constructed candidates for the current block. On the other hand, if mv0 and / or mv1 for the current block are not available, i.e., if at least one of mv0 and mv1 is not derived from a neighboring block of the current block, the encoding / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0271] Furthermore, when the affine motion model applied to the current block is a 6-affine motion model, if mv0, mv1, and mv2 for the current block are available, the encoding / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block. On the other hand, if mv0, mv1, and / or mv2 for the current block are not available, i.e., if at least one of mv0, mv1, and mv2 is not derived from a neighboring block of the current block, the encoding / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0272] The above-described proposed embodiment is a method in which a constructed candidate is considered only if all motion vectors of CPs for generating an affine motion model for the current block are available. Here, "available" may refer to the fact that the reference picture of a neighboring block is the same as the reference picture of the current block. That is, the constructed candidate can be derived only if there is a motion vector that satisfies the above condition among the motion vectors of neighboring blocks for each CP of the current block. Therefore, if the affine motion model applied to the current block is a 4-affine motion model, the constructed candidate can be considered only if the motion vectors of CP0 and CP1 of the current block (i.e., mv0 and mv1) are available. Also, if the affine motion model applied to the current block is a 6-affine motion model, the constructed candidate can be considered only if the motion vectors of CP0, CP1, and CP2 of the current block (i.e., mv0, mv1, and mv2) are available. Therefore, according to the proposed embodiment, an additional configuration for deriving a motion vector for a CP based on the above-described Equation 6 or Equation 7 may not be necessary. This can reduce the computational complexity for deriving the constructed candidate. Also, since the constructed candidate is determined only when a CPMVP candidate having the same reference picture is available, overall coding performance can be improved.

[0273] Meanwhile, a pruning check process between the derived inherited affine candidate and the constructed affine candidate may not be performed. The pruning check process may refer to a process of checking whether they are identical to each other, and if they are the same candidate, removing the candidate derived later.

[0274] The above-described embodiment can be illustrated as shown in FIGS.

[0275] FIG. 21 shows an example of deriving the constructed candidates when four affine motion models are applied to the current block.

[0276] 21, an encoding / decoding apparatus may determine whether mv0 and mv1 are available for the current block (S2100). That is, the encoding / decoding apparatus may determine whether mv0 and mv1 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, and mv1 may be a CPMVP candidate for CP1.

[0277] The encoding / decoding device can determine whether there is an available mv0 in the first group and whether there is an available mv1 in the second group.

[0278] Specifically, the neighboring blocks of the current block are divided into three groups, and the neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G. The first group may include the motion vectors of neighboring block A, neighboring block B, and neighboring block C, the second group may include the motion vectors of neighboring block D and neighboring block E, and the third group may include the motion vectors of neighboring block F and neighboring block G. The peripheral block A may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block B may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block C may indicate a peripheral block located in the left corner of the upper left sample position of the current block, the peripheral block D may indicate a peripheral block located in the upper right corner of the current block, the peripheral block E may indicate a peripheral block located in the upper right corner of the upper right sample position of the current block, the peripheral block F may indicate a peripheral block located in the left corner of the lower left sample position of the current block, and the peripheral block G may indicate a peripheral block located in the lower left corner of the lower left sample position of the current block.

[0279] The encoding / decoding apparatus may check whether motion vectors of neighboring blocks in the first group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no usable mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0280] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no usable mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0281] If the mv0 and mv1 for the current block are available, i.e., if the mv0 and mv1 for the current block are derived, the encoding / decoding device may derive the derived mv0 and mv1 as constructed candidates for the current block (S2110). On the other hand, if the mv0 and / or mv1 for the current block are not available, i.e., if at least one of mv0 and mv1 is not derived from a neighboring block of the current block, the encoding / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0282] Meanwhile, the pruning check process between the derived inherited affine candidate and the constructed affine candidate may not be performed. The pruning check process may check whether they are identical to each other, and if they are the same candidate, remove the candidate derived later.

[0283] FIG. 22 shows an example of deriving the constructed candidates when a 6-affine motion model is applied to the current block.

[0284] 22, an encoding / decoding apparatus may determine whether mv0, mv1, and mv2 are available for the current block (S2200). That is, the encoding / decoding apparatus may determine whether mv0, mv1, and mv2 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0285] The encoding device / decoding device can determine whether there is an available mv0 in the first group, whether there is an available mv1 in the second group, and whether there is an available mv2 in the third group.

[0286] Specifically, the neighboring blocks of the current block are divided into three groups, and the neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G. The first group may include the motion vectors of neighboring block A, neighboring block B, and neighboring block C, the second group may include the motion vectors of neighboring block D and neighboring block E, and the third group may include the motion vectors of neighboring block F and neighboring block G. The peripheral block A may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block B may indicate a peripheral block located in the upper left corner of the upper left sample position of the current block, the peripheral block C may indicate a peripheral block located in the left corner of the upper left sample position of the current block, the peripheral block D may indicate a peripheral block located in the upper right corner of the current block, the peripheral block E may indicate a peripheral block located in the upper right corner of the upper right sample position of the current block, the peripheral block F may indicate a peripheral block located in the left corner of the lower left sample position of the current block, and the peripheral block G may indicate a peripheral block located in the lower left corner of the lower left sample position of the current block.

[0287] The encoding / decoding apparatus may check whether motion vectors of neighboring blocks in the first group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no usable mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0288] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no usable mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0289] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no usable mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0290] If mv0, mv1, and mv2 for the current block are available, i.e., if mv0, mv1, and mv2 for the current block are derived, the encoding device / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block (S2210). On the other hand, if mv0, mv1, and / or mv2 for the current block are not available, i.e., if at least one of mv0, mv1, and mv2 is not derived from a neighboring block of the current block, the encoding device / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0291] Meanwhile, the pruning check process between the derived inherited affine candidate and the constructed affine candidate may not be performed.

[0292] On the other hand, if the number of derived affine candidates is less than two (i.e., if the number of inherited affine candidates and / or constructed affine candidates is less than two), an HEVC AMVP candidate may be added to the affine MVP list of the current block.

[0293] For example, the HEVC AMVP candidates may be derived in the following order:

[0294] Specifically, when the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is available, the CPMV0 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is available (i.e., when the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is derived), a first affine MVP candidate including the CPMV0 of the constructed affine candidate in CPMV0, CPMV1, and CPMV2 can be derived.

[0295] Furthermore, if the number of derived affine candidates is less than 2 and the constructed affine candidate CPMV1 is available, the CPMV1 can be used as the affine MVP candidate. That is, if the number of derived affine candidates is less than 2 and the constructed affine candidate CPMV1 is available (i.e., the number of derived affine candidates is less than 2 and the constructed affine candidate CPMV1 is derived), a second affine MVP candidate can be derived that includes the constructed affine candidate CPMV1 in CPMV0, CPMV1, and CPMV2.

[0296] Next, when the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is available, the CPMV2 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is available (i.e., when the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is derived), a third affine MVP candidate can be derived that includes the CPMV2 of the constructed affine candidate in CPMV0, CPMV1, and CPMV2.

[0297] Next, if the number of derived affine candidates is less than two, an HEVC Temporal Motion Vector Predictor (TMVP) can be used as the affine MVP candidate. The HEVC TMVP can be derived based on motion information of temporally neighboring blocks of the current block. That is, if the number of derived affine candidates is less than two, a third affine MVP candidate including motion vectors of temporally neighboring blocks of the current block as CPMV0, CPMV1, and CPMV2 can be derived. The temporal neighboring blocks can indicate collocated blocks in a collocated picture corresponding to the current block.

[0298] Next, if the number of derived affine candidates is less than 2, a zero motion vector (zero MV) can be used as the affine MVP candidate. That is, if the number of derived affine candidates is less than 2, a third affine MVP candidate including the zero motion vector as CPMV0, CPMV1, and CPMV2 can be derived. The zero motion vector can indicate a motion vector with a value of 0.

[0299] This is because the step of using the CPMV of the constructed affine candidate reuses the MV already considered for generating the constructed affine candidate, which can reduce complexity compared to existing methods of deriving HEVC AMVP candidates.

[0300] However, this document proposes another embodiment for deriving the inherited affine candidates.

[0301] In order to derive the inherited affine candidates, information on affine prediction of neighboring blocks is required. Specifically, the following affine prediction information is required.

[0302] 1) An affine flag (affine_flag) indicating whether affine prediction-based encoding of the neighboring blocks is applied.

[0303] 2) Motion information of the surrounding blocks

[0304] When a 4-affine motion model is applied to the surrounding blocks, the motion information of the surrounding blocks may include L0 motion information and L1 motion information for CP0 and L0 motion information and L1 motion information for CP1. When a 6-affine motion model is applied to the surrounding blocks, the motion information of the surrounding blocks may include L0 motion information and L1 motion information for CP0 and L0 motion information and L1 motion information for CP2. Here, the L0 motion information may indicate motion information for L0 (List 0), and the L1 motion information may indicate motion information for L1 (List 1). The L0 motion information may include an L0 reference picture index and an L0 motion vector, and the L1 motion information may include an L1 reference picture index and an L1 motion vector.

[0305] As described above, affine prediction requires a large amount of information to be stored, which can be a major cause of increased hardware costs in actual implementations of encoding / decoding devices. In particular, if a neighboring block is located above the current block and is a CTU boundary, a line buffer must be used to store affine prediction-related information for the neighboring block, which can lead to even greater cost problems. This problem may be referred to as a line buffer issue. Therefore, this document proposes an embodiment in which affine prediction-related information is not stored or is reduced in the line buffer, thereby minimizing hardware costs and deriving inherited affine candidates. The proposed embodiment can reduce computational complexity and improve coding performance when deriving the inherited affine candidates. For reference, the line buffer already stores motion information for 4x4 blocks. If the affine prediction-related information is further stored, the amount of information stored may increase by three times compared to the existing storage amount.

[0306] In this embodiment, no further information for affine prediction may be stored in the line buffer, and if information in the line buffer must be referenced to generate the inherited affine candidate, the generation of the inherited affine candidate may be restricted.

[0307] 23a-23b exemplarily show an embodiment for deriving the inherited affine candidates.

[0308] Referring to FIG. 23a, if a neighboring block B of the current block (i.e., a neighboring block above the current block) is not present in the same CTU as the current block (i.e., the current CTU), the neighboring block B may not be used to generate the inherited affine candidate. Meanwhile, although a neighboring block A is not present in the same CTU as the current block, it may be used to generate the inherited affine candidate because information about the neighboring block A is not stored in a line buffer. Therefore, in this embodiment, the neighboring block above the current block may be used to derive the inherited affine candidate only if it is included in the same CTU as the current block. Furthermore, if a neighboring block above the current block is not included in the same CTU as the current block, the neighboring block above may not be used to derive the inherited affine candidate.

[0309] 23b, a neighboring block B of the current block (i.e., a neighboring block above the current block) may be present in the same CTU as the current block, in which case the encoding / decoding device may refer to the neighboring block B to generate the inherited affine candidate.

[0310] FIG. 24 schematically illustrates an image encoding method by an encoding apparatus according to the present document. The method disclosed in FIG. 24 is performed by the encoding apparatus disclosed in FIG. 2. Specifically, for example, steps S2400 to S2430 of FIG. 24 may be performed by a prediction unit of the encoding apparatus, and step S2440 may be performed by an entropy encoding unit of the encoding apparatus. Also, although not shown, the step of deriving predicted samples for the current block based on the CPMV may be performed by a prediction unit of the encoding apparatus, the step of deriving residual samples for the current block based on original samples and predicted samples for the current block may be performed by a subtraction unit of the encoding apparatus, the step of generating information about the residual for the current block based on the residual samples may be performed by a transform unit of the encoding apparatus, and the step of encoding information about the residual may be performed by an entropy encoding unit of the encoding apparatus.

[0311] The encoding apparatus constructs an affine motion vector predictor (MVP) candidate list for a current block (S2400). The encoding apparatus may construct an affine MVP candidate list including affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list may be two.

[0312] Also, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The encoding apparatus may check whether the inherited affine MVP candidates of the current block are available, and if the inherited affine MVP candidates are available, the inherited affine MVP candidates may be derived. For example, the inherited affine MVP candidates may be derived based on neighboring blocks of the current block, and the maximum number of inherited affine MVP candidates may be two. The neighboring blocks may be checked for availability in a specific order, and the inherited affine MVP candidates may be derived based on the checked available neighboring blocks. That is, the neighboring blocks may be checked for availability in a specific order, and a first inherited affine MVP candidate may be derived based on the first checked available neighboring block, and a second inherited affine MVP candidate may be derived based on the second checked available neighboring block. The availability may indicate that the neighboring block is coded using an affine motion model and that the reference picture of the neighboring block is the same as the reference picture of the current block. That is, the available neighboring block may be coded using an affine motion model (i.e., affine prediction is applied) and that the reference picture is the same as the reference picture of the current block. Specifically, the encoding apparatus may derive a motion vector for the CP of the current block based on the affine motion model of the first checked available neighboring block, and derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. Also, the encoding apparatus may derive a motion vector for the CP of the current block based on the affine motion model of the second checked available neighboring block, and derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model may be derived as shown in Equation 1 or 3 above.

[0313] In other words, the neighboring blocks may be checked to see if they satisfy a specific condition in a specific order, and the inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the checked specific condition. That is, the neighboring blocks may be checked to see if they satisfy the specific condition in a specific order, and a first inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the specific condition that are checked first, and a second inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the specific condition that are checked second. Specifically, the encoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring blocks that satisfy the specific condition that are checked first, and may derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. Furthermore, the encoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring blocks that satisfy the specific condition that are checked second, and may derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model may be derived as shown in Equation 1 or 3. Meanwhile, the specific condition may indicate that the current block is coded using an affine motion model and that the reference picture of the neighboring block is the same as the reference picture of the current block. That is, the neighboring block satisfying the specific condition may be coded using an affine motion model (i.e., affine prediction is applied) and that the reference picture is the same as the reference picture of the current block.

[0314] Here, for example, the neighboring blocks may include a left neighboring block, an upper neighboring block, a right upper corner neighboring block, a lower left corner neighboring block, and an upper left corner neighboring block of the current block. In this case, the identifying order may be from the left neighboring block to the lower left corner neighboring block, the upper neighboring block, the right upper corner neighboring block, and the upper left corner neighboring block.

[0315] Alternatively, for example, the peripheral blocks may include only the left peripheral block and the upper peripheral block, in which case the specific order may be from the left peripheral block to the upper peripheral block.

[0316] Alternatively, for example, the peripheral blocks may include the left peripheral block, and if the upper peripheral block is included in a current CTU including the current block, the peripheral blocks may further include the upper peripheral block. In this case, the specified order may be from the left peripheral block to the upper peripheral block. Also, if the upper peripheral block is not included in the current CTU, the peripheral blocks may not include the upper peripheral block. In this case, only the left peripheral block may be checked.

[0317] On the other hand, if the size is WxH and the x component and y component of the top-left sample position of the current block are 0, the lower-left neighboring block may be a block including a sample at (-1, H) coordinates, the left neighboring block may be a block including a sample at (-1, H-1) coordinates, the upper-right neighboring block may be a block including a sample at (W, -1) coordinates, the upper neighboring block may be a block including a sample at (W-1, -1) coordinates, and the upper-left neighboring block may be a block including a sample at (-1, -1) coordinates. That is, the left neighboring block may be the left-most neighboring block among the left neighboring blocks of the current block, and the upper neighboring block may be the left-most neighboring block among the upper neighboring blocks of the current block.

[0318] Also, as an example, if a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. The encoding apparatus may check whether a constructed affine MVP candidate for the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate may be derived. Also, for example, the constructed affine MVP candidate may be derived after the inherited affine MVP candidate is derived. If the number of derived affine MVP candidates (i.e., the inherited affine MVP candidates) is less than two and the constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. Here, the constructed affine MVP candidate may include candidate motion vectors for the CP. The constructed affine MVP candidate may be available if all of the candidate motion vectors are available.

[0319] For example, if a 4-affine motion model is applied to the current block, the CPs of the current block may include CP0 and CP1. If a candidate motion vector for CP0 is available and a candidate motion vector for CP1 is available, the constructed affine MVP candidates may be available, and the affine MVP candidate list may include the constructed affine MVP candidates. Here, CP0 may indicate a position in the upper left corner of the current block, and CP1 may indicate a position in the upper right corner of the current block.

[0320] The constructed affine MVP candidates may include a candidate motion vector for the CP0 and a candidate motion vector for the CP1, where the candidate motion vector for the CP0 may be a motion vector of a first block, and the candidate motion vector for the CP1 may be a motion vector of a second block.

[0321] Furthermore, the first block may be determined by checking neighboring blocks in the first group in a first specific order, and the first identified reference picture may be the same as the reference picture of the current block. That is, the candidate motion vector for CP1 may be determined by checking neighboring blocks in the first group in a first order, and the first identified reference picture may be the same as the reference picture of the current block. The availability may indicate that the neighboring block exists and is coded using inter-prediction. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0322] Furthermore, the second block may be determined by checking neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0323] On the other hand, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at a (-1, -1) coordinate, the surrounding block B may be a block including a sample at a (0, -1) coordinate, the surrounding block C may be a block including a sample at a (-1, 0) coordinate, the surrounding block D may be a block including a sample at a (W-1, -1) coordinate, and the surrounding block E may be a block including a sample at a (W, -1) coordinate. That is, the peripheral block A may be the peripheral block in the upper left corner of the current block, the peripheral block B may be the upper peripheral block located on the leftmost side among the peripheral blocks above the current block, the peripheral block C may be the left peripheral block located on the uppermost side among the peripheral blocks on the left side of the current block, the peripheral block D may be the upper peripheral block located on the rightmost side among the peripheral blocks above the current block, and the peripheral block E may be the peripheral block in the upper right corner of the current block.

[0324] On the other hand, if at least one of the candidate motion vectors of CP0 and the candidate motion vectors of CP1 is unavailable, the constructed affine MVP candidate may not be available.

[0325] Alternatively, for example, if a 6-affine motion model is applied to the current block, the CPs of the current block may include CP0, CP1, and CP2. If a candidate motion vector for CP0 is available, a candidate motion vector for CP1 is available, and a candidate motion vector for CP2 is available, the constructed affine MVP candidates may be available, and the affine MVP candidate list may include the constructed affine MVP candidates. Here, CP0 may indicate a position in the upper left corner of the current block, CP1 may indicate a position in the upper right corner of the current block, and CP2 may indicate a position in the lower left corner of the current block.

[0326] The constructed affine MVP candidates may include a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, where the candidate motion vector for CP0 may be a motion vector of a first block, the candidate motion vector for CP1 may be a motion vector of a second block, and the candidate motion vector for CP2 may be a motion vector of a third block.

[0327] Furthermore, the first block may check neighboring blocks in the first group according to a first specific order, and a first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, a candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0328] Furthermore, the second block may check neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0329] Furthermore, the third block may be determined by checking neighboring blocks in the third group according to a third specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the third block in the third group is the same as the reference picture of the current block, a candidate motion vector for CP2 may be available. For example, the third group may include neighboring blocks F and G, and the third specific order may be from neighboring block F to neighboring block G.

[0330] On the other hand, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at a (-1, -1) coordinate, the surrounding block B may be a block including a sample at a (0, -1) coordinate, the surrounding block C may be a block including a sample at a (-1, 0) coordinate, the surrounding block D may be a block including a sample at a (W-1, -1) coordinate, the surrounding block E may be a block including a sample at a (W, -1) coordinate, the surrounding block F may be a block including a sample at a (-1, H-1) coordinate, and the surrounding block G may be a block including a sample at a (-1, H) coordinate. That is, the peripheral block A may be the peripheral block at the upper left corner of the current block, the peripheral block B may be the upper peripheral block located at the leftmost position among the peripheral blocks above the current block, the peripheral block C may be the left peripheral block located at the topmost position among the peripheral blocks to the left of the current block, the peripheral block D may be the upper peripheral block located at the rightmost position among the peripheral blocks above the current block, the peripheral block E may be the peripheral block at the upper right corner of the current block, the peripheral block F may be the left peripheral block located at the bottommost position among the peripheral blocks to the left of the current block, and the peripheral block G may be the peripheral block at the lower left corner of the current block.

[0331] On the other hand, if at least one of the candidate motion vectors of CP0, CP1, and CP2 is unavailable, the constructed affine MVP candidate may not be available.

[0332] Thereafter, the affine MVP candidate list can be derived based on the following sequential steps.

[0333] For example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP0 is available, the encoding apparatus may derive a first affine MVP candidate, which may be an affine MVP candidate that includes the motion vector for the CP0 in the candidate motion vector for the CP.

[0334] Also, for example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP1 is available, the encoding apparatus may derive a second affine MVP candidate, which may be an affine MVP candidate that includes the motion vector for the CP1 in the candidate motion vector for the CP.

[0335] Furthermore, for example, if the number of derived affine MVP candidates is less than two and a motion vector for CP2 is available, the encoding apparatus may derive a third affine MVP candidate, which may be an affine MVP candidate that includes the motion vector for CP2 in the candidate motion vector for the CP.

[0336] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the encoding apparatus may derive a fourth affine MVP candidate including a temporal MVP derived based on temporal neighboring blocks of the current block as a candidate motion vector for the CP. The temporal neighboring blocks may indicate collocated blocks in a collocated picture corresponding to the current block. The temporal MVP may be derived based on the motion vectors of the temporal neighboring blocks.

[0337] Also, for example, if the number of derived affine MVP candidates is less than two, the encoding apparatus may derive a fifth affine MVP candidate including a zero motion vector as a candidate motion vector for the CP. The zero motion vector may indicate a motion vector with a value of 0.

[0338] The encoding apparatus derives control point motion vector predictors (CPMVPs) for the control points (CPs) of the current block based on the affine MVP candidate list (S2410). The encoding apparatus may derive a CPMV for the CP of the current block having an optimal RD cost and may select an affine MVP candidate most similar to the CPMV from the affine MVP candidates as an affine MVP candidate for the current block. The encoding apparatus may derive control point motion vector predictors (CPMVPs) for the CP of the current block based on the selected affine MVP candidate from the affine MVP candidates included in the affine MVP candidate list. Specifically, if the affine MVP candidates include a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate may be derived as the CPMVP of CP0, and the candidate motion vector for CP1 of the affine MVP candidate may be derived as the CPMVP of CP1. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived as the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2.

[0339] The encoding apparatus may encode an affine MVP candidate index that indicates the selected affine MVP candidate from among the affine MVP candidates, and the affine MVP candidate index may indicate the one affine MVP candidate from among affine MVP candidates included in an affine motion vector predictor (MVP) candidate list for the current block.

[0340] The encoding apparatus derives a CPMV for the CP of the current block (S2420). The encoding apparatus can derive a CPMV for each of the CPs of the current block.

[0341] The encoding apparatus derives control point motion vector differences (CPMVDs) for the CPs of the current block based on the CPMVP and the CPMV (S2430). The encoding apparatus can derive CPMVDs for the CPs of the current block based on the CPMVP and the CPMV for each of the CPs.

[0342] The encoding apparatus encodes motion prediction information including information on the CPMVD (S2440). The encoding apparatus may output the motion prediction information including information on the CPMVD in the form of a bitstream. That is, the encoding apparatus may output image information including the motion prediction information in the form of a bitstream. The encoding apparatus may encode information on the CPMVD for each of the CPs, and the motion prediction information may include information on the CPMVD.

[0343] The motion prediction information may also include the affine MVP candidate index, which may indicate the selected affine MVP candidate from among affine MVP candidates included in an affine motion vector predictor (MVP) candidate list for the current block.

[0344] Meanwhile, as an example, an encoding device may derive predicted samples for the current block based on the CPMV, derive residual samples for the current block based on original samples and predicted samples for the current block, generate information about the residual for the current block based on the residual samples, and encode the information about the residual. The image information may include information about the residual. Meanwhile, the bitstream may be transmitted to a decoding device via a network or a (digital) storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0345] FIG. 25 schematically illustrates an encoding apparatus that performs the image encoding method according to the present document. The method disclosed in FIG. 24 is performed by the encoding apparatus disclosed in FIG. 25. Specifically, for example, a prediction unit of the encoding apparatus of FIG. 25 may perform steps S2400 to S2430 of FIG. 24, and an entropy encoding unit of the encoding apparatus of FIG. 25 may perform step S2440 of FIG. 24. Although not shown, the process of deriving prediction samples for the current block based on the CPMV may be performed by the prediction unit of the encoding apparatus of FIG. 25, the process of deriving residual samples for the current block based on original samples and prediction samples for the current block may be performed by a subtraction unit of the encoding apparatus of FIG. 25, the process of generating information about the residual for the current block based on the residual samples may be performed by a transform unit of the encoding apparatus of FIG. 25, and the process of encoding information about the residual may be performed by the entropy encoding unit of the encoding apparatus of FIG. 25.

[0346] Figure 26 outlines a method for decoding an image by a decoding device according to this document. The method disclosed in Figure 26 is performed by the decoding device disclosed in Figure 3. Specifically, for example, S2600 in Figure 26 may be performed by an entropy decoding unit of the decoding device, S2610 to S2650 may be performed by a prediction unit of the decoding device, and S2660 may be performed by an adder of the decoding device. Also, although not shown, the process of acquiring information about the residual of the current block via a bitstream may be performed by the entropy decoding unit of the decoding device, and the process of deriving the residual sample for the current block based on the residual information may be performed by an inverse transform unit of the decoding device.

[0347] The decoding apparatus acquires motion prediction information for a current block from a bitstream (S2600). The decoding apparatus can acquire image information including the motion prediction information from the bitstream.

[0348] Also, for example, the motion prediction information may include information on control point motion vector differences (CPMVDs) for control points (CPs) of the current block, i.e., the motion prediction information may include information on CPMVDs for each of the CPs of the current block.

[0349] Also, for example, the motion prediction information may include an affine MVP candidate index for the current block, which may indicate one of the affine MVP candidates included in an affine motion vector predictor (MVP) candidate list for the current block.

[0350] The decoding apparatus constructs an affine motion vector predictor (MVP) candidate list for the current block (S2610). The decoding apparatus may construct an affine MVP candidate list including affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list may be two.

[0351] Also, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The decoding device may check whether the inherited affine MVP candidates of the current block are available, and if the inherited affine MVP candidates are available, the inherited affine MVP candidates may be derived. For example, the inherited affine MVP candidates may be derived based on neighboring blocks of the current block, and the maximum number of inherited affine MVP candidates may be two. The neighboring blocks may be checked for availability in a specific order, and the inherited affine MVP candidates may be derived based on the checked available neighboring blocks. That is, the neighboring blocks may be checked for availability in a specific order, and a first inherited affine MVP candidate may be derived based on the first checked available neighboring block, and a second inherited affine MVP candidate may be derived based on the second checked available neighboring block. The availability may indicate that the neighboring block is coded using an affine motion model and that the reference picture of the neighboring block is the same as the reference picture of the current block. That is, the available neighboring block may be coded using an affine motion model (i.e., affine prediction is applied) and that the reference picture is the same as the reference picture of the current block. Specifically, the decoding apparatus may derive a motion vector for the CP of the current block based on the affine motion model of the first checked available neighboring block, and derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. Also, the decoding apparatus may derive a motion vector for the CP of the current block based on the affine motion model of the second checked available neighboring block, and derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model may be derived as shown in Equation 1 or 3 above.

[0352] In other words, the neighboring blocks may be checked to see if they satisfy a specific condition in a specific order, and the inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the checked specific condition. That is, the neighboring blocks may be checked to see if they satisfy the specific condition in a specific order, and a first inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the first checked specific condition, and a second inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the second checked specific condition. Specifically, the decoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring blocks that satisfy the first checked specific condition, and may derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. Furthermore, the decoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring blocks that satisfy the second checked specific condition, and may derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model may be derived as shown in Equation 1 or 3. Meanwhile, the specific condition may indicate that the current block is coded using an affine motion model and that the reference picture of the neighboring block is the same as the reference picture of the current block. That is, the neighboring block satisfying the specific condition may be coded using an affine motion model (i.e., affine prediction is applied) and that the reference picture is the same as the reference picture of the current block.

[0353] Here, for example, the neighboring blocks may include a left neighboring block, an upper neighboring block, a right upper corner neighboring block, a lower left corner neighboring block, and an upper left corner neighboring block of the current block. In this case, the identifying order may be from the left neighboring block to the lower left corner neighboring block, the upper neighboring block, the right upper corner neighboring block, and the upper left corner neighboring block.

[0354] Alternatively, for example, the peripheral blocks may include only the left peripheral block and the upper peripheral block, in which case the specific order may be from the left peripheral block to the upper peripheral block.

[0355] Alternatively, for example, the peripheral blocks may include the left peripheral block, and if the upper peripheral block is included in a current CTU including the current block, the peripheral blocks may further include the upper peripheral block. In this case, the specified order may be from the left peripheral block to the upper peripheral block. Also, if the upper peripheral block is not included in the current CTU, the peripheral blocks may not include the upper peripheral block. In this case, only the left peripheral block may be checked.

[0356] On the other hand, if the size is WxH and the x component and y component of the top-left sample position of the current block are 0, the lower-left neighboring block may be a block including a sample at coordinates (-1, H), the left neighboring block may be a block including a sample at coordinates (-1, H-1), the upper-right neighboring block may be a block including a sample at coordinates (W, -1), the upper neighboring block may be a block including a sample at coordinates (W-1, -1), and the upper-left neighboring block may be a block including a sample at coordinates (-1, -1). That is, the left neighboring block may be the leftmost neighboring block among the left neighboring blocks of the current block, and the upper neighboring block may be the leftmost neighboring block among the upper neighboring blocks of the current block.

[0357] Also, as an example, if a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. The decoding apparatus may check whether a constructed affine MVP candidate for the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate may be derived. Also, for example, the constructed affine MVP candidate may be derived after the inherited affine MVP candidate is derived. If the number of derived affine MVP candidates (i.e., the inherited affine MVP candidates) is less than two and the constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. Here, the constructed affine MVP candidate may include candidate motion vectors for the CP. The constructed affine MVP candidate may be available if all of the candidate motion vectors are available.

[0358] For example, if a 4-affine motion model is applied to the current block, the CPs of the current block may include CP0 and CP1. If a candidate motion vector for CP0 is available and a candidate motion vector for CP1 is available, the constructed affine MVP candidates may be available, and the affine MVP candidate list may include the constructed affine MVP candidates. Here, CP0 may indicate a position in the upper left corner of the current block, and CP1 may indicate a position in the upper right corner of the current block.

[0359] The constructed affine MVP candidates may include a candidate motion vector for the CP0 and a candidate motion vector for the CP1, where the candidate motion vector for the CP0 may be a motion vector of a first block, and the candidate motion vector for the CP1 may be a motion vector of a second block.

[0360] Furthermore, the first block may be determined by checking neighboring blocks in the first group in a first specific order, and the first identified reference picture may be the same as the reference picture of the current block. That is, the candidate motion vector for CP1 may be determined by checking neighboring blocks in the first group in a first order, and the first identified reference picture may be the same as the reference picture of the current block. The availability may indicate that the neighboring block exists and is coded using inter-prediction. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0361] Furthermore, the second block may be determined by checking neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0362] On the other hand, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at (-1, -1) coordinates, the surrounding block B may be a block including a sample at (0, -1) coordinates, the surrounding block C may be a block including a sample at (-1, 0) coordinates, the surrounding block D may be a block including a sample at (W-1, -1) coordinates, and the surrounding block E may be a block including a sample at (W, -1) coordinates. That is, the peripheral block A may be the peripheral block in the upper left corner of the current block, the peripheral block B may be the upper peripheral block located on the leftmost side among the peripheral blocks above the current block, the peripheral block C may be the left peripheral block located on the uppermost side among the peripheral blocks on the left side of the current block, the peripheral block D may be the upper peripheral block located on the rightmost side among the peripheral blocks above the current block, and the peripheral block E may be the peripheral block in the upper right corner of the current block.

[0363] On the other hand, if at least one of the candidate motion vectors of CP0 and the candidate motion vectors of CP1 is unavailable, the constructed affine MVP candidate may not be available.

[0364] Alternatively, for example, if a 6-affine motion model is applied to the current block, the CPs of the current block may include CP0, CP1, and CP2. If a candidate motion vector for CP0 is available, a candidate motion vector for CP1 is available, and a candidate motion vector for CP2 is available, the constructed affine MVP candidates may be available, and the affine MVP candidate list may include the constructed affine MVP candidates. Here, CP0 may indicate a position in the upper left corner of the current block, CP1 may indicate a position in the upper right corner of the current block, and CP2 may indicate a position in the lower left corner of the current block.

[0365] The constructed affine MVP candidates may include a candidate motion vector for the CP0, a candidate motion vector for the CP1, and a candidate motion vector for the CP2, where the candidate motion vector for the CP0 may be a motion vector of a first block, the candidate motion vector for the CP1 may be a motion vector of a second block, and the candidate motion vector for the CP2 may be a motion vector of a third block.

[0366] Furthermore, the first block may check neighboring blocks in the first group according to a first specific order, and a first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, a candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0367] Furthermore, the second block may be determined by checking neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0368] Furthermore, the third block may be determined by checking neighboring blocks in the third group according to a third specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the third block in the third group is the same as the reference picture of the current block, a candidate motion vector for CP2 may be available. For example, the third group may include neighboring blocks F and G, and the third specific order may be from neighboring block F to neighboring block G.

[0369] On the other hand, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at a (-1, -1) coordinate, the surrounding block B may be a block including a sample at a (0, -1) coordinate, the surrounding block C may be a block including a sample at a (-1, 0) coordinate, the surrounding block D may be a block including a sample at a (W-1, -1) coordinate, the surrounding block E may be a block including a sample at a (W, -1) coordinate, the surrounding block F may be a block including a sample at a (-1, H-1) coordinate, and the surrounding block G may be a block including a sample at a (-1, H) coordinate. That is, the peripheral block A may be the peripheral block at the upper left corner of the current block, the peripheral block B may be the upper peripheral block located at the leftmost position among the peripheral blocks above the current block, the peripheral block C may be the left peripheral block located at the topmost position among the peripheral blocks to the left of the current block, the peripheral block D may be the upper peripheral block located at the rightmost position among the peripheral blocks above the current block, the peripheral block E may be the peripheral block at the upper right corner of the current block, the peripheral block F may be the left peripheral block located at the bottommost position among the peripheral blocks to the left of the current block, and the peripheral block G may be the peripheral block at the lower left corner of the current block.

[0370] On the other hand, if at least one of the candidate motion vectors of CP0, CP1, and CP2 is unavailable, the constructed affine MVP candidate may not be available.

[0371] Thereafter, the affine MVP candidate list can be derived based on the following sequential steps.

[0372] For example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP0 is available, the decoding apparatus may derive a first affine MVP candidate, which may be an affine MVP candidate that includes the motion vector for the CP0 in the candidate motion vector for the CP.

[0373] Also, for example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP1 is available, the decoding apparatus may derive a second affine MVP candidate, where the second affine MVP candidate may be an affine MVP candidate that includes the motion vector for the CP1 in the candidate motion vector for the CP.

[0374] Also, for example, if the number of derived affine MVP candidates is less than two and a motion vector for CP2 is available, the decoding apparatus may derive a third affine MVP candidate, which may be an affine MVP candidate that includes the motion vector for CP2 in the candidate motion vector for the CP.

[0375] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the decoding apparatus may derive a fourth affine MVP candidate including a temporal MVP derived based on a temporal neighboring block of the current block as a candidate motion vector for the CP. The temporal neighboring block may indicate a collocated block in a collocated picture corresponding to the current block. The temporal MVP may be derived based on the motion vector of the temporal neighboring block.

[0376] Also, for example, if the number of derived affine MVP candidates is less than two, the decoding apparatus may derive a fifth affine MVP candidate including a zero motion vector as a candidate motion vector for the CP. The zero motion vector may indicate a motion vector with a value of 0.

[0377] The decoding apparatus derives Control Point Motion Vector Predictors (CPMVPs) for the Control Point (CP) of the current block based on the affine MVP candidate list (S2620).

[0378] A decoding apparatus may select a specific affine MVP candidate from the affine MVP candidates included in the affine MVP candidate list and derive the selected affine MVP candidate as a CPMVP for the CP of the current block. For example, the decoding apparatus may acquire the affine MVP candidate index for the current block from a bitstream and derive the affine MVP candidate pointed to by the affine MVP candidate index from the affine MVP candidate list as a CPMVP for the CP of the current block. Specifically, if affine MVP candidates include a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate may be derived as the CPMVP of CP0, and the candidate motion vector for CP1 of the affine MVP candidate may be derived as the CPMVP of CP1. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived as the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2.

[0379] The decoding apparatus derives control point motion vector differences (CPMVDs) for the CPs of the current block based on the motion prediction information (S2630). The motion prediction information may include information on CPMVDs for each of the CPs, and the decoding apparatus may derive the CPMVDs for each of the CPs of the current block based on the information on the CPMVDs for each of the CPs.

[0380] The decoding apparatus derives control point motion vectors (CPMV) for the CP of the current block based on the CPMVP and the CPMVD (S2640). The decoding apparatus may derive a CPMV for each CP based on the CPMVP and CPMVD for each CP. For example, the decoding apparatus may add the CPMVP and CPMVD for each CP to derive a CPMV for the CP.

[0381] The decoding apparatus derives prediction samples for the current block based on the CPMV (S2650). The decoding apparatus may derive motion vectors for sub-blocks or samples of the current block based on the CPMV. That is, the decoding apparatus may derive motion vectors for each sub-block or each sample of the current block based on the CPMV. The sub-block-based or sample-based motion vectors may be derived based on Equation 1 or Equation 3 above. The motion vectors may be represented as an affine motion vector field (MVF) or a motion vector array.

[0382] The decoding apparatus may derive prediction samples for the current block based on the motion vector in sub-block units or sample units, derive a reference area in a reference picture based on the motion vector in sub-block units or sample units, and generate prediction samples for the current block based on reconstructed samples in the reference area.

[0383] The decoding apparatus generates a reconstructed picture for the current block based on the derived prediction samples (S2660). The decoding apparatus may generate a reconstructed picture for the current block based on the derived prediction samples. Depending on the prediction mode, the decoding apparatus may directly use the prediction samples as reconstructed samples, or may generate reconstructed samples by adding residual samples to the prediction samples. If residual samples for the current block exist, the decoding apparatus may acquire information about the residual for the current block from the bitstream. The information about the residual may include transform coefficients related to the residual samples. The decoding apparatus may derive the residual samples (or residual sample array) for the current block based on the residual information. The decoding apparatus may generate reconstructed samples based on the prediction samples and the residual samples, and may derive a reconstructed block or picture based on the reconstructed samples. As described above, the decoding apparatus may apply an in-loop filtering procedure, such as deblocking filtering and / or an SAO procedure, to the reconstructed picture to improve subjective / objective image quality as needed.

[0384] Figure 27 schematically illustrates a decoding device that performs the image decoding method according to the present document. The method disclosed in Figure 26 is performed by the decoding device disclosed in Figure 27. Specifically, for example, the entropy decoding unit of the decoding device of Figure 27 may perform S2600 of Figure 26, the prediction unit of the decoding device of Figure 27 may perform S2610 to S2650 of Figure 26, and the adder of the decoding device of Figure 27 may perform S2660 of Figure 26. Also, although not shown, the process of obtaining image information including information about the residual of the current block via a bitstream may be performed by the entropy decoding unit of the decoding device of Figure 27, and the process of deriving the residual sample for the current block based on the residual information may be performed by an inverse transform unit of the decoding device of Figure 27.

[0385] According to the aforementioned document, it is possible to increase the efficiency of image coding based on affine motion prediction.

[0386] In addition, according to this document, when deriving an affine MVP candidate list, a constructed affine MVP candidate can be added only if all candidate motion vectors for the CP of the constructed affine MVP candidate are available, thereby reducing the complexity of the process of deriving a constructed affine MVP candidate and the process of constructing an affine MVP candidate list and improving coding efficiency.

[0387] In addition, according to this document, when deriving an affine MVP candidate list, in the process of deriving a constructed affine MVP candidate, further affine MVP candidates can be derived based on candidate motion vectors for the derived CP, thereby reducing the complexity of the process of constructing an affine MVP candidate list and improving coding efficiency.

[0388] In addition, according to this document, in the process of deriving an inherited affine MVP candidate, the upper surrounding block can be used only if the upper surrounding block is included in the current CTU, and the inherited affine MVP candidate can be derived. This can reduce the amount of line buffer storage for affine prediction and minimize hardware costs.

[0389] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but this document is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, one skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and that different steps may be included, or one or more steps of the flowcharts may be removed without affecting the scope of this document.

[0390] The embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in the figures may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored in a digital storage medium.

[0391] In addition, decoding devices and encoding devices to which this document is applied may be included in, and used to process video signals or data signals, devices for transmitting and receiving multimedia broadcasts, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, devices providing customized video (VoD) services, over-the-top (OTT) video devices, devices providing Internet streaming services, three-dimensional (3D) video devices, image telephone video devices, transportation terminals (e.g., vehicle terminals, airplane terminals, ship terminals, etc.), medical video devices, etc. For example, over-the-top (OTT) video devices may include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0392] In addition, the processing method to which this document is applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium may also include media realized in the form of a carrier wave (e.g., transmission via the Internet). The bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0393] Furthermore, the embodiments of the present document may be implemented in a computer program product by program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.

[0394] FIG. 28 illustrates an example of a content streaming system to which the embodiments disclosed herein can be applied.

[0395] Referring to FIG. 28, the content streaming system to which this document applies can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0396] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0397] The bitstream may be generated by an encoding method or a bitstream generation method to which this document applies, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0398] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.

[0399] The streaming server may receive content from a media storage and / or an encoding server. For example, if content is received from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.

[0400] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, a smart glass, an HMD (head mounted display), a digital TV, a desktop computer, and a digital signage.

[0401] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.

Claims

1. In a method for decoding an image by a decoding device, obtaining motion prediction information for the current block from the bitstream; constructing an affine motion vector predictor (MVP) candidate list for the current block; deriving control point motion vector predictors (CPMVPs) for control points (CPs) of the current block based on the affine MVP candidate list; deriving a control point motion vector differential (CPMVD) for the CP of the current block based on the motion prediction information; deriving a control point motion vector (CPMV) for the CP of the current block based on the CPMVP and the CPMVD; deriving a predicted sample for the current block based on the CPMV; generating a reconstructed picture for the current block based on the derived prediction samples; Including, The step of constructing the affine MVP candidate list comprises: checking whether a first affine MVP candidate is available, where the first affine MVP candidate is available based on the fact that a first block in a left group of blocks is coded with an affine motion model and the reference picture index of the first block is the same as the reference picture index of the current block; checking whether a second affine MVP candidate is available, where the second affine MVP candidate is available based on the fact that a second block in an upper group of blocks is coded with the affine motion model and the reference picture index of the second block is the same as the reference picture index of the current block; checking whether a third affine MVP candidate is available based on the number of available affine MVP candidates being less than two; the left block group includes a neighboring block at a lower left corner of the current block and a first left neighboring block adjacent to an upper side of the neighboring block at the lower left corner; the upper block group includes a neighboring block at an upper right corner of the current block, a first upper neighboring block adjacent to the left side of the neighboring block at the upper right corner, and a neighboring block at an upper left corner; A four-parameter affine model or a six-parameter affine model is used for inter prediction; When the four-parameter affine model is used for the inter prediction, the third affine MVP candidate is available based on the fact that a first motion vector for CP0 of the current block and a second motion vector for CP1 of the current block are derived from a block group to the upper left of the current block and a block group to the upper right of the current block, respectively; When the six-parameter affine model is used for the inter prediction, the third affine MVP candidate is available based on the fact that the first motion vector for the CP0, the second motion vector for the CP1, and the third motion vector for the CP2 of the current block are derived from a block group on the upper left side of the current block, a block group on the upper right side of the current block, and a block group on the left side of the current block, respectively; the upper left block group includes the upper left corner peripheral block of the current block, a second left peripheral block adjacent to the lower side of the upper left corner peripheral block, and a second upper peripheral block adjacent to the right side of the upper left corner peripheral block; the upper right block group includes the upper right corner peripheral block and the first upper peripheral block; and the lower left block group includes the lower left corner peripheral block and the first left peripheral block; deriving a fourth affine MVP candidate as the affine MVP candidate based on the number of the affine MVP candidates being less than two and the availability of a third motion vector for the CP2 included in the third affine MVP candidate; An image decoding method, wherein the fourth affine MVP candidate is a first motion vector for the CP0, a second motion vector for the CP1, and a third motion vector for the CP2, and includes the third motion vector for the CP2 included in the third affine MVP candidate.

2. In an image encoding method using an encoding device, constructing an affine motion vector predictor (MVP) candidate list for the current block; deriving control point motion vector predictors (CPMVPs) for control points (CPs) of the current block based on the affine MVP candidate list; deriving a control point motion vector (CPMV) for the CP of the current block; deriving a control point motion vector differential (CPMVD) for the CP of the current block based on the CPMVP and the CPMV; encoding motion prediction information including information about the CPMVD; Including, The step of constructing the affine MVP candidate list comprises: checking whether a first affine MVP candidate is available, where the first affine MVP candidate is available based on the fact that a first block in a left group of blocks is coded with an affine motion model and the reference picture index of the first block is the same as the reference picture index of the current block; checking whether a second affine MVP candidate is available, where the second affine MVP candidate is available based on the fact that a second block in an upper group of blocks is coded with the affine motion model and the reference picture index of the second block is the same as the reference picture index of the current block; checking whether a third affine MVP candidate is available based on the number of available affine MVP candidates being less than two; the left block group includes a neighboring block at a lower left corner of the current block and a first left neighboring block adjacent to an upper side of the neighboring block at the lower left corner; the upper block group includes a neighboring block at an upper right corner of the current block, a first upper neighboring block adjacent to the left side of the neighboring block at the upper right corner, and a neighboring block at an upper left corner; A four-parameter affine model or a six-parameter affine model is used for inter prediction; When the four-parameter affine model is used for the inter prediction, the third affine MVP candidate is available based on the fact that a first motion vector for CP0 of the current block and a second motion vector for CP1 of the current block are derived from a block group to the upper left of the current block and a block group to the upper right of the current block, respectively; When the six-parameter affine model is used for the inter prediction, the third affine MVP candidate is available based on the fact that the first motion vector for the CP0, the second motion vector for the CP1, and the third motion vector for the CP2 of the current block are derived from a block group on the upper left side of the current block, a block group on the upper right side of the current block, and a block group on the left side of the current block, respectively; the upper left block group includes the upper left corner peripheral block of the current block, a second left peripheral block adjacent to the lower side of the upper left corner peripheral block, and a second upper peripheral block adjacent to the right side of the upper left corner peripheral block; the upper right block group includes the upper right corner peripheral block and the first upper peripheral block; and the lower left block group includes the lower left corner peripheral block and the first left peripheral block; deriving a fourth affine MVP candidate as the affine MVP candidate based on the number of the affine MVP candidates being less than two and the availability of a third motion vector for the CP2 included in the third affine MVP candidate; An image encoding method, wherein the fourth affine MVP candidate is a first motion vector for the CP0, a second motion vector for the CP1, and a third motion vector for the CP2, and includes the third motion vector for the CP2 included in the third affine MVP candidate.

3. In a method for transmitting data for an image, generating a bitstream for the image, the bitstream being generated based on: constructing an affine motion vector predictor (MVP) candidate list for a current block; deriving control point motion vector predictors (CPMVPs) for control points (CPs) of the current block based on the affine MVP candidate list; deriving control point motion vectors (CPMVs) for the CPs of the current block; deriving control point motion vector differentials (CPMVDs) for the CPs of the current block based on the CPMVPs and the CPMVs; and encoding motion prediction information including information about the CPMVDs; transmitting the data including the bitstream; Including, constructing the affine MVP candidate list checking whether a first affine MVP candidate is available, and determining that the first affine MVP candidate is available based on the fact that a first block in a left group of blocks is coded with an affine motion model and the reference picture index of the first block is the same as the reference picture index of the current block; checking whether a second affine MVP candidate is available, and determining that the second affine MVP candidate is available based on the fact that a second block in an upper group of blocks is coded with an affine motion model and the reference picture index of the second block is the same as the reference picture index of the current block; checking whether a third affine MVP candidate is available based on the number of available affine MVP candidates being less than two; the left block group includes a neighboring block at a lower left corner of the current block and a first left neighboring block adjacent to an upper side of the neighboring block at the lower left corner; the upper block group includes a neighboring block at an upper right corner of the current block, a first upper neighboring block adjacent to the left side of the neighboring block at the upper right corner, and a neighboring block at an upper left corner; A four-parameter affine model or a six-parameter affine model is used for inter prediction; When the four-parameter affine model is used for the inter prediction, the third affine MVP candidate is available based on the fact that a first motion vector for CP0 of the current block and a second motion vector for CP1 of the current block are derived from a block group to the upper left of the current block and a block group to the upper right of the current block, respectively; When the six-parameter affine model is used for the inter prediction, the third affine MVP candidate is available based on the fact that the first motion vector for the CP0, the second motion vector for the CP1, and the third motion vector for the CP2 of the current block are derived from a block group on the upper left side of the current block, a block group on the upper right side of the current block, and a block group on the left side of the current block, respectively; the upper left block group includes the upper left corner peripheral block of the current block, a second left peripheral block adjacent to the lower side of the upper left corner peripheral block, and a second upper peripheral block adjacent to the right side of the upper left corner peripheral block; the upper right block group includes the upper right corner peripheral block and the first upper peripheral block; and the lower left block group includes the lower left corner peripheral block and the first left peripheral block; deriving a fourth affine MVP candidate as the affine MVP candidate based on the number of the affine MVP candidates being less than two and the availability of a third motion vector for the CP2 included in the third affine MVP candidate; A data transmission method in which the fourth affine MVP candidate is a first motion vector for the CP0, a second motion vector for the CP1, and a third motion vector for the CP2, and includes the third motion vector for the CP2 included in the third affine MVP candidate.

Citation Information

Patent Citations

  • Motion vector prediction for affine motion models in video coding

    US20180098063A1

  • Motion vector generation for affine motion model for video coding

    US20180192069A1

  • Method and apparatus for affine merge mode prediction for video coding system

    WO2017118409A1

  • Method and apparatus of video coding with affine motion compensation

    WO2017148345A1