IMAGE DECODING METHOD BASED ON AFFIN MOTION PREDICTION AND APPARATUS USING THE AFFIN MVP CANDIDATE LIST IN THE IMAGE CODING SYSTEM

MX435025BActive Publication Date: 2026-06-12LG ELECTRONICS INC

Patent Information

Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2021-03-10
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

High-resolution, high-quality video data requires efficient compression methods to reduce storage and transmission costs, as existing technologies struggle to effectively manage the increased information bits associated with HD and UHD video.

Method used

A video decoding method and device utilizing affine motion prediction, which constructs an affine MVP candidate list by deriving candidates from surrounding blocks and uses these candidates for prediction, even when the number of available candidates is less than the maximum, by selecting motion vectors from control points and adding temporal or zero motion vectors as needed.

Benefits of technology

This approach enhances video coding efficiency by reducing the complexity of deriving affine MVP candidates and improving coding performance, while minimizing hardware costs and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure MX435025B0
    Figure MX435025B0
Patent Text Reader

Abstract

The present invention relates to a method by which a decoding apparatus performs image decoding, in accordance with the present document, comprising the steps of: obtaining motion prediction information about a current block from a bit stream; generating a list of affine MVP candidates for the current block; deriving CPMVPs for CPs of the current block based on the list of affine MVP candidates; deriving CPMVDs for the CPs of the current block based on the motion prediction information; deriving CPMVs for the CPs of the current block based on the CPMVPs and the CPMVDs; and deriving prediction samples for the current block based on the CPMVs.
Need to check novelty before this filing date? Find Prior Art

Description

A method and device for video decoding based on affine motion prediction using an affine MVP candidate list in a video coding system

[0001] This document relates to video coding technology, and more specifically, to a video decoding method and device based on affine motion prediction in a video coding system.

[0002] Demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), is growing across various fields. As video data becomes higher in resolution and quality, the amount of information or bits transmitted increases relative to conventional video data. Therefore, transmitting video data using existing media, such as wired or wireless broadband lines, or storing it using existing storage media, increases transmission and storage costs.

[0003] Accordingly, high-efficiency image compression technology is required to effectively transmit, store, and reproduce high-resolution, high-quality image information.

[0004] The technical task of this document is to provide a method and device for improving video coding efficiency.

[0005] Another technical challenge of this document is to provide a video decoding method and device that constructs an affine MVP candidate list for the current block by deriving a constructed affine MVP candidate based on surrounding blocks only when all candidate motion vectors for CPs are available, and performs prediction for the current block based on the constructed affine MVP candidate list.

[0006] Another technical challenge of this document is to provide a video decoding method and device for deriving an affine MVP candidate using a candidate motion vector derived in the process of deriving the constructed affine MVP candidate as an added affine MVP candidate when the number of available inherited affine MVP candidates and constructed affine MVP candidates is less than the maximum number of candidates in the MVP candidate list, and for performing a prediction for the current block based on the constructed affine MVP candidate list.

[0007] According to one embodiment of the present document, a method of image decoding performed by a decoding device is provided. The method comprises the steps of obtaining motion prediction information for a current block from a bitstream, constructing an affine motion vector predictor (MVP) candidate list for the current block, deriving CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the affine MVP candidate list, deriving CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the motion prediction information, deriving CPMVs (Control Point Motion Vectors) for the CPs of the current block based on the CPMVPs and the CPMVDs, deriving prediction samples for the current block based on the CPMVs, and generating a reconstructed picture for the current block based on the derived prediction samples, wherein the step of constructing the affine MVP candidate list checks whether an inherited affine MVP candidate of the current block is available, and the inherited affine MVP candidate is If an MVP candidate is available, a step of deriving the inherited affine MVP candidate, a step of checking whether a constructed affine MVP candidate of the current block is available, and if the constructed affine MVP candidate is available, a step of deriving the constructed affine MVP candidate, and the constructed affine MVP candidate including a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block,If the number of derived affine MVP candidates is less than 2 and a motion vector for CP0 is available, a step of deriving a first affine MVP candidate, wherein the first affine MVP candidate is an affine MVP candidate that includes a motion vector for CP0 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than 2 and a motion vector for CP1 is available, a step of deriving a second affine MVP candidate, wherein the second affine MVP candidate is an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than 2 and a motion vector for CP2 is available, a step of deriving a third affine MVP candidate, wherein the third affine MVP candidate is an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs; If the number is less than 2, a step of deriving a fourth affine MVP candidate including temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs is characterized by including a step of deriving a fifth affine MVP candidate including a zero motion vector as candidate motion vectors for the CPs is characterized by including a step of deriving a fifth affine MVP candidate including a zero motion vector as candidate motion vectors for the CPs is characterized by including a step of deriving a fourth affine MVP candidate including temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs is characterized by including a step of deriving a fifth affine MVP candidate including temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs is characterized by including a step of deriving a fifth affine MVP candidate including temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs is characterized by including a step of deriving a fifth affine MVP candidate including temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs is characterized by including a step of deriving a fifth affine MVP candidate including temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs is characterized by including a step of deriving a fourth affine MVP candidate including temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs is characterized by including a step of deriving a fifth ...

[0008] According to another embodiment of the present document, a decoding device for performing image decoding is provided. The decoding device comprises an entropy decoding unit that obtains motion prediction information for a current block from a bitstream, a prediction unit that constructs an affine motion vector predictor (MVP) candidate list for the current block, derives CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the affine MVP candidate list, derives CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the motion prediction information, derives CPMVs (Control Point Motion Vectors) for the CPs of the current block based on the CPMVPs and the CPMVDs, and derives prediction samples for the current block based on the CPMVs, and an addition unit that generates a reconstructed picture for the current block based on the derived prediction samples, wherein the affine MVP candidate list is configured to determine whether an inherited affine MVP candidate of the current block is available. A step of checking, if the inherited affine MVP candidate is available, deriving the inherited affine MVP candidate; a step of checking whether the constructed affine MVP candidate of the current block is available; if the constructed affine MVP candidate is available, deriving the constructed affine MVP candidate, and the constructed affine MVP candidate including a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; the number of derived affine MVP candidates is less than 2;If a motion vector for CP0 is available, a step of deriving a first affine MVP candidate, wherein the first affine MVP candidate is an affine MVP candidate that includes a motion vector for CP0 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than two and a motion vector for CP1 is available, a step of deriving a second affine MVP candidate, wherein the second affine MVP candidate is an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than two and a motion vector for CP2 is available, a step of deriving a third affine MVP candidate, wherein the third affine MVP candidate is an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than two, a temporal neighborhood of the current block It is characterized by being configured based on a step of deriving a fourth affine MVP candidate that includes a temporal MVP derived based on a block as candidate motion vectors for the CPs, and a step of deriving a fifth affine MVP candidate that includes a zero motion vector as candidate motion vectors for the CPs when the number of derived affine MVP candidates is less than two.

[0009] According to another embodiment of the present document, a video encoding method performed by an encoding device is provided. The method comprises the steps of constructing an affine motion vector predictor (MVP) candidate list for a current block, deriving CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the affine MVP candidate list, deriving CPMVs for the CPs of the current block, deriving CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the CPMVPs and the CPMVs, and encoding motion prediction information including information about the CPMVDs, wherein the step of constructing the affine MVP candidate list comprises the steps of checking whether an inherited affine MVP candidate of the current block is available, and if the inherited affine MVP candidate is available, deriving the inherited affine MVP candidate, and determining whether a constructed affine MVP candidate of the current block is available. However, if the constructed affine MVP candidate is available, the constructed affine MVP candidate is derived, and the constructed affine MVP candidate includes a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block, and if the number of derived affine MVP candidates is less than 2 and the motion vector for CP0 is available, a first affine MVP candidate is derived.The first affine MVP candidate is an affine MVP candidate that includes a motion vector for CP0 as candidate motion vectors for the CPs, a step of deriving a second affine MVP candidate, wherein the second affine MVP candidate is an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs, a step of deriving a third affine MVP candidate, wherein the third affine MVP candidate is an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs, a step of deriving a fourth affine MVP, wherein the temporal MVP derived based on a temporal neighboring block of the current block is as candidate motion vectors for the CPs, if the number of derived affine MVP candidates is less than two and a motion vector for CP1 is available. It is characterized by including a step of deriving a candidate, and a step of deriving a fifth affine MVP candidate including a zero motion vector as candidate motion vectors for the CPs when the number of derived affine MVP candidates is less than two.

[0010] According to another embodiment of the present document, a video encoding device is provided. The encoding device comprises a prediction unit that constructs an affine motion vector predictor (MVP) candidate list for a current block, derives CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the affine MVP candidate list, and derives CPMVs for the CPs of the current block, a subtraction unit that derives CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the CPMVPs and the CPMVs, and an entropy encoding unit that encodes motion prediction information including information on the CPMVDs, wherein the affine MVP candidate list includes a step of checking whether an inherited affine MVP candidate of the current block is available, and if the inherited affine MVP candidate is available, deriving the inherited affine MVP candidate, and a step of deriving a constructed affine MVP candidate of the current block. A step of checking availability, and if the constructed affine MVP candidate is available, deriving the constructed affine MVP candidate, and the constructed affine MVP candidate including a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block, a step of deriving a first affine MVP candidate, and if the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, the first affine MVP candidate is an affine MVP candidate including a motion vector for CP0 as candidate motion vectors for the CPs,If the number of derived affine MVP candidates is less than 2 and a motion vector for CP1 is available, a step of deriving a second affine MVP candidate, wherein the second affine MVP candidate is an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than 2 and a motion vector for CP2 is available, a step of deriving a third affine MVP candidate, wherein the third affine MVP candidate is an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than 2, a step of deriving a fourth affine MVP candidate that includes a temporal MVP derived based on a temporal neighboring block of the current block as candidate motion vectors for the CPs; and If the number of derived affine MVP candidates is less than 2, a zero motion vector is generated. It is characterized by being configured based on a step of deriving a fifth affine MVP candidate including candidate motion vectors for the above CPs.

[0011] According to this document, the overall image / video compression efficiency can be improved.

[0012] According to this paper, the efficiency of image coding based on affine motion prediction can be improved.

[0013] According to this document, when deriving an affine MVP candidate list, the constructed affine MVP candidate can be added only when all candidate motion vectors for CPs of the constructed affine MVP candidate are available, thereby reducing the complexity of the process of deriving the constructed affine MVP candidate and the process of constructing the affine MVP candidate list and improving coding efficiency.

[0014] According to this document, in deriving an affine MVP candidate list, additional affine MVP candidates can be derived based on candidate motion vectors for CP derived in the process of deriving constructed affine MVP candidates, thereby reducing the complexity of the process of constructing an affine MVP candidate list and improving coding efficiency.

[0015] According to this document, in the process of deriving an inherited affine MVP candidate, the inherited affine MVP candidate can be derived using an upper neighboring block only when the upper neighboring block is included in the current CTU, thereby reducing the storage amount of the line buffer for affine prediction and minimizing hardware costs.

[0016] Figure 1 schematically illustrates an example of a video / image coding system to which embodiments of the present document can be applied.

[0017] FIG. 2 is a drawing schematically illustrating the configuration of a video / image encoding device to which embodiments of this document can be applied.

[0018] FIG. 3 is a drawing schematically illustrating the configuration of a video / image decoding device to which embodiments of this document can be applied.

[0019] Figure 4 illustrates an example of a movement expressed through the above affine movement model.

[0020] Figure 5 illustrates an example of the above affine motion model in which motion vectors for three control points are used.

[0021] Figure 6 illustrates an example of the above affine motion model in which motion vectors for two control points are used.

[0022] Figure 7 exemplarily shows a method for deriving a motion vector in sub-block units based on the above affine motion model.

[0023] Fig. 8 illustrates an exemplary flowchart of an affine motion prediction method according to one embodiment of the present document.

[0024] FIG. 9 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to one embodiment of the present document.

[0025] FIG. 10 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to one embodiment of the present document.

[0026] Figure 11 shows an example of affine prediction performed when a surrounding block A is selected as an affine merge candidate.

[0027] Figure 12 illustrates peripheral blocks for deriving the above-described inherited affine candidate.

[0028] Figure 13 illustrates an example of a spatial candidate for the above-mentioned constructed affine candidate.

[0029] Figure 14 illustrates an example of constructing an affine MVP list.

[0030] Figure 15 shows an example of deriving the above-mentioned constructed candidate.

[0031] Figure 16 shows an example of deriving the above-mentioned constructed candidate.

[0032] Figure 17 illustrates an example of the locations of surrounding blocks scanned to derive inherited affine candidates.

[0033] Figure 18 shows an example of deriving the constructed candidate when a 4-affine motion model is applied to the current block.

[0034] Figure 19 shows an example of deriving the constructed candidate when a 6-affine motion model is applied to the current block.

[0035] Figures 20a and 20b illustrate examples of deriving the inherited affine candidates.

[0036] Figure 21 schematically illustrates a method for encoding an image using an encoding device according to this document.

[0037] Fig. 22 schematically illustrates an encoding device that performs an image encoding method according to this document.

[0038] Figure 23 schematically illustrates a method of decoding an image by a decoding device according to this document.

[0039] Fig. 24 schematically illustrates a decoding device that performs an image decoding method according to this document.

[0040] Figure 25 illustrates an exemplary structure of a content streaming system to which embodiments of this document are applied.

[0041] This document may have various modifications and embodiments, and thus specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit this document to specific embodiments. The terminology used herein is only used to describe specific embodiments and is not intended to limit the technical idea of ​​this document. The singular expression includes plural expressions unless the context clearly indicates otherwise. It should be understood that the terms "comprises" or "has" in this specification specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0042] Meanwhile, each component in the drawings described in this document is depicted independently for the convenience of explaining their distinct functions. This does not imply that each component is implemented using separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of this document, as long as they do not deviate from the essence of this document.

[0043] Hereinafter, with reference to the attached drawings, a preferred embodiment of the present document will be described in more detail. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components may be omitted.

[0044] Figure 1 schematically illustrates an example of a video / image coding system to which embodiments of the present document can be applied.

[0045] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or streaming.

[0046] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.

[0047] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.

[0048] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0049] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0050] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.

[0051] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.

[0052] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the VVC (versatile video coding) standard, the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation of audio video coding standard), or the next generation video / image coding standard (e.g., H.267 or H.268).

[0053] This document presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.

[0054] In this document, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (coding tree units). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each of which consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan may represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in a CTU raster scan in a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture. A tile scan may represent a specific sequential ordering of CTUs partitioning a picture, wherein the CTUs are ordered consecutively in a CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit.A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably. For example, in this document, the terms tile group / tile group header may be referred to as slice / slice header.

[0055] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0056] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0057] In this document, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." (In this document, the terms " / " and "," should be interpreted to indicate "and / or." For instance, the expression "A / B" may mean "A and / or B." Further, "A, B" may mean "A and / or B." Further, "A / B / C" may mean "at least one of A, B, and / or C." Also, "A / B / C" may mean "at least one of A, B, and / or C.")

[0058] Additionally, in this document, "or" is interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) both "A and B." In other words, "or" in this document can mean "additionally or alternatively." (Furthermore, in the document, the term "or" should be interpreted to indicate "and / or." For instance, the expression "A or B" may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted to indicate "additionally or alternatively.")

[0059] Figure 2 is a drawing schematically illustrating the configuration of a video / image encoding device to which embodiments of this document may be applied. Hereinafter, the term "video encoding device" may include an image encoding device.

[0060] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0061] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit may be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present document may be performed based on the final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit for deriving a transformation coefficient and / or a unit for deriving a residual signal from a transformation coefficient.

[0062] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).

[0063] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0064] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail in the prediction direction. However, this is merely an example, and a greater or lesser number of directional prediction modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0065] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0066] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, sample values ​​within a picture can be signaled based on information about the palette table and palette index.

[0067] The prediction signal generated through the above prediction unit (including the inter prediction unit (221) and / or the intra prediction unit (222)) can be used to generate a restored signal or a residual signal. The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. Additionally, the transformation process can be applied to blocks of pixels of equal size, either square or variable size, non-square.

[0068] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit (240) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) may encode, together or separately, information necessary for video / image restoration (e.g., values ​​of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The above bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).

[0069] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (155) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.

[0070] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0071] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit the information to the entropy encoding unit (240), as described later in the description of each filtering method. The information regarding filtering may be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0072] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (100) and the decoding device, and can also improve encoding efficiency.

[0073] The memory (270) DPB can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks within the current picture and transfer them to the intra prediction unit (222).

[0074] FIG. 3 is a drawing schematically illustrating the configuration of a video / image decoding device to which embodiments of this document can be applied.

[0075] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (331) and an intra-prediction unit (332). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., decoder chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0076] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 0.2-1. For example, the decoding device (300) can derive units / blocks based on block division related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.

[0077] The decoding device (300) can receive a signal output from the encoding device of FIG. 0.2-1 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode a picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values ​​on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).

[0078] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0079] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).

[0080] The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit (310), the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.

[0081] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be signaled and included in the video / image information.

[0082] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit (331) can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0083] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information about the prediction can include information indicating the mode of inter prediction for the current block.

[0084] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.

[0085] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture.

[0086] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0087] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0088] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (260) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks within the current picture and transfer them to the intra prediction unit (331).

[0089] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (100) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.

[0090] Meanwhile, with regard to inter prediction, an inter prediction method that takes into account image distortion has been proposed. Specifically, an affine motion model has been proposed that efficiently derives motion vectors for sub-blocks or sample points of the current block and increases the accuracy of inter prediction despite deformations such as rotation, zoom-in, or zoom-out of the image. That is, an affine motion model that derives motion vectors for sub-blocks or sample points of the current block has been proposed. Prediction using the above affine motion model may be called affine inter prediction or affine motion prediction.

[0091] For example, the affine inter prediction using the above affine motion model can efficiently express four motions, i.e., four transformations, as described below.

[0092] Fig. 4 exemplarily shows a motion expressed through the affine motion model. Referring to Fig. 4, motions that can be expressed through the affine motion model may include translational motion, scale motion, rotational motion, and shear motion. That is, not only translational motion in which (a part of) an image moves planarly over time as shown in Fig. 4, but also scale motion in which (a part of) an image is scaled over time, rotational motion in which (a part of) an image is rotated over time, and shear motion in which (a part of) an image is deformed into an equilibrium quadrilateral over time can be efficiently expressed through the affine inter prediction.

[0093] The encoding device / decoding device can predict the distortion form of the image based on the motion vectors at the control points (CPs) of the current block through the affine inter prediction, thereby improving the compression performance of the image by increasing the accuracy of the prediction. In addition, since the motion vector for at least one control point of the current block can be derived using the motion vector of the surrounding blocks of the current block, the data burden for the additional information can be reduced, and the inter prediction efficiency can be significantly improved.

[0094] As an example of the above affine inter prediction, motion information at three control points, i.e. three reference points, may be required.

[0095] Figure 5 illustrates an example of the above affine motion model in which motion vectors for three control points are used.

[0096] If the top-left sample position within the current block (500) is (0,0), the sample positions (0,0), (w, 0), and (0, h) can be set as the control points as illustrated in FIG. 5. Hereinafter, the control point of the (0,0) sample position can be represented as CP0, the control point of the (w, 0) sample position can be represented as CP1, and the control point of the (0, h) sample position can be represented as CP2.

[0097] A mathematical formula for the affine motion model can be derived using the above-described control points and the motion vectors for the control points. The mathematical formula for the affine motion model can be expressed as follows.

[0098]

[0099] Here, w represents the width of the current block (500), h represents the height of the current block (500), and v 0x , v 0y represents the x-component and y-component of the motion vector of CP0, respectively, and v 1x , v 1y represent the x-component and y-component of the motion vector of CP1, respectively, and v 2x , v 2y Each represents the x-component and y-component of the motion vector of CP2. In addition, x represents the x-component of the position of the target sample within the current block (500), y represents the y-component of the position of the target sample within the current block (500), and v x is the x component of the motion vector of the target sample in the current block (500), v y represents the y component of the motion vector of the target sample in the current block (500).

[0100] Since the motion vector of the CP0, the motion vector of the CP1, and the motion vector of the CP2 are known, the motion vector according to the sample position in the current block can be derived based on the mathematical expression 1. That is, according to the affine motion model, the motion vectors v0(v) at the control points are derived based on the distance ratio between the coordinates (x, y) of the target sample and the three control points. 0x , v 0y ), v1(v 1x , v 1y ), v2(v 2x , v 2y ) can be scaled to derive the motion vector of the target sample according to the target sample location. That is, according to the affine motion model, the motion vector of each sample in the current block can be derived based on the motion vectors of the control points. Meanwhile, the set of motion vectors of the samples in the current block derived according to the affine motion model can be expressed as an affine motion vector field (MVF).

[0101] Meanwhile, the six parameters for the above mathematical expression 1 can be expressed as a, b, c, d, e, and f as in the following mathematical expression, and the mathematical expression for the affine motion model expressed by the six parameters can be as follows.

[0102]

[0103] Here, w represents the width of the current block (500), h represents the height of the current block (500), and v 0x , v 0y represents the x-component and y-component of the motion vector of CP0, respectively, and v 1x , v 1y represent the x-component and y-component of the motion vector of CP1, respectively, and v 2x , v 2yEach represents the x-component and y-component of the motion vector of CP2. In addition, x represents the x-component of the position of the target sample within the current block (500), y represents the y-component of the position of the target sample within the current block (500), and v x is the x component of the motion vector of the target sample in the current block (500), v y represents the y component of the motion vector of the target sample in the current block (500).

[0104] The above affine motion model or the above affine inter prediction using the above six parameters can be referred to as a six-parameter affine motion model or AF6.

[0105] Additionally, as an example of the above affine inter prediction, motion information at two control points, i.e., two reference points, may be required.

[0106] Figure 6 illustrates an example of the affine motion model using motion vectors for two control points. The affine motion model using two control points can express three motions, including translational motion, scale motion, and rotational motion. The affine motion model expressing three motions may also be referred to as a similarity affine motion model or a simplified affine motion model.

[0107] If the top-left sample position within the current block (600) is (0,0), the (0,0) and (w, 0) sample positions can be set as the control points as shown in FIG. 6. Hereinafter, the control point of the (0,0) sample position can be represented as CP0, and the control point of the (w, 0) sample position can be represented as CP1.

[0108] A mathematical formula for the affine motion model can be derived using the above-described control points and the motion vectors for the control points. The mathematical formula for the affine motion model can be expressed as follows.

[0109]

[0110] Here, w represents the width of the current block (600), and v 0x , v 0y represents the x-component and y-component of the motion vector of CP0, respectively, and v 1x , v 1y Each represents the x-component and y-component of the motion vector of CP1. In addition, x represents the x-component of the position of the target sample within the current block (600), y represents the y-component of the position of the target sample within the current block (600), and v x is the x component of the motion vector of the target sample in the current block (600), v y represents the y component of the motion vector of the target sample in the current block (600).

[0111] Meanwhile, the four parameters for the above mathematical expression 3 can be expressed as a, b, c, and d as in the following mathematical expression, and the mathematical expression for the affine motion model expressed by the four parameters can be as follows.

[0112]

[0113] Here, w represents the width of the current block (600), and v 0x , v 0y represents the x-component and y-component of the motion vector of CP0, respectively, and v 1x , v 1yEach represents the x-component and y-component of the motion vector of CP1. In addition, x represents the x-component of the position of the target sample within the current block (600), y represents the y-component of the position of the target sample within the current block (600), and v x is the x component of the motion vector of the target sample in the current block (600), v y represents the y component of the motion vector of the target sample in the current block (600). The affine motion model using the two control points can be expressed by four parameters a, b, c, d as in the mathematical expression 4, and the affine motion model or the affine inter prediction using the four parameters can be expressed as a four-parameter affine motion model or AF4. That is, according to the affine motion model, the motion vector of each sample in the current block can be derived based on the motion vectors of the control points. Meanwhile, a set of motion vectors of samples in the current block derived according to the affine motion model can be expressed as an affine motion vector field (MVF).

[0114] Meanwhile, as described above, a motion vector per sample can be derived using the affine motion model, thereby significantly improving the accuracy of inter prediction. However, this may significantly increase the complexity of the motion compensation process.

[0115] Accordingly, instead of deriving a motion vector in sample units, it is possible to restrict the derivation of a motion vector in sub-block units within the current block.

[0116] Fig. 7 exemplarily shows a method for deriving a motion vector in units of sub-blocks based on the affine motion model. Fig. 7 exemplarily shows a case where the size of the current block is 16×16 and a motion vector is derived in units of 4×4 sub-blocks. The sub-blocks can be set to various sizes, and for example, when the sub-blocks are set to an n×n size (n is a positive integer, e.g., n is 4), a motion vector can be derived in units of n×n sub-blocks within the current block based on the affine motion model, and various methods for deriving a motion vector representing each sub-block can be applied.

[0117] For example, referring to FIG. 7, the motion vector of each sub-block can be derived using the center or center lower right side sample position of each sub-block as a representative coordinate. Here, the center lower right position may refer to a sample position located on the lower right side among four samples located at the center of the sub-block. For example, when n is an odd number, one sample may be located at the exact center of the sub-block, in which case the center sample position may be used to derive the motion vector of the sub-block. However, when n is an even number, four samples may be located adjacently at the center of the sub-block, in which case the lower right sample position may be used to derive the motion vector. For example, referring to FIG. 7, the representative coordinates for each sub-block can be derived as (2, 2), (6, 2), (10, 2), ..., (14, 14), and the encoding device / decoding device can derive the motion vector of each sub-block by substituting each of the representative coordinates of the sub-blocks into the above-described mathematical expression 1 or 3. The motion vectors of the sub-blocks in the current block derived through the above-described affine motion model can be expressed as affine MVF.

[0118] Meanwhile, as an example, the size of a sub-block within the current block may be derived based on the following mathematical formula.

[0119]

[0120] Here, M represents the width of the sub-block, and N represents the height of the sub-block. Also, v 0x , v 0y represents the x-component and y-component of CPMV0 of the current block, respectively, and v 0x , v 0y represents the x-component and y-component of CPMV1 of the current block, respectively, w represents the width of the current block, h represents the height of the current block, and MvPre represents the motion vector fraction accuracy. For example, the motion vector fraction accuracy can be set to 1 / 16.

[0121] Meanwhile, inter prediction using the above-described affine motion model, i.e., affine motion prediction, may have an affine merge mode (AF_MERGE) and an affine inter mode (AF_INTER). Here, the affine inter mode may also be expressed as an affine MVP mode (affine motion vector prediction mode, AF_MVP).

[0122] The above affine merge mode is similar to the existing merge mode in that it does not transmit the MVD for the motion vector of the control points. That is, the above affine merge mode can represent an encoding / decoding method that performs prediction by deriving CPMV for each of two or three control points from neighboring blocks of the current block without coding the MVD (motion vector difference), similar to the existing skip / merge mode.

[0123] For example, when the AF_MRG mode is applied to the current block, MVs for CP0 and CP1 (i.e., CPMV0 and CPMV1) can be derived from the neighboring blocks of the current block to which the affine mode is applied. That is, CPMV0 and CPMV1 of the neighboring blocks to which the affine mode is applied can be derived as merge candidates, and CPMV0 and CPMV1 for the current block can be derived based on the merge candidates. An affine motion model can be derived based on CPMV0 and CPMV1 of the neighboring blocks indicated by the merge candidates, and the CPMV0 and CPMV1 for the current block can be derived based on the affine motion model.

[0124] The above affine inter mode may represent inter prediction that derives an MVP (motion vector predictor) for the motion vectors of the control points, derives the motion vectors of the control points based on the received MVD (motion vector difference) and the MVP, derives an affine MVF of the current block based on the motion vectors of the control points, and performs prediction based on the affine MVF. Here, the motion vector of the control point may be expressed as CPMV (Control Point Motion Vector), the MVP of the control point may be expressed as CPMVP (Control Point Motion Vector Predictor), and the MVD of the control point may be expressed as CPMVD (Control Point Motion Vector Difference). Specifically, for example, the encoding device can derive a control point point motion vector predictor (CPMVP) and a control point point motion vector (CPMV) for each of CP0 and CP1 (or CP0, CP1, and CP2), and transmit or store information about the CPMVP and / or a CPMVD which is a difference value between the CPMVP and CPMV.

[0125] Here, when the above affine inter mode is applied to the current block, the encoding device / decoding device can construct an affine MVP candidate list based on the surrounding blocks of the current block, and the affine MVP candidate can be referred to as a CPMVP pair candidate, and the affine MVP candidate list can also be referred to as a CPMVP candidate list.

[0126] Additionally, each affine MVP candidate can mean a combination of CPMVPs of CP0 and CP1 in a four-parameter affine motion model, and can mean a combination of CPMVPs of CP0, CP1, and CP2 in a six-parameter affine motion model.

[0127] Fig. 8 illustrates an exemplary flowchart of an affine motion prediction method according to one embodiment of the present document.

[0128] Referring to Fig. 8, the affine motion prediction method can be broadly expressed as follows. When the affine motion prediction method starts, a CPMV pair can first be acquired (S800). Here, the CPMV pair can include CPMV0 and CPMV1 when using a 4-parameter affine model.

[0129] Afterwards, affine motion compensation can be performed based on the CPMV pair (S810), and affine motion prediction can be terminated.

[0130] Additionally, there may be two affine prediction modes to determine the CPMV0 and the CPMV1. Here, the two affine prediction modes may include an affine inter mode and an affine merge mode. The affine inter mode can clearly determine CPMV0 and CPMV1 by signaling two pieces of motion vector difference (MVD) information for CPMV0 and CPMV1. On the other hand, the affine merge mode can derive CPMV pairs without signaling MVD information.

[0131] In other words, the affine merge mode can derive the CPMV of the current block using the CPMV of the surrounding blocks coded in the affine mode, and when the motion vector is determined in units of sub-blocks, the affine merge mode can also be referred to as a sub-block merge mode.

[0132] In affine merge mode, the encoding device can signal an index of a neighboring block coded in affine mode for deriving a CPMV of a current block to a decoding device, and can also signal a difference value between the CPMV of the neighboring block and the CPMV of the current block. Here, the affine merge mode can construct an affine merge candidate list based on the neighboring blocks, and the index of the neighboring block can indicate a neighboring block to be referenced for deriving the CPMV of the current block among the affine merge candidate list. The affine merge candidate list may also be referred to as a subblock merge candidate list.

[0133] The affine inter mode may also be referred to as the affine MVP mode. In the affine MVP mode, the CPMV of the current block can be derived based on the CPMVP (Control Point Motion Vector Predictor) and the CPMVD (Control Point Motion Vector Difference). In other words, the encoding device can determine the CPMVP for the CPMV of the current block, derive the CPMVD, which is the difference between the CPMV and the CPMVP of the current block, and signal information about the CPMVP and information about the CPMVD to the decoding device. Here, the affine MVP mode can construct an affine MVP candidate list based on neighboring blocks, and the information about the CPMVP can indicate neighboring blocks to be referenced to derive the CPMVP for the CPMV of the current block among the affine MVP candidate list. The affine MVP candidate list may also be referred to as a control point motion vector predictor candidate list.

[0134] For example, when the affine inter mode of the 6-parameter affine motion model is applied, the current block can be encoded as described below.

[0135] FIG. 9 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to one embodiment of the present document.

[0136] Referring to Fig. 9, the motion vector of CP0 of the current block can be expressed as v0, the motion vector of CP1 as v1, the motion vector of the control point at the bottom-left sample position as v2, and the motion vector of CP2 as v3. That is, v0 can represent the CPMVP of CP0, v1 can represent the CPMVP of CP1, and v2 can represent the CPMVP of CP2.

[0137] The Affine MVP candidate may be a combination of the CPMVP candidate of the CP0, the CPMVP candidate of the CP1, and the candidate of the CP2.

[0138] For example, the above affine MVP candidate can be derived as follows.

[0139] Specifically, up to 12 CPMVP candidate combinations can be determined as shown in the following mathematical formula.

[0140]

[0141] Here, v A is the motion vector of the surrounding block A, v B is the motion vector of the surrounding block B, v C is the motion vector of the surrounding block C, v D is the motion vector of the surrounding block D, v E is the motion vector of the surrounding block E, v F is the motion vector of the surrounding block F, v G can represent the motion vector of the surrounding block G.

[0142] In addition, the peripheral block A may represent a peripheral block located at the upper left of the upper left sample position of the current block, the peripheral block B may represent a peripheral block located above the upper left sample position of the current block, and the peripheral block C may represent a peripheral block located to the left of the upper left sample position of the current block. In addition, the peripheral block D may represent a peripheral block located above the upper right sample position of the current block, and the peripheral block E may represent a peripheral block located at the upper right of the upper right sample position of the current block. In addition, the peripheral block F may represent a peripheral block located to the left of the lower left sample position of the current block, and the peripheral block G may represent a peripheral block located at the lower left of the lower left sample position of the current block.

[0143] That is, referring to the above mathematical expression 6, the CPMVP candidate of the CP0 is the motion vector v of the surrounding block A. A , the motion vector v of the surrounding block B B and / or the motion vector v of the surrounding block C C , and the CPMVP candidate of CP1 may include the motion vector v of the surrounding block D. D , and / or the motion vector v of the surrounding block E E , and the CPMVP candidate of CP2 may include the motion vector v of the surrounding block F. F , and / or the motion vector v of the surrounding block G G may include.

[0144] In other words, the CPMVP v0 of the CP0 can be derived based on the motion vector of at least one of the surrounding blocks A, B, and C of the upper left sample position. Here, the surrounding block A may refer to a block located at the upper left of the upper left sample position of the current block, the surrounding block B may refer to a block located above the upper left sample position of the current block, and the surrounding block C may refer to a block located to the left of the upper left sample position of the current block.

[0145] Based on the motion vectors of the surrounding blocks, up to 12 CPMVP candidate combinations including the CPMVP candidate of CP0, the CPMVP candidate of CP1, and the CPMVP candidate of CP2 can be derived.

[0146] Afterwards, the derived CPMVP candidate combinations are sorted in order of smallest DV, and the top two CPMVP candidate combinations can be derived as the above-mentioned affine MVP candidates.

[0147] The DV of the CPMVP candidate combination can be derived using the following mathematical formula.

[0148]

[0149] Thereafter, the encoding device can determine CPMVs for each of the affine MVP candidates, compare RD (Rate Distortion) costs for the CPMVs, and select an affine MVP candidate with a smaller RD cost as the optimal affine MVP candidate for the current block. The encoding device can encode and signal an index and CPMVD indicating the optimal candidate.

[0150] Additionally, for example, when affine merge mode is applied, the current block can be encoded as described below.

[0151] FIG. 10 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to one embodiment of the present document.

[0152] An affine merge candidate list of the current block can be constructed based on the surrounding blocks of the current block illustrated in FIG. 10. The surrounding blocks can include surrounding block A, surrounding block B, surrounding block C, surrounding block D, and surrounding block E. The surrounding block A can represent a left surrounding block of the current block, the surrounding block B can represent an upper surrounding block of the current block, the surrounding block C can represent an upper-right corner surrounding block of the current block, the surrounding block D can represent a lower-left corner surrounding block of the current block, and the surrounding block E can represent an upper-left corner surrounding block of the current block.

[0153] For example, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the left peripheral block may be a block including a sample with coordinates (-1, H-1), the upper peripheral block may be a block including a sample with coordinates (W-1, -1), the upper-right corner peripheral block may be a block including a sample with coordinates (W, -1), the lower-left corner peripheral block may be a block including a sample with coordinates (-1, H), and the upper-left corner peripheral block may be a block including a sample with coordinates (-1, -1).

[0154] Specifically, for example, the encoding device can scan the surrounding blocks A, B, C, D, and E of the current block in a specific scanning order, and determine the surrounding block encoded in the affine prediction mode first in the scanning order as a candidate block of the affine merge mode, i.e., an affine merge candidate. Here, for example, the specific scanning order may be an alphabetical order. That is, the specific scanning order may be the order of surrounding block A, B, C, D, and E.

[0155] Thereafter, the encoding device can determine an affine motion model of the current block using the CPMV of the determined candidate block, determine a CPMV of the current block based on the affine motion model, and determine an affine MVF of the current block based on the CPMV.

[0156] For example, if a surrounding block A is determined to be a candidate block of the current block, it can be coded as described below.

[0157] Figure 11 shows an example of affine prediction performed when a surrounding block A is selected as an affine merge candidate.

[0158] Referring to FIG. 11, the encoding device can determine a neighboring block A of the current block as a candidate block, and derive an affine motion model of the current block based on CPMV, v2, and v3 of the neighboring block. Thereafter, the encoding device can determine CPMV, v0, and v1 of the current block based on the affine motion model. The encoding device can determine an affine MVF based on CPMV, v0, and v1 of the current block, and perform an encoding process for the current block based on the affine MVF.

[0159] Meanwhile, with regard to affine inter prediction, inherited affine candidates and constructed affine candidates are being considered for constructing the list of affine MVP candidates.

[0160] Here, the inherited affine candidates may be as follows.

[0161] For example, if a neighboring block of the current block is an affine block and a reference picture of the current block is the same as a reference picture of the neighboring block, an affine MVP pair of the current block can be determined from an affine motion model of the neighboring block. Here, the affine block can represent a block to which the affine inter prediction is applied. The inherited affine candidate can represent CPMVPs (e.g., the affine MVP pair) derived based on the affine motion model of the neighboring block.

[0162] Specifically, as an example, the inherited affine candidate can be derived as described below.

[0163] Figure 12 illustrates peripheral blocks for deriving the above-described inherited affine candidate.

[0164] Referring to FIG. 12, the surrounding blocks of the current block may include a left surrounding block A0 of the current block, a lower left corner surrounding block A1 of the current block, an upper surrounding block B0 of the current block, an upper right corner surrounding block B1 of the current block, and an upper left corner surrounding block B2 of the current block.

[0165] For example, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the left peripheral block may be a block including a sample with coordinates (-1, H-1), the upper peripheral block may be a block including a sample with coordinates (W-1, -1), the upper-right corner peripheral block may be a block including a sample with coordinates (W, -1), the lower-left corner peripheral block may be a block including a sample with coordinates (-1, H), and the upper-left corner peripheral block may be a block including a sample with coordinates (-1, -1).

[0166] An encoding device / decoding device can sequentially check neighboring blocks A0, A1, B0, B1, and B2, and if the neighboring blocks are coded using an affine motion model, and a reference picture of the current block and a reference picture of the neighboring block are the same, two CPMVs or three CPMVs of the current block can be derived based on the affine motion model of the neighboring blocks. The CPMVs can be derived as affine MVP candidates of the current block. The affine MVP candidates can represent the inherited affine candidates.

[0167] For example, up to two inherited affine candidates can be derived based on the above surrounding blocks.

[0168] For example, the encoding device / decoding device can derive a first affine MVP candidate of the current block based on a first block within the surrounding blocks. Here, the first block can be coded with an affine motion model, and the reference picture of the first block can be the same as the reference picture of the current block. That is, the first block can be a block that satisfies a condition that is first confirmed by checking the surrounding blocks in a specific order. The condition can be coded with an affine motion model, and the reference picture of the block can be the same as the reference picture of the current block.

[0169] Thereafter, the encoding device / decoding device can derive a second affine MVP candidate of the current block based on a second block within the surrounding blocks. Here, the second block can be coded with an affine motion model, and the reference picture of the second block can be the same as the reference picture of the current block. That is, the second block can be a block that satisfies a second confirmed condition by checking the surrounding blocks in a specific order. The condition can be coded with an affine motion model, and the reference picture of the block can be the same as the reference picture of the current block.

[0170] Meanwhile, for example, if the available number of inherited affine candidates is less than 2 (i.e., the number of derived inherited affine candidates is less than 2), a constructed affine candidate may be considered. The constructed affine candidate may be derived as follows.

[0171] Figure 13 illustrates an example of a spatial candidate for the above-mentioned constructed affine candidate.

[0172] As illustrated in Fig. 13, the motion vectors of the surrounding blocks of the current block can be divided into three groups. Referring to Fig. 13, the surrounding blocks can include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G.

[0173] The above-described peripheral block A may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block B may represent a peripheral block located at the upper left of the upper left sample position of the current block, and the above-described peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block. In addition, the above-described peripheral block D may represent a peripheral block located at the upper right of the upper right sample position of the current block, and the above-described peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block. In addition, the above-described peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the above-described peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0174] For example, the three groups above may include S0, S1, and S2, and the S0, the S1, and the S2 may be derived as shown in the following table.

[0175]

[0176] Here, mv A is the motion vector of the surrounding block A, mv B is the motion vector of the surrounding block B, mv C is the motion vector of the surrounding block C, mv D is the motion vector of the surrounding block D, mv E is the motion vector of the surrounding block E, mv Fis the motion vector of the surrounding block F, mv G represents the motion vector of the surrounding block G. S0 may be represented as the first group, S1 as the second group, and S2 as the third group.

[0177] The encoding device / decoding device can derive mv0 from the above S0, can derive mv1 from S1, can derive mv2 from S2, and can derive an affine MVP candidate including the above mv0, the above mv1, and the above mv2. The above affine MVP candidate can represent the above constructed affine candidate. In addition, the above mv0 can be a CPMVP candidate of CP0, the above mv1 can be a CPMVP candidate of CP1, and the above mv2 can be a CPMVP candidate of CP2.

[0178] Here, the reference picture for the mv0 may be the same as the reference picture of the current block. That is, the mv0 may be a motion vector that satisfies a condition that is first confirmed by checking the motion vectors in the S0 in a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be the neighboring block A → the neighboring block B → the neighboring block C in the S0. In addition, the process may be performed in an order other than the above-described order, and may not be limited to the above-described example.

[0179] In addition, the reference picture for the mv1 may be the same as the reference picture of the current block. That is, the mv1 may be a motion vector that satisfies a condition that is first confirmed by checking the motion vectors in the S1 in a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be from the neighboring block D to the neighboring block E in the S1. In addition, the process may be performed in an order other than the above-described order, and may not be limited to the above-described example.

[0180] In addition, the reference picture for the mv2 may be the same as the reference picture of the current block. That is, the mv2 may be a motion vector that satisfies a condition that is first confirmed by checking the motion vectors in the S2 in a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be from the neighboring block F to the neighboring block G in the S2. In addition, the process may be performed in an order other than the above-described order, and may not be limited to the above-described example.

[0181] Meanwhile, when only the above mv0 and the above mv1 are available, i.e., when only the above mv0 and the above mv1 are derived, the above mv2 can be derived as in the following mathematical formula.

[0182]

[0183] Here, mv2 x represents the x component of the above mv2, and mv2 y represents the y component of the above mv2, and mv0 x represents the x component of the above mv0, and mv0 y represents the y component of the above mv0, and mv1 x represents the x component of the above mv1, and mv1 yrepresents the y component of the above mv1. In addition, w represents the width of the current block, and h represents the height of the current block.

[0184] Meanwhile, when only the above mv0 and the above mv2 are derived, the above mv1 can be derived as in the following mathematical formula.

[0185]

[0186] Here, mv1 x represents the x component of the above mv1, and mv1 y represents the y component of the above mv1, and mv0 x represents the x component of the above mv0, and mv0 y represents the y component of the above mv0, and mv2 x represents the x component of the above mv2, and mv2 y represents the y component of the above mv2. In addition, w represents the width of the current block, and h represents the height of the current block.

[0187] Additionally, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the AMVP process of the existing HEVC standard may be applied to construct the affine MVP list. That is, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the process of constructing MVP candidates in the existing HEVC standard may be performed.

[0188] Meanwhile, the flowcharts of the examples that constitute the above-described Affine MVP list are as follows.

[0189] Figure 14 illustrates an example of constructing an affine MVP list.

[0190] Referring to FIG. 14, the encoding / decoding device may add an inherited candidate to the affine MVP list of the current block (S1400). The inherited candidate may represent the inherited affine candidate described above.

[0191] Specifically, the encoding device / decoding device can derive up to two inherited affine candidates from the surrounding blocks of the current block (S1405). Here, the surrounding blocks can include a left surrounding block A0, a lower left corner surrounding block A1, an upper surrounding block B0, an upper right corner surrounding block B1, and an upper left corner surrounding block B2 of the current block.

[0192] For example, the encoding device / decoding device can derive a first affine MVP candidate of the current block based on a first block within the surrounding blocks. Here, the first block can be coded with an affine motion model, and the reference picture of the first block can be the same as the reference picture of the current block. That is, the first block can be a block that satisfies a condition that is first confirmed by checking the surrounding blocks in a specific order. The condition can be coded with an affine motion model, and the reference picture of the block can be the same as the reference picture of the current block.

[0193] Thereafter, the encoding device / decoding device can derive a second affine MVP candidate of the current block based on a second block within the surrounding blocks. Here, the second block can be coded with an affine motion model, and the reference picture of the second block can be the same as the reference picture of the current block. That is, the second block can be a block that satisfies a second confirmed condition by checking the surrounding blocks in a specific order. The condition can be coded with an affine motion model, and the reference picture of the block can be the same as the reference picture of the current block.

[0194] Meanwhile, the specific order may be left peripheral block A0 → lower left corner peripheral block A1 → upper peripheral block B0 → upper right corner peripheral block B1 → upper left corner peripheral block B2. In addition, the order may be performed in a different order than the above-described order, and may not be limited to the above-described example.

[0195] The encoding device / decoding device may add a constructed candidate to the affine MVP list of the current block (S1410). The constructed candidate may represent the above-described constructed affine candidate. The constructed candidate may also be represented as a constructed affine MVP candidate. If the number of available inherited candidates is less than two, the encoding device / decoding device may add the constructed candidate to the affine MVP list of the current block. For example, the encoding device / decoding device may derive one constructed affine candidate.

[0196] Meanwhile, the method for deriving the constructed affine candidate may vary depending on whether the affine motion model applied to the current block is a 6-affine motion model or a 4-affine motion model. Specific details regarding the method for deriving the constructed candidate will be described below.

[0197] The encoding device / decoding device may add the HEVC AMVP candidate to the affine MVP list of the current block (S1420). If the number of available inherited candidates and / or constructed candidates is less than two, the encoding device / decoding device may add the HEVC AMVP candidate to the affine MVP list of the current block. That is, if the number of available inherited candidates and / or constructed candidates is less than two, the encoding device / decoding device may perform a process of constructing an MVP candidate in the existing HEVC standard.

[0198] Meanwhile, a method for deriving the above-mentioned constructed candidate may be as follows.

[0199] For example, if the affine motion model applied to the current block is a 6-affine motion model, the constructed candidate can be derived as in the example illustrated in FIG. 15.

[0200] Figure 15 shows an example of deriving the above-mentioned constructed candidate.

[0201] Referring to Fig. 15, the encoding device / decoding device can check mv0, mv1, and mv2 for the current block (S1500). That is, the encoding device / decoding device can check mv0, mv1, and mv2 available in the surrounding blocks of the current block. It can be determined whether this exists. Here, the mv0 may be a CPMVP candidate of CP0 of the current block, the mv1 may be a CPMVP candidate of CP1, and the mv2 may be a CPMVP candidate of CP2. In addition, the mv0, the mv1, and the mv2 can be represented as candidate motion vectors for the above CPs.

[0202] For example, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the first group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as the mv0. That is, the mv0 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the first group in a specific order. If the motion vectors of the surrounding blocks in the first group do not satisfy the specific condition, an available mv0 may not exist. Here, for example, the specific order may be an order from the surrounding block A to the surrounding block B to the surrounding block C in the first group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0203] In addition, for example, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the second group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as the mv1. That is, the mv1 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the surrounding blocks in the second group do not satisfy the specific condition, an available mv1 may not exist. Here, for example, the specific order may be an order from the surrounding block D to the surrounding block E in the second group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0204] In addition, for example, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as the mv2. That is, the mv2 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the third group in a specific order. If the motion vectors of the surrounding blocks in the third group do not satisfy the specific condition, an available mv2 may not exist. Here, for example, the specific order may be an order from the surrounding block F to the surrounding block G in the third group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0205] Meanwhile, the first group may include a motion vector of a surrounding block A, a motion vector of a surrounding block B, and a motion vector of a surrounding block C, the second group may include a motion vector of a surrounding block D, and a motion vector of a surrounding block E, and the third group may include a motion vector of a surrounding block F, and a motion vector of a surrounding block G. The above-described peripheral block A may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block B may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the above-described peripheral block D may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the above-described peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0206] If only mv0 and mv1 for the current block are available, i.e., if only mv0 and mv1 for the current block are derived, the encoding device / decoding device can derive mv2 for the current block based on the above-described mathematical expression 8 (S1510). The encoding device / decoding device can derive mv2 by substituting the derived mv0 and mv1 into the above-described mathematical expression 8.

[0207] If only mv0 and mv2 for the current block are available, i.e., if only mv0 and mv2 for the current block are derived, the encoding device / decoding device can derive mv1 for the current block based on the above-described mathematical expression 9 (S1520). The encoding device / decoding device can derive mv1 by substituting the derived mv0 and mv2 into the above-described mathematical expression 9.

[0208] The encoding device / decoding device is derived from the above mv0, mv1 and mv2 can be derived as a constructed candidate of the current block (S1530). If the mv0, the mv1, and the mv2 are available, that is, based on the surrounding blocks of the current block, the mv0, the mv1, and the mv2 In this derived case, the encoding device / decoding device is the derived mv0, the mv1 and the mv2 can be derived as a constructed candidate for the current block.

[0209] In addition, when only the mv0 and the mv1 for the current block are available, i.e., when only the mv0 and the mv1 for the current block are derived, the encoding device / decoding device can derive the derived mv0, the mv1, and the mv2 derived based on the above-described mathematical expression 8 as the constructed candidates for the current block.

[0210] In addition, when only the mv0 and the mv2 for the current block are available, i.e., when only the mv0 and the mv2 for the current block are derived, the encoding device / decoding device can derive the derived mv0, the mv2, and the mv1 derived based on the above-described mathematical expression 9 as the constructed candidates for the current block.

[0211] Additionally, for example, if the affine motion model applied to the current block is a 4-affine motion model, the constructed candidate can be derived as in the example illustrated in FIG. 15.

[0212] Figure 16 shows an example of deriving the above-mentioned constructed candidate.

[0213] Referring to Fig. 16, the encoding device / decoding device can check mv0, mv1, and mv2 for the current block (S1600). That is, the encoding device / decoding device can check mv0, mv1, and mv2 available in the surrounding blocks of the current block. It can be determined whether this exists. Here, the mv0 may be a CPMVP candidate of CP0 of the current block, the mv1 may be a CPMVP candidate of CP1, and the mv2 may be a CPMVP candidate of CP2.

[0214] For example, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the first group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as the mv0. That is, the mv0 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the first group in a specific order. If the motion vectors of the surrounding blocks in the first group do not satisfy the specific condition, an available mv0 may not exist. Here, for example, the specific order may be an order from the surrounding block A to the surrounding block B to the surrounding block C in the first group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0215] In addition, for example, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the second group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as the mv1. That is, the mv1 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the surrounding blocks in the second group do not satisfy the specific condition, an available mv1 may not exist. Here, for example, the specific order may be an order from the surrounding block D to the surrounding block E in the second group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0216] In addition, for example, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as the mv2. That is, the mv2 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the third group in a specific order. If the motion vectors of the surrounding blocks in the third group do not satisfy the specific condition, an available mv2 may not exist. Here, for example, the specific order may be an order from the surrounding block F to the surrounding block G in the third group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0217] Meanwhile, the first group may include a motion vector of a surrounding block A, a motion vector of a surrounding block B, and a motion vector of a surrounding block C, the second group may include a motion vector of a surrounding block D, and a motion vector of a surrounding block E, and the third group may include a motion vector of a surrounding block F, and a motion vector of a surrounding block G. The above-described peripheral block A may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block B may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the above-described peripheral block D may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the above-described peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0218] When only mv0 and mv1 for the current block are available, or when mv0, mv1, and mv2 for the current block are available, i.e., when only mv0 and mv1 for the current block are derived, or when mv0, mv1, and mv2 for the current block are derived, the encoding device / decoding device can derive the derived mv0 and mv1 as constructed candidates for the current block (S1610).

[0219] Meanwhile, if only mv0 and mv2 for the current block are available, i.e., if only mv0 and mv2 for the current block are derived, the encoding device / decoding device can derive mv1 for the current block based on the above-described mathematical expression 9 (S1620). The encoding device / decoding device can derive mv1 by substituting the derived mv0 and mv2 into the above-described mathematical expression 9.

[0220] Thereafter, the encoding device / decoding device can derive the derived mv0 and mv1 as constructed candidates of the current block (S1610).

[0221] Meanwhile, this document proposes another embodiment for deriving the inherited affine candidates. The proposed embodiment can improve coding performance by reducing computational complexity in deriving the inherited affine candidates.

[0222] Figure 17 illustrates an example of the locations of surrounding blocks scanned to derive inherited affine candidates.

[0223] The encoding device / decoding device can derive up to two inherited affine candidates from the neighboring blocks of the current block. Fig. 17 may represent the neighboring blocks for the inherited affine candidates. For example, the neighboring blocks may include neighboring block A and neighboring block B as illustrated in Fig. 17. The neighboring block A may represent the left neighboring block A0 described above, and the neighboring block B may represent the upper neighboring block B0 described above.

[0224] For example, the encoding device / decoding device can check whether the neighboring blocks are available in a specific order, and can derive an inherited affine candidate of the current block based on the first identified available neighboring block. That is, the encoding device / decoding device can check whether the neighboring blocks satisfy a specific condition in a specific order, and can derive an inherited affine candidate of the current block based on the first identified available neighboring block. In addition, the encoding device / decoding device can derive an inherited affine candidate of the current block based on the second identified neighboring block that satisfies a specific condition. That is, the encoding device / decoding device can derive an inherited affine candidate of the current block based on the second identified neighboring block that satisfies a specific condition. Here, the availability may be coded with an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. That is, the specific condition may be coded as an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. Also, for example, the specific order may be the neighboring block A → the neighboring block B. Meanwhile, a pruning check process between two inherited affine candidates (i.e., derived inherited affine candidates) may not be performed. The pruning check process may refer to a process of checking whether they are identical and, if they are identical candidates, removing the candidate derived in a later order.

[0225] The above-described embodiment proposes a method of deriving the inherited affine candidate by checking only two surrounding blocks (i.e., surrounding block A, surrounding block B), instead of deriving the inherited affine candidate by checking all of the existing surrounding blocks (i.e., surrounding block A, surrounding block B). Here, the surrounding block C may represent the above-described upper-right corner surrounding block B1, the surrounding block D may represent the above-described lower-left corner surrounding block A1, and the surrounding block E may represent the above-described upper-left corner surrounding block B2.

[0226] In order to analyze the spatial correlation between the surrounding blocks and the current block according to affine inter prediction, the probability that the affine prediction is applied to the current block when the affine prediction is applied to each surrounding block can be referenced. The probability that the affine prediction is applied to the current block when the affine prediction is applied to each surrounding block can be derived as shown in the following table.

[0227]

[0228] Referring to Table 2 above, it can be confirmed that among the surrounding blocks, the surrounding blocks A and B have high spatial correlations with respect to the current block. Therefore, through an embodiment of deriving the inherited affine candidates using only the surrounding blocks A and B with high spatial correlations, it is possible to obtain the effect of deriving high decoding performance while reducing the processing time.

[0229] Meanwhile, the pruning check process can be performed to prevent the existence of identical candidates in the candidate list. The pruning check process can eliminate redundancy, which may result in an advantage in encoding efficiency; however, it has the disadvantage of increasing computational complexity by performing the pruning check process. In particular, the pruning check process for affine candidates has very high computational complexity because it must be performed on the affine type (e.g., whether the affine motion model is a 4-affine motion model or a 6-affine motion model), the reference picture (or reference picture index), and the MVs of CP0, CP1, and CP2. Therefore, the present embodiment proposes a method of not performing the pruning check process between the inherited affine candidate derived based on the neighboring block A (e.g., inherited_A) and the inherited affine candidate derived based on the neighboring block B (e.g., inherited_B). For the surrounding blocks A and B, the distance is far and therefore the spatial correlation is low, so the likelihood that inherited_A and inherited_B are identical is low. Therefore, it may be reasonable not to perform the pruning check process between the inherited affine candidates.

[0230] Alternatively, a method for performing a minimal pruning check process based on the above-mentioned basis may be proposed. For example, the encoding / decoding device may perform the pruning check process by comparing only the MVs of CP0 of the inherited affine candidates.

[0231] Furthermore, this document proposes a method for deriving constructed candidates different from the aforementioned embodiments. The proposed embodiment can improve coding performance by reducing complexity compared to the aforementioned embodiments for deriving constructed candidates. The proposed embodiment is described below. Furthermore, when the available number of inherited affine candidates is less than 2 (i.e., when the number of derived inherited affine candidates is less than 2), a constructed affine candidate may be considered.

[0232] For example, the encoding device / decoding device can check mv0, mv1, mv2 for the current block. That is, the encoding device / decoding device can check mv0, mv1, mv2 available in the surrounding blocks of the current block. It can be determined whether this exists. Here, the mv0 may be a CPMVP candidate of CP0 of the current block, the mv1 may be a CPMVP candidate of CP1, and the mv2 may be a CPMVP candidate of CP2.

[0233] Specifically, the surrounding blocks of the current block can be divided into three groups, and the surrounding blocks can include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G. The first group can include a motion vector of surrounding block A, a motion vector of surrounding block B, and a motion vector of surrounding block C, the second group can include a motion vector of surrounding block D, a motion vector of surrounding block E, and the third group can include a motion vector of surrounding block F, a motion vector of surrounding block G. The above-described peripheral block A may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block B may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the above-described peripheral block D may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the above-described peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0234] The encoding device / decoding device can determine whether there is an available mv0 in the first group, whether there is an available mv1 in the second group, and whether there is an available mv2 in the third group. can determine whether it exists.

[0235] Specifically, for example, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the first group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as the mv0. That is, the mv0 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the first group in a specific order. If the motion vectors of the surrounding blocks in the first group do not satisfy the specific condition, an available mv0 may not exist. Here, for example, the specific order may be an order from the surrounding block A to the surrounding block B to the surrounding block C in the first group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0236] In addition, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the second group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as mv1. That is, the mv1 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the surrounding blocks in the second group do not satisfy the specific condition, an available mv1 may not exist. Here, for example, the specific order may be an order from the surrounding block D to the surrounding block E in the second group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0237] In addition, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as mv2. That is, the mv2 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the third group in a specific order. If the motion vectors of the surrounding blocks in the third group do not satisfy the specific condition, an available mv2 may not exist. Here, for example, the specific order may be an order from the surrounding block F to the surrounding block G in the third group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0238] Afterwards, if the affine motion model applied to the current block is a 4-affine motion model, mv0 for the current block and if mv1 is available, the encoding device / decoding device derives the above mv0 and mv1 can be derived as constructed candidates for the current block. Meanwhile, mv0 for the current block and / or if mv1 is not available, i.e. mv0 from a neighboring block of the current block. and mv1 If at least one of them is not derived, the encoding device / decoding device may not add the constructed candidate to the affine MVP list of the current block.

[0239] In addition, if the affine motion model applied to the current block is a 6-affine motion model, mv0 and mv1 for the current block and mv2 are available, the encoding device / decoding device derives the above mv0, mv1 and mv2 can be derived as constructed candidates for the current block. Meanwhile, mv0 and mv1 for the current block and / or mv2 is not available, i.e., mv0, mv1 from the surrounding blocks of the current block. and mv2 If at least one of them is not derived, the encoding device / decoding device may not add the constructed candidate to the affine MVP list of the current block.

[0240] The proposed embodiment described above is a method for considering as a constructed candidate only when all motion vectors of CPs for generating an affine motion model of the current block are available. Here, the meaning of available may indicate that the reference picture of the neighboring block and the reference picture of the current block are the same. In other words, the constructed candidate can be derived only when a motion vector satisfying the above condition exists among the motion vectors of the neighboring blocks for each of the CPs of the current block. Accordingly, when the affine motion model applied to the current block is a 4-affine motion model, the constructed candidate can be considered only when the MVs of CP0 and CP1 of the current block (i.e., mv0 and mv1) are available. In addition, when the affine motion model applied to the current block is a 6-affine motion model, the constructed candidate can be considered only when the MVs of CP0, CP1, and CP2 of the current block (i.e., mv0, mv1, and mv2) are available. Therefore, according to the proposed embodiment, an additional configuration for deriving a motion vector for CP based on the above-described mathematical expression 8 or mathematical expression 9 may not be required. This can reduce the computational complexity for deriving the constructed candidate. In addition, since the constructed candidate is determined only when a CPMVP candidate having the same reference picture is available, the overall coding performance can be improved.

[0241] Meanwhile, a pruning check process may not be performed between the derived inherited affine candidates and the constructed affine candidates. The pruning check process may refer to a process of checking whether the candidates are identical and, if so, eliminating the candidate derived later.

[0242] The above-described embodiment can be represented as in FIG. 18 and FIG. 19.

[0243] Figure 18 shows an example of deriving the constructed candidate when a 4-affine motion model is applied to the current block.

[0244] Referring to FIG. 18, the encoding device / decoding device can determine whether mv0 and mv1 are available for the current block (S1800). That is, the encoding device / decoding device can determine whether mv0 and mv1 are available in the surrounding blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, and mv1 may be a CPMVP candidate for CP1.

[0245] The encoding device / decoding device can determine whether there is an available mv0 in the first group and whether there is an available mv1 in the second group.

[0246] Specifically, the surrounding blocks of the current block can be divided into three groups, and the surrounding blocks can include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G. The first group can include a motion vector of surrounding block A, a motion vector of surrounding block B, and a motion vector of surrounding block C, the second group can include a motion vector of surrounding block D, a motion vector of surrounding block E, and the third group can include a motion vector of surrounding block F, a motion vector of surrounding block G. The above-described peripheral block A may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block B may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the above-described peripheral block D may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the above-described peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0247] The encoding device / decoding device can check whether the motion vectors of the neighboring blocks in the first group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the neighboring block that satisfies the condition that is first confirmed in the checking process as the mv0. That is, the mv0 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, an available mv0 may not exist. Here, for example, the specific order may be an order from neighboring block A to neighboring block B to neighboring block C in the first group. In addition, for example, the specific condition may be that a reference picture for a motion vector of a neighboring block is the same as a reference picture of the current block.

[0248] In addition, the encoding device / decoding device can check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the neighboring block that satisfies the condition that is first confirmed in the checking process as the mv1. That is, the mv1 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, an available mv1 may not exist. Here, for example, the specific order may be an order from the neighboring block D to the neighboring block E in the second group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0249] The above mv0 for the current block and if the above mv1 is available, i.e., the above mv0 for the current block And if the above mv1 is derived, the encoding device / decoding device derives the above mv0 and mv1 can be derived as constructed candidates of the current block (S1810). Meanwhile, if mv0 and / or mv1 for the current block are not available, i.e., mv0 and mv1 are derived from neighboring blocks of the current block. If at least one of them is not derived, the encoding device / decoding device may not add the constructed candidate to the affine MVP list of the current block.

[0250] Meanwhile, a pruning check process may not be performed between the derived inherited affine candidates and the constructed affine candidates. The pruning check process may refer to a process of checking whether the candidates are identical and, if so, eliminating the candidate derived later.

[0251] Figure 19 shows an example of deriving the constructed candidate when a 6-affine motion model is applied to the current block.

[0252] Referring to FIG. 19, the encoding device / decoding device can determine whether mv0, mv1, and mv2 are available for the current block (S1900). That is, the encoding device / decoding device can determine whether mv0, mv1, and mv2 are available in the surrounding blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0253] The encoding device / decoding device can determine whether there is an available mv0 in the first group, whether there is an available mv1 in the second group, and whether there is an available mv2 in the third group.

[0254] Specifically, the surrounding blocks of the current block can be divided into three groups, and the surrounding blocks can include surrounding block A, surrounding block B, surrounding block C, surrounding block D, surrounding block E, surrounding block F, and surrounding block G. The first group can include a motion vector of surrounding block A, a motion vector of surrounding block B, and a motion vector of surrounding block C, the second group can include a motion vector of surrounding block D, a motion vector of surrounding block E, and the third group can include a motion vector of surrounding block F, a motion vector of surrounding block G. The above-described peripheral block A may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block B may represent a peripheral block located at the upper left of the upper left sample position of the current block, the above-described peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the above-described peripheral block D may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the above-described peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the above-described peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0255] The encoding device / decoding device can check whether the motion vectors of the neighboring blocks in the first group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the neighboring block that satisfies the condition that is first confirmed in the checking process as the mv0. That is, the mv0 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, an available mv0 may not exist. Here, for example, the specific order may be an order from neighboring block A to neighboring block B to neighboring block C in the first group. In addition, for example, the specific condition may be that a reference picture for a motion vector of a neighboring block is the same as a reference picture of the current block.

[0256] In addition, the encoding device / decoding device can check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the neighboring block that satisfies the condition that is first confirmed in the checking process as the mv1. That is, the mv1 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, an available mv1 may not exist. Here, for example, the specific order may be an order from the neighboring block D to the neighboring block E in the second group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0257] In addition, the encoding device / decoding device can check whether the motion vectors of the surrounding blocks in the third group satisfy a specific condition in a specific order. The encoding device / decoding device can derive the motion vector of the surrounding block that satisfies the condition that is first confirmed in the checking process as mv2. That is, the mv2 may be a motion vector that satisfies the specific condition that is first confirmed by checking the motion vectors in the third group in a specific order. If the motion vectors of the surrounding blocks in the third group do not satisfy the specific condition, an available mv2 may not exist. Here, for example, the specific order may be an order from the surrounding block F to the surrounding block G in the third group. In addition, for example, the specific condition may be that the reference picture for the motion vector of the surrounding block is the same as the reference picture of the current block.

[0258] The above mv0, the above mv1 for the above current block and if the above mv2 is available, i.e., the above mv0, the above mv1 for the current block And when the above mv2 is derived, the encoding device / decoding device derives the above mv0, mv1 and mv2 can be derived as constructed candidates of the current block (S1910). Meanwhile, if mv0, mv1 and / or mv2 for the current block are not available, i.e., mv0, mv1 and mv2 are available from the surrounding blocks of the current block. If at least one of them is not derived, the encoding device / decoding device may not add the constructed candidate to the affine MVP list of the current block.

[0259] Meanwhile, the pruning check process between the derived inherited affine candidate and the above-mentioned constructed affine candidate may not be performed.

[0260] Meanwhile, if the number of derived affine candidates is less than 2 (i.e., the number of inherited affine candidates and / or constructed affine candidates is less than 2), the HEVC AMVP candidate may be added to the affine MVP list of the current block.

[0261] For example, the HEVC AMVP candidate can be derived in the following order.

[0262] Specifically, if the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is available, the CPMV0 can be used as the affine MVP candidate. That is, if the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is available (i.e., the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is derived), a first affine MVP candidate including the CPMV0 of the constructed affine candidate as CPMV0, CPMV1, and CPMV2 can be derived.

[0263] In addition, next, if the number of derived affine candidates is less than 2 and the CPMV1 of the constructed affine candidate is available, the CPMV1 can be used as the affine MVP candidate. That is, if the number of derived affine candidates is less than 2 and the CPMV1 of the constructed affine candidate is available (i.e., the number of derived affine candidates is less than 2 and the CPMV1 of the constructed affine candidate is derived), a second affine MVP candidate including the CPMV1 of the constructed affine candidate as CPMV0, CPMV1, and CPMV2 can be derived.

[0264] In addition, next, if the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is available, the CPMV2 can be used as the affine MVP candidate. That is, if the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is available (i.e., the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is derived), a third affine MVP candidate including the CPMV2 of the constructed affine candidate as CPMV0, CPMV1, and CPMV2 can be derived.

[0265] In addition, next, when the number of derived affine candidates is less than 2, HEVC TMVP (Temporal Motion vector predictor) can be used as the affine MVP candidate. The HEVC TMVP can be derived based on motion information of temporal neighboring blocks of the current block. That is, when the number of derived affine candidates is less than 2, a third affine MVP candidate including motion vectors of temporal neighboring blocks of the current block as CPMV0, CPMV1, and CPMV2 can be derived. The temporal neighboring blocks can represent collocated blocks in a collocated picture corresponding to the current block.

[0266] In addition, next, if the number of derived affine candidates is less than 2, a zero motion vector (zero MV) can be used as the affine MVP candidate. That is, if the number of derived affine candidates is less than 2, a third affine MVP candidate including the zero motion vector as CPMV0, CPMV1, and CPMV2 can be derived. The zero motion vector can represent a motion vector whose value is 0.

[0267] This can reduce complexity compared to the existing method of deriving HEVC AMVP candidates because the steps of using the CPMV of the constructed affine candidate reuse the MV already considered for generating the constructed affine candidate.

[0268] Meanwhile, this document proposes another embodiment of deriving the above-described inherited affine candidate.

[0269] In order to derive the above-mentioned inherited affine candidates, affine prediction information of surrounding blocks is required, and specifically, the following affine prediction information is required.

[0270] 1) Affine flag (affine_flag) indicating whether affine prediction-based encoding of the above surrounding blocks is applied.

[0271] 2) Movement information of the surrounding blocks

[0272] When a 4-affine motion model is applied to the surrounding block, the motion information of the surrounding block may include L0 motion information and L1 motion information for CP0, and L0 motion information and L1 motion information for CP1. In addition, when a 6-affine motion model is applied to the surrounding block, the motion information of the surrounding block may include L0 motion information and L1 motion information for CP0, and L0 motion information and L1 motion information for CP2. Here, the L0 motion information may indicate motion information for L0 (List 0), and the L1 motion information may indicate motion information for L1 (List 1). The L0 motion information may include an L0 reference picture index and an L0 motion vector, and the L1 motion information may include an L1 reference picture index and an L1 motion vector.

[0273] As described above, affine prediction requires a large amount of information to be stored, which can be a major factor in increasing hardware costs in actual implementations in encoding / decoding devices. In particular, when a neighboring block is located above the current block and is a CTU boundary, a line buffer must be used to store affine prediction-related information for the neighboring block, which can lead to even greater cost issues. This problem may be referred to as the "line buffer issue." Therefore, this document proposes an embodiment that minimizes hardware costs by eliminating or reducing the storage of affine prediction-related information in the line buffer, thereby deriving inherited affine candidates. The proposed embodiment can improve coding performance by reducing the computational complexity in deriving the inherited affine candidates. Meanwhile, for reference, the line buffer already stores motion information for 4x4 blocks, and if the affine prediction-related information is additionally stored, the amount of information stored can increase by three times compared to the existing storage amount.

[0274] In this embodiment, no additional information about affine prediction may be stored in the line buffer, and if information in the line buffer must be referenced for generation of the inherited affine candidate, generation of the inherited affine candidate may be restricted.

[0275] Figures 20a and 20b illustrate examples of deriving the inherited affine candidates.

[0276] Referring to FIG. 20a, if the neighboring block B of the current block (i.e., the upper neighboring block of the current block) does not exist in the same CTU as the current block (i.e., the current CTU), the neighboring block B may not be used to generate the inherited affine candidate. On the other hand, although the neighboring block A does not exist in the same CTU as the current block, the information about the neighboring block A is not stored in the line buffer, and thus may be used to generate the inherited affine candidate. Therefore, in the present embodiment, the upper neighboring block of the current block may be used to derive the inherited affine candidate only when it is included in the same CTU as the current block. In addition, if the upper neighboring block of the current block is not included in the same CTU as the current block, the upper neighboring block may not be used to derive the inherited affine candidate.

[0277] Referring to FIG. 20b, a neighboring block B of the current block (i.e., an upper neighboring block of the current block) may exist in the same CTU as the current block. In this case, the encoding device / decoding device may generate the inherited affine candidate by referring to the neighboring block B.

[0278] Fig. 21 schematically illustrates a video encoding method by an encoding device according to the present document. The method disclosed in Fig. 21 can be performed by the encoding device disclosed in Fig. 2. Specifically, for example, S2100 to S2120 of Fig. 21 can be performed by a prediction unit of the encoding device, S2130 can be performed by a subtraction unit of the encoding device, and S2140 can be performed by an entropy encoding unit of the encoding device. In addition, although not illustrated, a process of deriving prediction samples for the current block based on the CPMVs can be performed by a prediction unit of the encoding device, a process of deriving a residual sample for the current block based on the original sample and the prediction sample for the current block can be performed by a subtraction unit of the encoding device, a process of generating information about the residual for the current block based on the residual sample can be performed by a transformation unit of the encoding device, and a process of encoding information about the residual can be performed by an entropy encoding unit of the encoding device.

[0279] The encoding device constructs an affine motion vector predictor (MVP) candidate list for the current block (S2100). The encoding device can construct an affine MVP candidate list that includes affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list can be 2.

[0280] In addition, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The encoding device may check whether the inherited affine MVP candidate of the current block is available, and if the inherited affine MVP candidate is available, the inherited affine MVP candidate may be derived. For example, the inherited affine MVP candidates may be derived based on neighboring blocks of the current block, and the maximum number of the inherited affine MVP candidates may be 2. The neighboring blocks may be checked for availability in a specific order, and the inherited affine MVP candidate may be derived based on the checked available neighboring blocks. That is, the surrounding blocks can be checked for availability in a specific order, and a first inherited affine MVP candidate can be derived based on the first checked available surrounding block, and a second inherited affine MVP candidate can be derived based on the second checked available surrounding block. The availability can be coded with an affine motion model, and the reference picture of the surrounding block can indicate that it is the same as the reference picture of the current block. That is, the available surrounding blocks can be coded with an affine motion model (i.e., affine prediction is applied), and the reference picture can be a surrounding block that is the same as the reference picture of the current block. Specifically, the encoding device can derive motion vectors for CPs of the current block based on the affine motion model of the first checked available surrounding block, and can derive the first inherited affine MVP candidate that includes the motion vectors as CPMVP candidates. Additionally, the encoding device can derive motion vectors for CPs of the current block based on the affine motion model of the second checked available surrounding block, and can derive the second inherited affine MVP candidate including the motion vectors as CPMVP candidates.The above affine motion model can be derived as in the above mathematical expression 1 or mathematical expression 3.

[0281] In addition, in other words, the neighboring blocks can be checked in a specific order to see whether they satisfy a specific condition, and the inherited affine MVP candidate can be derived based on the neighboring blocks that satisfy the checked specific condition. That is, the neighboring blocks can be checked in a specific order to see whether they satisfy the specific condition, and a first inherited affine MVP candidate can be derived based on the neighboring block that satisfies the specific condition that is checked for the first time, and a second inherited affine MVP candidate can be derived based on the neighboring block that satisfies the specific condition that is checked for the second time. Specifically, the encoding device can derive motion vectors for CPs of the current block based on the affine motion model of the neighboring block that satisfies the specific condition that is checked for the first time, and can derive the first inherited affine MVP candidate that includes the motion vectors as CPMVP candidates. In addition, the encoding device can derive motion vectors for CPs of the current block based on the affine motion model of the surrounding block that satisfies the specific condition checked for the second time, and can derive the second inherited affine MVP candidate that includes the motion vectors as CPMVP candidates. The affine motion model can be derived as in the above-described mathematical expression 1 or mathematical expression 3. Meanwhile, the specific condition can indicate that the affine motion model is coded, and the reference picture of the surrounding block is the same as the reference picture of the current block. That is, the surrounding block that satisfies the specific condition is coded with the affine motion model (i.e., affine prediction is applied), and the reference picture can be a surrounding block that is the same as the reference picture of the current block.

[0282] Here, for example, the surrounding blocks may include a left surrounding block, an upper surrounding block, an upper right corner surrounding block, a lower left corner surrounding block, and an upper left corner surrounding block of the current block. In this case, the specific order may be from the left surrounding block to the lower left corner surrounding block, the upper surrounding block, the upper right corner surrounding block, and the upper left corner surrounding block.

[0283] Alternatively, for example, the peripheral blocks may include only the left peripheral block and the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block.

[0284] Alternatively, for example, the peripheral blocks may include the left peripheral block, and if the upper peripheral block is included in a current CTU that includes the current block, the peripheral blocks may further include the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block. Furthermore, if the upper peripheral block is not included in the current CTU, the peripheral blocks may not include the upper peripheral block. In this case, only the left peripheral block may be checked.

[0285] Meanwhile, when the size is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the lower-left corner surrounding block may be a block including a sample with coordinates (-1, H), the left surrounding block may be a block including a sample with coordinates (-1, H-1), the upper-right corner surrounding block may be a block including a sample with coordinates (W, -1), the upper surrounding block may be a block including a sample with coordinates (W-1, -1), and the upper-left corner surrounding block may be a block including a sample with coordinates (-1, -1). That is, the left surrounding block may be a left surrounding block located at the lowermost position among the left surrounding blocks of the current block, and the upper surrounding block may be an upper surrounding block located at the leftmost position among the upper surrounding blocks of the current block.

[0286] In addition, for example, if a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. The encoding device may check whether the constructed affine MVP candidate of the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate may be derived. In addition, for example, the constructed affine MVP candidate may be derived after the inherited affine MVP candidate is derived. If the number of derived affine MPV candidates (i.e., the inherited affine MVP candidates) is less than two and the constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. Here, the constructed affine MVP candidate may include candidate motion vectors for the CPs. The constructed affine MVP candidate may be available when all of the candidate motion vectors are available.

[0287] For example, when a 4 affine motion model is applied to the current block, the CPs of the current block may include CP0 and CP1. When a candidate motion vector for CP0 is available and a candidate motion vector for CP1 is available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, CP0 may indicate an upper left position of the current block, and CP1 may indicate an upper right position of the current block.

[0288] The above-mentioned constructed affine MVP candidate may include a candidate motion vector for CP0 and a candidate motion vector for CP1. The candidate motion vector for CP0 may be a motion vector of a first block, and the candidate motion vector for CP1 may be a motion vector of a second block.

[0289] In addition, the first block may check the neighboring blocks within the first group according to a first specific order, and the first identified reference picture may be the same block as the reference picture of the current block. That is, the candidate motion vector for CP1 may be the motion vector of the block whose first identified reference picture is the same as the reference picture of the current block by checking the neighboring blocks within the first group according to a first specific order. The availability may indicate that the neighboring block exists and that the neighboring block is coded with inter prediction. Here, when the reference picture of the first block within the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. In addition, for example, the first group may include neighboring block A, neighboring block B, and neighboring block C, and the first specific order may be the order from neighboring block A to neighboring block B to neighboring block C.

[0290] In addition, the second block may check the neighboring blocks within the second group according to a second specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the second block within the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. In addition, for example, the second group may include neighboring blocks D and E, and the second specific order may be from the neighboring blocks D to E.

[0291] Meanwhile, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample of coordinates (-1, -1), the surrounding block B may be a block including a sample of coordinates (0, -1), the surrounding block C may be a block including a sample of coordinates (-1, 0), the surrounding block D may be a block including a sample of coordinates (W-1, -1), and the surrounding block E may be a block including a sample of coordinates (W, -1). That is, the peripheral block A may be a block surrounding the upper left corner of the current block, the peripheral block B may be an upper peripheral block located at the leftmost position among the upper peripheral blocks of the current block, the peripheral block C may be a left peripheral block located at the uppermost position among the left peripheral blocks of the current block, the peripheral block D may be an upper peripheral block located at the rightmost position among the upper peripheral blocks of the current block, and the peripheral block E may be a block surrounding the upper right corner of the current block.

[0292] Meanwhile, if at least one of the candidate motion vectors of CP0 and CP1 is not available, the constructed affine MVP candidate may not be available.

[0293] Alternatively, for example, when a 6 affine motion model is applied to the current block, the CPs of the current block may include CP0, CP1, and CP2. When a candidate motion vector for CP0 is available, a candidate motion vector for CP1 is available, and a candidate motion vector for CP2 is available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, the CP0 may represent an upper-left position of the current block, the CP1 may represent an upper-right position of the current block, and the CP2 may represent a lower-left position of the current block.

[0294] The above-described constructed affine MVP candidate may include a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2. The candidate motion vector for CP0 may be a motion vector of a first block, the candidate motion vector for CP1 may be a motion vector of a second block, and the candidate motion vector for CP2 may be a motion vector of a third block.

[0295] In addition, the first block may check the neighboring blocks within the first group according to a first specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the first block within the first group is the same as the reference picture of the current block, a candidate motion vector for CP0 may be available. In addition, for example, the first group may include neighboring block A, neighboring block B, and neighboring block C, and the first specific order may be an order from neighboring block A to neighboring block B to neighboring block C.

[0296] In addition, the second block may check the neighboring blocks within the second group according to a second specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the second block within the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. In addition, for example, the second group may include neighboring blocks D and E, and the second specific order may be from the neighboring blocks D to E.

[0297] In addition, the third block may check the neighboring blocks within the third group according to a third specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the third block within the third group is the same as the reference picture of the current block, a candidate motion vector for CP2 may be available. In addition, for example, the third group may include neighboring blocks F and G, and the third specific order may be an order from the neighboring block F to the neighboring block G.

[0298] Meanwhile, when the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample of coordinates (-1, -1), the surrounding block B may be a block including a sample of coordinates (0, -1), the surrounding block C may be a block including a sample of coordinates (-1, 0), the surrounding block D may be a block including a sample of coordinates (W-1, -1), the surrounding block E may be a block including a sample of coordinates (W, -1), the surrounding block F may be a block including a sample of coordinates (-1, H-1), and the surrounding block G may be a block including a sample of coordinates (-1, H). That is, the surrounding block A may be a block surrounding the upper left corner of the current block, the surrounding block B may be a block located at the leftmost side among the upper surrounding blocks of the current block, the surrounding block C may be a block located at the uppermost side among the left surrounding blocks of the current block, the surrounding block D may be a block located at the rightmost side among the upper surrounding blocks of the current block, the surrounding block E may be a block surrounding the upper right corner of the current block, the surrounding block F may be a block located at the lowermost side among the left surrounding blocks of the current block, and the surrounding block G may be a block surrounding the lower left corner of the current block.

[0299] Meanwhile, if at least one of the candidate motion vector of CP0, the candidate motion vector of CP1, and the candidate motion vector of CP2 is not available, the constructed affine MVP candidate may not be available.

[0300] Thereafter, the above-mentioned Affine MVP candidate list can be derived based on the steps in the sequence described below.

[0301] For example, if the number of derived affine MVP candidates is less than two and a motion vector for CP0 is available, the encoding device can derive a first affine MVP candidate. Here, the first affine MVP candidate may be an affine MVP candidate that includes a motion vector for CP0 as candidate motion vectors for the CPs.

[0302] Additionally, for example, if the number of derived affine MVP candidates is less than two and a motion vector for CP1 is available, the encoding device may derive a second affine MVP candidate. Here, the second affine MVP candidate may be an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs.

[0303] Additionally, for example, if the number of derived affine MVP candidates is less than two and a motion vector for CP2 is available, the encoding device may derive a third affine MVP candidate. Here, the third affine MVP candidate may be an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs.

[0304] In addition, for example, when the number of derived affine MVP candidates is less than two, the encoding device can derive a fourth affine MVP candidate that includes a temporal MVP derived based on a temporal neighboring block of the current block as candidate motion vectors for the CPs. The temporal neighboring block can represent a collocated block in a collocated picture corresponding to the current block. The temporal MVP can be derived based on the motion vector of the temporal neighboring block.

[0305] Additionally, for example, if the number of derived affine MVP candidates is less than two, the encoding device may derive a fifth affine MVP candidate that includes a zero motion vector as candidate motion vectors for the CPs. The zero motion vector may represent a motion vector with a value of 0.

[0306] The encoding device derives CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the affine MVP candidate list (S2110). The encoding device can derive CPMVs for the CPs of the current block having optimal RD costs, and select an affine MVP candidate most similar to the CPMVs among the affine MVP candidates as the affine MVP candidate for the current block. The encoding device can derive CPMVPs (Control Point Motion Vector Predictors) for the CPs of the current block based on the selected affine MVP candidate among the affine MVP candidates included in the affine MVP candidate list. Specifically, when an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, and the candidate motion vector for CP1 of the affine MVP candidate can be derived as the CPMVP of CP1. In addition, when an affine MVP candidate includes a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived as the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2.In addition, if the affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2.

[0307] The encoding device can encode an affine MVP candidate index that points to the selected affine MVP candidate among the affine MVP candidates. The affine MVP candidate index can point to one affine MVP candidate among the affine MVP candidates included in the affine motion vector predictor (MVP) candidate list for the current block.

[0308] The encoding device derives CPMVs for the CPs of the current block (S2120). The encoding device can derive CPMVs for each of the CPs of the current block.

[0309] The encoding device derives CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the CPMVPs and CPMVs (S2130). The encoding device can derive CPMVDs for the CPs of the current block based on the CPMVPs and CPMVs for each of the CPs.

[0310] The encoding device encodes motion prediction information including information about the CPMVDs (S2140). The encoding device can output the motion prediction information including information about the CPMVDs in the form of a bitstream. That is, the encoding device can output image information including the motion prediction information in the form of a bitstream. The encoding device can encode information about the CPMVD for each of the CPs, and the motion prediction information can include information about the CPMVDs.

[0311] Additionally, the motion prediction information may include the affine MVP candidate index. The affine MVP candidate index may point to the selected affine MVP candidate among the affine MVP candidates included in the affine motion vector predictor (MVP) candidate list for the current block.

[0312] Meanwhile, as an example, the encoding device can derive prediction samples for the current block based on the CPMVs, derive a residual sample for the current block based on the original sample and the prediction sample for the current block, generate information about the residual for the current block based on the residual sample, and encode the information about the residual. The image information can include information about the residual.

[0313] Meanwhile, the bitstream may be transmitted to a decoding device via a network or (digital) storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0314] Fig. 22 schematically illustrates an encoding device that performs an image encoding method according to the present document. The method disclosed in Fig. 21 can be performed by the encoding device disclosed in Fig. 22. Specifically, for example, the prediction unit of the encoding device of Fig. 22 can perform S2100 to S2130 of Fig. 21, and the entropy encoding unit of the encoding device of Fig. 22 can perform S2140 of Fig. 21. In addition, although not shown, the process of deriving prediction samples for the current block based on the CPMVs can be performed by the prediction unit of the encoding device of FIG. 22, the process of deriving a residual sample for the current block based on the original sample and the prediction sample for the current block can be performed by the subtraction unit of the encoding device of FIG. 22, the process of generating information about the residual for the current block based on the residual sample can be performed by the conversion unit of the encoding device of FIG. 22, and the process of encoding information about the residual can be performed by the entropy encoding unit of the encoding device of FIG. 22.

[0315] Fig. 23 schematically illustrates a video decoding method by a decoding device according to the present document. The method disclosed in Fig. 23 can be performed by the decoding device disclosed in Fig. 3. Specifically, for example, S2300 of Fig. 23 can be performed by an entropy decoding unit of the decoding device, S2310 to S2350 can be performed by a prediction unit of the decoding device, and S2360 can be performed by an addition unit of the decoding device. In addition, although not illustrated, a process of obtaining information about the residual of the current block through a bitstream can be performed by an entropy decoding unit of the decoding device, and a process of deriving the residual sample for the current block based on the residual information can be performed by an inverse transform unit of the decoding device.

[0316] The decoding device obtains motion prediction information for the current block from the bitstream (S2300). The decoding device can obtain image information including the motion prediction information from the bitstream.

[0317] Additionally, for example, the motion prediction information may include information about CPMVDs (Control Point Motion Vector Differences) for CPs (control points) of the current block. That is, the motion prediction information may include information about CPMVDs for each of the CPs of the current block.

[0318] Additionally, for example, the motion prediction information may include an affine MVP candidate index for the current block. The affine MVP candidate index may point to one of the affine MVP candidates included in the affine motion vector predictor (MVP) candidate list for the current block.

[0319] The decoding device constructs an affine motion vector predictor (MVP) candidate list for the current block (S2310). The decoding device may construct an affine MVP candidate list that includes affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list may be 2.

[0320] In addition, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The decoding device may check whether the inherited affine MVP candidate of the current block is available, and if the inherited affine MVP candidate is available, the inherited affine MVP candidate may be derived. For example, the inherited affine MVP candidates may be derived based on neighboring blocks of the current block, and the maximum number of the inherited affine MVP candidates may be 2. The neighboring blocks may be checked for availability in a specific order, and the inherited affine MVP candidate may be derived based on the checked available neighboring blocks. That is, the surrounding blocks can be checked for availability in a specific order, and a first inherited affine MVP candidate can be derived based on the first checked available surrounding block, and a second inherited affine MVP candidate can be derived based on the second checked available surrounding block. The availability can be coded with an affine motion model, and the reference picture of the surrounding block can indicate that it is the same as the reference picture of the current block. That is, the available surrounding blocks can be coded with an affine motion model (i.e., affine prediction is applied), and the reference picture can be a surrounding block that is the same as the reference picture of the current block. Specifically, the decoding device can derive motion vectors for CPs of the current block based on the affine motion model of the first checked available surrounding block, and can derive the first inherited affine MVP candidate that includes the motion vectors as CPMVP candidates. Additionally, the decoding device can derive motion vectors for CPs of the current block based on the affine motion model of the second checked available surrounding block, and can derive the second inherited affine MVP candidate including the motion vectors as CPMVP candidates.The above affine motion model can be derived as in the above mathematical expression 1 or mathematical expression 3.

[0321] In addition, in other words, the neighboring blocks can be checked in a specific order to see whether they satisfy a specific condition, and the inherited affine MVP candidate can be derived based on the neighboring blocks that satisfy the checked specific condition. That is, the neighboring blocks can be checked in a specific order to see whether they satisfy the specific condition, and a first inherited affine MVP candidate can be derived based on the neighboring block that satisfies the specific condition that is checked for the first time, and a second inherited affine MVP candidate can be derived based on the neighboring block that satisfies the specific condition that is checked for the second time. Specifically, the decoding device can derive motion vectors for CPs of the current block based on the affine motion model of the neighboring block that satisfies the specific condition that is checked for the first time, and can derive the first inherited affine MVP candidate that includes the motion vectors as CPMVP candidates. In addition, the decoding device can derive motion vectors for CPs of the current block based on the affine motion model of the surrounding block that satisfies the specific condition checked for the second time, and can derive the second inherited affine MVP candidate that includes the motion vectors as CPMVP candidates. The affine motion model can be derived as in the above-described mathematical expression 1 or mathematical expression 3. Meanwhile, the specific condition can indicate that the affine motion model is coded, and the reference picture of the surrounding block is the same as the reference picture of the current block. That is, the surrounding block that satisfies the specific condition is coded with the affine motion model (i.e., affine prediction is applied), and the reference picture can be a surrounding block that is the same as the reference picture of the current block.

[0322] Here, for example, the surrounding blocks may include a left surrounding block, an upper surrounding block, an upper right corner surrounding block, a lower left corner surrounding block, and an upper left corner surrounding block of the current block. In this case, the specific order may be from the left surrounding block to the lower left corner surrounding block, the upper surrounding block, the upper right corner surrounding block, and the upper left corner surrounding block.

[0323] Alternatively, for example, the peripheral blocks may include only the left peripheral block and the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block.

[0324] Alternatively, for example, the peripheral blocks may include the left peripheral block, and if the upper peripheral block is included in a current CTU including the current block, the peripheral blocks may further include the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block. Additionally, if the upper peripheral block is not included in the current CTU, the peripheral blocks may not include the upper peripheral block. In this case, only the left peripheral block may be checked. That is, if the upper peripheral block of the current block is included in a current CTU (Coding Tree Unit) including the current block, the upper peripheral block may be used to derive the inherited affine MVP candidate, and if the upper peripheral block of the current block is not included in the current CTU, the upper peripheral block may not be used to derive the inherited affine MVP candidate.

[0325] Meanwhile, when the size is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the lower-left corner surrounding block may be a block including a sample with coordinates (-1, H), the left surrounding block may be a block including a sample with coordinates (-1, H-1), the upper-right corner surrounding block may be a block including a sample with coordinates (W, -1), the upper surrounding block may be a block including a sample with coordinates (W-1, -1), and the upper-left corner surrounding block may be a block including a sample with coordinates (-1, -1). That is, the left surrounding block may be a left surrounding block located at the lowermost position among the left surrounding blocks of the current block, and the upper surrounding block may be an upper surrounding block located at the leftmost position among the upper surrounding blocks of the current block.

[0326] In addition, for example, if a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. The decoding device may check whether the constructed affine MVP candidate of the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate may be derived. In addition, for example, the constructed affine MVP candidate may be derived after the inherited affine MVP candidate is derived. If the number of derived affine MPV candidates (i.e., the inherited affine MVP candidates) is less than two and the constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. Here, the constructed affine MVP candidate may include candidate motion vectors for the CPs. The constructed affine MVP candidate may be available when all of the candidate motion vectors are available.

[0327] For example, when a 4 affine motion model is applied to the current block, the CPs of the current block may include CP0 and CP1. When a candidate motion vector for CP0 is available and a candidate motion vector for CP1 is available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, CP0 may indicate an upper left position of the current block, and CP1 may indicate an upper right position of the current block.

[0328] The above-mentioned constructed affine MVP candidate may include a candidate motion vector for CP0 and a candidate motion vector for CP1. The candidate motion vector for CP0 may be a motion vector of a first block, and the candidate motion vector for CP1 may be a motion vector of a second block.

[0329] In addition, the first block may check the neighboring blocks within the first group according to a first specific order, and the first identified reference picture may be the same block as the reference picture of the current block. That is, the candidate motion vector for CP1 may be the motion vector of the block whose first identified reference picture is the same as the reference picture of the current block by checking the neighboring blocks within the first group according to a first specific order. The availability may indicate that the neighboring block exists and that the neighboring block is coded with inter prediction. Here, when the reference picture of the first block within the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. In addition, for example, the first group may include neighboring block A, neighboring block B, and neighboring block C, and the first specific order may be the order from neighboring block A to neighboring block B to neighboring block C.

[0330] In addition, the second block may check the neighboring blocks within the second group according to a second specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the second block within the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. In addition, for example, the second group may include neighboring blocks D and E, and the second specific order may be from the neighboring blocks D to E.

[0331] Meanwhile, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample of coordinates (-1, -1), the surrounding block B may be a block including a sample of coordinates (0, -1), the surrounding block C may be a block including a sample of coordinates (-1, 0), the surrounding block D may be a block including a sample of coordinates (W-1, -1), and the surrounding block E may be a block including a sample of coordinates (W, -1). That is, the peripheral block A may be a block surrounding the upper left corner of the current block, the peripheral block B may be an upper peripheral block located at the leftmost position among the upper peripheral blocks of the current block, the peripheral block C may be a left peripheral block located at the uppermost position among the left peripheral blocks of the current block, the peripheral block D may be an upper peripheral block located at the rightmost position among the upper peripheral blocks of the current block, and the peripheral block E may be a block surrounding the upper right corner of the current block.

[0332] Meanwhile, if at least one of the candidate motion vectors of CP0 and CP1 is not available, the constructed affine MVP candidate may not be available.

[0333] Alternatively, for example, when a 6 affine motion model is applied to the current block, the CPs of the current block may include CP0, CP1, and CP2. When a candidate motion vector for CP0 is available, a candidate motion vector for CP1 is available, and a candidate motion vector for CP2 is available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, the CP0 may represent an upper-left position of the current block, the CP1 may represent an upper-right position of the current block, and the CP2 may represent a lower-left position of the current block.

[0334] The above-described constructed affine MVP candidate may include a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2. The candidate motion vector for CP0 may be a motion vector of a first block, the candidate motion vector for CP1 may be a motion vector of a second block, and the candidate motion vector for CP2 may be a motion vector of a third block.

[0335] In addition, the first block may check the neighboring blocks within the first group according to a first specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the first block within the first group is the same as the reference picture of the current block, a candidate motion vector for CP0 may be available. In addition, for example, the first group may include neighboring block A, neighboring block B, and neighboring block C, and the first specific order may be an order from neighboring block A to neighboring block B to neighboring block C.

[0336] In addition, the second block may check the neighboring blocks within the second group according to a second specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the second block within the second group is the same as the reference picture of the current block, a candidate motion vector for CP1 may be available. In addition, for example, the second group may include neighboring blocks D and E, and the second specific order may be from the neighboring blocks D to E.

[0337] In addition, the third block may check the neighboring blocks within the third group according to a third specific order, and the first identified reference picture may be the same block as the reference picture of the current block. Here, if the reference picture of the third block within the third group is the same as the reference picture of the current block, a candidate motion vector for CP2 may be available. In addition, for example, the third group may include neighboring blocks F and G, and the third specific order may be an order from the neighboring block F to the neighboring block G.

[0338] Meanwhile, when the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample of coordinates (-1, -1), the surrounding block B may be a block including a sample of coordinates (0, -1), the surrounding block C may be a block including a sample of coordinates (-1, 0), the surrounding block D may be a block including a sample of coordinates (W-1, -1), the surrounding block E may be a block including a sample of coordinates (W, -1), the surrounding block F may be a block including a sample of coordinates (-1, H-1), and the surrounding block G may be a block including a sample of coordinates (-1, H). That is, the surrounding block A may be a block surrounding the upper left corner of the current block, the surrounding block B may be a block located at the leftmost side among the upper surrounding blocks of the current block, the surrounding block C may be a block located at the uppermost side among the left surrounding blocks of the current block, the surrounding block D may be a block located at the rightmost side among the upper surrounding blocks of the current block, the surrounding block E may be a block surrounding the upper right corner of the current block, the surrounding block F may be a block located at the lowermost side among the left surrounding blocks of the current block, and the surrounding block G may be a block surrounding the lower left corner of the current block.

[0339] Meanwhile, if at least one of the candidate motion vector of CP0, the candidate motion vector of CP1, and the candidate motion vector of CP2 is not available, the constructed affine MVP candidate may not be available.

[0340] Meanwhile, a pruning check may not be performed between the inherited affine MVP candidate and the constructed affine MVP candidate. The pruning check may refer to a process of checking whether the constructed affine MVP candidate is identical to the inherited affine MVP candidate, and if so, not deriving the constructed affine MVP candidate.

[0341] Thereafter, the above-mentioned Affine MVP candidate list can be derived based on the steps in the sequence described below.

[0342] For example, if the number of derived affine MVP candidates is less than two and a motion vector for CP0 is available, the decoding device can derive a first affine MVP candidate. Here, the first affine MVP candidate may be an affine MVP candidate that includes a motion vector for CP0 as candidate motion vectors for the CPs.

[0343] Additionally, for example, if the number of derived affine MVP candidates is less than two and a motion vector for CP1 is available, the decoding device may derive a second affine MVP candidate. Here, the second affine MVP candidate may be an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs.

[0344] Additionally, for example, if the number of derived affine MVP candidates is less than two and a motion vector for CP2 is available, the decoding device may derive a third affine MVP candidate. Here, the third affine MVP candidate may be an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs.

[0345] In addition, for example, when the number of derived affine MVP candidates is less than two, the decoding device can derive a fourth affine MVP candidate that includes temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs. The temporal neighboring blocks can represent collocated blocks in collocated pictures corresponding to the current block. The temporal MVPs can be derived based on motion vectors of the temporal neighboring blocks.

[0346] Additionally, for example, if the number of derived affine MVP candidates is less than two, the decoding device may derive a fifth affine MVP candidate that includes a zero motion vector as candidate motion vectors for the CPs. The zero motion vector may represent a motion vector with a value of 0.

[0347] The decoding device derives CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the above-mentioned affine MVP candidate list (S2320).

[0348] The decoding device can select a specific affine MVP candidate from among the affine MVP candidates included in the affine MVP candidate list, and derive the selected affine MVP candidate as CPMVPs for the CPs of the current block. For example, the decoding device can obtain the affine MVP candidate index for the current block from the bitstream, and derive the affine MVP candidate indicated by the affine MVP candidate index from among the affine MVP candidates included in the affine MVP candidate list, as CPMVPs for the CPs of the current block. Specifically, when the affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, and the candidate motion vector for CP1 of the affine MVP candidate can be derived as the CPMVP of CP1. In addition, when an affine MVP candidate includes a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived as the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2. In addition, when an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived as the CPMVP of CP0, and the candidate motion vector for CP2 of the affine MVP candidate can be derived as the CPMVP of CP2.

[0349] The decoding device derives CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the motion prediction information (S2330). The motion prediction information may include information on the CPMVD for each of the CPs, and the decoding device may derive the CPMVD for each of the CPs of the current block based on the information on the CPMVD for each of the CPs.

[0350] The decoding device derives CPMVs (Control Point Motion Vectors) for the CPs of the current block based on the CPMVPs and CPMVDs (S2340). The decoding device can derive a CPMV for each CP based on the CPMVP and CPMVD for each CP. For example, the decoding device can derive a CPMV for each CP by adding the CPMVP and CPMVD for each CP.

[0351] The decoding device derives prediction samples for the current block based on the CPMVs (S2350). The decoding device can derive motion vectors of the current block in units of subblocks or samples based on the CPMVs. That is, the decoding device can derive motion vectors of each subblock or each sample of the current block based on the CPMVs. The motion vectors of the subblock or sample can be derived based on the above-described mathematical expression 1 or mathematical expression 3. The motion vectors can be expressed as an affine motion vector field (MVF) or a motion vector array.

[0352] The decoding device can derive prediction samples for the current block based on the motion vectors of the sub-block unit or the sample unit. The decoding device can derive a reference region within a reference picture based on the motion vectors of the sub-block unit or the sample unit, and can generate a prediction sample of the current block based on the reconstructed sample within the reference region.

[0353] The decoding device generates a reconstructed picture for the current block based on the derived prediction samples (S2360). The decoding device can generate a reconstructed picture for the current block based on the derived prediction samples. Depending on the prediction mode, the decoding device may directly use the prediction sample as a reconstructed sample, or may generate a reconstructed sample by adding a residual sample to the prediction sample. If a residual sample for the current block exists, the decoding device can obtain information about the residual for the current block from the bitstream. The information about the residual may include a transform coefficient for the residual sample. The decoding device can derive the residual sample (or residual sample array) for the current block based on the residual information. The decoding device can generate a reconstructed sample based on the prediction sample and the residual sample, and can derive a reconstructed block or a reconstructed picture based on the reconstructed sample. As described above, the decoding device may then apply an in-loop filtering procedure, such as deblocking filtering and / or SAO procedure, to the restored picture to improve subjective / objective picture quality as needed.

[0354] Fig. 24 schematically illustrates a decoding device that performs an image decoding method according to the present document. The method disclosed in Fig. 23 can be performed by the decoding device disclosed in Fig. 24. Specifically, for example, the entropy decoding unit of the decoding device of Fig. 24 can perform S2300 of Fig. 23, the prediction unit of the decoding device of Fig. 24 can perform S2310 to S2350 of Fig. 23, and the adding unit of the decoding device of Fig. 24 can perform S2360 of Fig. 23. In addition, although not illustrated, a process of obtaining image information including information about the residual of the current block through a bitstream can be performed by the entropy decoding unit of the decoding device of Fig. 24, and a process of deriving the residual sample for the current block based on the residual information can be performed by the inverse transform unit of the decoding device of Fig. 24.

[0355] According to the above-described document, the efficiency of image coding based on affine motion prediction can be improved.

[0356] In addition, according to this document, when deriving an affine MVP candidate list, the constructed affine MVP candidate can be added only when all candidate motion vectors for CPs of the constructed affine MVP candidate are available, thereby reducing the complexity of the process of deriving the constructed affine MVP candidate and the process of constructing the affine MVP candidate list and improving coding efficiency.

[0357] In addition, according to this document, in deriving an affine MVP candidate list, additional affine MVP candidates can be derived based on candidate motion vectors for CP derived in the process of deriving constructed affine MVP candidates, thereby reducing the complexity of the process of constructing an affine MVP candidate list and improving coding efficiency.

[0358] In addition, according to this document, in the process of deriving an inherited affine MVP candidate, the inherited affine MVP candidate can be derived using an upper neighboring block only when the upper neighboring block is included in the current CTU, thereby reducing the storage amount of the line buffer for affine prediction and minimizing hardware costs.

[0359] While the methods described in the above-described embodiments are described based on a flowchart as a series of steps or blocks, this document is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will appreciate that the steps depicted in the flowchart are not exclusive, and that other steps may be included, or one or more steps in the flowchart may be deleted, without affecting the scope of this document.

[0360] The embodiments described in this document may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units depicted in each drawing may be implemented and performed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information on instructions) or algorithms may be stored on a digital storage medium.

[0361] In addition, the decoding device and encoding device to which the embodiments of the present document are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a transportation terminal (e.g., a vehicle terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT video (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0362] In addition, the processing method to which the embodiments of the present document are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0363] Additionally, the embodiments of the present document can be implemented as a computer program product by program code, and the program code can be executed on a computer by the embodiments of the present document. The program code can be stored on a computer-readable carrier.

[0364] Figure 25 illustrates an exemplary structure of a content streaming system to which embodiments of this document are applied.

[0365] A content streaming system to which the embodiments of this document are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0366] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.

[0367] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiments of the present document are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0368] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.

[0369] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0370] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system may be operated as a distributed server, in which case data received from each server may be distributedly processed.

Claims

1. In a video decoding method performed by a decoding device, A step of obtaining motion prediction information for a current block from a bitstream; A step of constructing a list of affine motion vector predictor (MVP) candidates for the current block; A step of deriving CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the above affine MVP candidate list; A step of deriving CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the motion prediction information; A step of deriving CPMVs (Control Point Motion Vectors) for the CPs of the current block based on the CPMVPs and CPMVDs; A step of deriving prediction samples for the current block based on the CPMVs; and A step of generating a restoration picture for the current block based on the above-described predicted samples, The steps for forming the above-mentioned Affine MVP candidate list are: A step of checking whether an inherited affine MVP candidate of the current block is available, and if the inherited affine MVP candidate is available, deriving the inherited affine MVP candidate; A step of checking whether a constructed affine MVP candidate of the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate is derived, and the constructed affine MVP candidate includes a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; A step of deriving a first affine MVP candidate when the number of derived affine MVP candidates is less than two and a motion vector for CP0 is available, wherein the first affine MVP candidate is an affine MVP candidate that includes a motion vector for CP0 as candidate motion vectors for the CPs; A step of deriving a second affine MVP candidate when the number of derived affine MVP candidates is less than two and a motion vector for CP1 is available, wherein the second affine MVP candidate is an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs; A step of deriving a third affine MVP candidate when the number of derived affine MVP candidates is less than two and a motion vector for CP2 is available, wherein the third affine MVP candidate is an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than two, a step of deriving a fourth affine MVP candidate that includes temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs; and An image decoding method characterized by comprising a step of deriving a fifth affine MVP candidate including a zero motion vector as candidate motion vectors for the CPs when the number of derived affine MVP candidates is less than two.

2. In paragraph 1, The CP0 above represents the upper left position of the current block, the CP1 above represents the upper right position of the current block, and the CP2 above represents the lower left position of the current block. An image decoding method characterized in that the above-mentioned constructed affine MVP candidate is available when the above-mentioned candidate motion vectors are available.

3. In paragraph 2, If the reference picture of the first block in the first group is identical to the reference picture of the current block, a candidate motion vector for CP0 is available, If the reference picture of the second block in the second group is identical to the reference picture of the current block, a candidate motion vector for CP1 is available, If the reference picture of the third block in the third group is identical to the reference picture of the current block, a candidate motion vector for CP2 is available, A video decoding method, characterized in that when the candidate motion vector for the CP0 is available, the candidate motion vector for the CP1 is available, and the candidate motion vector for the CP2 is available, the affine MVP candidate list includes the constructed affine MVP candidate.

4. In paragraph 3, The above first group includes peripheral block A, peripheral block B, and peripheral block C, The second group includes peripheral block D and peripheral block E, The third group includes peripheral block F and peripheral block G, A video decoding method, characterized in that when the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A is a block including a sample of coordinates (-1, -1), the surrounding block B is a block including a sample of coordinates (0, -1), the surrounding block C is a block including a sample of coordinates (-1, 0), the surrounding block D is a block including a sample of coordinates (W-1, -1), the surrounding block E is a block including a sample of coordinates (W, -1), the surrounding block F is a block including a sample of coordinates (-1, H-1), and the surrounding block G is a block including a sample of coordinates (-1, H).

5. In paragraph 4, The first block checks the surrounding blocks within the first group according to the first specific order, and the first confirmed reference picture is the same block as the reference picture of the current block, The second block checks the surrounding blocks within the second group according to the second specific order, and the first confirmed reference picture is the same block as the reference picture of the current block. A video decoding method characterized in that the third block checks surrounding blocks within the third group according to a third specific order and the first confirmed reference picture is the same block as the reference picture of the current block.

6. In paragraph 5, The above first specific order is the order from the peripheral block A to the peripheral block B to the peripheral block C, The second specific order is the order from the peripheral block D to the peripheral block E, An image decoding method, characterized in that the third specific order is an order from the peripheral block F to the peripheral block G.

7. In paragraph 1, Checking whether the surrounding blocks of the current block are available in a specific order, An image decoding method characterized in that the inherited affine MVP candidate is derived based on the checked available surrounding blocks.

8. In paragraph 7, A video decoding method, characterized in that the above available surrounding blocks are coded with an affine motion model, and the reference picture is a surrounding block identical to the reference picture of the current block.

9. In paragraph 8, An image decoding method, characterized in that the above-mentioned surrounding blocks include a left-side surrounding block and an upper-side surrounding block of the current block.

10. In paragraph 8, The above surrounding blocks include the left surrounding blocks of the current block, An image decoding method, characterized in that when an upper peripheral block of the current block is included in a current CTU (Coding Tree Unit) including the current block, the peripheral blocks include an upper peripheral block of the current block.

11. In paragraph 9, An image decoding method, characterized in that when the peripheral blocks include the left peripheral block and the upper peripheral block, the specific order is an order from the left peripheral block to the upper peripheral block.

12. In paragraph 11, An image decoding method, characterized in that when the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the left peripheral block is a block including a sample at coordinates (-1, H-1), and the upper peripheral block is a block including a sample at coordinates (W-1, -1).

13. In paragraph 1, An image decoding method characterized in that a pruning check is not performed between the inherited affine MVP candidate and the constructed affine MVP candidate.

14. In paragraph 1, If the upper surrounding block of the current block is included in the current CTU (Coding Tree Unit) that includes the current block, the upper surrounding block is used to derive the inherited affine MVP candidate, An image decoding method, characterized in that if the upper surrounding block of the current block is not included in the current CTU, the upper surrounding block is not used for deriving the inherited affine MVP candidate.

15. In a video encoding method performed by an encoding device, A step of constructing a list of affine motion vector predictor (MVP) candidates for the current block; A step of deriving CPMVPs (Control Point Motion Vector Predictors) for CPs (Control Points) of the current block based on the above affine MVP candidate list; A step of deriving CPMVs for the CPs of the current block; A step of deriving CPMVDs (Control Point Motion Vector Differences) for the CPs of the current block based on the CPMVPs and CPMVs; and A step of encoding motion prediction information including information about the CPMVDs, The steps for forming the above-mentioned Affine MVP candidate list are: A step of checking whether an inherited affine MVP candidate of the current block is available, and if the inherited affine MVP candidate is available, deriving the inherited affine MVP candidate; A step of checking whether a constructed affine MVP candidate of the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate is derived, and the constructed affine MVP candidate includes a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; A step of deriving a first affine MVP candidate when the number of derived affine MVP candidates is less than two and a motion vector for CP0 is available, wherein the first affine MVP candidate is an affine MVP candidate that includes a motion vector for CP0 as candidate motion vectors for the CPs; A step of deriving a second affine MVP candidate when the number of derived affine MVP candidates is less than two and a motion vector for CP1 is available, wherein the second affine MVP candidate is an affine MVP candidate that includes a motion vector for CP1 as candidate motion vectors for the CPs; A step of deriving a third affine MVP candidate when the number of derived affine MVP candidates is less than two and a motion vector for CP2 is available, wherein the third affine MVP candidate is an affine MVP candidate that includes a motion vector for CP2 as candidate motion vectors for the CPs; If the number of derived affine MVP candidates is less than two, a step of deriving a fourth affine MVP candidate that includes temporal MVPs derived based on temporal neighboring blocks of the current block as candidate motion vectors for the CPs; and A video encoding method characterized by comprising a step of deriving a fifth affine MVP candidate including a zero motion vector as candidate motion vectors for the CPs when the number of derived affine MVP candidates is less than two.