Affine motion prediction-based image decoding method and apparatus using affine MVP candidate list in image coding system

The image decoding method constructs affine MVP candidate lists to enhance image coding efficiency, addressing the increased costs of high-resolution image transmission and storage by deriving CPMVP and CPMVD, thus improving compression efficiency and reducing hardware costs.

JP2025161863AActive Publication Date: 2025-10-24LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025134967
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-09-10
Filing Date
2025-08-14
Publication Date
2025-10-24
Estimated Expiration
2039-09-10

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images has led to a rise in transmission and storage costs due to the increased amount of information, necessitating the development of highly efficient image compression techniques.

Method used

An image decoding method and apparatus that constructs affine MVP candidate lists based on available candidate motion vectors, deriving Control Point Motion Vector Predictors (CPMVP) and Differences (CPMVD) to improve image coding efficiency.

Benefits of technology

This approach enhances image/video compression efficiency by reducing the complexity of deriving affine MVP candidates and minimizing hardware costs, while maintaining high-quality image transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025161863000001_ABST
    Figure 2025161863000001_ABST
Patent Text Reader

Abstract

To provide a video decoding method.SOLUTION: A method comprises the steps of: obtaining motion prediction information for a current block from a bitstream; and constructing an affine MVP candidate list for the current block. The step of constructing the affine MVP candidate list comprises, when the number of derived affine MVP candidates including an inherited affine MVP candidate and a constructed affine MVP candidate is less than 2, deriving a first affine MVP candidate, where the first affine MVP candidate is an affine MVP candidate including a specific motion vector as candidate motion vectors for CPs and the specific motion vector is an available motion vector among a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2.SELECTED DRAWING: Figure 23
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This document relates to image coding techniques, and more particularly to an image decoding method and apparatus based on affine motion prediction in an image coding system. [Background technology]

[0002] In recent years, the demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the resolution and quality of image data increases, the amount of information or bits to be transmitted increases relatively compared to existing image data. Therefore, when image data is transmitted using a medium such as an existing wired / wireless broadband line or stored using an existing storage medium, transmission costs and storage costs increase.

[0003] This necessitates the development of highly efficient image compression techniques for the effective transmission, storage, and playback of high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] The technical problem of this document is to provide a method and apparatus for increasing image coding efficiency.

[0005] Another technical problem of this document is to provide an image decoding method and apparatus that derives constructed affine MVP candidates based on surrounding blocks only when all candidate motion vectors for a CP are available, constructs an affine MVP candidate list for the current block, and performs prediction for the current block based on the constructed affine MVP candidate list.

[0006] Another technical problem of this document is to provide an image decoding method and apparatus that derives affine MVP candidates using candidate motion vectors derived in the constructed affine MVP candidate derivation process as affine MVP candidates added when the number of available inherited affine MVP candidates and constructed affine MVP candidates is less than the maximum number of candidates in the MVP candidate list, and performs prediction for the current block based on the constructed affine MVP candidate list. [Means for solving the problem]

[0007] According to an embodiment of the present document, there is provided an image decoding method performed by a decoding device, the method including the steps of obtaining motion prediction information for a current block from a bitstream, constructing an affine motion vector predictor (MVP) candidate list for the current block, deriving Control Point Motion Vector Predictors (CPMVP) for a Control Point (CP) of the current block based on the affine MVP candidate list, deriving Control Point Motion Vector Differences (CPMVD) for the CP of the current block based on the motion prediction information, and deriving Control Point Motion Vector Differences (CPMVD) for the CP of the current block based on the CPMVP and the CPMVD. and generating a reconstructed picture for the current block based on the derived predicted samples, wherein the constructing the affine MVP candidate list includes checking whether an inherited affine MVP candidate for the current block is available, and if the inherited affine MVP candidate is available, deriving the inherited affine MVP candidate; constructing a constructed affine MVP candidate list for the current block. checking whether constructed affine MVP candidates are available, and if the constructed affine MVP candidates are available, deriving the constructed affine MVP candidates, the constructed affine MVP candidates including a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; if the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, deriving a first affine MVP candidate, the first affine MVP candidate being:a step of deriving a second affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two and the motion vector for CP1 is available, the second affine MVP candidate being an affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP; a step of deriving a third affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the third affine MVP candidate being an affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two, deriving a fourth affine MVP candidate including a temporal MVP derived based on temporal neighboring blocks of the current block as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two, and a zero motion vector (Zero Motion Vector) is used if the number of derived affine MVP candidates is less than two. vector) as a candidate motion vector for the CP.

[0008] According to another embodiment of the present document, there is provided a decoding device for performing image decoding, the decoding device including an entropy decoding unit that obtains motion prediction information for a current block from a bitstream, constructs an affine motion vector predictor (MVP) candidate list for the current block, derives control point motion vector predictors (CPMVP) for a control point (CP) of the current block based on the affine MVP candidate list, derives control point motion vector differences (CPMVD) for the CP of the current block based on the motion prediction information, and calculates control point motion vector differences (CPMVD) for the CP of the current block based on the CPMVP and the CPMVD. a prediction unit that derives CPMV vectors and derives predicted samples for the current block based on the CPMV; and an adder that generates a reconstructed picture for the current block based on the derived predicted samples, wherein the affine MVP candidate list includes a step of checking whether an inherited affine MVP candidate for the current block is available, and if the inherited affine MVP candidate is available, the inherited affine MVP candidate is derived; and a constructed affine MVP candidate for the current block is checking whether the constructed affine MVP candidates are available, and if the constructed affine MVP candidates are available, deriving the constructed affine MVP candidates, the constructed affine MVP candidates including a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; if the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, deriving a first affine MVP candidate, the first affine MVP candidate being:a step of deriving a second affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two and the motion vector for CP1 is available, the second affine MVP candidate being an affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP; a step of deriving a third affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the third affine MVP candidate being an affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two, deriving a fourth affine MVP candidate including a temporal MVP derived based on temporal neighboring blocks of the current block as a candidate motion vector for the CP if the number of derived affine MVP candidates is less than two, and a zero motion vector (Zero Motion Vector) is used if the number of derived affine MVP candidates is less than two. and deriving a fifth affine MVP candidate that includes the motion vector (i.e., the first affine MVP candidate) as a candidate motion vector for the CP.

[0009] According to another embodiment of the present document, there is provided a video encoding method performed by an encoding apparatus, the method including the steps of: constructing an affine motion vector predictor (MVP) candidate list for a current block; deriving control point motion vector predictors (CPMVPs) for a control point (CP) of the current block based on the affine MVP candidate list; deriving CPMVs for the CPs of the current block; deriving control point motion vector differences (CPMVDs) for the CPs of the current block based on the CPMVPs and the CPMVs; and generating motion prediction information (motion prediction information) including information on the CPMVDs.The step of constructing the affine MVP candidate list includes the steps of: checking whether inherited affine MVP candidates of the current block are available; and deriving the inherited affine MVP candidates if the inherited affine MVP candidates are available; checking whether constructed affine MVP candidates of the current block are available; and deriving the constructed affine MVP candidates if the constructed affine MVP candidates are available, the constructed affine MVP candidates including a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; and if the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, deriving a first affine MVP candidate, the first affine MVP candidate being the motion vector for CP0. a step of deriving a second affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP, if the number of derived affine MVP candidates is less than two and the motion vector for CP1 is available, the second affine MVP candidate being an affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP; a step of deriving a third affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP, if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the third affine MVP candidate being an affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP; a step of deriving a fourth affine MVP candidate including a temporal MVP derived based on temporal neighboring blocks of the current block as a candidate motion vector for the CP, if the number of derived affine MVP candidates is less than two, andvector) as a candidate motion vector for the CP.

[0010] According to another embodiment of the present document, there is provided a video encoding device, the encoding device including: a prediction unit that constructs an affine motion vector predictor (MVP) candidate list for a current block, derives control point motion vector predictors (CPMVPs) for a control point (CP) of the current block based on the affine MVP candidate list, and derives a CPMV for the CP of the current block; a subtraction unit that derives control point motion vector differences (CPMVDs) for the CP of the current block based on the CPMVPs and the CPMVs; and motion prediction information (MPI) including information on the CPMVDs.The affine MVP candidate list includes an entropy encoding unit that encodes motion vector information (motion vector information), and the affine MVP candidate list includes a step of checking whether an inherited affine MVP candidate of the current block is available, and if the inherited affine MVP candidate is available, deriving the inherited affine MVP candidate; a step of checking whether a constructed affine MVP candidate of the current block is available, and if the constructed affine MVP candidate is available, deriving the constructed affine MVP candidate, the constructed affine MVP candidate including a candidate motion vector for CP0 of the current block, a candidate motion vector for CP1 of the current block, and a candidate motion vector for CP2 of the current block; if the number of derived affine MVP candidates is less than two and the motion vector for CP0 is available, deriving a first affine MVP candidate, the first affine MVP candidate being a motion vector for CP0. a step of deriving a second affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP, if the number of derived affine MVP candidates is less than two and the motion vector for CP1 is available, the second affine MVP candidate being an affine MVP candidate including the motion vector for CP1 as a candidate motion vector for the CP; a step of deriving a third affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP, if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the third affine MVP candidate being an affine MVP candidate including the motion vector for CP2 as a candidate motion vector for the CP; a step of deriving a fourth affine MVP candidate including a temporal MVP derived based on temporal neighboring blocks of the current block as a candidate motion vector for the CP, if the number of derived affine MVP candidates is less than two, andvector) as a candidate motion vector for the CP. [Effects of the Invention]

[0011] According to this document, it is possible to improve the overall image / video compression efficiency.

[0012] According to this document, the efficiency of image coding based on affine motion prediction can be increased.

[0013] According to this document, when deriving an affine MVP candidate list, a constructed affine MVP candidate can be added only if all candidate motion vectors for the CP of the constructed affine MVP candidate are available, thereby reducing the complexity of the process of deriving a constructed affine MVP candidate and the process of constructing an affine MVP candidate list and improving coding efficiency.

[0014] According to this document, when deriving an affine MVP candidate list, additional affine MVP candidates can be derived based on candidate motion vectors for CPs derived in the process of deriving constructed affine MVP candidates, thereby reducing the complexity of the process of constructing an affine MVP candidate list and improving coding efficiency.

[0015] According to this document, in the process of deriving an inherited affine MVP candidate, the inherited affine MVP candidate can be derived using the upper surrounding block only if the upper surrounding block is included in the current CTU, thereby reducing the amount of storage in the line buffer for affine prediction and minimizing hardware costs. [Brief explanation of the drawings]

[0016] [Figure 1] 1 illustrates an example of a video / image coding system in which embodiments of the present document may be applied; [Figure 2] 1 is a diagram illustrating an outline of the configuration of a video / image encoding device to which an embodiment of this document can be applied. [Figure 3] FIG. 1 is a diagram illustrating an outline of the configuration of a video / image decoding device to which an embodiment of the present document can be applied. [Figure 4] 10 illustrates an example of motion represented through the affine motion model. [Figure 5] 1 shows an example of the affine motion model in which motion vectors for three control points are used. [Figure 6] 1 exemplarily illustrates the affine motion model in which motion vectors for two control points are used. [Figure 7] A method for deriving a motion vector in sub-block units based on the affine motion model will now be described as an example. [Figure 8] 1 illustrates an exemplary flow diagram of an affine motion prediction method according to one embodiment of the present document. [Figure 9] FIG. 10 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to an embodiment of the present document. [Figure 10] FIG. 10 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to an embodiment of the present document. [Figure 11] 10 shows an example of affine prediction performed when a neighboring block A is selected as an affine merge candidate. [Figure 12] 10 exemplarily shows surrounding blocks for deriving the inherited affine candidates. [Figure 13] 10 shows an example of spatial candidates for the constructed affine candidates. [Figure 14] An example of constructing an affine MVP list is shown below. [Figure 15] An example of deriving the constructed candidates will be shown below. [Figure 16] An example of deriving the constructed candidates will be shown below. [Figure 17]10 exemplarily shows the neighboring block positions scanned to derive the inherited affine candidates. [Figure 18] An example of deriving the constructed candidates when four affine motion models are applied to the current block will be described below. [Figure 19] An example of deriving the constructed candidates when a 6-affine motion model is applied to the current block will be described below. [Figure 20a] 10 illustrates an exemplary embodiment for deriving the inherited affine candidates. [Figure 20b] 10 illustrates an exemplary embodiment for deriving the inherited affine candidates. [Figure 21] 1 shows an outline of an image encoding method using an encoding device according to this document. [Figure 22] 1 shows a schematic diagram of an encoding device that performs the image encoding method according to this document. [Figure 23] 1 shows an outline of an image decoding method using a decoding device according to this document. [Figure 24] 1 shows an outline of a decoding device that performs the image decoding method according to this document. [Figure 25] 1 exemplarily illustrates a structural diagram of a content streaming system to which an embodiment of the present document is applied. DETAILED DESCRIPTION OF THE INVENTION

[0017] Although this document may be modified in various ways and may have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to the specific embodiment. Common terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. A singular expression includes a plural expression unless the context clearly dictates otherwise. In this specification, terms such as "comprise" or "have" are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should not be understood as precluding the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0018] Meanwhile, each component in the drawings described in this document is illustrated independently for the convenience of explaining the different characteristic functions, etc., and does not mean that each component is implemented as separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also within the scope of this document as long as they do not deviate from the essence of this document.

[0019] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and duplicate descriptions of the same components will be omitted.

[0020] FIG. 1 shows an overview of an example of a video / image coding system in which embodiments of this document may be applied.

[0021] As shown in Figure 1, a video / image encoding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.

[0022] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.

[0023] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which the associated data is generated.

[0024] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0025] The transmitter may transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver may receive / extract the bitstream and transmit it to a decoding device.

[0026] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, prediction, etc., which correspond to the operations of the encoding device.

[0027] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.

[0028] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the versatile video coding (VVC) standard, the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding (AVS2) standard, or next generation video / image coding standards (e.g., H.267 or H.268).

[0029] In this document, various embodiments relating to video / image coding are presented, which, unless otherwise stated, may also be performed in combination with one another.

[0030] In this document, video can refer to a collection of a series of images over time. A picture generally refers to a unit that shows an image at a specific time period, and a slice / tile is a unit that constitutes part of a picture during encoding. A slice / tile can include one or more coding tree units (CTUs). A picture can be composed of one or more slices / tiles. A picture can be composed of one or more tile groups. A tile group can include one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile can be partitioned into multiple bricks, each consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan can represent a specific sequential ordering of CTUs partitioning a picture, where the CTUs can be ordered consecutively in a CTU raster scan within a brick, the bricks within a tile can be ordered consecutively in a raster scan of the bricks of the tile, and the tiles within a picture can be ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the height of the picture. A tile scan may represent a specific sequential ordering of CTUs partitioning a picture, where the CTUs may be consecutively ordered in a CTU raster scan within a tile, and tiles within a picture may be consecutively ordered in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit.A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably. For example, in this document, tile group / tile group header may be referred to as slice / slice header.

[0031] A pixel or a pel can refer to the smallest unit that makes up a picture (or an image). A "sample" can also be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luma component, or can represent only a pixel / pixel value of a chroma component.

[0032] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients (or transform coefficients) consisting of M columns and N rows.

[0033] In this document, the terms " / " and "," should be interpreted as "and / or." For example, "A / B" means "A and / or B," and "A, B" means "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." (In this document, the terms " / " and "," should be interpreted to indicate "and / or." For instance, the expression "A / B" may mean "A and / or B." Further, "A, B" may mean "A and / or B." Further, "A / B / C" may mean "at least one of A, B, and / or C." Also, "A / B / C" may mean "at least one of A, B, and / or C.")

[0034] Additionally, in this document, "or" should be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." In other words, in this document, the term "or" should be interpreted to indicate "and / or." For instance, the expression "A or B" may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted to indicate "additionally or alternatively."

[0035] 2 is a diagram illustrating an outline of the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the term "video encoding device" may include an image encoding device.

[0036] As shown in FIG. 2, the encoding device 200 may include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor 230 may further include a subtractor (231). The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image divider 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0037] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) using a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding procedure described herein may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be immediately used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of deeper depth as needed, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0038] The term "unit" can be used interchangeably with terms such as "block" or "area." In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample generally represents a pixel or pixel value, and can represent only a pixel / pixel value of the luma component or only a pixel / pixel value of the chroma component. A sample can be used as a term corresponding to one pixel or pel of a picture (or image).

[0039] The encoding apparatus 200 may subtract a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as the subtraction unit 231. The prediction unit may perform prediction on a current block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0040] The intra prediction unit 222 can predict the current block by referring to samples in the current picture. The referenced samples can be located in the neighborhood of the current block or can be located far away, depending on the prediction mode. In intra prediction, prediction modes can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, DC mode and planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0041] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidates are used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information for the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of a neighboring block as a motion vector predictor and signaling the motion vector difference.

[0042] The predictor 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use intra block copy (IBC) prediction mode or palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be seen as an example of intra coding or intra prediction. When palette mode is applied, sample values ​​within a picture may be signaled based on information about a palette table and a palette index.

[0043] The prediction signal generated by the previous prediction unit (including the inter prediction unit 221 and / or the previous intra prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transform unit 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process can be applied to square pixel blocks of the same size, or to non-square, variable-sized blocks.

[0044] The quantizer 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoder 240. The entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition to the quantized transform coefficients, the entropy encoder 240 may also encode information required for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. Information and / or syntax elements transmitted / signaled from an encoding device to a decoding device in this document may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal can be configured as internal / external elements of the encoding device 200, or the transmitter can be included in the entropy encoding unit 240.

[0045] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter predictor 221 or the intra predictor 222. When there is no residual for the current block, such as when a skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.

[0046] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.

[0047] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0048] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatch between the encoding device 100 and the decoding device, and can also improve coding efficiency.

[0049] The memory 270DPB may store modified reconstructed pictures for use as reference pictures in the inter predictor 221. The memory 270 may store motion information of blocks from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in already reconstructed pictures. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.

[0050] FIG. 3 is a diagram illustrating an outline of the configuration of a video / image decoding device to which the embodiments of this document can be applied.

[0051] As shown in FIG. 3, the decoding device 300 may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor 330 may include an inter predictor 332 and an intra predictor 331. The residual processor 320 may include a dequantizer (321) and an inverse transformer (322). Depending on the embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). The memory 360 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.

[0052] When a bitstream including video / image information is input, the decoding device 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using a processing unit applied by the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a playback device.

[0053] The decoding device 300 may receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding device may further decode pictures based on the information on the parameter sets and / or the general constraint information. Signaling / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on an encoding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration, quantized values ​​of transform coefficients related to residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbols / bins for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, information related to prediction is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual values ​​entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). Furthermore, among the information decoded by the entropy decoding unit 310, information related to filtering may be provided to the filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.

[0054] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0055] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0056] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.

[0057] The predictor 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in deriving a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / image information and signaled.

[0058] The intra prediction unit 331 can predict a current block by referring to samples in a current picture. The referenced samples can be located in the neighborhood of the current block or far away from it depending on the prediction mode. In intra prediction, prediction modes can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.

[0059] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted from the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.

[0060] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block.

[0061] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.

[0062] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during the picture decoding process.

[0063] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0064] The (modified) reconstructed picture stored in the DPB of the memory 360 may be used as a reference picture in the inter predictor 332. The memory 360 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.

[0065] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 100 can be similarly or correspondingly applied to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.

[0066] Meanwhile, in relation to inter prediction, an inter prediction method that takes image distortion into consideration has been proposed. Specifically, an affine motion model has been proposed that efficiently derives motion vectors for sub-blocks or sample points of a current block and improves the accuracy of inter prediction despite deformations such as image rotation, zoom-in, or zoom-out. That is, an affine motion model has been proposed that derives motion vectors for sub-blocks or sample points of a current block. Prediction using the affine motion model may be called affine inter prediction or affine motion prediction.

[0067] For example, the affine inter-prediction using the affine motion model can efficiently represent four motions, ie, four transformations, as described below.

[0068] 4 exemplarily illustrates motions represented by the affine motion model. As shown in FIG. 4, motions that can be represented by the affine motion model include translational motion, scaled motion, rotational motion, and sheared motion. That is, not only the translational motion in which an image (or a part thereof) moves across a plane over time as shown in FIG. 4, but also the scaled motion in which an image (or a part thereof) scales over time, the rotational motion in which an image (or a part thereof) rotates over time, and the sheared motion in which an image (or a part thereof) is transformed into a parallelogram over time can be efficiently represented through the affine inter-prediction.

[0069] The encoding / decoding device can predict the distortion type of the image based on the motion vector at the control point (CP) of the current block through the affine inter prediction, thereby improving the accuracy of the prediction and thereby improving the image compression performance. Also, since the motion vector for at least one control point of the current block can be derived using the motion vectors of the neighboring blocks of the current block, the burden of the amount of data for added additional information can be reduced and the efficiency of inter prediction can be significantly improved.

[0070] As an example of the affine inter prediction, three control points, i.e., motion information at three reference points may be required.

[0071] FIG. 5 exemplarily shows the affine motion model in which motion vectors for three control points are used.

[0072] If the top-left sample position in the current block 500 is (0,0), the (0,0), (w,0), and (0,h) sample positions can be determined as the control points as shown in Figure 5. Hereinafter, the control point at the (0,0) sample position can be represented as CP0, the control point at the (w,0) sample position as CP1, and the control point at the (0,h) sample position as CP2.

[0073] Using the above-described control points and the motion vectors for the control points, a mathematical formula for the affine motion model can be derived. The mathematical formula for the affine motion model can be expressed as follows:

[0074]

number

[0075] Here, w represents the width of the current block 500, h represents the height of the current block 500, and v represents the width of the current block 500. 0x , v 0y indicate the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively, and v 2x , v 2y indicate the x and y components of the motion vector of CP2, respectively. Also, x indicates the x component of the position of the target sample within the current block 500, y indicates the y component of the position of the target sample within the current block 500, and v x is the x-component of the motion vector of the target sample in the current block 500, v y denotes the y component of the motion vector of the target sample in the current block 500.

[0076] Since the motion vectors of CP0, CP1, and CP2 are known, a motion vector according to the sample position in the current block can be derived based on Equation 1. That is, according to the affine motion model, the motion vector v0 (v 0x , v 0y ), v1(v 1x , v 1y ), v2(v 2x , v 2y ) may be scaled to derive a motion vector of the target sample according to the target sample position. That is, according to the affine motion model, a motion vector of each sample in the current block may be derived based on the motion vector of the control point. Meanwhile, a set of motion vectors of the samples in the current block derived according to the affine motion model may be represented as an affine motion vector field (MVF).

[0077] Meanwhile, the six parameters for Equation 1 can be expressed as a, b, c, d, e, and f as in the following equation, and the equation for the affine motion model expressed by the six parameters can be as follows:

[0078]

number

[0079] Here, w represents the width of the current block 500, h represents the height of the current block 500, and v represents the width of the current block 500. 0x , v 0y indicate the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively, and v 2x , v 2yindicate the x and y components of the motion vector of CP2, respectively. Also, x indicates the x component of the position of the target sample within the current block 500, y indicates the y component of the position of the target sample within the current block 500, and v x is the x-component of the motion vector of the target sample in the current block 500, v y denotes the y component of the motion vector of the target sample in the current block 500.

[0080] The affine motion model or the affine inter-prediction using the six parameters can be referred to as a six-parameter affine motion model or AF6.

[0081] Also, as an example of the affine inter-prediction, two control points, i.e., motion information at two reference points may be required.

[0082] 6 exemplarily illustrates the affine motion model in which motion vectors for two control points are used. The affine motion model using two control points can express three types of motion, including translational motion, scale motion, and rotational motion. The affine motion model expressing three types of motion can also be referred to as a similarity affine motion model or a simplified affine motion model.

[0083] If the top-left sample position in the current block 600 is (0,0), the (0,0) and (w,0) sample positions can be determined as the control points as shown in Figure 6. Hereinafter, the control point at the (0,0) sample position can be represented as CP0, and the control point at the (w,0) sample position can be represented as CP1.

[0084] Using the above-described control points and the motion vectors for the control points, a mathematical formula for the affine motion model can be derived. The mathematical formula for the affine motion model can be expressed as follows:

[0085]

number

[0086] Here, w represents the width of the current block 600, and v 0x , v 0y indicate the x and y components of the motion vector of CP0, respectively, and v 1x , v 1y indicate the x and y components of the motion vector of CP1, respectively. Also, x indicates the x component of the position of the target sample within the current block 600, y indicates the y component of the position of the target sample within the current block 600, and v x is the x-component of the motion vector of the target sample in the current block 600, v y denotes the y component of the motion vector of the target sample in the current block 600.

[0087] Meanwhile, the four parameters for Equation 3 can be expressed as a, b, c, and d as in the following equation, and the equation for the affine motion model expressed by the four parameters can be as follows:

[0088]

number

[0089] Here, w represents the width of the current block 600, and v 0x , v 0y indicate the x and y components of the motion vector of CP0, respectively, and v 1x , v 1yindicate the x and y components of the motion vector of CP1, respectively. Also, x indicates the x component of the position of the target sample within the current block 600, y indicates the y component of the position of the target sample within the current block 600, and v x is the x-component of the motion vector of the target sample in the current block 600, v y indicates the y-component of the motion vector of the target sample in the current block 600. The affine motion model using the two control points can be expressed by four parameters a, b, c, and d as in Equation 4, and the affine motion model or the affine inter-prediction using the four parameters can be expressed as a four-parameter affine motion model or AF4. That is, according to the affine motion model, a motion vector for each sample in the current block can be derived based on the motion vector of the control point. Meanwhile, a collection of motion vectors of samples in the current block derived according to the affine motion model can be expressed as an affine motion vector field (MVF).

[0090] Meanwhile, as described above, a motion vector per sample can be derived through the affine motion model, thereby significantly improving the accuracy of inter prediction, although in this case, the complexity of the motion compensation process may be significantly increased.

[0091] Accordingly, instead of deriving a motion vector in units of samples, it is possible to restrict the motion vector to be derived in units of sub-blocks within the current block.

[0092] 7 exemplarily illustrates a method for deriving a motion vector in units of sub-blocks based on the affine motion model. Figure 7 exemplarily illustrates a case where the size of the current block is 16x16 and motion vectors are derived in units of 4x4 sub-blocks. The sub-blocks can be set to various sizes. For example, if the sub-blocks are set to an nxn size (n is a positive integer, e.g., 4), motion vectors can be derived in units of nxn sub-blocks within the current block based on the affine motion model, and various methods can be applied to derive a motion vector representing each sub-block.

[0093] For example, as shown in FIG. 7, a motion vector for each sub-block may be derived using the center or lower right sample position of each sub-block as a representative coordinate. Here, the center lower right position may refer to the sample position located at the lower right of four samples located at the center of the sub-block. For example, if n is an odd number, one sample may be located in the center of the sub-block, and in this case, the center sample position may be used to derive the motion vector for the sub-block. However, if n is an even number, four samples may be located adjacent to each other in the center of the sub-block, and in this case, the lower right sample position may be used to derive the motion vector. For example, as shown in FIG. 7, the representative coordinates for each sub-block may be derived as (2, 2), (6, 2), (10, 2), ..., (14, 14), and the encoding / decoding device may derive the motion vector for each sub-block by substituting each of the representative coordinates of the sub-block into Equation 1 or 3 above. The motion vector of a sub-block within the current block derived through the affine motion model can be expressed as an affine MVF.

[0094] Meanwhile, as an example, the size of the sub-block within the current block may be derived based on the following equation:

[0095]

number

[0096] Here, M indicates the width of the sub-block, and N indicates the height of the sub-block. 0x , v 0y indicate the x and y components of the CPMV0 of the current block, respectively, and v 0x , v 0y where MvPre denotes the x and y components of the CPMV1 of the current block, w denotes the width of the current block, h denotes the height of the current block, and MvPre denotes the motion vector fraction accuracy. For example, the motion vector fraction accuracy may be set to 1 / 16.

[0097] Meanwhile, inter prediction using the above-mentioned affine motion model, i.e., affine motion prediction, may include an affine merge mode (AF_MERGE) and an affine inter mode (AF_INTER). Here, the affine inter mode may also be referred to as an affine motion vector prediction mode (AF_MVP).

[0098] The affine merge mode is similar to the existing merge mode in that MVDs for the motion vectors of the control points are not transmitted. That is, like the existing skip / merge mode, the affine merge mode can represent an encoding / decoding method in which CPMVs for each of two or three control points from neighboring blocks of the current block are derived and prediction is performed without encoding motion vector differences (MVDs).

[0099] For example, when the AF_MRG mode is applied to the current block, MVs for CP0 and CP1 (i.e., CPMV0 and CPMV1) can be derived from neighboring blocks of the current block to which affine mode is applied. That is, CPMV0 and CPMV1 of the neighboring blocks to which the affine mode is applied can be derived as merge candidates, and CPMV0 and CPMV1 for the current block can be derived based on the merge candidates. An affine motion model can be derived based on CPMV0 and CPMV1 of the neighboring blocks represented by the merge candidates, and the CPMV0 and CPMV1 for the current block can be derived based on the affine motion model.

[0100] The affine inter mode may represent inter prediction in which an MVP (Motion Vector Predictor) for the motion vector of the control point is derived, the motion vector of the control point is derived based on a received MVD (motion vector difference) and the MVP, an affine MVF of the current block is derived based on the motion vector of the control point, and prediction is performed based on the affine MVF. Here, the motion vector of the control point may be expressed as a CPMV (Control Point Motion Vector), the MVP of the control point may be expressed as a CPMVP (Control Point Motion Vector Predictor), and the MVD of the control point may be expressed as a CPMVD (Control Point Motion Vector Difference). Specifically, for example, an encoding device may derive a CPMVP (Control Point Motion Vector Predictor) and a CPMV (Control Point Motion Vector) for each of CP0 and CP1 (or CP0, CP1, and CP2), and may transmit or store information about the CPMVP and / or CPMVD, which is a difference between the CPMVP and CPMV.

[0101] Here, when the affine inter mode is applied to the current block, the encoding device / decoding device can construct an affine MVP candidate list based on the neighboring blocks of the current block, and the affine MVP candidates can be referred to as CPMVP pair candidates, and the affine MVP candidate list can also be referred to as a CPMVP candidate list.

[0102] In addition, each affine MVP candidate can mean a combination of CPMVPs of CP0 and CP1 in a four parameter affine motion model, and can mean a combination of CPMVPs of CP0, CP1, and CP2 in a six parameter affine motion model.

[0103] FIG. 8 exemplarily illustrates a flow diagram of an affine motion prediction method according to one embodiment of the present document.

[0104] As shown in Figure 8, the affine motion prediction method can be roughly divided into the following: When the affine motion prediction method starts, first, a CPMV pair can be obtained (S800). Here, when a four-parameter affine model is used, the CPMV pair can include CPMV0 and CPMV1.

[0105] Affine motion compensation may then be performed based on the CPMV pair (S810), and affine motion prediction may be completed.

[0106] Furthermore, two affine prediction modes may exist for determining the CPMV0 and CPMV1. Here, the two affine prediction modes may include an affine inter mode and an affine merge mode. The affine inter mode can explicitly determine the CPMV0 and CPMV1 by signaling two motion vector difference (MVD) information for the CPMV0 and CPMV1. In contrast, the affine merge mode can derive a CPMV pair without signaling MVD information.

[0107] In other words, the affine merge mode can derive the CPMV of the current block using the CPMV of surrounding blocks coded in affine mode, and if the motion vector is determined on a sub-block basis, the affine merge mode can also be called the sub-block merge mode.

[0108] In the affine merge mode, the encoding device may signal to the decoding device indexes of neighboring blocks coded in the affine mode for deriving the CPMV of the current block, and may further signal a difference value between the CPMV of the neighboring blocks and the CPMV of the current block. Here, the affine merge mode may construct an affine merge candidate list based on neighboring blocks, and the indexes of the neighboring blocks may represent neighboring blocks in the affine merge candidate list that are referenced to derive the CPMV of the current block. The affine merge candidate list may also be referred to as a sub-block merge candidate list.

[0109] The affine inter mode may also be referred to as an affine MVP mode. In the affine MVP mode, the CPMV of the current block may be derived based on a control point motion vector predictor (CPMVP) and a control point motion vector difference (CPMVD). In other words, the encoding device may determine a CPMVP for the CPMV of the current block, derive a CPMVD, which is a difference between the CPMV of the current block and the CPMVP, and signal information about the CPMVP and information about the CPMVD to the decoding device. Here, the affine MVP mode may construct an affine MVP candidate list based on neighboring blocks, and the information about the CPMVP may indicate neighboring blocks in the affine MVP candidate list that are referenced to derive a CPMVP for the CPMV of the current block. The affine MVP candidate list may also be referred to as a control point motion vector predictor candidate list.

[0110] For example, when the affine inter mode of the six-parameter affine motion model is applied, the current block can be encoded as described below.

[0111] FIG. 9 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to one embodiment of the present document.

[0112] As shown in Figure 9, the motion vector of CP0 of the current block can be expressed as v0, the motion vector of CP1 as v1, the motion vector of the control point at the bottom-left sample position as v2, and the motion vector of CP2 as v3. That is, v0 represents the CPMVP of CP0, v1 represents the CPMVP of CP1, and v2 represents the CPMVP of CP2.

[0113] The affine MVP candidate may be a combination of the CPMVP candidate for CP0, the CPMVP candidate for CP1, and the candidate for CP2.

[0114] For example, the affine MVP candidates can be derived as follows:

[0115] Specifically, a combination of up to 12 CPMVP candidates can be determined as follows:

[0116]

number

[0117] where v A is the motion vector of the surrounding block A, v B is the motion vector of the surrounding block B, v C is the motion vector of the surrounding block C, v D is the motion vector of the surrounding block D, v E is the motion vector of the surrounding block E, v F is the motion vector of the surrounding block F, v G can represent the motion vector of the surrounding block G.

[0118] The neighboring block A may represent a neighboring block located at the upper left corner of the upper left sample position of the current block, the neighboring block B may represent a neighboring block located at the top of the upper left sample position of the current block, and the neighboring block C may represent a neighboring block located at the left of the upper left sample position of the current block. The neighboring block D may represent a neighboring block located at the top of the upper right sample position of the current block, and the neighboring block E may represent a neighboring block located at the upper right corner of the upper right sample position of the current block. The neighboring block F may represent a neighboring block located at the left of the lower left sample position of the current block, and the neighboring block G may represent a neighboring block located at the lower left corner of the lower left sample position of the current block.

[0119] That is, referring to Equation 6, the CPMVP candidate of CP0 is the motion vector v of the neighboring block A. A , the motion vector v of the surrounding block B B , and / or the motion vector v of the surrounding block C C and the CPMVP candidates for CP1 may include the motion vector v of the neighboring block D. D , and / or the motion vector v of the surrounding block E E and the CPMVP candidates for CP2 may include the motion vector v of the neighboring block F. F , and / or the motion vector v of the surrounding block G G may include:

[0120] In other words, the CPMVP v0 of CP0 may be derived based on the motion vector of at least one of neighboring blocks A, B, and C at the top left sample position, where neighboring block A may refer to the block located at the top left of the top left sample position of the current block, neighboring block B may refer to the block located at the top of the top left sample position of the current block, and neighboring block C may refer to the block located to the left of the top left sample position of the current block.

[0121] A combination of up to 12 CPMVP candidates including the CPMVP candidate for CP0, the CPMVP candidate for CP1, and the CPMVP candidate for CP2 may be derived based on the motion vectors of the surrounding blocks.

[0122] Then, the derived combinations of CPMVP candidates are sorted in ascending order of DV, and the top two combinations of CPMVP candidates can be derived as the affine MVP candidates.

[0123] The DV of a combination of CPMVP candidates can be derived as follows:

[0124]

number

[0125] The encoding device can then determine a CPMV for each of the affine MVP candidates, compare RD (Rate Distortion) costs for the CPMV, and select the affine MVP candidate with the smallest RD cost as the optimal affine MVP candidate for the current block. The encoding device can encode and signal an index pointing to the optimal candidate and the CPMVD.

[0126] Also, for example, if the affine merge mode is applied, the current block can be encoded as described below.

[0127] FIG. 10 is a diagram illustrating a method for deriving a motion vector predictor at a control point according to one embodiment of the present document.

[0128] An affine merge candidate list for a current block may be constructed based on the neighboring blocks of the current block shown in Fig. 10. The neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, and neighboring block E. The neighboring block A may represent a left neighboring block of the current block, the neighboring block B may represent an upper neighboring block of the current block, the neighboring block C may represent a right upper corner neighboring block of the current block, the neighboring block D may represent a lower left corner neighboring block of the current block, and the neighboring block E may represent an upper left corner neighboring block of the current block.

[0129] For example, if the size of the current block is W×H and the x component of the top-left sample position of the current block is 0 and the y component is 0, the left peripheral block may be a block including a sample at a (-1, H-1) coordinate, the upper peripheral block may be a block including a sample at a (W-1, -1) coordinate, the upper right corner peripheral block may be a block including a sample at a (W, -1) coordinate, the lower left corner peripheral block may be a block including a sample at a (-1, H) coordinate, and the upper left corner peripheral block may be a block including a sample at a (-1, -1) coordinate.

[0130] Specifically, for example, the encoding apparatus may scan neighboring blocks A, B, C, D, and E of the current block in a specific scanning order, and determine the neighboring block encoded in the affine prediction mode first in the scanning order as a candidate block for the affine merge mode, i.e., an affine merge candidate. Here, for example, the specific scanning order may be alphabetical order. That is, the specific scanning order may be the order of neighboring blocks A, B, C, D, and E.

[0131] The encoding device can then determine an affine motion model for the current block using the CPMV of the determined candidate block, determine the CPMV of the current block based on the affine motion model, and determine the affine MVF of the current block based on the CPMV.

[0132] For example, if neighboring block A is determined as a candidate block for the current block, it can be coded as described below.

[0133] FIG. 11 shows an example of affine prediction performed when surrounding block A is selected as an affine merge candidate.

[0134] 11, the encoding apparatus may determine a neighboring block A of the current block as a candidate block and derive an affine motion model for the current block based on the CPMV, v2, and v3 of the neighboring blocks. Then, the encoding apparatus may determine the CPMV, v0, and v1 of the current block based on the affine motion model. The encoding apparatus may determine an affine MVF based on the CPMV, v0, and v1 of the current block and perform an encoding process for the current block based on the affine MVF.

[0135] Meanwhile, in relation to affine inter-prediction, inherited affine candidates and constructed affine candidates are considered for constructing an affine MVP candidate list.

[0136] Here, the inherited affine candidates may be as follows:

[0137] For example, if a neighboring block of the current block is an affine block and the reference picture of the current block is the same as the reference picture of the neighboring block, an affine MVP pair of the current block may be determined from the affine motion model of the neighboring block. Here, the affine block may represent a block to which the affine inter-prediction is applied. The inherited affine candidate may represent a CPMVP (e.g., the affine MVP pair) derived based on the affine motion model of the neighboring block.

[0138] Specifically, as an example, the inherited affine candidates can be derived as described below.

[0139] FIG. 12 exemplarily shows surrounding blocks for deriving the inherited affine candidates.

[0140] As shown in FIG. 12, the peripheral blocks of the current block may include a left peripheral block A0 of the current block, a lower left corner peripheral block A1 of the current block, an upper peripheral block B0 of the current block, a top right corner peripheral block B1 of the current block, and an upper left corner peripheral block B2 of the current block.

[0141] For example, if the size of the current block is W×H and the x component of the top-left sample position of the current block is 0 and the y component is 0, the left peripheral block may be a block including a sample at a (-1, H-1) coordinate, the upper peripheral block may be a block including a sample at a (W-1, -1) coordinate, the upper right corner peripheral block may be a block including a sample at a (W, -1) coordinate, the lower left corner peripheral block may be a block including a sample at a (-1, H) coordinate, and the upper left corner peripheral block may be a block including a sample at a (-1, -1) coordinate.

[0142] The encoding / decoding device may sequentially check neighboring blocks A0, A1, B0, B1, and B2. If the neighboring blocks are coded using an affine motion model and the reference picture of the current block is the same as the reference picture of the neighboring blocks, the encoding / decoding device may derive two or three CPMVs of the current block based on the affine motion models of the neighboring blocks. The CPMVs may be derived as affine MVP candidates of the current block. The affine MVP candidates may represent the inherited affine candidates.

[0143] As an example, up to two inherited affine candidates may be derived based on the surrounding blocks.

[0144] For example, the encoding / decoding device may derive a first affine MVP candidate for the current block based on a first block among neighboring blocks. Here, the first block may be coded using an affine motion model, and the reference picture of the first block may be the same as the reference picture of the current block. That is, the first block may be a block that satisfies a condition that is first confirmed when checking the neighboring blocks in a specific order. The condition may be that the block is coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block.

[0145] The encoding / decoding device may then derive a second affine MVP candidate for the current block based on a second block among the neighboring blocks. Here, the second block may be coded using an affine motion model, and the reference picture of the second block may be the same as the reference picture of the current block. That is, the second block may be the block that satisfies a second confirmed condition when checking the neighboring blocks in a specific order. The condition may be that the block is coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block.

[0146] On the other hand, for example, if the available number of inherited affine candidates is less than 2 (i.e., if the number of derived inherited affine candidates is less than 2), a constructed affine candidate may be considered. The constructed affine candidate may be derived as follows:

[0147] FIG. 13 exemplarily shows spatial candidates for the constructed affine candidates.

[0148] 13, the motion vectors of the neighboring blocks of the current block can be divided into three groups: neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G.

[0149] The neighboring block A may represent a neighboring block located at the upper left corner of the upper left sample position of the current block, the neighboring block B may represent a neighboring block located at the upper end of the upper left sample position of the current block, and the neighboring block C may represent a neighboring block located at the left end of the upper left sample position of the current block. Also, the neighboring block D may represent a neighboring block located at the upper end of the upper right sample position of the current block, and the neighboring block E may represent a neighboring block located at the upper right corner of the upper right sample position of the current block. Also, the neighboring block F may represent a neighboring block located at the left end of the lower left sample position of the current block, and the neighboring block G may represent a neighboring block located at the lower left corner of the lower left sample position of the current block.

[0150] For example, the three groups may include S0, S1, and S2, and S0, S1, and S2 may be derived as shown in the following table.

[0151] [Table 1]

[0152] Here, mv A is the motion vector of the surrounding block A, mv B is the motion vector of the surrounding block B, mv C is the motion vector of the surrounding block C, mv D is the motion vector of the surrounding block D, mv E is the motion vector of the surrounding block E, mvF is the motion vector of the surrounding block F, mv G represents the motion vector of the surrounding block G. S0 can also be expressed as the first group, S1 as the second group, and S2 as the third group.

[0153] The encoding device / decoding device can derive mv0 from S0, mv1 from S1, and mv2 from S2, and can derive affine MVP candidates including mv0, mv1, and mv2. The affine MVP candidates can represent the constructed affine candidates. Furthermore, mv0 can be a CPMVP candidate for CP0, mv1 can be a CPMVP candidate for CP1, and mv2 can be a CPMVP candidate for CP2.

[0154] Here, the reference picture for mv0 may be the same as the reference picture of the current block. That is, mv0 may be a motion vector that satisfies a condition that is first confirmed by checking motion vectors in S0 according to a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be neighboring block A → neighboring block B → neighboring block C in S0. Orders other than the above order may be used, and are not limited to the above example.

[0155] Furthermore, the reference picture for mv1 may be the same as the reference picture of the current block. That is, mv1 may be a motion vector that satisfies a condition that is first confirmed when motion vectors in S1 are checked according to a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be neighboring block D → neighboring block E in S1. The order may be other than the above and is not limited to the above example.

[0156] Furthermore, the reference picture for mv2 may be the same as the reference picture of the current block. That is, mv2 may be a motion vector that satisfies a condition that is first confirmed when motion vectors in S2 are checked according to a specific order. The condition may be that the reference picture for the motion vector is the same as the reference picture of the current block. The specific order may be neighboring block F→neighboring block G in S2. The order may be other than the above-mentioned order and is not limited to the above-mentioned example.

[0157] On the other hand, when only the mv0 and the mv1 are available, that is, when only the mv0 and the mv1 are derived, the mv2 can be derived as follows:

[0158]

number

[0159] where mv2 x indicates the x-component of mv2, and mv2 y indicates the y component of mv2, and mv0 x indicates the x-component of mv0, and mv0 y indicates the y component of mv0, and mv1 x indicates the x-component of mv1, and mv1 y indicates the y component of mv1, w indicates the width of the current block, and h indicates the height of the current block.

[0160] Meanwhile, when only the mv0 and the mv2 are derived, the mv1 can be derived as follows:

[0161]

number

[0162] where mv1 x indicates the x-component of mv1, and mv1 yindicates the y component of mv1, and mv0 x indicates the x-component of mv0, and mv0 y indicates the y component of mv0, and mv2 x indicates the x-component of mv2, and mv2 y indicates the y component of mv2, w indicates the width of the current block, and h indicates the height of the current block.

[0163] Also, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the AMVP process of the existing HEVC standard can be applied to construct the affine MVP list. That is, if the number of available inherited affine candidates and / or constructed affine candidates is less than 2, the process of constructing MVP candidates in the existing HEVC standard can be performed.

[0164] Meanwhile, a flowchart of an embodiment of constructing the above-mentioned affine MVP list will be described later.

[0165] FIG. 14 exemplarily shows an example of constructing an affine MVP list.

[0166] As shown in Figure 14, the encoding / decoding device can add an inherited candidate to the affine MVP list of the current block (S1400). The inherited candidate can represent the inherited affine candidate described above.

[0167] Specifically, the encoding / decoding apparatus may derive up to two inherited affine candidates from the neighboring blocks of the current block (S1405), where the neighboring blocks may include a left neighboring block A0, a lower left corner neighboring block A1, an upper neighboring block B0, a right upper corner neighboring block B1, and an upper left corner neighboring block B2 of the current block.

[0168] For example, the encoding / decoding device may derive a first affine MVP candidate for the current block based on a first block among neighboring blocks. Here, the first block may be coded using an affine motion model, and the reference picture of the first block may be the same as the reference picture of the current block. That is, the first block may be a block that satisfies a condition that is first confirmed when checking the neighboring blocks in a specific order. The condition may be that the block is coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block.

[0169] The encoding / decoding device may then derive a second affine MVP candidate for the current block based on a second block among the neighboring blocks. Here, the second block may be coded using an affine motion model, and the reference picture of the second block may be the same as the reference picture of the current block. That is, the second block may be the block that satisfies a second confirmed condition when checking the neighboring blocks in a specific order. The condition may be that the block is coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block.

[0170] Meanwhile, the specific order may be the left peripheral block A0 → the lower left corner peripheral block A1 → the upper peripheral block B0 → the upper right corner peripheral block B1 → the upper left corner peripheral block B2, etc. Also, the order may be other than the above-mentioned order and is not limited to the above-mentioned example.

[0171] The encoding / decoding device may add a constructed candidate to the affine MVP list of the current block (S1410). The constructed candidate may represent the above-mentioned constructed affine candidate. The constructed candidate may also be referred to as a constructed affine MVP candidate. If the number of available inherited candidates is less than two, the encoding / decoding device may add a constructed candidate to the affine MVP list of the current block. For example, the encoding / decoding device may derive one constructed affine candidate.

[0172] Meanwhile, the method of deriving the constructed affine candidate may differ depending on whether the affine motion model applied to the current block is a 6-affine motion model or a 4-affine motion model. The method of deriving the constructed affine candidate will be described in detail later.

[0173] The encoding device / decoding device may add an HEVC AMVP candidate to the affine MVP list of the current block (S1420). If the number of available inherited and / or constructed candidates is less than two, the encoding device / decoding device may add an HEVC AMVP candidate to the affine MVP list of the current block. That is, if the number of available inherited and / or constructed candidates is less than two, the encoding device / decoding device may perform a process of configuring an MVP candidate in the existing HEVC standard.

[0174] Meanwhile, the constructed candidate may be derived as follows.

[0175] For example, if the affine motion model applied to the current block is a 6-affine motion model, the constructed candidate may be derived as in the embodiment shown in FIG.

[0176] FIG. 15 shows an example of deriving the constructed candidates.

[0177] As shown in FIG. 15, the encoding / decoding device may check mv0, mv1, and mv2 for the current block (S1500). That is, the encoding / decoding device may determine whether mv0, mv1, and mv2 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2. Also, mv0, mv1, and mv2 may be represented as candidate motion vectors for the CP.

[0178] For example, the encoding / decoding device may check whether motion vectors of neighboring blocks in a first group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of a neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no available mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0179] Furthermore, for example, the encoding / decoding device may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of the neighboring block that satisfies the condition first confirmed in the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no available mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0180] Furthermore, for example, the encoding / decoding device may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no available mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0181] Meanwhile, the first group may include motion vectors of neighboring blocks A, B, and C, the second group may include motion vectors of neighboring blocks D and E, and the third group may include motion vectors of neighboring blocks F and G. The neighboring block A may represent a neighboring block located at the upper left corner of the upper left sample position of the current block, the neighboring block B may represent a neighboring block located at the upper end of the upper left sample position of the current block, the neighboring block C may represent a neighboring block located at the left end of the upper left sample position of the current block, the neighboring block D may represent a neighboring block located at the upper end of the upper right sample position of the current block, the neighboring block E may represent a neighboring block located at the upper right corner of the upper right sample position of the current block, the neighboring block F may represent a neighboring block located at the left end of the lower left sample position of the current block, and the neighboring block G may represent a neighboring block located at the lower left corner of the lower left sample position of the current block.

[0182] If only the mv0 and mv1 for the current block are available, i.e., if only the mv0 and mv1 for the current block are derived, the encoding / decoding device can derive mv2 for the current block based on the above-described Equation 8 (S1510). The encoding / decoding device can derive mv2 by substituting the derived mv0 and mv1 into the above-described Equation 8.

[0183] If only the mv0 and mv2 for the current block are available, i.e., if only the mv0 and mv2 for the current block are derived, the encoding / decoding device can derive mv1 for the current block based on the above-described Equation 9 (S1520). The encoding / decoding device can derive mv1 by substituting the derived mv0 and mv2 into the above-described Equation 9.

[0184] The encoding device / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block (S1530). If the mv0, mv1, and mv2 are available, i.e., if the mv0, mv1, and mv2 are derived based on neighboring blocks of the current block, the encoding device / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block.

[0185] Also, if only mv0 and mv1 for the current block are available, i.e., if only mv0 and mv1 for the current block are derived, the encoding device / decoding device can derive mv2 derived based on the derived mv0, mv1 and the above-mentioned Equation 8 as the constructed candidate for the current block.

[0186] Also, if only mv0 and mv2 for the current block are available, i.e., if only mv0 and mv2 for the current block are derived, the encoding device / decoding device can derive mv1 derived based on the derived mv0, mv2 and the above-mentioned Equation 9 as the constructed candidate for the current block.

[0187] Also, for example, if the affine motion model applied to the current block is a 4-affine motion model, the constructed candidate can be derived as in the embodiment shown in FIG.

[0188] FIG. 16 shows an example of deriving the constructed candidates.

[0189] 16, the encoding / decoding device may check mv0, mv1, and mv2 for the current block (S1600). That is, the encoding / decoding device may determine whether mv0, mv1, and mv2 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0190] For example, the encoding / decoding device may check whether motion vectors of neighboring blocks in a first group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of a neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no available mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0191] Furthermore, for example, the encoding / decoding device may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of the neighboring block that satisfies the condition first confirmed in the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no available mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0192] Furthermore, for example, the encoding / decoding device may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of the neighboring block that satisfies the condition first confirmed during the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no available mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0193] Meanwhile, the first group may include motion vectors of neighboring blocks A, B, and C, the second group may include motion vectors of neighboring blocks D and E, and the third group may include motion vectors of neighboring blocks F and G. The neighboring block A may represent a neighboring block located at the upper left corner of the upper left sample position of the current block, the neighboring block B may represent a neighboring block located at the upper end of the upper left sample position of the current block, the neighboring block C may represent a neighboring block located at the left end of the upper left sample position of the current block, the neighboring block D may represent a neighboring block located at the upper end of the upper right sample position of the current block, the neighboring block E may represent a neighboring block located at the upper right corner of the upper right sample position of the current block, the neighboring block F may represent a neighboring block located at the left end of the lower left sample position of the current block, and the neighboring block G may represent a neighboring block located at the lower left corner of the lower left sample position of the current block.

[0194] If only mv0 and mv1 for the current block are available or if mv0, mv1, and mv2 for the current block are available, i.e., if only mv0 and mv1 for the current block are derived or if mv0, mv1, and mv2 for the current block are derived, the encoding device / decoding device can derive the derived mv0 and mv1 as constructed candidates for the current block (S1610).

[0195] On the other hand, if only the mv0 and mv2 for the current block are available, i.e., if only the mv0 and mv2 for the current block are derived, the encoding / decoding device can derive mv1 for the current block based on the above-described Equation 9 (S1620). The encoding / decoding device can derive mv1 by substituting the derived mv0 and mv2 into the above-described Equation 9.

[0196] Thereafter, the encoding / decoding apparatus may derive the derived mv0 and mv1 as constructed candidates for the current block (S1610).

[0197] Meanwhile, this document proposes another embodiment for deriving the inherited affine candidates, which can reduce the computational complexity and improve coding performance when deriving the inherited affine candidates.

[0198] FIG. 17 exemplarily shows the surrounding block positions scanned to derive the inherited affine candidates.

[0199] The encoding / decoding device may derive up to two inherited affine candidates from the neighboring blocks of the current block. Figure 17 may represent the neighboring blocks for the inherited affine candidates. For example, the neighboring blocks may include neighboring block A and neighboring block B shown in Figure 17. The neighboring block A may represent the left neighboring block A0 described above, and the neighboring block B may represent the upper neighboring block B0 described above.

[0200] For example, the encoding / decoding device may check whether the neighboring blocks are available in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. That is, the encoding / decoding device may check whether the neighboring blocks satisfy a specific condition in a specific order and derive the inherited affine candidate for the current block based on the first identified available neighboring block. The encoding / decoding device may also derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies the specific condition. That is, the encoding / decoding device may derive the inherited affine candidate for the current block based on the second identified neighboring block that satisfies the specific condition. Here, the availability may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. That is, the specific condition may be coded using an affine motion model, and the reference picture of the block may be the same as the reference picture of the current block. For example, the specific order may be neighboring block A → neighboring block B. Meanwhile, a pruning check process between two inherited affine candidates (i.e., derived inherited affine candidates, etc.) may not be performed. The pruning check process may refer to a process of checking whether the candidates are identical to each other, and if they are identical, removing the candidate derived in the final order.

[0201] The above-described embodiment proposes a method of deriving the inherited affine candidate by checking only two neighboring blocks (i.e., neighboring block A and neighboring block B) instead of deriving the inherited affine candidate by checking all existing neighboring blocks (i.e., neighboring block A, neighboring block B, neighboring block C, neighboring block D, and neighboring block E). Here, the neighboring block C may represent the above-described upper right corner neighboring block B1, the neighboring block D may represent the above-described lower left corner neighboring block A1, and the neighboring block E may represent the above-described upper left corner neighboring block B2.

[0202] In order to analyze the spatial correlation between the neighboring blocks and the current block through affine inter-prediction, the probability that affine prediction is applied to the current block when affine prediction is applied to each neighboring block may be referenced. The probability that affine prediction is applied to the current block when affine prediction is applied to each neighboring block may be derived as shown in the following table.

[0203] [Table 2]

[0204] Referring to Table 2, it can be seen that among the neighboring blocks, neighboring block A and neighboring block B have high spatial correlation with the current block. Therefore, by deriving the inherited affine candidate using only neighboring block A and neighboring block B having high spatial correlation, it is possible to obtain an effect of deriving high decoding performance while reducing processing time.

[0205] Meanwhile, the pruning check process may be performed to prevent the same candidate from existing in the candidate list. The pruning check process may be advantageous in terms of encoding efficiency because it can eliminate redundancy, but has the disadvantage of increasing computational complexity. In particular, the pruning check process for affine candidates must be performed on the affine type (e.g., whether the affine motion model is a 4-affine motion model or a 6-affine motion model), reference pictures (or reference picture indexes), and MVs of CP0, CP1, and CP2, resulting in very high computational complexity. Therefore, this embodiment proposes a method of not performing a pruning check process between an inherited affine candidate (e.g., inherited_A) derived based on the neighboring block A and an inherited affine candidate (e.g., inherited_B) derived based on the neighboring block B. In the case of neighboring blocks A and B, the distance between them is large, and therefore the spatial correlation is low, so it is unlikely that inherited_A and inherited_B are identical. Therefore, it may be appropriate not to perform a pruning check process between the inherited affine candidates.

[0206] Alternatively, a method of performing a minimal pruning check process based on the above-mentioned reasons may be proposed. For example, the encoding / decoding device may perform a pruning check process by comparing only the MVs of the CP0s of the inherited affine candidates.

[0207] This document also proposes a method for deriving constructed candidates that is different from the above-described embodiment. The proposed embodiment can reduce complexity and improve coding performance compared to the above-described embodiment for deriving constructed candidates. The proposed embodiment is described below. Furthermore, if the available number of inherited affine candidates is less than two (i.e., if the number of derived inherited affine candidates is less than two), a constructed affine candidate may be considered.

[0208] For example, the encoding / decoding device can check mv0, mv1, and mv2 for the current block. That is, the encoding / decoding device can determine whether mv0, mv1, and mv2 are available in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0209] Specifically, the neighboring blocks of the current block may be divided into three groups, and the neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G. The first group may include the motion vectors of neighboring block A, neighboring block B, and neighboring block C, the second group may include the motion vectors of neighboring block D and neighboring block E, and the third group may include the motion vectors of neighboring block F and neighboring block G. The peripheral block A may represent a peripheral block located at the upper left end of the upper left sample position of the current block, the peripheral block B may represent a peripheral block located at the upper end of the upper left sample position of the current block, the peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the peripheral block D may represent a peripheral block located at the upper end of the upper right sample position of the current block, the peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0210] The encoding device / decoding device can determine whether mv0 is available in the first group, can determine whether mv1 is available in the second group, and can determine whether mv2 is available in the third group.

[0211] Specifically, for example, the encoding / decoding device may check whether motion vectors of neighboring blocks in a first group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of a neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no available mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0212] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed in the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no available mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0213] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed in the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no available mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. For example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0214] Thereafter, when the affine motion model applied to the current block is a 4-affine motion model, if mv0 and mv1 for the current block are available, the encoding / decoding device may derive the derived mv0 and mv1 as constructed candidates for the current block. On the other hand, if mv0 and / or mv1 for the current block are not available, i.e., if at least one of mv0 and mv1 is not derived from a neighboring block of the current block, the encoding / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0215] Furthermore, when the affine motion model applied to the current block is a 6-affine motion model, if mv0, mv1, and mv2 for the current block are available, the encoding / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block. On the other hand, if mv0, mv1, and / or mv2 for the current block are not available, i.e., if at least one of mv0, mv1, and mv2 is not derived from a neighboring block of the current block, the encoding / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0216] The above-described proposed embodiment is a method of considering a current block as a constructed candidate only if all motion vectors of CPs for generating an affine motion model are available. Here, "available" may mean that the reference picture of a neighboring block is the same as the reference picture of the current block. That is, the constructed candidate can be derived only if there is a motion vector that satisfies the above condition among the motion vectors of neighboring blocks for each CP of the current block. Therefore, if the affine motion model applied to the current block is a 4-affine motion model, the constructed candidate can be considered only if the motion vectors of CP0 and CP1 of the current block (i.e., mv0 and mv1) are available. Also, if the affine motion model applied to the current block is a 6-affine motion model, the constructed candidate can be considered only if the motion vectors of CP0, CP1, and CP2 of the current block (i.e., mv0, mv1, and mv2) are available. Therefore, according to the proposed embodiment, an additional configuration for deriving a motion vector for a CP based on the above-described Equation 8 or Equation 9 may not be necessary. This reduces the computational complexity for deriving the constructed candidate. Also, since the constructed candidate is determined only when a CPMVP candidate having the same reference picture is available, overall coding performance can be improved.

[0217] Meanwhile, a pruning check process between the derived inherited affine candidate and the constructed affine candidate may not be performed. The pruning check process may refer to a process of checking whether the candidates are identical to each other, and if they are identical, removing the candidate derived in the final order.

[0218] The above-described embodiment can be illustrated as in FIGS.

[0219] FIG. 18 shows an example of deriving the constructed candidates when four affine motion models are applied to the current block.

[0220] 18, the encoding / decoding device may determine whether mv0 and mv1 are available for the current block (S1800). That is, the encoding / decoding device may determine whether mv0 and mv1 are available for neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, and mv1 may be a CPMVP candidate for CP1.

[0221] The encoding / decoding device can determine if there is mv0 available in the first group and can determine if there is mv1 available in the second group.

[0222] Specifically, the neighboring blocks of the current block may be divided into three groups, and the neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G. The first group may include the motion vectors of neighboring block A, neighboring block B, and neighboring block C, the second group may include the motion vectors of neighboring block D and neighboring block E, and the third group may include the motion vectors of neighboring block F and neighboring block G. The peripheral block A may represent a peripheral block located at the upper left end of the upper left sample position of the current block, the peripheral block B may represent a peripheral block located at the upper end of the upper left sample position of the current block, the peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the peripheral block D may represent a peripheral block located at the upper end of the upper right sample position of the current block, the peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0223] The encoding / decoding device may check whether motion vectors of neighboring blocks in the first group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of a neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no available mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0224] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed in the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no available mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0225] If the mv0 and mv1 for the current block are available, i.e., if the mv0 and mv1 for the current block are derived, the encoding / decoding device may derive the derived mv0 and mv1 as constructed candidates for the current block (S1810). On the other hand, if the mv0 and / or mv1 for the current block are not available, i.e., if at least one of mv0 and mv1 is not derived from a neighboring block of the current block, the encoding / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0226] Meanwhile, a pruning check process between the derived inherited affine candidate and the constructed affine candidate may not be performed. The pruning check process may refer to a process of checking whether the candidates are identical to each other, and if they are identical, removing the candidate derived in the final order.

[0227] FIG. 19 shows an example of deriving the constructed candidates when a 6-affine motion model is applied to the current block.

[0228] 19, the encoding / decoding device may determine whether mv0, mv1, and mv2 are available for the current block (S1900). That is, the encoding / decoding device may determine whether mv0, mv1, and mv2 are available for use in neighboring blocks of the current block. Here, mv0 may be a CPMVP candidate for CP0 of the current block, mv1 may be a CPMVP candidate for CP1, and mv2 may be a CPMVP candidate for CP2.

[0229] The encoding / decoding device can determine if there is an mv0 available in the first group, can determine if there is an mv1 available in the second group, and can determine if there is an mv2 available in the third group.

[0230] Specifically, the neighboring blocks of the current block may be divided into three groups, and the neighboring blocks may include neighboring block A, neighboring block B, neighboring block C, neighboring block D, neighboring block E, neighboring block F, and neighboring block G. The first group may include the motion vectors of neighboring block A, neighboring block B, and neighboring block C, the second group may include the motion vectors of neighboring block D and neighboring block E, and the third group may include the motion vectors of neighboring block F and neighboring block G. The peripheral block A may represent a peripheral block located at the upper left end of the upper left sample position of the current block, the peripheral block B may represent a peripheral block located at the upper end of the upper left sample position of the current block, the peripheral block C may represent a peripheral block located at the left end of the upper left sample position of the current block, the peripheral block D may represent a peripheral block located at the upper end of the upper right sample position of the current block, the peripheral block E may represent a peripheral block located at the upper right end of the upper right sample position of the current block, the peripheral block F may represent a peripheral block located at the left end of the lower left sample position of the current block, and the peripheral block G may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0231] The encoding / decoding device may check whether motion vectors of neighboring blocks in the first group satisfy a specific condition in a specific order. The encoding / decoding device may derive the motion vector of a neighboring block that satisfies the condition first confirmed during the checking process as mv0. That is, mv0 may be the motion vector that satisfies the specific condition first confirmed after checking motion vectors in the first group in a specific order. If the motion vectors of the neighboring blocks in the first group do not satisfy the specific condition, there may be no available mv0. Here, for example, the specific order may be from neighboring block A to neighboring block B to neighboring block C in the first group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0232] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the second group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed in the checking process as mv1. That is, mv1 may be the motion vector that satisfies the specific condition first confirmed by checking the motion vectors in the second group in a specific order. If the motion vectors of the neighboring blocks in the second group do not satisfy the specific condition, there may be no available mv1. Here, for example, the specific order may be from neighboring block D to neighboring block E in the second group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0233] Furthermore, the encoding / decoding apparatus may check whether the motion vectors of the neighboring blocks in the third group satisfy a specific condition in a specific order. The encoding / decoding apparatus may derive the motion vector of the neighboring block that satisfies the condition first confirmed in the checking process as mv2. That is, mv2 may be the motion vector that satisfies the specific condition first confirmed after checking the motion vectors in the third group in a specific order. If the motion vectors of the neighboring blocks in the third group do not satisfy the specific condition, there may be no available mv2. Here, for example, the specific order may be from neighboring block F to neighboring block G in the third group. Also, for example, the specific condition may be that the reference picture for the motion vector of the neighboring block is the same as the reference picture of the current block.

[0234] If the mv0, mv1, and mv2 for the current block are available, i.e., if the mv0, mv1, and mv2 for the current block are derived, the encoding / decoding device may derive the derived mv0, mv1, and mv2 as constructed candidates for the current block (S1910). On the other hand, if the mv0, mv1, and / or mv2 for the current block are not available, i.e., if at least one of mv0, mv1, and mv2 is not derived from a neighboring block of the current block, the encoding / decoding device may not add a constructed candidate to the affine MVP list of the current block.

[0235] Meanwhile, a pruning check process between the derived inherited affine candidate and the constructed affine candidate may not be performed.

[0236] On the other hand, if the number of derived affine candidates is less than two (i.e., if the number of inherited affine candidates and / or constructed affine candidates is less than two), an HEVC AMVP candidate may be added to the affine MVP list of the current block.

[0237] For example, the HEVC AMVP candidates can be derived in the following order:

[0238] Specifically, when the number of derived affine candidates is less than 2, if the CPMV0 of the constructed affine candidate is available, the CPMV0 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2, if the CPMV0 of the constructed affine candidate is available (i.e., when the number of derived affine candidates is less than 2 and the CPMV0 of the constructed affine candidate is derived), a first affine MVP candidate including the CPMV0 of the constructed affine candidate in CPMV0, CPMV1, and CPMV2 can be derived.

[0239] Next, when the number of derived affine candidates is less than 2 and CPMV1 of the constructed affine candidate is available, the CPMV1 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2 and CPMV1 of the constructed affine candidate is available (i.e., when the number of derived affine candidates is less than 2 and CPMV1 of the constructed affine candidate is derived), a second affine MVP candidate can be derived that includes CPMV1 of the constructed affine candidate in CPMV0, CPMV1, and CPMV2.

[0240] Next, when the number of derived affine candidates is less than 2, if the CPMV2 of the constructed affine candidate is available, the CPMV2 can be used as the affine MVP candidate. That is, when the number of derived affine candidates is less than 2, if the CPMV2 of the constructed affine candidate is available (i.e., when the number of derived affine candidates is less than 2 and the CPMV2 of the constructed affine candidate is derived), a third affine MVP candidate can be derived that includes the CPMV2 of the constructed affine candidate in CPMV0, CPMV1, and CPMV2.

[0241] Next, if the number of derived affine candidates is less than two, an HEVC Temporal Motion Vector Predictor (TMVP) may be used as the affine MVP candidate. The HEVC TMVP may be derived based on motion information of temporal neighboring blocks of the current block. That is, if the number of derived affine candidates is less than two, a third affine MVP candidate including motion vectors of temporal neighboring blocks of the current block as CPMV0, CPMV1, and CPMV2 may be derived. The temporal neighboring blocks may represent collocated blocks in a collocated picture corresponding to the current block.

[0242] Next, if the number of derived affine candidates is less than 2, a zero motion vector (zero MV) can be used as the affine MVP candidate. That is, if the number of derived affine candidates is less than 2, a third affine MVP candidate including the zero motion vector as CPMV0, CPMV1, and CPMV2 can be derived. The zero motion vector can represent a motion vector with a value of 0.

[0243] This is because the step of using the CPMV of the constructed affine candidate reuses the MV that has already been considered for generating the constructed affine candidate, thereby reducing the complexity compared to existing methods of deriving HEVC AMVP candidates.

[0244] However, this document proposes another embodiment for deriving the inherited affine candidates.

[0245] In order to derive the inherited affine candidates, affine prediction information of the surrounding blocks is required. Specifically, the following affine prediction information is required:

[0246] 1) An affine flag (affine_flag) indicating whether affine prediction-based encoding of the neighboring blocks is applied.

[0247] 2) Motion information of the surrounding blocks

[0248] When a 4-affine motion model is applied to the surrounding blocks, the motion information of the surrounding blocks may include L0 motion information and L1 motion information for CP0, and L0 motion information and L1 motion information for CP1. When a 6-affine motion model is applied to the surrounding blocks, the motion information of the surrounding blocks may include L0 motion information and L1 motion information for CP0, and L0 motion information and L1 motion information for CP2. Here, the L0 motion information may represent motion information for L0 (List 0), and the L1 motion information may represent motion information for L1 (List 1). The L0 motion information may include an L0 reference picture index and an L0 motion vector, and the L1 motion information may include an L1 reference picture index and an L1 motion vector.

[0249] As described above, affine prediction requires a large amount of information to be stored, which can be a major cause of increased hardware costs in actual implementations of encoding / decoding devices. In particular, if a neighboring block is located above the current block and is a CTU boundary, a line buffer must be used to store affine prediction-related information for the neighboring block, which can result in greater cost problems. This problem can be referred to as a line buffer issue. In response, this document proposes an embodiment in which affine prediction-related information is not stored or is reduced in the line buffer, thereby minimizing hardware costs and deriving inherited affine candidates. The proposed embodiment can improve encoding performance by reducing the computational complexity when deriving the inherited affine candidates. Meanwhile, for reference, if motion information for a 4x4 size block is already stored in the line buffer and the affine prediction-related information is additionally stored, the amount of information stored can increase by three times compared to the existing storage amount.

[0250] In this embodiment, no additional information for affine prediction may be stored in the line buffer, and if information in the line buffer must be referenced to generate the inherited affine candidate, the generation of the inherited affine candidate may be restricted.

[0251] 20a-20b exemplarily illustrate an embodiment for deriving the inherited affine candidates.

[0252] As shown in FIG. 20a, if a neighboring block B of the current block (i.e., an upper neighboring block of the current block) is not present in the same CTU as the current block (i.e., the current CTU), the neighboring block B may not be used to generate the inherited affine candidate. Meanwhile, although a neighboring block A is not present in the same CTU as the current block, information about the neighboring block A is not stored in a line buffer and may be used to generate the inherited affine candidate. Therefore, in this embodiment, the upper neighboring block of the current block may be used to derive the inherited affine candidate only if it is included in the same CTU as the current block. Furthermore, if an upper neighboring block of the current block is not included in the same CTU as the current block, the upper neighboring block may not be used to derive the inherited affine candidate.

[0253] As shown in Figure 20b, a neighboring block B of the current block (i.e., an upper neighboring block of the current block) may be present in the same CTU as the current block. In this case, the encoding / decoding device may generate the inherited affine candidate by referring to the neighboring block B.

[0254] FIG. 21 illustrates an outline of an image encoding method by an encoding device according to the present document. The method disclosed in FIG. 21 may be performed by the encoding device disclosed in FIG. 2. Specifically, for example, steps S2100 to S2120 in FIG. 21 may be performed by a prediction unit of the encoding device, step S2130 may be performed by a subtraction unit of the encoding device, and step S2140 may be performed by an entropy encoding unit of the encoding device. Also, for example, although not shown, a step of deriving predicted samples for the current block based on the CPMV may be performed by a prediction unit of the encoding device, a step of deriving residual samples for the current block based on original samples and predicted samples for the current block may be performed by a subtraction unit of the encoding device, a step of generating information about the residual for the current block based on the residual samples may be performed by a transformation unit of the encoding device, and a step of encoding information about the residual may be performed by an entropy encoding unit of the encoding device.

[0255] The encoding apparatus constructs an affine motion vector predictor (MVP) candidate list for a current block (S2100). The encoding apparatus may construct an affine MVP candidate list including affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list may be two.

[0256] Also, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The encoding apparatus may check whether the inherited affine MVP candidates of the current block are available, and if the inherited affine MVP candidates are available, the inherited affine MVP candidates may be derived. For example, the inherited affine MVP candidates may be derived based on neighboring blocks of the current block, and the maximum number of inherited affine MVP candidates may be two. The neighboring blocks may be checked for availability in a specific order, and the inherited affine MVP candidates may be derived based on the checked available neighboring blocks. That is, the neighboring blocks may be checked for availability in a specific order, and a first inherited affine MVP candidate may be derived based on the first checked available neighboring block, and a second inherited affine MVP candidate may be derived based on the second checked available neighboring block. The available neighboring blocks may be coded using an affine motion model, and the reference picture of the neighboring blocks may be the same as the reference picture of the current block. That is, the available neighboring blocks may be coded using an affine motion model (i.e., affine prediction is applied), and the reference picture may be the same as the reference picture of the current block. Specifically, the encoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the first checked available neighboring block, and derive the first inherited affine MVP candidate including the motion vector as a CPMVP candidate. Also, the encoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the second checked available neighboring block, and derive the second inherited affine MVP candidate including the motion vector as a CPMVP candidate. The affine motion model may be derived as shown in Equation 1 or 3 above.

[0257] In other words, the neighboring blocks may be checked to see if they satisfy a specific condition in a specific order, and the inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the checked specific condition. That is, the neighboring blocks may be checked to see if they satisfy the specific condition in a specific order, and a first inherited affine MVP candidate may be derived based on the neighboring block that satisfies the specific condition that is checked first, and a second inherited affine MVP candidate may be derived based on the neighboring block that satisfies the specific condition that is checked second. Specifically, the encoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring block that satisfies the specific condition that is checked first, and may derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. Furthermore, the encoding apparatus may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring block that satisfies the specific condition that is checked second, and may derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model can be derived as shown in Equation 1 or 3. Meanwhile, the specific condition may represent that the current block is coded using an affine motion model and the reference picture of the neighboring block is the same as the reference picture of the current block. That is, the neighboring block satisfying the specific condition may be coded using an affine motion model (i.e., affine prediction is applied) and the reference picture may be the same as the reference picture of the current block.

[0258] Here, for example, the peripheral blocks may include a left peripheral block, an upper peripheral block, a right upper corner peripheral block, a lower left corner peripheral block, and an upper left corner peripheral block of the current block, and the specific order may be from the left peripheral block to the lower left corner peripheral block, the upper peripheral block, the right upper corner peripheral block, and the upper left corner peripheral block.

[0259] Alternatively, for example, the peripheral blocks may include only the left peripheral block and the top peripheral block, in which case the specific order may be from the left peripheral block to the top peripheral block.

[0260] Alternatively, for example, the peripheral blocks may include the left peripheral block, and if the upper peripheral block is included in a current CTU including the current block, the peripheral blocks may further include the upper peripheral block. In this case, the specific order may be from the left peripheral block to the upper peripheral block. Also, if the upper peripheral block is not included in the current CTU, the peripheral blocks may not include the upper peripheral block. In this case, only the left peripheral block may be checked.

[0261] On the other hand, if the size is W×H and the x component and y component of the top-left sample position of the current block are 0, the lower-left corner peripheral block may be a block including a sample at coordinates (-1, H), the left peripheral block may be a block including a sample at coordinates (-1, H-1), the upper-right corner peripheral block may be a block including a sample at coordinates (W, -1), the upper peripheral block may be a block including a sample at coordinates (W-1, -1), and the upper-left corner peripheral block may be a block including a sample at coordinates (-1, -1). That is, the left peripheral block may be the leftmost peripheral block among the left peripheral blocks of the current block, and the upper peripheral block may be the leftmost peripheral block among the upper peripheral blocks of the current block.

[0262] Also, as an example, if a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. The encoding apparatus may check whether a constructed affine MVP candidate for the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate may be derived. Also, for example, the constructed affine MVP candidate may be derived after the inherited affine MVP candidate is derived. If the number of derived affine MVP candidates (i.e., the inherited affine MVP candidates) is less than two and the constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. Here, the constructed affine MVP candidate may include candidate motion vectors for the CP. The constructed affine MVP candidate may be available if all of the candidate motion vectors are available.

[0263] For example, if a 4-affine motion model is applied to the current block, the CPs of the current block may include CP0 and CP1. If a candidate motion vector for CP0 is available and a candidate motion vector for CP1 is available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, CP0 may represent the upper left corner of the current block, and CP1 may represent the upper right corner of the current block.

[0264] The constructed affine MVP candidates may include a candidate motion vector for the CP0 and a candidate motion vector for the CP1. The candidate motion vector for the CP0 may be a motion vector of a first block, and the candidate motion vector for the CP1 may be a motion vector of a second block.

[0265] Furthermore, the first block may check neighboring blocks in the first group according to a first specific order, and the first identified reference picture may be the same as the reference picture of the current block. That is, the candidate motion vector for CP1 may be the motion vector of the block whose first identified reference picture is the same as the reference picture of the current block when neighboring blocks in the first group are checked according to a first order. The availability may indicate that the neighboring blocks exist and are coded using inter-prediction. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0266] Furthermore, the second block may check neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, the candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0267] On the other hand, if the size of the current block is W×H and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at (-1, -1) coordinates, the surrounding block B may be a block including a sample at (0, -1) coordinates, the surrounding block C may be a block including a sample at (-1, 0) coordinates, the surrounding block D may be a block including a sample at (W-1, -1) coordinates, and the surrounding block E may be a block including a sample at (W, -1) coordinates. That is, the peripheral block A may be the upper left corner peripheral block of the current block, the peripheral block B may be the upper peripheral block located on the leftmost side among the upper peripheral blocks of the current block, the peripheral block C may be the left peripheral block located on the topmost side among the left peripheral blocks of the current block, the peripheral block D may be the upper peripheral block located on the rightmost side among the upper peripheral blocks of the current block, and the peripheral block E may be the upper right corner peripheral block of the current block.

[0268] On the other hand, if at least one of the candidate motion vectors of CP0 and the candidate motion vectors of CP1 is unavailable, the constructed affine MVP candidate may be unavailable.

[0269] Alternatively, for example, if a 6-affine motion model is applied to the current block, the CPs of the current block may include CP0, CP1, and CP2. If a candidate motion vector for CP0 is available, a candidate motion vector for CP1 is available, and a candidate motion vector for CP2 is available, the constructed affine MVP candidates may be available, and the affine MVP candidate list may include the constructed affine MVP candidates. Here, CP0 may represent the upper left corner of the current block, CP1 may represent the upper right corner of the current block, and CP2 may represent the lower left corner of the current block.

[0270] The constructed affine MVP candidates may include a candidate motion vector for the CP0, a candidate motion vector for the CP1, and a candidate motion vector for the CP2. The candidate motion vector for the CP0 may be a motion vector of a first block, the candidate motion vector for the CP1 may be a motion vector of a second block, and the candidate motion vector for the CP2 may be a motion vector of a third block.

[0271] Furthermore, the first block may check neighboring blocks in the first group according to a first specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0272] Furthermore, the second block may check neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, the candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0273] Furthermore, the third block may check neighboring blocks in the third group according to a third specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the third block in the third group is the same as the reference picture of the current block, the candidate motion vector for CP2 may be available. For example, the third group may include neighboring blocks F and G, and the third specific order may be from neighboring block F to neighboring block G.

[0274] On the other hand, if the size of the current block is W×H and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at a (-1, -1) coordinate, the surrounding block B may be a block including a sample at a (0, -1) coordinate, the surrounding block C may be a block including a sample at a (-1, 0) coordinate, the surrounding block D may be a block including a sample at a (W-1, -1) coordinate, the surrounding block E may be a block including a sample at a (W, -1) coordinate, the surrounding block F may be a block including a sample at a (-1, H-1) coordinate, and the surrounding block G may be a block including a sample at a (-1, H) coordinate. That is, the peripheral block A may be the upper left corner peripheral block of the current block, the peripheral block B may be the upper leftmost peripheral block among the upper peripheral blocks of the current block, the peripheral block C may be the leftmost peripheral block among the left peripheral blocks of the current block, the peripheral block D may be the upper rightmost peripheral block among the upper peripheral blocks of the current block, the peripheral block E may be the upper right corner peripheral block of the current block, the peripheral block F may be the leftmost peripheral block among the left peripheral blocks of the current block, and the peripheral block G may be the lower left corner peripheral block of the current block.

[0275] On the other hand, if at least one of the candidate motion vectors of CP0, CP1, and CP2 is unavailable, the constructed affine MVP candidate may be unavailable.

[0276] The affine MVP candidate list can then be derived based on the following sequential steps:

[0277] For example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP0 is available, the encoding apparatus may derive a first affine MVP candidate, where the first affine MVP candidate may be an affine MVP candidate that includes the motion vector for the CP0 as a candidate motion vector for the CP.

[0278] Also, for example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP1 is available, the encoding device can derive a second affine MVP candidate, where the second affine MVP candidate can be an affine MVP candidate that includes the motion vector for the CP1 as a candidate motion vector for the CP.

[0279] Also, for example, if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the encoding apparatus may derive a third affine MVP candidate, where the third affine MVP candidate may be an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for the CP.

[0280] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the encoding apparatus may derive a fourth affine MVP candidate including a temporal MVP derived based on temporal neighboring blocks of the current block as a candidate motion vector for the CP. The temporal neighboring blocks may represent collocated blocks in a collocated picture corresponding to the current block. The temporal MVP may be derived based on the motion vectors of the temporal neighboring blocks.

[0281] Also, for example, if the number of derived affine MVP candidates is less than two, the encoding apparatus may derive a fifth affine MVP candidate that includes a zero motion vector as a candidate motion vector for the CP. The zero motion vector may represent a motion vector with a value of 0.

[0282] The encoding apparatus derives control point motion vector predictors (CPMVPs) for the control points (CPs) of the current block based on the affine MVP candidate list (S2110). The encoding apparatus may derive a CPMV for the CP of the current block having an optimal RD cost and may select an affine MVP candidate most similar to the CPMV from the affine MVP candidates as an affine MVP candidate for the current block. The encoding apparatus may derive control point motion vector predictors (CPMVPs) for the CP of the current block based on the selected affine MVP candidate from the affine MVP candidates included in the affine MVP candidate list. Specifically, if the affine MVP candidates include a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate may be derived from the CPMVP of CP0, and the candidate motion vector for CP1 of the affine MVP candidate may be derived from the CPMVP of CP1. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived using the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived using the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived using the CPMVP of CP2. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived using the CPMVP of CP0, and the candidate motion vector for CP2 of the affine MVP candidate can be derived using the CPMVP of CP2.

[0283] The encoding apparatus may encode an affine MVP candidate index indicating the selected affine MVP candidate from the affine MVP candidates, and the affine MVP candidate index may indicate the one affine MVP candidate from among affine MVP candidates included in an affine motion vector predictor (MVP) candidate list for the current block.

[0284] The encoding apparatus derives a CPMV for the CP of the current block (S2120). The encoding apparatus can derive a CPMV for each of the CPs of the current block.

[0285] The encoding apparatus derives control point motion vector differences (CPMVDs) for the CPs of the current block based on the CPMVP and the CPMV (S2130). The encoding apparatus can derive CPMVDs for the CPs of the current block based on the CPMVP and the CPMV for each of the CPs.

[0286] The encoding apparatus encodes motion prediction information including information about the CPMVD (S2140). The encoding apparatus can output the motion prediction information including information about the CPMVD in the form of a bitstream. That is, the encoding apparatus can output image information including the motion prediction information in the form of a bitstream. The encoding apparatus can encode information about CPMVD for each of the CPs, and the motion prediction information can include information about the CPMVD.

[0287] The motion prediction information may also include the affine MVP candidate index, which may indicate the selected affine MVP candidate from among affine MVP candidates included in an affine motion vector predictor (MVP) candidate list for the current block.

[0288] Meanwhile, as an example, the encoding device may derive predicted samples for the current block based on the CPMV, derive residual samples for the current block based on original samples and predicted samples for the current block, generate information about the residual for the current block based on the residual samples, and encode the information about the residual. The image information may include information about the residual.

[0289] Meanwhile, the bitstream can be transmitted to the decoding device via a network or a (digital) storage medium, where the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0290] FIG. 22 illustrates an outline of an encoding device that performs the image encoding method according to the present document. The method disclosed in FIG. 21 can be performed by the encoding device disclosed in FIG. 22. Specifically, for example, a prediction unit of the encoding device of FIG. 22 can perform steps S2100 to S2130 of FIG. 21, and an entropy encoding unit of the encoding device of FIG. 22 can perform step S2140 of FIG. 21. Also, for example, although not shown, a process of deriving predicted samples for the current block based on the CPMV can be performed by a prediction unit of the encoding device of FIG. 22, a process of deriving residual samples for the current block based on original samples and predicted samples for the current block can be performed by a subtraction unit of the encoding device of FIG. 22, a process of generating information about the residual for the current block based on the residual samples can be performed by a transformation unit of the encoding device of FIG. 22, and a process of encoding information about the residual can be performed by an entropy encoding unit of the encoding device of FIG. 22.

[0291] Figure 23 shows an outline of an image decoding method by a decoding device according to the present document. The method disclosed in Figure 23 may be performed by the decoding device disclosed in Figure 3. Specifically, for example, S2300 in Figure 23 may be performed by an entropy decoding unit of the decoding device, S2310 to S2350 may be performed by a prediction unit of the decoding device, and S2360 may be performed by an adder unit of the decoding device. Also, for example, although not shown, a step of obtaining information about the residual of a current block via a bitstream may be performed by an entropy decoding unit of the decoding device, and a step of deriving the residual sample for the current block based on the residual information may be performed by an inverse transform unit of the decoding device.

[0292] The decoding device obtains motion prediction information for a current block from a bitstream (S2300). The decoding device can obtain image information including the motion prediction information from the bitstream.

[0293] Also, for example, the motion prediction information may include information on Control Point Motion Vector Differences (CPMVD) for the Control Points (CPs) of the current block. That is, the motion prediction information may include information on CPMVD for each of the CPs of the current block.

[0294] Also, for example, the motion prediction information may include an affine MVP candidate index for the current block, which may indicate one of affine MVP candidates included in an affine motion vector predictor (MVP) candidate list for the current block.

[0295] The decoding apparatus constructs an affine motion vector predictor (MVP) candidate list for the current block (S2310). The decoding apparatus may construct an affine MVP candidate list including affine MVP candidates for the current block. The maximum number of affine MVP candidates in the affine MVP candidate list may be two.

[0296] Also, as an example, the affine MVP candidate list may include inherited affine MVP candidates. The decoding device may check whether the inherited affine MVP candidates of the current block are available, and if the inherited affine MVP candidates are available, the inherited affine MVP candidates may be derived. For example, the inherited affine MVP candidates may be derived based on neighboring blocks of the current block, and the maximum number of inherited affine MVP candidates may be two. The neighboring blocks may be checked for availability in a specific order, and the inherited affine MVP candidates may be derived based on the checked available neighboring blocks. That is, the neighboring blocks may be checked for availability in a specific order, and a first inherited affine MVP candidate may be derived based on the first checked available neighboring block, and a second inherited affine MVP candidate may be derived based on the second checked available neighboring block. The available neighboring blocks may be coded using an affine motion model, and the reference picture of the neighboring blocks may be the same as the reference picture of the current block. That is, the available neighboring blocks may be coded using an affine motion model (i.e., affine prediction is applied), and the reference picture may be the same as the reference picture of the current block. Specifically, the decoding device may derive a motion vector for the CP of the current block based on the affine motion model of the first checked available neighboring block, and derive the first inherited affine MVP candidate including the motion vector as a CPMVP candidate. Also, the decoding device may derive a motion vector for the CP of the current block based on the affine motion model of the second checked available neighboring block, and derive the second inherited affine MVP candidate including the motion vector as a CPMVP candidate. The affine motion model may be derived as shown in Equation 1 or 3 above.

[0297] In other words, the neighboring blocks may be checked to see if they satisfy a specific condition in a specific order, and the inherited affine MVP candidate may be derived based on the neighboring blocks that satisfy the checked specific condition. That is, the neighboring blocks may be checked to see if they satisfy the specific condition in a specific order, and a first inherited affine MVP candidate may be derived based on the neighboring block that satisfies the specific condition that is checked first, and a second inherited affine MVP candidate may be derived based on the neighboring block that satisfies the specific condition that is checked second. Specifically, the decoding device may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring block that satisfies the specific condition that is checked first, and may derive the first inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. Furthermore, the decoding device may derive a motion vector for a CP of the current block based on the affine motion model of the neighboring block that satisfies the specific condition that is checked second, and may derive the second inherited affine MVP candidate that includes the motion vector as a CPMVP candidate. The affine motion model can be derived as shown in Equation 1 or 3. Meanwhile, the specific condition may represent that the current block is coded using an affine motion model and the reference picture of the neighboring block is the same as the reference picture of the current block. That is, the neighboring block satisfying the specific condition may be coded using an affine motion model (i.e., affine prediction is applied) and the reference picture may be the same as the reference picture of the current block.

[0298] Here, for example, the peripheral blocks may include a left peripheral block, an upper peripheral block, a right upper corner peripheral block, a lower left corner peripheral block, and an upper left corner peripheral block of the current block, and the specific order may be from the left peripheral block to the lower left corner peripheral block, the upper peripheral block, the right upper corner peripheral block, and the upper left corner peripheral block.

[0299] Alternatively, for example, the peripheral blocks may include only the left peripheral block and the top peripheral block, in which case the specific order may be from the left peripheral block to the top peripheral block.

[0300] Alternatively, for example, the neighboring blocks may include the left neighboring block, and if the upper neighboring block is included in a current coding tree unit (CTU) including the current block, the neighboring blocks may further include the upper neighboring block. In this case, the specific order may be from the left neighboring block to the upper neighboring block. Also, if the upper neighboring block is not included in the current CTU, the neighboring blocks may not include the upper neighboring block. In this case, only the left neighboring block may be checked. That is, if the upper neighboring block of the current block is included in a current coding tree unit (CTU) including the current block, the upper neighboring block may be used for deriving the inherited affine MVP candidate, and if the upper neighboring block of the current block is not included in the current CTU, the upper neighboring block may not be used for deriving the inherited affine MVP candidate.

[0301] On the other hand, if the size is W×H and the x component and y component of the top-left sample position of the current block are 0, the lower-left corner peripheral block may be a block including a sample at coordinates (-1, H), the left peripheral block may be a block including a sample at coordinates (-1, H-1), the upper-right corner peripheral block may be a block including a sample at coordinates (W, -1), the upper peripheral block may be a block including a sample at coordinates (W-1, -1), and the upper-left corner peripheral block may be a block including a sample at coordinates (-1, -1). That is, the left peripheral block may be the leftmost peripheral block among the left peripheral blocks of the current block, and the upper peripheral block may be the leftmost peripheral block among the upper peripheral blocks of the current block.

[0302] Also, as an example, if a constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. The decoding apparatus may check whether a constructed affine MVP candidate for the current block is available, and if the constructed affine MVP candidate is available, the constructed affine MVP candidate may be derived. Also, for example, the constructed affine MVP candidate may be derived after the inherited affine MVP candidate is derived. If the number of derived affine MVP candidates (i.e., the inherited affine MVP candidates) is less than two and the constructed affine MVP candidate is available, the affine MVP candidate list may include the constructed affine MVP candidate. Here, the constructed affine MVP candidate may include candidate motion vectors for the CP. The constructed affine MVP candidate may be available if all of the candidate motion vectors are available.

[0303] For example, if a 4-affine motion model is applied to the current block, the CPs of the current block may include CP0 and CP1. If a candidate motion vector for CP0 is available and a candidate motion vector for CP1 is available, the constructed affine MVP candidate may be available, and the affine MVP candidate list may include the constructed affine MVP candidate. Here, CP0 may represent the upper left corner of the current block, and CP1 may represent the upper right corner of the current block.

[0304] The constructed affine MVP candidates may include a candidate motion vector for the CP0 and a candidate motion vector for the CP1. The candidate motion vector for the CP0 may be a motion vector of a first block, and the candidate motion vector for the CP1 may be a motion vector of a second block.

[0305] Furthermore, the first block may check neighboring blocks in the first group according to a first specific order, and the first identified reference picture may be the same as the reference picture of the current block. That is, the candidate motion vector for CP1 may be the motion vector of the block whose first identified reference picture is the same as the reference picture of the current block when neighboring blocks in the first group are checked according to a first order. The availability may indicate that the neighboring blocks exist and are coded using inter-prediction. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0306] Furthermore, the second block may check neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, the candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0307] On the other hand, if the size of the current block is W×H and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at (-1, -1) coordinates, the surrounding block B may be a block including a sample at (0, -1) coordinates, the surrounding block C may be a block including a sample at (-1, 0) coordinates, the surrounding block D may be a block including a sample at (W-1, -1) coordinates, and the surrounding block E may be a block including a sample at (W, -1) coordinates. That is, the peripheral block A may be the upper left corner peripheral block of the current block, the peripheral block B may be the upper peripheral block located on the leftmost side among the upper peripheral blocks of the current block, the peripheral block C may be the left peripheral block located on the topmost side among the left peripheral blocks of the current block, the peripheral block D may be the upper peripheral block located on the rightmost side among the upper peripheral blocks of the current block, and the peripheral block E may be the upper right corner peripheral block of the current block.

[0308] On the other hand, if at least one of the candidate motion vectors of CP0 and the candidate motion vectors of CP1 is unavailable, the constructed affine MVP candidate may be unavailable.

[0309] Alternatively, for example, if a 6-affine motion model is applied to the current block, the CPs of the current block may include CP0, CP1, and CP2. If a candidate motion vector for CP0 is available, a candidate motion vector for CP1 is available, and a candidate motion vector for CP2 is available, the constructed affine MVP candidates may be available, and the affine MVP candidate list may include the constructed affine MVP candidates. Here, CP0 may represent the upper left corner of the current block, CP1 may represent the upper right corner of the current block, and CP2 may represent the lower left corner of the current block.

[0310] The constructed affine MVP candidates may include a candidate motion vector for the CP0, a candidate motion vector for the CP1, and a candidate motion vector for the CP2. The candidate motion vector for the CP0 may be a motion vector of a first block, the candidate motion vector for the CP1 may be a motion vector of a second block, and the candidate motion vector for the CP2 may be a motion vector of a third block.

[0311] Furthermore, the first block may check neighboring blocks in the first group according to a first specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the first block in the first group is the same as the reference picture of the current block, the candidate motion vector for CP0 may be available. For example, the first group may include neighboring blocks A, B, and C, and the first specific order may be from neighboring block A to neighboring block B to neighboring block C.

[0312] Furthermore, the second block may check neighboring blocks in the second group according to a second specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the second block in the second group is the same as the reference picture of the current block, the candidate motion vector for CP1 may be available. For example, the second group may include neighboring blocks D and E, and the second specific order may be from neighboring block D to neighboring block E.

[0313] Furthermore, the third block may check neighboring blocks in the third group according to a third specific order, and the first identified reference picture may be the same as the reference picture of the current block. Here, if the reference picture of the third block in the third group is the same as the reference picture of the current block, the candidate motion vector for CP2 may be available. For example, the third group may include neighboring blocks F and G, and the third specific order may be from neighboring block F to neighboring block G.

[0314] On the other hand, if the size of the current block is W×H and the x component of the top-left sample position of the current block is 0 and the y component is 0, the surrounding block A may be a block including a sample at a (-1, -1) coordinate, the surrounding block B may be a block including a sample at a (0, -1) coordinate, the surrounding block C may be a block including a sample at a (-1, 0) coordinate, the surrounding block D may be a block including a sample at a (W-1, -1) coordinate, the surrounding block E may be a block including a sample at a (W, -1) coordinate, the surrounding block F may be a block including a sample at a (-1, H-1) coordinate, and the surrounding block G may be a block including a sample at a (-1, H) coordinate. That is, the peripheral block A may be the upper left corner peripheral block of the current block, the peripheral block B may be the upper leftmost peripheral block among the upper peripheral blocks of the current block, the peripheral block C may be the leftmost peripheral block among the left peripheral blocks of the current block, the peripheral block D may be the upper rightmost peripheral block among the upper peripheral blocks of the current block, the peripheral block E may be the upper right corner peripheral block of the current block, the peripheral block F may be the leftmost peripheral block among the left peripheral blocks of the current block, and the peripheral block G may be the lower left corner peripheral block of the current block.

[0315] On the other hand, if at least one of the candidate motion vectors of CP0, CP1, and CP2 is unavailable, the constructed affine MVP candidate may be unavailable.

[0316] Meanwhile, a pruning check between the inherited affine MVP candidate and the constructed affine MVP candidate may not be performed. The pruning check may represent a process of checking whether the constructed affine MVP candidate is the same as the inherited affine MVP candidate, and not deriving the constructed affine MVP candidate if they are the same.

[0317] The affine MVP candidate list can then be derived based on the following sequential steps:

[0318] For example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP0 is available, the decoding device may derive a first affine MVP candidate, where the first affine MVP candidate may be an affine MVP candidate that includes the motion vector for the CP0 as a candidate motion vector for the CP.

[0319] Also, for example, if the number of derived affine MVP candidates is less than two and a motion vector for the CP1 is available, the decoding device can derive a second affine MVP candidate, where the second affine MVP candidate may be an affine MVP candidate that includes the motion vector for the CP1 as a candidate motion vector for the CP.

[0320] Also, for example, if the number of derived affine MVP candidates is less than two and the motion vector for CP2 is available, the decoding device can derive a third affine MVP candidate, where the third affine MVP candidate can be an affine MVP candidate that includes the motion vector for CP2 as a candidate motion vector for the CP.

[0321] Furthermore, for example, if the number of derived affine MVP candidates is less than two, the decoding apparatus may derive a fourth affine MVP candidate including a temporal MVP derived based on a temporal neighboring block of the current block as a candidate motion vector for the CP. The temporal neighboring block may represent a collocated block in a collocated picture corresponding to the current block. The temporal MVP may be derived based on the motion vector of the temporal neighboring block.

[0322] Also, for example, if the number of derived affine MVP candidates is less than two, the decoding device may derive a fifth affine MVP candidate that includes a zero motion vector as a candidate motion vector for the CP. The zero motion vector may represent a motion vector with a value of 0.

[0323] The decoding apparatus derives Control Point Motion Vector Predictors (CPMVPs) for the Control Point (CP) of the current block based on the affine MVP candidate list (S2320).

[0324] The decoding device may select a specific affine MVP candidate from the affine MVP candidates included in the affine MVP candidate list and derive the selected affine MVP candidate using the CPMVP for the CP of the current block. For example, the decoding device may obtain the affine MVP candidate index for the current block from a bitstream and derive the affine MVP candidate pointed to by the affine MVP candidate index from the affine MVP candidate list using the CPMVP for the CP of the current block. Specifically, if the affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP1, the candidate motion vector for CP0 of the affine MVP candidate may be derived using the CPMVP for CP0, and the candidate motion vector for CP1 of the affine MVP candidate may be derived using the CPMVP for CP1. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived using the CPMVP of CP0, the candidate motion vector for CP1 of the affine MVP candidate can be derived using the CPMVP of CP1, and the candidate motion vector for CP2 of the affine MVP candidate can be derived using the CPMVP of CP2. Furthermore, when an affine MVP candidate includes a candidate motion vector for CP0 and a candidate motion vector for CP2, the candidate motion vector for CP0 of the affine MVP candidate can be derived using the CPMVP of CP0, and the candidate motion vector for CP2 of the affine MVP candidate can be derived using the CPMVP of CP2.

[0325] The decoding apparatus derives control point motion vector differences (CPMVDs) for the CPs of the current block based on the motion prediction information (S2330). The motion prediction information may include information about the CPMVDs for each of the CPs, and the decoding apparatus may derive the CPMVDs for each of the CPs of the current block based on the information about the CPMVDs for each of the CPs.

[0326] The decoding device derives control point motion vectors (CPMV) for the CP of the current block based on the CPMVP and the CPMVD (S2340). The decoding device can derive a CPMV for each CP based on the CPMVP and CPMVD for each CP. For example, the decoding device can derive a CPMV for the CP by adding the CPMVP and CPMVD for each CP.

[0327] The decoding apparatus derives prediction samples for the current block based on the CPMV (S2350). The decoding apparatus can derive sub-block-based or sample-based motion vectors for the current block based on the CPMV. That is, the decoding apparatus can derive motion vectors for each sub-block or each sample of the current block based on the CPMV. The sub-block-based or sample-based motion vectors can be derived based on Equation 1 or Equation 3 above. The motion vectors can be expressed as an affine motion vector field (MVF) or a motion vector array.

[0328] The decoding device may derive predicted samples for the current block based on the motion vector in sub-block units or sample units, derive a reference area in a reference picture based on the motion vector in sub-block units or sample units, and generate predicted samples for the current block based on reconstructed samples in the reference area.

[0329] The decoding device generates a reconstructed picture for the current block based on the derived predicted samples (S2360). The decoding device can generate a reconstructed picture for the current block based on the derived predicted samples. The decoding device can directly use the predicted samples as reconstructed samples depending on the prediction mode, or can generate reconstructed samples by adding residual samples to the predicted samples. If residual samples for the current block exist, the decoding device can obtain information about the residuals for the current block from the bitstream. The information about the residuals may include transform coefficients related to the residual samples. The decoding device can derive the residual samples (or a residual sample array) for the current block based on the residual information. The decoding device can generate reconstructed samples based on the predicted samples and the residual samples, and can derive a reconstructed block or picture based on the reconstructed samples. As described above, the decoding device can then apply in-loop filtering procedures, such as deblocking filtering and / or SAO procedures, to the reconstructed picture to improve subjective / objective image quality as needed.

[0330] FIG. 24 illustrates an outline of a decoding device that performs the image decoding method according to the present document. The method disclosed in FIG. 23 can be performed by the decoding device disclosed in FIG. 24. Specifically, for example, an entropy decoding unit of the decoding device of FIG. 24 can perform S2300 of FIG. 23, a prediction unit of the decoding device of FIG. 24 can perform S2310 to S2350 of FIG. 23, and an adder of the decoding device of FIG. 24 can perform S2360 of FIG. 23. Also, for example, although not shown, a process of obtaining image information including information about the residual of a current block via a bitstream can be performed by the entropy decoding unit of the decoding device of FIG. 24, and a process of deriving the residual sample for the current block based on the residual information can be performed by an inverse transform unit of the decoding device of FIG. 24.

[0331] According to the above-mentioned document, it is possible to increase the efficiency of image coding based on affine motion prediction.

[0332] In addition, according to this document, when deriving an affine MVP candidate list, a constructed affine MVP candidate can be added only if all candidate motion vectors for the CP of the constructed affine MVP candidate are available, thereby reducing the complexity of the process of deriving a constructed affine MVP candidate and the process of constructing an affine MVP candidate list and improving coding efficiency.

[0333] In addition, according to this document, when deriving an affine MVP candidate list, additional affine MVP candidates can be derived based on candidate motion vectors for CPs derived in the process of deriving constructed affine MVP candidates, thereby reducing the complexity of the process of constructing an affine MVP candidate list and improving coding efficiency.

[0334] In addition, according to this document, in the process of deriving an inherited affine MVP candidate, the inherited affine MVP candidate can be derived using the upper surrounding block only if the upper surrounding block is included in the current CTU, thereby reducing the amount of storage in the line buffer for affine prediction and minimizing hardware costs.

[0335] In the above-described embodiments, methods and the like are described based on flowcharts as a series of steps or blocks, but this document is not limited to the order of steps and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps in the flowcharts may be deleted without affecting the scope of this document.

[0336] The embodiments described in this document may be implemented on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the drawings may be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0337] In addition, the decoding device and encoding device to which the embodiments of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video device, an internet streaming service providing device, a three-dimensional (3D) video device, an image telephone video device, a vehicle terminal (e.g., a vehicle terminal, an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0338] In addition, a processing method to which an embodiment of this document is applied can be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to this document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes any type of storage device or distributed storage device on which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium can also include media implemented in the form of a carrier wave (e.g., transmission via the Internet). The bitstream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0339] Furthermore, the embodiments of the present document may be implemented in a computer program product by program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.

[0340] FIG. 25 exemplarily illustrates a structural diagram of a content streaming system to which the embodiments of this document are applied.

[0341] A content streaming system to which the embodiments of this document are applied can broadly include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0342] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0343] The bitstream can be generated by an encoding method or a bitstream generation method to which an embodiment of this document is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0344] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.

[0345] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.

[0346] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.

Claims

1. An image decoding method by a decoding device, comprising: obtaining motion prediction information for a current block from the bitstream; constructing an affine motion vector predictor (mvp) candidate list for the current block; deriving a control point motion vector predictor (CPMVP) for a control point (CP) of the current block based on the affine MVP candidate list; deriving a control point motion vector difference (CPMVD) for the CP of the current block based on the motion prediction information; deriving a control point motion vector (CPMV) for the CP of the current block based on the CPMVP and the CPMVD; deriving a predicted sample for the current block based on the CPMV; generating a reconstructed picture for the current block based on the derived prediction samples; The step of constructing the affine MVP candidate list comprises: checking whether inherited affine mvp candidates are available, The inherited affine mvp candidate is an affine mvp candidate that configures a motion vector derived from an affine model of a neighboring block of the current block as a candidate motion vector for the CP, the inherited affine mvp candidates are derived based on the availability of the inherited affine mvp candidates; The availability condition of the inherited affine MVP candidate is whether the reference picture of the neighboring block is the same as the reference picture of the current block; checking whether constructed affine mvp candidates are available, The constructed affine mvp candidate is a motion vector derived from a first neighboring block in a first neighboring block group of the current block as a candidate motion vector for CP0; a motion vector derived from a second neighboring block in a second neighboring block group of the current block as a candidate motion vector for CP1; a motion vector derived from a third neighboring block in a third neighboring block group of the current block is configured as a candidate motion vector for CP2; and the constructed affine mvp candidates are derived based on the availability of the constructed affine mvp candidates; a condition for the availability of the constructed affine MVP candidate is whether or not all of the motion vector derived from the first peripheral block, the motion vector derived from the second peripheral block, and the motion vector derived from the third peripheral block are available; a condition for using the motion vector derived from the first peripheral block is whether a reference picture of the first peripheral block is the same as the reference picture of the current block; a condition for using the motion vector derived from the second peripheral block is whether a reference picture of the second peripheral block is the same as the reference picture of the current block; a condition for using the motion vector derived from the third peripheral block is whether a reference picture of the third peripheral block is the same as the reference picture of the current block; deriving a first affine mvp candidate based on the number of derived affine mvp candidates including the inherited affine mvp candidate and the constructed affine mvp candidate being less than two, the first affine mvp candidate is an affine mvp candidate that includes a specific motion vector as a candidate motion vector for the CP, and a step in which the specific motion vector is an available motion vector among the motion vector derived from the first peripheral block, the motion vector derived from the second peripheral block, and the motion vector derived from the third peripheral block.

2. CP0 represents the top left corner position of the current block, CP1 represents the upper right corner position of the current block, The image decoding method of claim 1 , wherein the CP2 represents a bottom left corner position of the current block.

3. An image encoding method using an encoding device, comprising: constructing an affine motion vector predictor (mvp) candidate list for the current block; deriving a control point motion vector predictor (CPMVP) for a control point (CP) of the current block based on the affine MVP candidate list; deriving a control point motion vector (CPMV) for the CP of the current block; deriving a control point motion vector difference (CPMVD) for the CP of the current block based on the CPMVP and the CPMV; encoding motion prediction information including information about the CPMVD; The step of constructing the affine MVP candidate list comprises: checking whether inherited affine mvp candidates are available, The inherited affine mvp candidate is an affine mvp candidate that configures a motion vector derived from an affine model of a neighboring block of the current block as a candidate motion vector for the CP, the inherited affine mvp candidates are derived based on the availability of the inherited affine mvp candidates; The availability condition of the inherited affine MVP candidate is whether the reference picture of the neighboring block is the same as the reference picture of the current block; checking whether constructed affine mvp candidates are available, The constructed affine mvp candidate is a motion vector derived from a first neighboring block in a first neighboring block group of the current block as a candidate motion vector for CP0; a motion vector derived from a second neighboring block in a second neighboring block group of the current block as a candidate motion vector for CP1; a motion vector derived from a third neighboring block in a third neighboring block group of the current block is configured as a candidate motion vector for CP2; and the constructed affine mvp candidates are derived based on the availability of the constructed affine mvp candidates; a condition for the availability of the constructed affine MVP candidate is whether or not all of the motion vector derived from the first peripheral block, the motion vector derived from the second peripheral block, and the motion vector derived from the third peripheral block are available; a condition for using the motion vector derived from the first peripheral block is whether a reference picture of the first peripheral block is the same as the reference picture of the current block; a condition for using the motion vector derived from the second peripheral block is whether a reference picture of the second peripheral block is the same as the reference picture of the current block; a condition for using the motion vector derived from the third peripheral block is whether a reference picture of the third peripheral block is the same as the reference picture of the current block; deriving a first affine mvp candidate based on the number of derived affine mvp candidates including the inherited affine mvp candidate and the constructed affine mvp candidate being less than two, the first affine mvp candidate is an affine mvp candidate that includes a specific motion vector as a candidate motion vector for the CP, and a step in which the specific motion vector is an available motion vector among the motion vector derived from the first surrounding block, the motion vector derived from the second surrounding block, and the motion vector derived from the third surrounding block.

4. 1. A method of transmitting data for an image, comprising: obtaining a bitstream of image information, The bitstream comprises: constructing an affine motion vector predictor (mvp) candidate list for the current block; deriving a control point motion vector predictor (CPMVP) for a control point (CP) of the current block based on the affine MVP candidate list; deriving a control point motion vector (CPMV) for the CP of the current block; deriving a control point motion vector difference (CPMVD) for the CP of the current block based on the CPMVP and the CPMV; encoding motion prediction information including information about the CPMVD; transmitting the data including the bitstream of the image information; The step of constructing the affine MVP candidate list comprises: checking whether inherited affine mvp candidates are available, The inherited affine mvp candidate is an affine mvp candidate that configures a motion vector derived from an affine model of a neighboring block of the current block as a candidate motion vector for the CP, the inherited affine mvp candidates are derived based on the availability of the inherited affine mvp candidates; The availability condition of the inherited affine MVP candidate is whether the reference picture of the neighboring block is the same as the reference picture of the current block; checking whether constructed affine mvp candidates are available, The constructed affine mvp candidate is a motion vector derived from a first neighboring block in a first neighboring block group of the current block as a candidate motion vector for CP0; a motion vector derived from a second neighboring block in a second neighboring block group of the current block as a candidate motion vector for CP1; a motion vector derived from a third neighboring block in a third neighboring block group of the current block is configured as a candidate motion vector for CP2; and the constructed affine mvp candidates are derived based on the availability of the constructed affine mvp candidates; a condition for the availability of the constructed affine MVP candidate is whether or not all of the motion vector derived from the first peripheral block, the motion vector derived from the second peripheral block, and the motion vector derived from the third peripheral block are available; a condition for using the motion vector derived from the first peripheral block is whether a reference picture of the first peripheral block is the same as the reference picture of the current block; a condition for using the motion vector derived from the second peripheral block is whether a reference picture of the second peripheral block is the same as the reference picture of the current block; a condition for using the motion vector derived from the third peripheral block is whether a reference picture of the third peripheral block is the same as the reference picture of the current block; deriving a first affine mvp candidate based on the number of derived affine mvp candidates including the inherited affine mvp candidate and the constructed affine mvp candidate being less than two, the first affine mvp candidate is an affine mvp candidate that includes a specific motion vector as a candidate motion vector for the CP, a step in which the specific motion vector is an available motion vector among the motion vector derived from the first surrounding block, the motion vector derived from the second surrounding block, and the motion vector derived from the third surrounding block.

Citation Information

Patent Citations

  • Motion vector prediction for affine motion models in video coding

    US20180098063A1

  • Motion vector generation for affine motion model for video coding

    US20180192069A1

  • Method and apparatus for affine merge mode prediction for video coding system

    WO2017118409A1

  • Method and apparatus of video coding with affine motion compensation

    WO2017148345A1

  • Affine prediction for video coding

    WO2017156705A1