Method and apparatus for using inter-frame prediction information

By deriving combined inter prediction information in the video encoding and decoding device using processing units to deduce and perform inter prediction on the target block, the problem of low encoding efficiency of high resolution and high definition images in the prior art is solved, and more efficient video encoding and decoding is achieved.

CN111567045BActive Publication Date: 2025-05-09ELECTRONICS & TELECOMM RES INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201880076042.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-06-20
Filing Date
2018-10-10
Publication Date
2025-05-09
Estimated Expiration
2038-10-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize inter prediction information in video encoding and decoding, resulting in low encoding efficiency for high resolution and high definition images.

Method used

By introducing a processing unit in the encoding device and the decoding device, inter prediction is derived and performed on the target block using combined inter prediction information. The processing unit generates combined inter prediction information by combining inter prediction information of adjacent blocks of the target block.

Benefits of technology

Improves the efficiency of video encoding and decoding, especially when processing high-resolution and high-definition images, inter prediction can be performed more accurately and efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111567045B_ABST
    Figure CN111567045B_ABST
Patent Text Reader

Abstract

Disclosed are a method and apparatus for video decoding and a method and apparatus for video encoding. When performing video encoding or decoding, inter-frame prediction information of a block to be encoded or decoded may be derived, and the derived inter-frame prediction information may be used to perform inter-frame prediction of the block to be encoded or decoded. Combined inter-frame prediction information may be generated by combining multiple pieces of inter-frame prediction information, and the generated combined inter-frame prediction information may be added as a candidate to a list for inter-frame prediction. One of the candidates in the list may be selected for inter-frame prediction of the block to be encoded or decoded, and the selected candidate may be used to perform inter-frame prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following embodiments generally relate to a video decoding method and apparatus and a video encoding method and apparatus, and more particularly, to a method and apparatus for using inter-frame prediction information in video encoding and decoding. Background Art

[0002] As the information and communication industry continues to develop, broadcast services supporting high definition (HD) resolution have become popular throughout the world. Through this popularity, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.

[0003] In order to meet the user's demand for high definition, a large number of organizations have accelerated the development of next-generation imaging devices. In addition to high-definition TV (HDTV) and full high-definition (FHD) TV, user interest in UHD TV has also increased, where UHD TV has a resolution more than four times that of full high-definition (FHD) TV. As its interest increases, image encoding / decoding technology for images with higher resolution and higher definition continues to be required.

[0004] The image encoding / decoding apparatus and method may use inter-frame prediction technology, intra-frame prediction technology, entropy coding technology, etc., so as to perform encoding / decoding on high-resolution and high-definition images. The inter-frame prediction technology may be a technology for predicting the value of a pixel included in a target picture using a temporally previous picture and / or a temporally subsequent picture. The intra-frame prediction technology may be a technology for predicting the value of a pixel included in a target picture using information about the pixel in the target picture. The entropy coding technology may be a technology for assigning short codewords to frequently occurring symbols and assigning long codewords to rarely occurring symbols.

[0005] Various prediction methods have been developed to improve the efficiency and accuracy of intra prediction and / or inter prediction. The prediction efficiency may vary greatly depending on which prediction method among various applicable prediction methods is used to encode and / or decode a block. Summary of the invention

[0006] Technical issues

[0007] The embodiments are directed to providing an encoding device and method and a decoding device and method for performing inter-frame prediction on a target block.

[0008] The embodiments are directed to providing an encoding device and method and a decoding device and method for deriving combined inter-frame prediction information for a target block and performing inter-frame prediction using the derived combined inter-frame prediction information.

[0009] Solution

[0010] According to one aspect, a coding device is provided, comprising: a processing unit for deriving inter-frame prediction information for a target block and using the derived inter-frame prediction information to perform inter-frame prediction for the target block, wherein the processing unit uses the combined inter-frame prediction information to configure a list for the target block, and wherein the processing unit generates the combined inter-frame prediction information by combining two or more inter-frame prediction information of neighboring blocks of the target block.

[0011] According to another aspect, a decoding device is provided, comprising: a processing unit for deriving inter-frame prediction information for a target block and performing inter-frame prediction for the target block using the derived inter-frame prediction information, wherein the processing unit uses the combined inter-frame prediction information to configure a list for the target block, and wherein the processing unit generates the combined inter-frame prediction information by combining two or more inter-frame prediction information of neighboring blocks of the target block.

[0012] According to another aspect, a decoding method is provided, comprising: deriving inter-frame prediction information for a target block; and using the derived inter-frame prediction information to perform inter-frame prediction for the target block, wherein a list for the target block is configured using combined inter-frame prediction information, and wherein the combined inter-frame prediction information is generated by combining two or more inter-frame prediction information of neighboring blocks of the target block.

[0013] The inter prediction information may include at least one of an illumination compensation (IC) flag and an overlapped block motion compensation (OBMC) flag.

[0014] The list may be a merge list or an Advanced Motion Vector Prediction (AMVP) list.

[0015] The neighboring blocks may include spatial neighboring blocks and temporal neighboring blocks of the target block.

[0016] When inter prediction information of one of the neighboring blocks is unavailable, combined inter prediction information for the one neighboring block may be derived.

[0017] When inter prediction information of one of the neighboring blocks is not added to the list, combined inter prediction information derived for the one neighboring block may be added to the list.

[0018] The motion vector combining the inter prediction information may be a result of a formula using motion vectors of neighboring blocks.

[0019] The motion vector of the combined inter prediction information may be a weighted average of motion vectors of neighboring blocks.

[0020] The motion vector of the combined inter prediction information may be a result of weighted combination of motion vectors of neighboring blocks based on block size.

[0021] The motion vector combining the inter prediction information may be a result of a picture order count (POC) weighted combination of motion vectors of neighboring blocks.

[0022] The motion vector combining the inter prediction information may be a result of extrapolation-based combination of motion vectors of neighboring blocks.

[0023] First inter prediction information related to a left position of the specific block and second inter prediction information related to a right position of the specific block may be derived, and combined inter prediction information may be generated by combining the first inter prediction information with the second inter prediction information.

[0024] A scheme for configuring the list may be determined based on the shape of the target block.

[0025] A scheme for configuring the list may be determined based on a partition state of the target block.

[0026] A scheme for configuring the list may be determined based on the location of the target block.

[0027] The combined inter prediction information may be added to the list at a priority lower than that of pieces of inter prediction information of neighboring blocks.

[0028] The combined inter prediction information may be added to a position in the list after the inter prediction information of the spatially neighboring blocks and before the inter prediction information of the temporally neighboring blocks.

[0029] A scheme for configuring the list may be determined based on a depth of a target block.

[0030] Beneficial Effects

[0031] Provided are an encoding device and method and a decoding device and method for performing inter-frame prediction on a target block.

[0032] Provided are an encoding apparatus and method and a decoding apparatus and method for deriving combined inter-frame prediction information for a target block and performing inter-frame prediction using the derived combined inter-frame prediction information. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present disclosure is applied;

[0034] Figure 2 is a block diagram showing a configuration of an embodiment of a decoding device to which the present disclosure is applied;

[0035] Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded;

[0036] Figure 4 is a diagram showing a form of a prediction unit (PU) that a coding unit (CU) can include;

[0037] Figure 5 is a diagram showing a form of a transform unit (TU) that can be included in a CU;

[0038] Figure 6 is a diagram for explaining an embodiment of an intra prediction process;

[0039] Figure 7 is a diagram for explaining the positions of reference samples used in the intra prediction process;

[0040] Figure 8 is a diagram for explaining an embodiment of an inter-frame prediction process;

[0041] Fig. 9 shows spatial candidates according to an embodiment;

[0042] Fig.10 shows the order in which motion information of spatial candidates is added to a merge list according to an embodiment;

[0043] Fig.11 shows a transform and quantization process according to an example;

[0044] Fig.12 is a configuration diagram of an encoding device according to an embodiment;

[0045] Fig.13 is a configuration diagram of a decoding device according to an embodiment;

[0046] Fig.14 is a flowchart illustrating an inter-frame prediction method according to an embodiment;

[0047] Fig.15 showing spatial neighboring blocks of a target block according to an example;

[0048] Fig.16 showing temporal neighboring blocks of a target block according to an example;

[0049] Fig.17 shows generation of combined inter prediction information for an upper right neighboring block according to an example;

[0050] Fig.18 shows generation of combined inter prediction information of upper neighboring blocks according to an example;

[0051] Fig.19 shows generation of combined inter prediction information for neighboring blocks according to an example;

[0052] Fig. 20shows generation of inter prediction information for block AL according to an example;

[0053] Fig.21 shows generation of inter prediction information for block AR according to an example;

[0054] Fig. 22 The generation of inter prediction information of a target CU according to an example is shown;

[0055] Fig.23 Shows the case where a CU with the same width and height is divided vertically;

[0056] Fig.24 Shows the case where a CU with the same width and height is split horizontally;

[0057] Fig.25 Shows the case where a CU whose width is greater than its height is split vertically;

[0058] Fig.26 Shows the case where a CU whose height is greater than its width is split horizontally;

[0059] Fig. 27 showing sub-blocks of temporally neighboring blocks and sub-blocks of a target block according to an example;

[0060] Fig.28 showing spatial neighboring blocks of a target block and sub-blocks of the target block according to an example;

[0061] Fig.29 Derivation of inter-frame prediction information using bilateral matching according to an example is shown;

[0062] Fig.30 The use of template matching mode according to an example to derive inter-frame prediction information is shown;

[0063] Fig.31 illustrates the application of OBMC according to an example;

[0064] Fig.32 shows a sub-PU in ATMVP mode according to an example;

[0065] Fig.33 is a flowchart illustrating a target block prediction method and a bit stream generation method according to an embodiment; and

[0066] Fig.34 is a flowchart illustrating a target block prediction method using a bitstream according to an embodiment.

[0067] Best Mode for Carrying Out the Invention

[0068] The present invention can be variously changed and can have various embodiments, and specific embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the present invention to a specific disclosed form, and they include all changes, equivalent forms or modified forms included in the spirit and scope of the present invention.

[0069] The following exemplary embodiments will be described in detail with reference to the accompanying drawings showing specific embodiments. These embodiments are described so that those of ordinary skill in the art to which the present disclosure belongs can easily practice these embodiments. It should be noted that the various embodiments are different from each other, but do not need to be mutually exclusive. For example, the specific shapes, structures and characteristics described herein can be implemented as other embodiments without departing from the spirit and scope of multiple embodiments associated with one embodiment. In addition, it should be understood that the position or arrangement of each component in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiments. Therefore, the attached detailed description is not intended to limit the scope of the present disclosure, and the scope of the exemplary embodiments is limited only by the attached claims and their equivalents (as long as they are properly described).

[0070] In the drawings, like reference numerals are used to designate the same or similar functions in various aspects. The shapes, sizes, etc. of components in the drawings may be exaggerated to make the description clear.

[0071] Terms such as "first" and "second" may be used to describe various components, but the components are not limited by the terms. The terms are only used to distinguish one component from another component. For example, without departing from the scope of this specification, a first component may be referred to as a second component. Similarly, a second component may be referred to as a first component. The term "and / or" may include a combination of multiple related description items or any one of multiple related description items.

[0072] It will be understood that when a component is referred to as being "connected" or "coupled" to another component, the two components may be directly connected or coupled to each other, or an intermediate component may exist between the two components. It will be understood that when a component is referred to as being "directly connected or coupled", there are no intermediate components between the two components.

[0073] In addition, the components described in the embodiments are shown independently to represent different characteristic functions, but this does not mean that each component is formed by a separate hardware or software. That is, for the convenience of description, multiple components are arranged and included separately. For example, at least two components in the multiple components can be integrated into a single component. On the contrary, a component can be divided into multiple components. As long as it does not depart from the essence of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included in the scope of this specification.

[0074] Furthermore, it should be noted that, in the exemplary embodiments, the expression describing that components “include” specific components means that additional components may be included within the scope of practice or technical spirit of the exemplary embodiments, but does not exclude the existence of components other than the specific components.

[0075] The terms used in this specification are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context specifically indicates the opposite description. In this specification, it should be understood that terms such as "including" or "having" are only intended to indicate the presence of features, numbers, steps, operations, components, parts or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0076] The embodiments will be described in detail below with reference to the accompanying drawings so that a person skilled in the art can easily practice the embodiments. In the following description of the embodiments, a detailed description of well-known functions or configurations that are considered to obscure the main points of this specification will be omitted. In addition, the same reference numerals are used throughout the drawings to designate the same components, and repeated descriptions of the same components will be omitted.

[0077] Hereinafter, "image" may refer to a single picture constituting a video, or may refer to the video itself. For example, "encoding and / or decoding an image" may refer to "encoding and / or decoding a video", and may also refer to "encoding and / or decoding any one of a plurality of images constituting a video".

[0078] Hereinafter, the terms "video" and "moving picture" may be used to have the same meaning and may be used interchangeably with each other.

[0079] Hereinafter, the target image may be an encoding target image as a target to be encoded and / or a decoding target image as a target to be decoded. In addition, the target image may be an input image input to an encoding device or an input image input to a decoding device.

[0080] Hereinafter, the terms “image”, “picture”, “frame”, and “screen” may be used to have the same meaning and may be used interchangeably with each other.

[0081] Hereinafter, the target block may be an encoding target block (i.e., a target to be encoded) and / or a decoding target block (i.e., a target to be decoded). In addition, the target block may be a current block, i.e., a target to be currently encoded and / or decoded. Here, the terms "target block" and "current block" may be used to have the same meaning and may be used interchangeably with each other.

[0082] Hereinafter, the terms "block" and "unit" may be used to have the same meaning and may be used interchangeably with each other. Alternatively, a "block" may refer to a specific unit.

[0083] Hereinafter, the terms "region" and "segment" may be used interchangeably with each other.

[0084] Hereinafter, a specific signal may be a signal indicating a specific block. For example, an original signal may be a signal indicating a target block. A prediction signal may be a signal indicating a prediction block. A residual signal may be a signal indicating a residual block.

[0085] In the following embodiments, specific information, data, flags, elements, and attributes may have their respective values. The value "0" corresponding to each of the information, data, flags, elements, and attributes may indicate a logical false or a first predefined value. In other words, the value "0", false, logical false, and the first predefined value may be used interchangeably. The value "1" corresponding to each of the information, data, flags, elements, and attributes may indicate a logical true or a second predefined value. In other words, the value "1", true, logical true, and the second predefined value may be used interchangeably.

[0086] When a variable such as i or j is used to indicate a row, column, or index, the value i may be an integer 0 or an integer greater than 0, or may be an integer 1 or an integer greater than 1. In other words, in an embodiment, each of the row, column, and index may be counted starting from 0, or may be counted starting from 1.

[0087] Hereinafter, terms to be used in the embodiments will be described.

[0088] Encoder: An encoder refers to a device used to perform encoding.

[0089] Decoder: A decoder refers to a device used to perform decoding.

[0090] Unit: A "unit" may refer to a unit of image encoding and decoding. The terms "unit" and "block" may be used to have the same meaning and may be used interchangeably with each other.

[0091] – A “cell” may be an M×N array of samples. M and N may be positive integers respectively. The term “cell” may generally refer to a two-dimensional (2D) array of samples.

[0092] – In the encoding and decoding process of an image, a "unit" may be a region generated by partitioning one image. A single image may be partitioned into a plurality of units. Alternatively, one image may be partitioned into sub-parts, and a unit may represent each partitioned sub-part when encoding or decoding is performed on the partitioned sub-parts.

[0093] – During encoding and decoding of an image, predefined processing can be performed on each unit depending on the type of unit.

[0094] – According to functions, the unit type may be classified into a macro unit, a coding unit (CU), a prediction unit (PU), a residual unit, a transform unit (TU), etc. Alternatively, according to functions, the unit may represent a block, a macro block, a coding tree unit (CTU), a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc.

[0095] – The term “unit” may mean information including a luma component block, a chroma component block corresponding to the luma component block, and syntax elements for each block such that the unit is designated to be distinguished from the block.

[0096] – The size and shape of the unit may be implemented differently. In addition, the unit may have any of various sizes and shapes. Specifically, the shape of the unit may include not only a square but also a geometric shape that can be represented in two dimensions (2D), such as a rectangle, a trapezoid, a triangle, and a pentagon.

[0097] – In addition, the unit information may include one or more of a type of unit (indicating a coding unit, a prediction unit, a residual unit, or a transform unit), a size of the unit, a depth of the unit, an order of encoding and decoding of the unit, and the like.

[0098] – A cell can be partitioned into sub-cells, each sub-cell having a smaller size than the size of the associated cell.

[0099] – Depth: The depth may indicate the degree to which a unit is partitioned. Also, the depth may indicate the level at which a corresponding unit exists when the unit is represented in a tree structure.

[0100] – The cell partition information may include a depth indicating the depth of the cell. The depth may indicate the number of times the cell is partitioned and / or the extent to which the cell is partitioned.

[0101] – In a tree structure, the root node can be considered to have the smallest depth and the leaf nodes can be considered to have the largest depth.

[0102] – A single unit may be hierarchically partitioned into a plurality of subunits, and the single unit has depth information based on a tree structure. In other words, a unit and a subunit generated by partitioning the unit may correspond to a node and a subnode of the node, respectively. Each partitioned subunit may have a depth. Since the depth indicates the number of times a unit is partitioned and / or the degree to which a unit is partitioned, the partition information of the subunit may include information about the size of the subunit.

[0103] – In the tree structure, the top node may correspond to the initial node before partitioning. The top node may be referred to as a "root node". In addition, the root node may have the smallest depth value. Here, the depth of the top node may be level "0".

[0104] – Nodes at depth level “1” may represent cells generated when the initial cell is partitioned once. Nodes at depth level “2” may represent cells generated when the initial cell is partitioned twice.

[0105] – Leaf nodes at depth level “n” may represent cells generated when the initial cell is partitioned n times.

[0106] – A leaf node may be a bottom node that cannot be partitioned further. The depth of a leaf node may be a maximum level. For example, a predefined value for the maximum level may be 3.

[0107] – QT depth can represent the depth of four partitions. BT depth can represent the depth of two partitions. TT depth can represent the depth of three partitions.

[0108] – Sample: Sample can be the basic unit of building blocks. Available from 0 to 2 according to the bit depth (Bd). Bd- The value of 1 represents the sample point.

[0109] – Samples can be pixels or pixel values.

[0110] – In the following, the terms “pixel” and “sample” may be used to have the same meaning and may be used interchangeably with each other.

[0111] Coding Tree Unit (CTU): A CTU may be composed of a single luma component (Y) coding tree block and two chroma component (Cb, Cr) coding tree blocks associated with the luma component coding tree block. In addition, a CTU may represent information including the above blocks and syntax elements for each block.

[0112] – Each coding tree unit (CTU) may be partitioned using one or more partitioning methods such as quadtree (QT), binary tree (BT), and ternary tree (TT) to configure sub-units such as coding units, prediction units, and transform units.

[0113] – “CTU” may be used as a term to designate a pixel block that is a processing unit in image decoding and encoding processes (e.g., in the case of partitioning an input image).

[0114] Coding Tree Block (CTB): “CTB” may be used as a term to designate any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.

[0115] – Neighboring (adjacent) block: A neighboring block refers to a block that is adjacent to a target block. The term “neighboring block” may also refer to a reconstructed neighboring block.

[0116] Hereinafter, the terms “neighboring block” and “adjacent block” may be used to have the same meaning and may be used interchangeably with each other.

[0117] – Spatial neighboring blocks: Spatial neighboring blocks may be blocks that are spatially adjacent to the target block. Neighboring blocks may include spatial neighboring blocks.

[0118] – The target block and spatially neighboring blocks may be included in the target picture.

[0119] – A spatially neighboring block may be a block whose boundary touches the target block or a block located a predetermined distance from the target block.

[0120] – The spatially adjacent block may be a block adjacent to a vertex of the target block. Here, the block adjacent to a vertex of the target block may be a block vertically adjacent to a neighboring block horizontally adjacent to the target block or a block horizontally adjacent to a neighboring block vertically adjacent to the target block.

[0121] – Temporal neighboring blocks: Temporal neighboring blocks may be blocks that are temporally adjacent to the target block. Neighboring blocks may include temporal neighboring blocks.

[0122] – Temporally neighboring blocks may include co-located blocks (col blocks).

[0123] – The col block may be a block in a previously reconstructed co-located picture (col picture). The position of the col block in the col picture may correspond to the position of the target block in the target picture. The col picture may be a picture included in the reference picture list.

[0124] – The temporal neighboring blocks may be spatial neighboring blocks of the target block.

[0125] Prediction unit: A prediction unit may be a basic unit for prediction such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation.

[0126] – A single prediction unit can be divided into multiple partitions or sub-prediction units with smaller sizes. The multiple partitions can also be basic units when performing prediction or compensation. Partitions generated by dividing a prediction unit can also be prediction units.

[0127] Prediction unit partition: A prediction unit partition may be a shape into which a prediction unit is divided.

[0128] Reconstructed neighboring cells: Reconstructed neighboring cells may be cells around the target cell that have been decoded and reconstructed.

[0129] – The reconstructed neighboring cells can be cells that are spatially adjacent to the target cell or temporally adjacent to the target cell.

[0130] – The reconstructed spatially neighboring unit may be a unit included in the target picture that has been reconstructed through encoding and / or decoding.

[0131] – The reconstructed temporal neighboring unit may be a unit included in the reference picture and has been reconstructed by encoding and / or decoding. The position of the reconstructed temporal neighboring unit in the reference picture may be the same as the position of the target unit in the target picture, or may correspond to the position of the target unit in the target picture.

[0132] Parameter set: Parameter set can be header information in the structure of a bitstream. For example, parameter set can include sequence parameter set, picture parameter set, adaptation parameter set, etc.

[0133] Rate-distortion optimization: The encoding device may use rate-distortion optimization in order to provide high encoding efficiency by utilizing a combination of the following items: the size of a coding unit (CU), a prediction mode, the size of a prediction unit (PU), motion information, and the size of a transform unit (TU).

[0134] – The rate-distortion optimization scheme may calculate the rate-distortion cost of each combination to select the best combination from these combinations. The rate-distortion cost may be calculated using the following equation 1. In general, the combination that minimizes the rate-distortion cost may be selected as the best combination under the rate-distortion optimization scheme.

[0135] [Equation 1]

[0136]

[0137] – D may represent distortion. D may be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (ie, the mean square error).

[0138] – R may represent the rate, which may use relevant context information to represent the bit rate.

[0139] – R may include not only encoding parameter information such as prediction mode, motion information, and coding block flag, but also bits generated by encoding transform coefficients.

[0140] – The encoding device may perform processes such as inter-frame prediction and / or intra-frame prediction, transform, quantization, entropy coding, inverse quantization (dequantization), and inverse transform in order to calculate accurate D and R. These processes greatly increase the complexity of the encoding device.

[0141] – Bitstream: A bitstream may refer to a stream of bits comprising coded image information.

[0142] – Parameter set: Parameter set can be header information in the structure of bitstream.

[0143] The parameter set may include at least one of a video parameter set, a sequence parameter set, a picture parameter set, and an adaptation parameter set. In addition, the parameter set may include information about a slice header and information about a tile header.

[0144] Parsing: Parsing may be the decision on the value of a syntax element made by performing entropy decoding on the bitstream. Alternatively, the term "parsing" may refer to such entropy decoding itself.

[0145] Symbol: A symbol may be at least one of a syntax element, a coding parameter, and a transform coefficient of a coding target unit and / or a decoding target unit. In addition, a symbol may be a target of entropy coding or a result of entropy decoding.

[0146] Reference picture: A reference picture may be an image referenced by a unit in order to perform inter-frame prediction or motion compensation. Alternatively, a reference picture may be an image including a reference unit referenced by a target unit in order to perform inter-frame prediction or motion compensation.

[0147] Hereinafter, the terms 'reference picture' and 'reference image' may be used to have the same meaning and may be used interchangeably with each other.

[0148] Reference picture list: A reference picture list may be a list including one or more reference images used for inter prediction or motion compensation.

[0149] – The types of reference picture lists may include merged list (LC), list 0 (L0), list 1 (L1), list 2 (L3), list 3 (L3), etc.

[0150] – For inter prediction, one or more reference picture lists may be used.

[0151] Inter prediction indicator: The inter prediction indicator may indicate the inter prediction direction of the target unit. The inter prediction may be one of unidirectional prediction and bidirectional prediction. Optionally, the inter prediction indicator may indicate the number of reference images used to generate the prediction unit of the target unit. Optionally, the inter prediction indicator may indicate the number of prediction blocks used for inter prediction or motion compensation of the target unit.

[0152] Reference picture index: The reference picture index may be an index indicating a specific reference picture in a reference picture list.

[0153] Motion Vector (MV): A motion vector may be a 2D vector used for inter-frame prediction or motion compensation. A motion vector may represent an offset between an encoding target image / decoding target image and a reference image.

[0154] – For example, you can use (mv x , mv y ) to represent MV. x Can indicate horizontal component, mv y May indicate the vertical component.

[0155] – Search range: The search range may be a 2D area where a search for an MV is performed during inter prediction. For example, the size of the search range may be M×N. M and N may be positive integers, respectively.

[0156] Motion vector candidate: A motion vector candidate may be a block that is a prediction candidate when a motion vector is predicted or a motion vector of a block that is a prediction candidate.

[0157] – Motion vector candidates may be included in a motion vector candidate list.

[0158] Motion vector candidate list: A motion vector candidate list may be a list configured using one or more motion vector candidates.

[0159] Motion vector candidate index: The motion vector candidate index may be an indicator for indicating a motion vector candidate in a motion vector candidate list. Alternatively, the motion vector candidate index may be an index of a motion vector predictor.

[0160] Motion information: The motion information may be information including at least one of a reference picture list, a reference image, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index, as well as a motion vector, a reference picture index, and an inter prediction indicator.

[0161] Merge candidate list: The merge candidate list may be a list using a merge candidate configuration.

[0162] Merge candidate: The merge candidate may be a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-predictive merge candidate, a zero merge candidate, etc. The merge candidate may include motion information such as an inter prediction indicator, a reference picture index for each list, and a motion vector.

[0163] Merge index: The merge index may be an indicator for indicating a merge candidate in the merge candidate list.

[0164] – The merge index may indicate a reconstructed unit for deriving a merge candidate between a reconstructed unit that is spatially adjacent to the target unit and a reconstructed unit that is temporally adjacent to the target unit.

[0165] – The merge index may indicate at least one of a plurality of pieces of motion information of a merge candidate.

[0166] Transform unit: A transform unit may be a basic unit of residual signal encoding and / or residual signal decoding (such as transformation, inverse transformation, quantization, inverse quantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit may be partitioned into multiple transform units with smaller sizes.

[0167] Scaling: Scaling can be referred to as the process of multiplying the factors by the levels of the transform coefficients.

[0168] – Transform coefficients may be generated as a result of scaling the transform coefficient levels. Scaling may also be referred to as "inverse quantization".

[0169] Quantization parameter (QP): A quantization parameter may be a value used to generate transform coefficient levels for transform coefficients in quantization. Alternatively, a quantization parameter may also be a value used to generate transform coefficients by scaling the transform coefficient levels in inverse quantization. Alternatively, a quantization parameter may be a value mapped to a quantization step size.

[0170] Delta quantization parameter: Delta quantization parameter is the difference between the quantization parameter of the encoding / decoding target unit and the predicted quantization parameter.

[0171] Scanning: Scanning may refer to a method of arranging the order of coefficients in a cell, block, or matrix. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array may be referred to as "scanning". Alternatively, a method for arranging a 1D array in the form of a 2D array may also be referred to as "scanning" or "inverse scanning".

[0172] Transform coefficient: The transform coefficient may be a coefficient value generated when the encoding device performs transformation. Alternatively, the transform coefficient may be a coefficient value generated when the decoding device performs at least one of entropy decoding and inverse quantization.

[0173] – The level at which quantization is applied to the transform coefficients or the quantization of the residual signal or the level of quantized transform coefficients may also be included in the meaning of the term "transform coefficient".

[0174] Quantization level: The quantization level may be a value generated when the encoding device performs quantization on the transform coefficient or the residual signal. Alternatively, the quantization level may be a value that is a target of inverse quantization when the decoding device performs inverse quantization.

[0175] – The quantized transform coefficient level as a result of transformation and quantization may also be included in the meaning of the quantization level.

[0176] Non-zero transform coefficient: A non-zero transform coefficient may be a transform coefficient having a value other than 0, or may be a transform coefficient level having a value other than 0. Alternatively, a non-zero transform coefficient may be a transform coefficient having a value with a magnitude other than 0, or may be a transform coefficient level having a value with a magnitude other than 0.

[0177] Quantization matrix: A quantization matrix may be a matrix used in a quantization or inverse quantization process in order to improve the subjective or objective image quality of an image. A quantization matrix may also be referred to as a "scaling list".

[0178] Quantization matrix coefficient: A quantization matrix coefficient can be each element in the quantization matrix. A quantization matrix coefficient can also be referred to as a "matrix coefficient".

[0179] Default matrix: The default matrix may be a quantization matrix predefined by an encoding device and a decoding device.

[0180] Non-default matrix: A non-default matrix may be a quantization matrix that is not pre-defined by the encoding device and the decoding device. The non-default matrix may be signaled by the encoding device to the decoding device.

[0181] Signaling: Signaling may indicate that information is sent from an encoding device to a decoding device. Alternatively, signaling may mean that the information is included in a bitstream or a storage medium. Information signaled by an encoding device may be used by a decoding device.

[0182] Figure 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present disclosure is applied.

[0183] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images (pictures). The encoding device 100 may sequentially encode one or more images of a video.

[0184] Reference Figure 1 , the encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization (inverse quantization) unit 160, an inverse transform unit 170, an adder 175, a filter unit 180 and a reference picture buffer 190.

[0185] The encoding apparatus 100 may perform encoding on a target image using an intra mode and / or an inter mode.

[0186] In addition, the encoding apparatus 100 may generate a bitstream including information about encoding by encoding the target image, and may output the generated bitstream. The generated bitstream may be stored in a computer-readable storage medium and may be streamed through a wireless / wired transmission medium.

[0187] When the intra mode is used as the prediction mode, the switch 115 may switch to the intra mode. When the inter mode is used as the prediction mode, the switch 115 may switch to the inter mode.

[0188] The encoding apparatus 100 may generate a prediction block of the target block. In addition, after having generated the prediction block, the encoding apparatus 100 may encode a residual between the target block and the prediction block.

[0189] When the prediction mode is the intra mode, the intra prediction unit 120 may use pixels of a previously encoded / decoded neighboring block around the target block as reference samples. The intra prediction unit 120 may perform spatial prediction on the target block using the reference samples, and may generate prediction samples for the target block via spatial prediction.

[0190] The inter prediction unit 110 may include a motion prediction unit and a motion compensation unit.

[0191] When the prediction mode is the inter mode, the motion prediction unit may search the reference image for a region that best matches the target block during the motion prediction process, and may derive a motion vector for the target block and the found region based on the found region.

[0192] The reference image may be stored in the reference picture buffer 190. More specifically, when encoding and / or decoding of the reference image has been processed, the reference image may be stored in the reference picture buffer 190.

[0193] The motion compensation unit may generate a prediction block for the target block by performing motion compensation using a motion vector. Here, the motion vector may be a two-dimensional (2D) vector for inter-frame prediction. In addition, the motion vector may represent an offset between a target image and a reference image.

[0194] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit may generate a prediction block by applying an interpolation filter to a partial area of ​​the reference image. In order to perform inter prediction or motion compensation, it may be determined which mode of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to a method for predicting and compensating the motion of a PU included in the CU based on the CU, and inter prediction or motion compensation may be performed according to the mode.

[0195] The subtractor 125 may generate a residual block, which is a difference between the target block and the prediction block. The residual block may also be referred to as a "residual signal".

[0196] The residual signal may be the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming or quantizing the difference between the original signal and the predicted signal or a signal generated by transforming and quantizing the difference. The residual block may be a residual signal for a block unit.

[0197] The transform unit 130 may generate a transform coefficient by transforming the residual block, and may output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by transforming the residual block.

[0198] When the transform skip mode is used, the transform unit 130 may omit the operation of transforming the residual block.

[0199] By performing quantization on the transform coefficient, a quantized transform coefficient level or a quantized level may be generated. Hereinafter, in an embodiment, each of the quantized transform coefficient level and the quantized level may also be referred to as a 'transform coefficient'.

[0200] The quantization unit 140 may generate a quantized transform coefficient level or a quantized level by quantizing the transform coefficient according to the quantization parameter. The quantization unit 140 may output the generated quantized transform coefficient level or the quantized level. In this case, the quantization unit 140 may quantize the transform coefficient using a quantization matrix.

[0201] The entropy encoding unit 150 may generate a bitstream by performing entropy encoding based on probability distribution based on the value calculated by the quantization unit 140 and / or the encoding parameter value calculated during the encoding process. The entropy encoding unit 150 may output the generated bitstream.

[0202] The entropy encoding unit 150 may perform entropy encoding on information about pixels of an image and information required for decoding the image. For example, the information required for decoding the image may include syntax elements and the like.

[0203] The coding parameters may be information required for encoding and / or decoding. The coding parameters may include information encoded by the encoding device 100 and transmitted from the encoding device 100 to the decoding device, and may also include information derived during the encoding or decoding process. For example, the information transmitted to the decoding device may include syntax elements.

[0204] For example, the coding parameters may include values ​​or statistical information such as prediction mode, motion vector, reference picture index, coding block pattern, presence or absence of residual signal, transform coefficient, quantized transform coefficient, quantization parameter, block size, and block partition information. The prediction mode may be an intra-frame prediction mode or an inter-frame prediction mode.

[0205] The residual signal may represent the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the prediction signal.

[0206] When entropy coding is applied, fewer bits can be allocated to symbols that appear more frequently, and more bits can be allocated to symbols that appear less frequently. Since the symbols are represented by this allocation, the size of the bit string for the target symbol to be encoded can be reduced. Therefore, the compression performance of video coding can be improved by entropy coding.

[0207] In addition, in order to perform entropy coding, the entropy coding unit 150 may use a coding method such as exponential Golomb, context adaptive variable length coding (CAVLC), or context adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 may use a variable length coding / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 may derive a binarization method for a target symbol. In addition, the entropy coding unit 150 may derive a probability model for a target symbol / binary bit. The entropy coding unit 150 may perform arithmetic coding using the derived binarization method, probability model, and context model.

[0208] The entropy encoding unit 150 may transform coefficients in a 2D block form into a 1D vector form through a transform coefficient scanning method in order to encode transform coefficient levels.

[0209] The coding coefficient may include not only information (or flag or index) such as syntax elements encoded by the encoding device and transmitted by the encoding device to the decoding device in a signal, but also information derived in the encoding or decoding process. In addition, the coding parameter may include information required for encoding or decoding an image.For example, the encoding parameters may include at least one of the following items or a combination of the following items: the size of the unit / block, the depth of the unit / block, the partition information of the unit / block, the partition structure of the unit / block, information indicating whether the unit / block is partitioned in a quadtree structure, information indicating whether the unit / block is partitioned in a binary tree (BT) structure, the partition direction of the binary tree structure (horizontal or vertical), the partition form of the binary tree structure (symmetric partitioning or asymmetric partitioning), information indicating whether the unit / block is partitioned in a ternary tree structure, the partition direction of the ternary tree structure (horizontal or vertical), the prediction scheme (intra-frame prediction or inter-frame prediction), the intra-frame prediction mode / direction, the reference sample filtering method, the prediction block filtering method method, prediction block boundary filtering method, filter taps for filtering, filter coefficients for filtering, inter prediction mode, motion information, motion vector, reference picture index, inter prediction direction, inter prediction indicator, reference picture list, reference image, motion vector predictor, motion vector prediction candidate, motion vector candidate list, information indicating whether merge mode is used, merge candidate, merge candidate list, information indicating whether skip mode is used, type of interpolation filter, taps of interpolation filter, filter coefficients of interpolation filter, size of motion vector, accuracy of motion vector representation, transform type, transform size, information indicating whether primary transform is used, information indicating whether additional (secondary) transform is used information indicating whether a residual signal is used, a first transform index, a second transform index, information indicating whether a residual signal exists or not, a coding block pattern, a coding block flag, a quantization parameter, a quantization matrix, information about an intra-loop filter, information indicating whether an intra-loop filter is applied, coefficients of the intra-loop filter, taps of the intra-loop filter, a shape / form of the intra-loop filter, information indicating whether a deblocking filter is applied, coefficients of the deblocking filter, taps of the deblocking filter, deblocking filter strength, shape / form of the deblocking filter, information indicating whether an adaptive sample offset is applied, a value of an adaptive sample offset, a category of an adaptive sample offset, a value indicating whether an adaptive loop filter is applied applied information, coefficients of an adaptive loop filter, taps of an adaptive loop filter, shape / form of an adaptive loop filter, a binarization / debinarization method, a context model, a context model decision method, a context model update method, information indicating whether a normal mode is performed, information indicating whether a bypass mode is performed, context bits, bypass bits, transform coefficients, transform coefficient levels, transform coefficient level scanning methods, image display / output order, slice identification information, slice type, slice partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth, information about a luminance signal, and information about a chrominance signal.

[0210] Here, transmitting a flag or an index with a signal may indicate that the encoding device 100 includes an entropy-coded flag or an entropy-coded index generated by performing entropy encoding on the flag or the index in the bitstream, and may indicate that the decoding device 200 obtains the flag or the index by performing entropy decoding on the entropy-coded flag or the entropy-coded index extracted from the bitstream.

[0211] Since the encoding apparatus 100 performs encoding via inter-frame prediction, the encoded target image can be used as a reference image for another image to be subsequently processed. Therefore, the encoding apparatus 100 may reconstruct or decode the encoded target image and store the reconstructed or decoded image as a reference image in the reference picture buffer 190. For decoding, inverse quantization and inverse transformation of the encoded target image may be performed.

[0212] The quantized level may be dequantized by the dequantization unit 160, and may be inversely transformed by the inverse transform unit 170. The dequantized and / or inversely transformed coefficient may be added to the prediction block by the adder 175. The dequantized and / or inversely transformed coefficient and the prediction block are added, and then a reconstructed block may be generated. Here, the dequantized and / or inversely transformed coefficient may represent a coefficient on which one or more of dequantization and inverse transformation are performed, and may also represent a reconstructed residual block.

[0213] The reconstructed block may be filtered by the filter unit 180. The filter unit 180 may apply one or more filters of a deblocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF) to the reconstructed block or the reconstructed picture. The filter unit 180 may also be referred to as a "loop filter".

[0214] The deblocking filter can eliminate block distortion that occurs at the boundary between blocks. In order to determine whether to apply the deblocking filter, the number of columns or rows of pixels included in the block and including the number of columns or rows based on which the determination of whether to apply the deblocking filter to the target block is based can be determined. When the deblocking filter is applied to the target block, the filter applied can be different depending on the strength of the deblocking filter required. In other words, among the different filters, the filter determined in consideration of the strength of the deblocking filter can be applied to the target block.

[0215] SAO may add an appropriate offset to a pixel value in order to compensate for a coding error. SAO may perform correction on a pixel-based basis on an image to which deblocking is applied, wherein the correction uses an offset of a difference between an original image and an image to which deblocking is applied. A method for dividing pixels included in an image into a certain number of regions, determining a region to which an offset is applied among the divided regions, and applying the offset to the determined region may be used, and a method for applying an offset in consideration of edge information of each pixel may also be used.

[0216] The ALF may perform filtering based on a value obtained by comparing a reconstructed image with an original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to the group may be determined, and filtering may be performed differently for each group. Information related to whether an adaptive loop filter is applied may be transmitted with a signal for each CU. The shape and filter coefficients of the ALF to be applied to each block may be different for each block.

[0217] The reconstructed block or the reconstructed image filtered by the filter unit 180 may be stored in the reference picture buffer 190. The reconstructed block filtered by the filter unit 180 may be a part of a reference picture. In other words, the reference picture may be a reconstructed picture composed of the reconstructed blocks filtered by the filter unit 180. The stored reference picture may then be used for inter prediction.

[0218] Figure 2 is a block diagram showing a configuration of an embodiment of a decoding device to which the present disclosure is applied.

[0219] The decoding device 200 may be a decoder, a video decoding device, or an image decoding device.

[0220] Reference Figure 2 , the decoding apparatus 200 may include an entropy decoding unit 210 , an inverse quantization (dequantization) unit 220 , an inverse transform unit 230 , an intra prediction unit 240 , an inter prediction unit 250 , an adder 255 , a filter unit 260 , and a reference picture buffer 270 .

[0221] The decoding apparatus 200 may receive a bitstream output from the encoding apparatus 100. The decoding apparatus 200 may receive a bitstream stored in a computer-readable storage medium, and may receive a bitstream streamed through a wired / wireless transmission medium.

[0222] The decoding apparatus 200 may perform decoding on a bitstream in an intra mode and / or an inter mode. In addition, the decoding apparatus 200 may generate a reconstructed image or a decoded image through decoding, and may output the reconstructed image or the decoded image.

[0223] For example, the switcher can be used to switch to the intra-frame mode or the inter-frame mode based on the prediction mode for decoding. When the prediction mode for decoding is the intra-frame mode, the switcher can be operated to switch to the intra-frame mode. When the prediction mode for decoding is the inter-frame mode, the switcher can be operated to switch to the inter-frame mode.

[0224] The decoding device 200 can obtain a reconstructed residual block by decoding the input bit stream, and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block as a decoding target by adding the reconstructed residual block to the prediction block.

[0225] The entropy decoding unit 210 may generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols may include quantized hierarchical format symbols. Here, the entropy decoding method may be similar to the entropy encoding method described above. That is, the entropy decoding method may be the inverse process of the entropy encoding method described above.

[0226] The quantized coefficients may be dequantized by the dequantization unit 220. The dequantization unit 220 may generate dequantized coefficients by performing dequantization on the quantized coefficients. In addition, the dequantized coefficients may be inversely transformed by the inverse transform unit 230. The inverse transform unit 230 may generate a reconstructed residual block by performing inverse transformation on the dequantized coefficients. As a result of performing dequantization and inverse transformation on the quantized coefficients, a reconstructed residual block may be generated. Here, when generating the reconstructed residual block, the dequantization unit 220 may apply a quantization matrix to the quantized coefficients.

[0227] When the intra mode is used, the intra prediction unit 240 may generate a prediction block by performing spatial prediction using pixel values ​​of previously decoded neighboring blocks around a target block.

[0228] The inter prediction unit 250 may include a motion compensation unit. Alternatively, the inter prediction unit 250 may be designated as a "motion compensation unit."

[0229] When the inter mode is used, the motion compensation unit may generate a prediction block by performing motion compensation using a motion vector and a reference image stored in the reference picture buffer 270 .

[0230] The motion compensation unit may apply an interpolation filter to a partial area of ​​a reference image when a motion vector has a value other than an integer, and may generate a prediction block using the reference image to which the interpolation filter is applied. In order to perform motion compensation, the motion compensation unit may determine which mode of a skip mode, a merge mode, an advanced motion vector prediction (AMVP) mode, and a current picture reference mode corresponds to a motion compensation method for a PU included in the CU based on the CU, and may perform motion compensation according to the determined mode.

[0231] The reconstructed residual block and the prediction block may be added to each other by the adder 255. The adder 255 may generate a reconstructed block by adding the reconstructed residual block and the prediction block.

[0232] The reconstructed block may be filtered by the filter unit 260. The filter unit 260 may apply at least one of a deblocking filter, an SAO filter, and an ALF to the reconstructed block or the reconstructed picture.

[0233] The reconstructed block filtered by the filter unit 260 may be stored in the reference picture buffer 270. The reconstructed block filtered by the filter unit 260 may be a part of the reference picture. In other words, the reference picture may be an image composed of the reconstructed blocks filtered by the filter unit 260. The stored reference image may then be used for inter prediction.

[0234] Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded.

[0235] Figure 3 An example where a single unit is partitioned into a plurality of sub-units may be schematically shown.

[0236] In order to efficiently partition an image, a coding unit (CU) may be used in encoding and decoding. The term "unit" may be used to collectively specify 1) a block including image samples and 2) a syntax element. For example, "partition of a unit" may mean "partition of a block corresponding to a unit".

[0237] CU may be used as a basic unit for image encoding / decoding. CU may be used as a unit to which a mode selected from intra mode and inter mode is applied in image encoding / decoding. In other words, in image encoding / decoding, it may be determined which mode of intra mode and inter mode will be applied to each CU.

[0238] Also, a CU may be a basic unit for predicting, transforming, quantizing, inversely transforming, dequantizing, and encoding / decoding a transform coefficient.

[0239] Reference Figure 3 , the image 300 may be sequentially partitioned into units corresponding to a maximum coding unit (LCU), and a partition structure of the image 300 may be determined according to the LCU. Here, the LCU may be used to have the same meaning as a coding tree unit (CTU).

[0240] Partitioning a unit may mean partitioning a block corresponding to the unit. Block partition information may include depth information about the depth of the unit. The depth information may indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned. A single unit may be hierarchically partitioned into sub-units, and the single unit has depth information based on a tree structure. Each partitioned sub-unit may have depth information. The depth information may be information indicating the size of a CU. Depth information may be stored for each CU. Each CU may have depth information.

[0241] The partition structure may represent the distribution of coding units (CUs) in the LCU 310 for efficiently encoding an image. Such distribution may be determined according to whether a single CU is to be partitioned into a plurality of CUs. The number of CUs generated by partitioning may be a positive integer of 2 or more, including 2, 4, 8, 16, and the like. Depending on the number of CUs generated by partitioning, the horizontal size and vertical size of each CU generated by partitioning may be smaller than the horizontal size and vertical size of the CU before partitioning.

[0242] Each partitioned CU may be recursively partitioned into four CUs in the same manner. Compared to at least one of the horizontal size and the vertical size of the CU before partitioning, at least one of the horizontal size and the vertical size of each partitioned CU may be reduced through recursive partitioning.

[0243] The partitioning of the CU may be performed recursively until a predefined depth or a predefined size. For example, the depth of the LCU may be 0, and the depth of the smallest coding unit (SCU) may be a predefined maximum depth. Here, as described above, the LCU may be a CU having a maximum coding unit size, and the SCU may be a CU having a minimum coding unit size.

[0244] Partitioning may begin at the LCU 310, and each time the horizontal and / or vertical dimensions of the CU are reduced by partitioning, the depth of the CU may increase by one.

[0245] For example, for each depth, a non-partitioned CU may have a size of 2N×2N. Also, in the case where the CU is partitioned, a CU of size 2N×2N may be partitioned into four CUs each of size N×N. Whenever the depth increases by 1, the value of N may be halved.

[0246] Reference Figure 3 , an LCU with a depth of 0 may have 64×64 pixels or a 64×64 block. 0 may be the minimum depth. An SCU with a depth of 3 may have 8×8 pixels or a 8×8 block. 3 may be the maximum depth. Here, a CU with a 64×64 block as an LCU may be represented by a depth of 0. A CU with a 32×32 block may be represented by a depth of 1. A CU with a 16×16 block may be represented by a depth of 2. A CU with an 8×8 block as an SCU may be represented by a depth of 3.

[0247] Information about whether the corresponding CU is partitioned can be represented by the partition information of the CU. The partition information can be 1-bit information. All CUs except the SCU may include partition information. For example, the value of the partition information of a non-partitioned CU may be 0. The value of the partition information of a partitioned CU may be 1.

[0248] For example, when a single CU is partitioned into four CUs, the horizontal size and vertical size of each of the four CUs generated by partitioning may be half the horizontal size and vertical size of the CU before partitioning. When a CU having a size of 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs may be 16×16. When a single CU is partitioned into four CUs, it may be considered that the CU has been partitioned in a quadtree structure.

[0249] For example, when a single CU is partitioned into two CUs, the horizontal size or vertical size of each of the two CUs generated by partitioning may be half the horizontal size or vertical size of the CU before partitioning. When a CU having a size of 32×32 is partitioned vertically into two CUs, the size of each of the two partitioned CUs may be 16×32. When a single CU is partitioned into two CUs, it may be considered that the CU has been partitioned in a binary tree structure.

[0250] Both quadtree partitioning and binary tree partitioning can be applied to Figure 3 LCU 310.

[0251] Figure 4 is a diagram illustrating a form of a prediction unit (PU) that a coding unit (CU) can include.

[0252] In a CU partitioned from an LCU, the CU that is no longer partitioned may be divided into one or more prediction units (PUs). This division is also referred to as "partitioning".

[0253] PU can be a basic unit for prediction. PU can be encoded and decoded in any one of skip mode, inter mode and intra mode. PU can be partitioned into various shapes according to each mode. For example, Figure 1 The target block described above refers to Figure 2 The target blocks described may all be PUs.

[0254] In skip mode, partitions may not exist in a CU.In skip mode, a 2N×2N mode 410 may be supported without partitioning, wherein in the 2N×2N mode 410, the size of the PU and the size of the CU are identical to each other.

[0255] In inter mode, there may be eight types of partition shapes in a CU. For example, in inter mode, 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440, and nR×2N mode 445 may be supported.

[0256] In intra mode, 2N×2N mode 410 and N×N mode 425 may be supported.

[0257] In 2N×2N mode 410, a PU of size 2N×2N may be encoded. A PU of size 2N×2N may represent a PU of the same size as a CU. For example, a PU of size 2N×2N may have a size of 64×64, 32×32, 16×16, or 8×8.

[0258] In N×N mode 425 , a PU of size N×N may be encoded.

[0259] For example, in intra prediction, when the size of a PU is 8×8, four partitioned PUs may be encoded. The size of each partitioned PU may be 4×4.

[0260] When a PU is encoded in intra mode, any one of multiple intra prediction modes may be used to encode the PU. For example, HEVC technology may provide 35 intra prediction modes, and a PU may be encoded in any one of the 35 intra prediction modes.

[0261] Which mode of the 2Nx2N mode 410 and the NxN mode 425 is to be used to encode the PU may be determined based on the rate-distortion cost.

[0262] The encoding device 100 may perform an encoding operation on a PU of size 2N×2N. Here, the encoding operation may be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, an optimal intra-prediction mode for a PU of size 2N×2N may be derived. The optimal intra-prediction mode may be an intra-prediction mode that has a minimum rate-distortion cost when encoding a PU of size 2N×2N among a plurality of intra-prediction modes that can be used by the encoding device 100.

[0263] In addition, the encoding device 100 may sequentially perform encoding operations on each PU obtained by performing N×N partitioning. Here, the encoding operation may be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, an optimal intra-prediction mode for a PU of size N×N may be derived. The optimal intra-prediction mode may be an intra-prediction mode that has a minimum rate-distortion cost when encoding a PU of size N×N among a plurality of intra-prediction modes that can be used by the encoding device 100.

[0264] The encoding apparatus 100 may determine which one of the PU having a size of 2N×2N and the PU having a size of N×N is to be encoded, based on a comparison between a rate-distortion cost of a PU having a size of 2N×2N and a rate-distortion cost of a PU having a size of N×N.

[0265] Figure 5 is a diagram illustrating a form of a transform unit (TU) that can be included in a CU.

[0266] A transform unit (TU) may be a basic unit used for processes such as transform, quantization, inverse transform, inverse quantization, entropy encoding, and entropy decoding in a CU. A TU may have a square or rectangular shape.

[0267] In the CU partitioned from the LCU, the CU that is no longer partitioned into a CU may be partitioned into one or more TUs. Here, the partition structure of the TU may be a quadtree structure. For example, Figure 5 As shown in , a single CU 510 may be partitioned one or more times according to a quadtree structure. Through such partitioning, a single CU 510 may be composed of TUs of various sizes.

[0268] In the encoding apparatus 100, a coding tree unit (CTU) having a size of 64×64 may be partitioned into a plurality of smaller CUs in a recursive quadtree structure. A single CU may be partitioned into four CUs having the same size. Each CU may be recursively divided and may have a quadtree structure.

[0269] A CU may have a given depth. When a CU is partitioned, a CU generated by partitioning may have a depth increased by 1 from the depth of the partitioned CU.

[0270] For example, the depth of the CU may have a value ranging from 0 to 3. According to the depth of the CU, the size of the CU may range from a size of 64×64 to a size of 8×8.

[0271] By recursively partitioning the CU, the best partitioning method that produces the minimum rate-distortion cost can be selected.

[0272] Figure 6 is a diagram for explaining an embodiment of an intra prediction process.

[0273] from Figure 6 An arrow extending radially from the center of the diagram in represents a prediction direction of the intra prediction mode. In addition, a number appearing near the arrow may represent an example of a mode value assigned to the intra prediction mode or a prediction direction of the intra prediction mode.

[0274] Intra-frame encoding and / or decoding may be performed using reference samples of blocks adjacent to the target block. The adjacent blocks may be adjacent reconstructed blocks. For example, intra-frame encoding and / or decoding may be performed using the values ​​of reference samples included in each adjacent reconstructed block or encoding parameters of the adjacent reconstructed blocks.

[0275] The encoding device 100 and / or the decoding device 200 may generate a prediction block by performing intra prediction on the target block based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may generate a prediction block for the target block by performing intra prediction based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.

[0276] The prediction block may be a block generated as a result of performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU.

[0277] The unit of the prediction block may have a size corresponding to at least one of a CU, a PU, and a TU. The prediction block may have a square shape having a size of 2N×2N or N×N. The size N×N may include a size of 4×4, 8×8, 16×16, 32×32, 64×64, etc.

[0278] Alternatively, the prediction block may be a square block of size 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc. or a rectangular block of size 2×8, 4×8, 2×16, 4×16, 8×16, etc.

[0279] Intra prediction may be performed considering the intra prediction mode for the target block. The number of intra prediction modes that the target block may have may be a predefined fixed value, and may be a value determined differently according to the properties of the prediction block. For example, the properties of the prediction block may include the size of the prediction block, the type of the prediction block, etc.

[0280] For example, regardless of the size of the prediction block, the number of intra prediction modes may be fixed to 35. Alternatively, the number of intra prediction modes may be 3, 5, 9, 17, 34, 35, or 36, for example.

[0281] The intra prediction mode can be a non-directional mode or a directional mode. Figure 6 As shown in , the intra prediction modes may include two non-directional modes and 33 directional modes.

[0282] The two non-directional modes may include a DC mode and a planar mode.

[0283] The direction pattern may be a pattern having a specific direction or a specific angle.

[0284] Each of the intra prediction modes may be represented by at least one of a mode number, a mode value, and a mode angle. The number of intra prediction modes may be M. The value of M may be 1 or greater. In other words, the number of intra prediction modes may be M, where M includes the number of non-directional modes and the number of directional modes.

[0285] The number of intra-frame prediction modes may be fixed to M regardless of the size of the block. Optionally, the number of intra-frame prediction modes may be different depending on the size of the block and / or the type of color component. For example, the number of prediction modes may be different depending on whether the color component is a luminance signal or a chrominance signal. For example, the larger the size of the block, the larger the number of intra-frame prediction modes. Optionally, the number of intra-frame prediction modes corresponding to the luminance component block may be greater than the number of intra-frame prediction modes corresponding to the chrominance component block.

[0286] For example, in a vertical mode with a mode value of 26, prediction may be performed in a vertical direction based on pixel values ​​of reference samples. For example, in a horizontal mode with a mode value of 10, prediction may be performed in a horizontal direction based on pixel values ​​of reference samples.

[0287] Even in a directional mode other than the above-described modes, the encoding apparatus 100 and the decoding apparatus 200 may perform intra prediction on a target unit using reference samples according to an angle corresponding to the directional mode.

[0288] An intra-frame prediction mode located to the right relative to the vertical mode may be referred to as a "vertical-right mode". An intra-frame prediction mode located below the horizontal mode may be referred to as a "horizontal-below mode". For example, in Figure 6 , an intra prediction mode whose mode value is one of 27, 28, 29, 30, 31, 32, 33, and 34 may be a vertical-right mode 613. An intra prediction mode whose mode value is one of 2, 3, 4, 5, 6, 7, 8, and 9 may be a horizontal-bottom mode 616.

[0289] The non-directional mode may include a DC mode and a planar mode. For example, the mode value of the DC mode may be 1. The mode value of the planar mode may be 0.

[0290] The directional mode may include an angular mode. Among the plurality of intra prediction modes, the remaining modes except the DC mode and the planar mode may be directional modes.

[0291] When the intra prediction mode is the DC mode, the prediction block may be generated based on an average value of pixel values ​​of a plurality of reference pixels. For example, the value of a pixel of the prediction block may be determined based on the average value of pixel values ​​of a plurality of reference pixels.

[0292] The number of intra prediction modes and the mode values ​​of each intra prediction mode described above are merely exemplary and may be defined differently according to embodiments, implementations and / or requirements.

[0293] In order to perform intra prediction on a target block, a step of checking whether a sample included in a reconstructed neighboring block can be used as a reference sample of the target block may be performed. When there is a sample that cannot be used as a reference sample of the target block among the samples in the neighboring block, a value generated by interpolation and / or duplication using at least one sample value among the samples included in the reconstructed neighboring block may replace the sample value of the sample that cannot be used as a reference sample. When the value generated by duplication and / or interpolation replaces the sample value of an existing sample, the sample may be used as a reference sample of the target block.

[0294] In the intra prediction, a filter may be applied to at least one of the reference samples and the prediction samples based on at least one of the intra prediction mode and the size of the target block.

[0295] When the intra prediction mode is the planar mode, the sample value of the predicted target block can be generated by using the weighted sum of the upper reference sample of the target block, the left reference sample of the target block, the upper right reference sample of the target block, and the lower left reference sample of the target block according to the position of the predicted target sample in the prediction block when generating the prediction block of the target block.

[0296] When the intra prediction mode is the DC mode, an average value of reference samples above the target block and reference samples on the left side of the target block may be used when generating a prediction block of the target block.

[0297] When the intra prediction mode is a directional mode, an upper reference sample, a left reference sample, an upper right reference sample, and / or a lower left reference sample of the target block may be used to generate a prediction block.

[0298] In order to generate the above-mentioned prediction samples, real number-based interpolation may be performed.

[0299] The intra prediction mode of the target block may perform prediction from intra prediction of neighboring blocks adjacent to the target block, and information used for the prediction may be entropy encoded / decoded.

[0300] For example, when the intra prediction modes of the target block and the neighboring block are identical to each other, a predefined flag may be used to signal that the intra prediction modes of the target block and the neighboring block are identical.

[0301] For example, an indicator indicating an intra prediction mode that is the same as the intra prediction mode of the target block among intra prediction modes of a plurality of neighboring blocks may be signaled.

[0302] When intra prediction modes of a target block and a neighboring block are different from each other, intra prediction mode information of the target block may be entropy encoded / decoded based on the intra prediction mode of the neighboring block.

[0303] Figure 7 is a diagram for explaining the positions of reference samples used in an intra prediction process.

[0304] Figure 7 The position of the reference sample points used for intra-frame prediction of the target block is shown. Figure 7 , the reconstructed reference samples used for intra-frame prediction of the target block may include a lower left reference sample 731 , a left reference sample 733 , an upper left corner reference sample 735 , an upper reference sample 737 , and an upper right reference sample 739 .

[0305] For example, the left reference sample 733 may represent a reconstructed reference pixel adjacent to the left side of the target block. The upper reference sample 737 may represent a reconstructed reference pixel adjacent to the top of the target block. The upper left corner reference sample 735 may represent a reconstructed reference pixel located at the upper left corner of the target block. The lower left reference sample 731 may represent a reference sample located below the left sample line among the samples located on the same line as the left sample line composed of the left reference sample 733. The upper right reference sample 739 may represent a reference sample located on the right side of the upper sample line among the samples located on the same line as the upper sample line composed of the upper reference sample 737.

[0306] When the size of the target block is N×N, the numbers of the lower left reference sample 731 , the left reference sample 733 , the upper reference sample 737 , and the upper right reference sample 739 may all be N.

[0307] By performing intra prediction on the target block, a prediction block may be generated. The process of generating the prediction block may include determining the values ​​of the pixels in the prediction block. The size of the target block and the prediction block may be the same.

[0308] The reference samples used for intra-frame prediction of the target block may change according to the intra-frame prediction mode of the target block. The direction of the intra-frame prediction mode may represent the dependency between the reference samples and the pixels of the prediction block. For example, the value of the specified reference sample may be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample and the one or more specified pixels in the prediction block may be samples and pixels located on a straight line along the direction of the intra-frame prediction mode. In other words, the value of the specified reference sample may be copied as the value of a pixel located in a direction opposite to the direction of the intra-frame prediction mode. Alternatively, the value of a pixel in the prediction block may be the value of a reference sample located in the direction of the intra-frame prediction mode relative to the position of the pixel.

[0309] In an example, when the intra prediction mode of the target block is a vertical mode with a mode value of 26, the upper reference sample 737 may be used for intra prediction. When the intra prediction mode is a vertical mode, the value of a pixel in the prediction block may be the value of a reference sample located vertically above the position of the pixel. Therefore, the upper reference sample 737 adjacent to the top of the target block may be used for intra prediction. In addition, the value of a pixel in a row of the prediction block may be the same as the value of the pixel of the upper reference sample 737.

[0310] In an example, when the intra prediction mode of the target block is a horizontal mode with a mode value of 10, the left reference sample 733 may be used for intra prediction. When the intra prediction mode is a horizontal mode, the value of a pixel in the prediction block may be a value of a reference sample located horizontally to the left of the position of the pixel. Therefore, the left reference sample 733 adjacent to the left side of the target block may be used for intra prediction. In addition, the value of a pixel in a column of the prediction block may be the same as the value of a pixel of the left reference sample 733.

[0311] In an example, when the mode value of the intra prediction mode of the current block is 18, at least some of the left reference samples 733, the upper left corner reference samples 735, and at least some of the upper reference samples 737 may be used for intra prediction. When the mode value of the intra prediction mode is 18, the value of a pixel in the prediction block may be a value of a reference sample diagonally located at the upper left corner of the pixel.

[0312] In addition, in a case where an intra prediction mode having a mode value of 27, 28, 29, 30, 31, 32, 33, or 34 is used, at least a portion of the upper right reference sample 739 may be used for intra prediction.

[0313] In addition, in the case where an intra prediction mode having a mode value of 2, 3, 4, 5, 6, 7, 8, or 9 is used, at least a portion of the lower left reference sample 731 may be used for intra prediction.

[0314] Also, in the case of an intra prediction mode in which the mode value is a value ranging from 11 to 25, the upper left corner reference sample 735 may be used for intra prediction.

[0315] The number of reference samples used to determine the pixel value of one pixel in the prediction block may be 1 or 2 or more.

[0316] As described above, the pixel value of the pixel in the prediction block may be determined according to the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode. When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are integer positions, the value of one reference sample indicated by the integer position may be used to determine the pixel value of the pixel in the prediction block.

[0317] When the position of a pixel and the position of a reference sample indicated by the direction of an intra prediction mode are not integer positions, an interpolated reference sample based on two reference samples closest to the position of the reference sample may be generated. The value of the interpolated reference sample may be used to determine a pixel value of a pixel in a prediction block. In other words, when the position of a pixel in a prediction block and the position of a reference sample indicated by the direction of an intra prediction mode indicate a position between two reference samples, an interpolated value based on the values ​​of the two samples may be generated.

[0318] The prediction block generated via prediction may be different from the original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.

[0319] Hereinafter, the terms “difference”, “error” and “residual” may be used to have the same meaning and may be used interchangeably with each other.

[0320] For example, in the case of directional intra prediction, the longer the distance between the pixels of the prediction block and the reference samples, the greater the prediction error that may occur. Such prediction error may lead to discontinuity between the generated prediction block and neighboring blocks.

[0321] In order to reduce the prediction error, a filtering operation for the prediction block may be used. The filtering operation may be configured to adaptively apply a filter to an area in the prediction block that is considered to have a large prediction error. For example, the area that is considered to have a large prediction error may be a boundary of the prediction block. In addition, the area that is considered to have a large prediction error in the prediction block may be different depending on the intra-frame prediction mode, and the characteristics of the filter may also be different depending on the intra-frame prediction mode.

[0322] Figure 8 is a diagram for explaining an embodiment of an inter-frame prediction process.

[0323] Figure 8 The rectangle shown in can represent an image (or picture). Figure 8 In FIG. 1 , an arrow may indicate a prediction direction. That is, each image may be encoded and / or decoded according to the prediction direction.

[0324] Images may be classified into intra-pictures (I pictures), unidirectional predictive pictures or predictive coded pictures (P pictures), and bidirectional predictive pictures or bidirectional predictive coded pictures (B pictures) according to encoding types. Each picture may be encoded according to its encoding type.

[0325] When a target image to be encoded is an I picture, the target image can be encoded using data contained in the image itself without performing inter-frame prediction with reference to other images. For example, the I picture can be encoded only via intra-frame prediction.

[0326] When the target image is a P picture, the target image may be encoded via inter prediction using a reference picture existing in one direction. Here, the one direction may be a forward direction or a backward direction.

[0327] When the target image is a B picture, the image may be encoded via inter-frame prediction using reference pictures existing in both directions, or may be encoded via inter-frame prediction using reference pictures existing in one of a forward direction and a backward direction. Here, the two directions may be a forward direction and a backward direction.

[0328] P pictures and B pictures encoded and / or decoded using reference pictures may be regarded as images using inter-frame prediction.

[0329] Hereinafter, inter prediction in the inter mode according to an embodiment will be described in detail.

[0330] Inter prediction may be performed using motion information.

[0331] In the inter mode, the encoding apparatus 100 may perform inter prediction and / or motion compensation on the target block. The decoding apparatus 200 may perform inter prediction and / or motion compensation corresponding to the inter prediction and / or motion compensation performed by the encoding apparatus 100 on the target block.

[0332] The motion information of the target block may be separately derived during inter prediction by the encoding apparatus 100 and the decoding apparatus 200. The motion information may be derived using motion information of a reconstructed neighboring block, motion information of a col block, and / or motion information of a block adjacent to the col block.

[0333] For example, the encoding apparatus 100 or the decoding apparatus 200 may perform prediction and / or motion compensation by using motion information of the spatial candidate and / or the temporal candidate as motion information of the target block. The target block may represent a PU and / or a PU partition.

[0334] The spatial candidate may be a reconstructed block that is spatially adjacent to the target block.

[0335] The temporal candidate may be a reconstructed block corresponding to the target block in a previously reconstructed co-located picture (col picture).

[0336] In inter-frame prediction, the encoding device 100 and the decoding device 200 can improve encoding efficiency and decoding efficiency by using motion information of spatial candidates and / or temporal candidates. The motion information of the spatial candidate may be referred to as "spatial motion information". The motion information of the temporal candidate may be referred to as "temporal motion information".

[0337] Hereinafter, the motion information of the spatial candidate may be the motion information of the PU including the spatial candidate. The motion information of the temporal candidate may be the motion information of the PU including the temporal candidate. The motion information of the candidate block may be the motion information of the PU including the candidate block.

[0338] Inter prediction may be performed using reference pictures.

[0339] The reference picture may be at least one of a picture before the target picture and a picture after the target picture. The reference picture may be an image used for prediction of the target block.

[0340] In inter prediction, a region in a reference picture may be specified using a reference picture index (or refIdx) indicating a reference picture, a motion vector to be described later, etc. Here, the region specified in the reference picture may indicate a reference block.

[0341] Inter prediction can select a reference picture, or select a reference block corresponding to a target block from the reference picture. In addition, inter prediction can use the selected reference block to generate a prediction block for the target block.

[0342] Motion information may be derived by each of the encoding apparatus 100 and the decoding apparatus 200 during inter prediction.

[0343] A spatial candidate may be a block that 1) exists in the target picture 2) has been previously reconstructed via encoding and / or decoding and 3) is adjacent to the target block or is located at a corner of the target block. Here, a "block located at a corner of a target block" may be a block that is vertically adjacent to a neighboring block that is horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighboring block that is vertically adjacent to the target block. In addition, a "block located at a corner of a target block" may have the same meaning as a "block adjacent to a corner of a target block." The meaning of a "block located at a corner of a target block" may be included in the meaning of a "block adjacent to a target block."

[0344] For example, the spatial candidate may be a reconstructed block located to the left of the target block, a reconstructed block located above the target block, a reconstructed block located at the lower left corner of the target block, a reconstructed block located at the upper right corner of the target block, or a target block located at the upper left corner of the target block.

[0345] Each of the encoding apparatus 100 and the decoding apparatus 200 may identify a block existing in a position spatially corresponding to the target block in the col picture. The position of the target block in the target picture and the position of the identified block in the col picture may correspond to each other.

[0346] Each of the encoding apparatus 100 and the decoding apparatus 200 may determine the col block existing at a predefined relevant position with respect to the identified block as a temporal candidate. The predefined relevant position may be a position existing inside and / or outside the identified block.

[0347] For example, the col block may include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block may be a block located at the coordinates (xP+nPSW, yP+nPSH). The second col block may be a block located at the coordinates (xP+(nPSW>>1), yP+(nPSH>>1)). When the first col block is not available, the second col block may be selectively used.

[0348] The motion vector of the target block may be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 may scale the motion vector of the col block. The scaled motion vector of the col block may be used as the motion vector of the target block. In addition, the motion vector of the running information of the temporal candidate stored in the list may be a scaled motion vector.

[0349] The ratio of the motion vector of the target block to the motion vector of the col block may be the same as the ratio of the first distance to the second distance. The first distance may be the distance between the reference picture and the target picture of the target block. The second distance may be the distance between the reference picture and the col picture of the col block.

[0350] The scheme for deriving motion information may change according to the inter prediction mode of the target block. For example, as the inter prediction mode applied to inter prediction, there may be an advanced motion vector predictor (AMVP) mode, a merge mode, a skip mode, a current picture reference mode, etc. The merge mode may also be referred to as a "motion merge mode". Each mode will be described in detail below.

[0351] 1) AMVP model

[0352] When the AMVP mode is used, the encoding device 100 may search for a similar block in a neighboring area of ​​the target block. The encoding device 100 may obtain a prediction block by performing prediction on the target block using motion information of the found similar block. The encoding device 100 may encode a residual block which is a difference between the target block and the prediction block.

[0353] 1-1) Create a list of predicted motion vector candidates

[0354] When the AMVP mode is used as the prediction mode, each of the encoding device 100 and the decoding device 200 may create a list of prediction motion vector candidates using a motion vector of a spatial candidate, a motion vector of a temporal candidate, and a zero vector. The prediction motion vector candidate list may include one or more prediction motion vector candidates. At least one of the motion vector of the spatial candidate, the motion vector of the temporal candidate, and the zero vector may be determined and used as the prediction motion vector candidate.

[0355] Hereinafter, the terms “prediction motion vector (candidate)” and “motion vector (candidate)” may be used to have the same meaning and may be used interchangeably with each other.

[0356] Hereinafter, the terms 'prediction motion vector candidate' and 'AMVP candidate' may be used to have the same meaning and may be used interchangeably with each other.

[0357] Hereinafter, the terms 'prediction motion vector candidate list' and 'AMVP candidate list' may be used to have the same meaning and may be used interchangeably with each other.

[0358] The spatial motion candidate may include a reconstructed spatial neighboring block. In other words, the motion vector of the reconstructed neighboring block may be referred to as a "spatial prediction motion vector candidate".

[0359] The temporal motion candidate may include the col block and blocks adjacent to the col block. In other words, the motion vector of the col block or the motion vector of the block adjacent to the col block may be referred to as a "temporal prediction motion vector candidate".

[0360] The zero vector may be the (0,0) motion vector.

[0361] The predicted motion vector candidate may be a motion vector predictor for predicting a motion vector. In addition, in the encoding apparatus 100, each predicted motion vector candidate may be an initial search position for a motion vector.

[0362] 1-2) Searching for motion vector using a list of predicted motion vector candidates

[0363] The encoding device 100 may determine a motion vector to be used for encoding a target block within a search range using the list of predicted motion vector candidates. In addition, the encoding device 100 may determine a predicted motion vector candidate to be used as a predicted motion vector of a target block among the predicted motion vector candidates present in the predicted motion vector candidate list.

[0364] The motion vector to be used for encoding the target block may be a motion vector that can be encoded at a minimum cost.

[0365] Also, the encoding apparatus 100 may determine whether to encode the target block using the AMVP mode.

[0366] 1-3) Transmission of inter-frame prediction information

[0367] The encoding apparatus 100 may generate a bitstream including inter prediction information required for inter prediction, and the decoding apparatus 200 may perform inter prediction on a target block using the inter prediction information of the bitstream.

[0368] The inter prediction information may include 1) mode information indicating whether AMVP is used, 2) a prediction motion vector index, 3) a motion vector difference (MVD), 4) a reference direction, and 5) a reference picture index.

[0369] Hereinafter, the terms "prediction motion vector index" and "AMVP index" may be used to have the same meaning and may be used interchangeably with each other. In addition, the inter prediction information may include a residual signal.

[0370] When the mode information indicates that the AMVP mode is used, the decoding apparatus 200 may acquire a prediction motion vector index, an MVD, a reference direction, and a reference picture index from a bitstream through entropy decoding.

[0371] The predicted motion vector index may indicate a predicted motion vector candidate to be used for predicting the target block among predicted motion vector candidates included in the predicted motion vector candidate list.

[0372] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information

[0373] The decoding apparatus 200 may derive a prediction motion vector candidate using the prediction motion vector candidate list, and may determine motion information of a target block based on the derived prediction motion vector candidate.

[0374] The decoding device 200 may determine a motion vector candidate for the target block among the predicted motion vector candidates included in the predicted motion vector candidate list using the predicted motion vector index. The decoding device 200 may select the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list as the predicted motion vector of the target block.

[0375] The motion vector actually used for inter-frame prediction of the target block may not match the predicted motion vector. In order to indicate the difference between the motion vector actually used for inter-frame prediction of the target block and the predicted motion vector, MVD may be used. The encoding apparatus 100 may derive a predicted motion vector similar to the motion vector actually used for inter-frame prediction of the target block so as to use as small an MVD as possible.

[0376] The MVD may be a difference between a motion vector of a target block and a predicted motion vector. The encoding apparatus 100 may calculate the MVD, and may entropy encode the MVD.

[0377] The MVD may be transmitted from the encoding device 100 to the decoding device 200 through a bitstream. The decoding device 200 may decode the received MVD. The decoding device 200 may derive a motion vector of the target block by summing the decoded MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 may be the sum of the entropy-decoded MVD and the motion vector candidate.

[0378] The reference direction may indicate a list of reference pictures to be used for prediction of the target block. For example, the reference direction may indicate one of reference picture list L0 and reference picture list L1.

[0379] The reference direction only indicates a reference picture list to be used for predicting a target block, and may not mean that the direction of the reference picture is limited to a forward direction or a backward direction. In other words, each of the reference picture lists L0 and L1 may include pictures in a forward direction and / or a backward direction.

[0380] The reference direction being unidirectional may mean that a single reference picture list is used. The reference direction being bidirectional may mean that two reference picture lists are used. In other words, the reference direction may indicate one of the following cases: a case where only reference picture list L0 is used, a case where only reference picture list L1 is used, and a case where two reference picture lists are used.

[0381] The reference picture index may indicate a reference picture to be used for predicting the target block among reference pictures in the reference picture list. The reference picture index may be entropy encoded by the encoding apparatus 100. The entropy encoded reference picture index may be signaled by the encoding apparatus 100 to the decoding apparatus 200 through a bitstream.

[0382] When two reference picture lists are used to predict the target block, a single reference picture index and a single motion vector may be used for each of the reference picture lists. In addition, when two reference picture lists are used to predict the target block, two prediction blocks may be specified for the target block. For example, an average or weighted sum of the two prediction blocks for the target block may be used to generate the (final) prediction block for the target block.

[0383] The motion vector of the target block may be derived by predicting the motion vector index, MVD, reference direction, and reference picture index.

[0384] The decoding apparatus 200 may generate a prediction block for the target block based on the derived motion vector and the reference picture index. For example, the prediction block may be a reference block indicated by the derived motion vector in the reference picture indicated by the reference picture index.

[0385] Since the predicted motion vector index and the MVD are encoded but the motion vector itself of the target block is not encoded, the number of bits transmitted from the encoding apparatus 100 to the decoding apparatus 200 may be reduced and encoding efficiency may be improved.

[0386] The motion information of the reconstructed neighboring blocks may be used for the target block. In a specific inter-frame prediction mode, the encoding device 100 may not encode the actual motion information of the target block separately. The motion information of the target block is not encoded, but additional information may be encoded, wherein the additional information enables the motion information of the target block to be derived using the motion information of the reconstructed neighboring blocks. Since the additional information is encoded, the number of bits sent to the decoding device 200 can be reduced, and the encoding efficiency can be improved.

[0387] For example, as an inter-prediction mode in which the motion information of the target block is not directly encoded, there may be a skip mode and / or a merge mode. Here, each of the encoding device 100 and the decoding device 200 may use an indicator and / or an index indicating a unit whose motion information is to be used as the motion information of the target unit among the reconstructed neighboring units.

[0388] 2) Merge mode

[0389] As a scheme for deriving motion information of a target block, there is merging. The term "merge" may mean merging the motions of multiple blocks. "Merge" may mean that the motion information of one block is also applied to other blocks. In other words, the merge mode may be a mode for deriving the motion information of the target block from the motion information of the neighboring blocks.

[0390] When using merge mode, the encoding device 100 may use the motion information of the spatial candidate and / or the motion information of the temporal candidate to predict the motion information of the target block. The spatial candidate may include a reconstructed spatially adjacent block that is spatially adjacent to the target block. Spatially adjacent blocks may include left adjacent blocks and upper adjacent blocks. Temporal candidates may include col blocks. The terms "spatial candidate" and "spatial merge candidate" may be used to have the same meaning and may be used interchangeably with each other. The terms "temporal candidate" and "temporal merge candidate" may be used to have the same meaning and may be used interchangeably with each other.

[0391] The encoding apparatus 100 may obtain a prediction block through prediction. The encoding apparatus 100 may encode a residual block which is a difference between a target block and a prediction block.

[0392] 2-1) Create a merge candidate list

[0393] When using the merge mode, each of the encoding device 100 and the decoding device 200 may create a merge candidate list using motion information of a spatial candidate and / or motion information of a temporal candidate. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may be unidirectional or bidirectional.

[0394] The merge candidate list may include a merge candidate. The merge candidate may be motion information. In other words, the merge candidate list may be a list storing multiple pieces of motion information.

[0395] The merge candidate may be motion information of multiple temporal candidates and / or spatial candidates. In addition, the merge candidate list may include new merge candidates generated by combining merge candidates already present in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple motion information previously present in the merge candidate list.

[0396] The merge candidate may be a specific mode for deriving inter-frame prediction information. The merge candidate may be information indicating a specific mode for deriving inter-frame prediction information. The inter-frame prediction information of the target block may be derived according to the specific mode indicated by the merge candidate. In addition, the specific mode may include a process for deriving a series of inter-frame prediction information. This specific mode may be an inter-frame prediction information derivation mode or a motion information derivation mode.

[0397] Inter prediction information of the target block may be derived according to a mode indicated by a merge candidate selected from among merge candidates in the merge candidate list by a merge index.

[0398] For example, the motion information derivation mode in the merge candidate list may be at least one of the following modes: 1) a motion information derivation mode for a sub-block unit; 2) an affine motion information derivation mode. In addition, the merge candidate list may include motion information of a zero vector. A zero vector may also be referred to as a "zero merge candidate".

[0399] In other words, the multiple motion information in the merge candidate list can be at least one of the following information: 1) motion information of spatial candidates, 2) motion information of temporal candidates, 3) motion information generated by combining multiple motion information previously existing in the merge candidate list, and 4) a zero vector.

[0400] The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an "inter prediction indicator". The reference direction may be unidirectional or bidirectional. A unidirectional reference direction may indicate L0 prediction or L1 prediction.

[0401] A merge candidate list may be created before performing prediction in merge mode.

[0402] The number of merge candidates in the merge candidate list may be predefined. Each of the encoding device 100 and the decoding device 200 may add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority so that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 may be made identical to each other using a predefined scheme and a predefined priority.

[0403] Merging may be applied on a CU or PU basis. When merging is performed on a CU or PU basis, the encoding device 100 may transmit a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merging is performed for each block partition, and 2) information about a block on which merging is to be performed among blocks that are spatial candidates and / or temporal candidates for a target block.

[0404] 2-2) Searching for motion vectors using the merge candidate list

[0405] The encoding device 100 may determine a merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using a merge candidate in the merge candidate list, and may generate a residual block for the merge candidate. The encoding device 100 may encode the target block using a merge candidate that generates the minimum cost in encoding the prediction and residual block.

[0406] Also, the encoding apparatus 100 may determine whether to encode the target block using the merge mode.

[0407] 2-3) Transmission of inter-frame prediction information

[0408] The encoding apparatus 100 may generate a bitstream including inter-frame prediction information required for inter-frame prediction. The encoding apparatus 100 may generate entropy-coded inter-frame prediction information by performing entropy encoding on the inter-frame prediction information, and may transmit the bitstream including the entropy-coded inter-frame prediction information to the decoding apparatus 200. The entropy-coded inter-frame prediction information may be signaled by the encoding apparatus 100 to the decoding apparatus 200 through a bitstream.

[0409] The decoding apparatus 200 may perform inter prediction on the target block using inter prediction information of the bitstream.

[0410] The inter prediction information may include 1) mode information indicating whether the merge mode is used and 2) a merge index.

[0411] In addition, the inter prediction information may include a residual signal.

[0412] The decoding apparatus 200 may acquire a merge index from a bitstream only when the mode information indicates that the merge mode is used.

[0413] The mode information may be a merge flag. The unit of the mode information may be a block. The information about the block may include the mode information, and the mode information may indicate whether the merge mode is applied to the block.

[0414] The merge index may indicate a merge candidate to be used for predicting the target block among the merge candidates included in the merge candidate list. Alternatively, the merge index may indicate a block to be merged with the target block among neighboring blocks spatially or temporally adjacent to the target block.

[0415] The encoding apparatus 100 may select a merge candidate having the highest encoding performance from among the merge candidates included in the merge candidate list, and may set a value of the merge index so that the merge index indicates the selected merge candidate.

[0416] 2-4) Inter-frame prediction using merge mode of inter-frame prediction information

[0417] The decoding apparatus 200 may perform prediction on the target block using a merge candidate indicated by a merge index among merge candidates included in the merge candidate list.

[0418] The motion vector of the target block may be specified by the motion vector of the merge candidate indicated by the merge index, the reference picture index, and the reference direction.

[0419] 3) Skip mode

[0420] The skip mode may be a mode in which the motion information of the spatial candidate or the motion information of the temporal candidate is applied to the target block without change. In addition, the skip mode may be a mode in which the residual signal is not used. In other words, when the skip mode is used, the reconstructed block may be a predicted block.

[0421] The difference between the merge mode and the skip mode is whether to transmit or use a residual signal. That is, the skip mode may be similar to the merge mode except that the residual signal is not transmitted or used.

[0422] When the skip mode is used, the encoding device 100 may transmit information about a block, among blocks that are spatial candidates or temporal candidates, whose motion information is to be used as the motion information of the target block, to the decoding device 200 through a bitstream. The encoding device 100 may generate entropy-coded information by performing entropy encoding on the information, and may signal the entropy-coded information to the decoding device 200 through a bitstream.

[0423] In addition, when the skip mode is used, the encoding device 100 may not transmit other syntax information (such as MVD) to the decoding device 200. For example, when the skip mode is used, the encoding device 100 may not signal a syntax element related to at least one of the MVC, the coded block flag, and the transform coefficient level to the decoding device 200.

[0424] 3-1) Create a merge candidate list

[0425] The merge candidate list may also be used in the skip mode. In other words, the merge candidate list may be used in both the merge mode and the skip mode. In this regard, the merge candidate list may also be referred to as a "skip candidate list" or a "merge / skip candidate list".

[0426] Optionally, the skip mode may use an additional candidate list different from the candidate list of the merge mode. In this case, in the following description, the merge candidate list and the merge candidate may be replaced with the skip candidate list and the skip candidate, respectively.

[0427] The merge candidate list may be created before performing prediction in skip mode.

[0428] 3-2) Searching for motion vectors using the merge candidate list

[0429] The encoding device 100 may determine a merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using a merge candidate in the merge candidate list. The encoding device 100 may encode the target block using a merge candidate that generates the minimum cost in prediction.

[0430] Also, the encoding apparatus 100 may determine whether to encode the target block using the skip mode.

[0431] 3-3) Transmission of inter-frame prediction information

[0432] The encoding apparatus 100 may generate a bitstream including inter prediction information required for inter prediction, and the decoding apparatus 200 may perform inter prediction on a target block using the inter prediction information of the bitstream.

[0433] The inter prediction information may include 1) mode information indicating whether the skip mode is used and 2) a skip index.

[0434] The skip index may be the same as the merge index described above.

[0435] When the skip mode is used, the target block may be encoded without using the residual signal. The inter prediction information may not include the residual signal. Alternatively, the bitstream may not include the residual signal.

[0436] The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the skip mode is used. As described above, the merge index and the skip index may be the same as each other. The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.

[0437] The skip index may indicate a merge candidate to be used for prediction of a target block among merge candidates included in the merge candidate list.

[0438] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information

[0439] The decoding apparatus 200 may perform prediction on the target block using a merge candidate indicated by a skip index among merge candidates included in the merge candidate list.

[0440] The motion vector of the target block may be specified by the motion vector of the merge candidate indicated by the skip index, the reference picture index, and the reference direction.

[0441] 4) Current picture reference mode

[0442] The current picture reference mode may denote a prediction mode that uses a previously reconstructed area in a target picture to which the target block belongs.

[0443] A vector for specifying a previously reconstructed region may be defined.A reference picture index of a target block may be used to determine whether the target block has been encoded in a current picture reference mode.

[0444] A flag or index indicating whether the target block is a block encoded in the current picture reference mode may be signaled by the encoding apparatus 100 to the decoding apparatus 200. Alternatively, whether the target block is a block encoded in the current picture reference mode may be inferred through a reference picture index of the target block.

[0445] When the target block is encoded in the current picture reference mode, the current picture may be added to a fixed position or an arbitrary position in the reference picture list for the target block.

[0446] For example, the fixed position may be a position where the reference picture index is 0 or a last position.

[0447] When a target picture is added to an arbitrary position in the reference picture list, an additional reference picture index indicating such arbitrary position may be signaled by the encoding apparatus 100 to the decoding apparatus 200 .

[0448] In the AMVP mode, merge mode, and skip mode described above, the index of the list may be used to specify motion information to be used for prediction of a target block among a plurality of pieces of motion information in the list.

[0449] In order to improve encoding efficiency, the encoding apparatus 100 may signal only the index of the element that generates the minimum cost in inter prediction of the target block among the elements in the list. The encoding apparatus 100 may encode the index and may signal the encoded index.

[0450] Therefore, it is necessary to be able to derive the above-described lists (i.e., the predicted motion vector candidate list and the merge candidate list) based on the same data using the same scheme by the encoding device 100 and the decoding device 200. Here, the same data may include a reconstructed picture and a reconstructed block. In addition, in order to specify an element using an index, the order of the elements in the list must be fixed.

[0451] Fig. 9 Spatial candidates according to an embodiment are shown.

[0452] exist Fig. 9 , the positions of the spatial candidates are shown.

[0453] The large block in the center of the figure may represent the target block. The five small blocks may represent spatial candidates.

[0454] The coordinates of the target block may be (xP, yP), and the size of the target block may be represented by (nPSW, nPSH).

[0455] The spatial candidate A0 may be a block adjacent to the lower left corner of the target block. A0 may be a block occupying a pixel located at coordinates (xP-1, yP+nPSH+1).

[0456] The spatial coordinate A1 may be a block adjacent to the left side of the target block. A1 may be the lowest block among the blocks adjacent to the left side of the target block. Alternatively, A1 may be a block adjacent to the top of A0. A1 may be a block occupying a pixel located at the coordinate (xP-1, yP+nPSH).

[0457] The spatial candidate B0 may be a block adjacent to the upper right corner of the target block. B0 may be a block occupying a pixel located at coordinates (xP+nPSW+1, yP-1).

[0458] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Alternatively, B1 may be a block adjacent to the left side of B0. B1 may be a block occupying a pixel located at coordinates (xP+nPSW, yP-1).

[0459] The spatial candidate B2 may be a block adjacent to the upper left corner of the target block. B2 may be a block occupying a pixel located at coordinates (xP-1, yP-1).

[0460] Determination of availability of spatial and temporal candidates

[0461] In order to include the motion information of the spatial candidate or the motion information of the temporal candidate in the list, it must be determined whether the motion information of the spatial candidate or the motion information of the temporal candidate is available.

[0462] Hereinafter, the candidate blocks may include spatial candidates and temporal candidates.

[0463] For example, the determination may be performed by sequentially applying the following steps 1) to 4) below.

[0464] Step 1) When the PU including the candidate block is located outside the boundary of the picture, the availability of the candidate block may be set to “false.” The expression “availability is set to false” may have the same meaning as “set to unavailable.”

[0465] Step 2) When the PU including the candidate block is located outside the boundary of the slice, the availability of the candidate block may be set to “false”. When the target block and the candidate block are located in different slices, the availability of the candidate block may be set to “false”.

[0466] Step 3) When the PU including the candidate block is located outside the boundary of the tile, the availability of the candidate block may be set to “false”. When the target block and the candidate block are located in different tiles, the availability of the candidate block may be set to “false”.

[0467] Step 4) When the prediction mode of the PU including the candidate block is an intra prediction mode, the availability of the candidate block may be set to “false”. When the PU including the candidate block does not use inter prediction, the availability of the candidate block may be set to “false”.

[0468] Fig.10 An order in which motion information of spatial candidates is added to a merge list according to an embodiment is shown.

[0469] like Fig.10 As shown in , when multiple motion information of spatial candidates are added to the merge list, the order of A1, B1, B0, A0 and B2 can be used. That is, multiple motion information of available spatial candidates can be added to the merge list in the order of A1, B1, B0, A0 and B2.

[0470] Methods for deriving merge lists in merge mode and skip mode

[0471] As described above, the maximum number of merge candidates in the merge list may be set. The set maximum number may be indicated by "N". The set number may be sent from the encoding device 100 to the decoding device 200. The slice header of the slice may include N. In other words, the maximum number of merge candidates in the merge list for the target block of the slice may be set by the slice header. For example, the value of N may be substantially 5.

[0472] A plurality of pieces of motion information (ie, merge candidates) may be added to the merge list in the order of the following steps 1) to 4).

[0473] Step 1) Among the space candidates, available space candidates can be added to the merge list. Fig.10 The multiple motion information of the available spatial candidates are added to the merge list in the order shown in . Here, when the motion information of the available spatial candidate overlaps with other motion information already in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information existing in the list may be simply referred to as "overlap check".

[0474] The maximum number of pieces of motion information added may be N.

[0475] Step 2) When the number of motion information in the merge list is less than N and the temporal candidate is available, the motion information of the temporal candidate may be added to the merge list. Here, when the motion information of the available temporal candidate overlaps with other motion information already in the merge list, the motion information of the temporal candidate may not be added to the merge list.

[0476] Step 3) When the number of pieces of motion information in the merge list is less than N and the type of the target slice is 'B', combined motion information generated by combined bidirectional prediction (bi-prediction) may be added to the merge list.

[0477] The target slice may be a slice including the target block.

[0478] The combined motion information may be a combination of L0 motion information and L1 motion information. The L0 motion information may be motion information referring only to the reference picture list L0. The L1 motion information may be motion information referring only to the reference picture list L1.

[0479] In the merge list, there may be one or more pieces of L0 motion information. In addition, in the merge list, there may be one or more pieces of L1 motion information.

[0480] The combined motion information may include one or more pieces of combined motion information. When generating the combined motion information, L0 motion information and L1 motion information to be used in the step of generating the combined motion information among the one or more pieces of L0 motion information and the one or more pieces of L1 motion information may be predefined. The one or more pieces of combined motion information may be generated in a predefined order via bidirectional prediction using a combination of a pair of different motion information in a merge list. One piece of motion information in the pair of different motion information may be L0 motion information, and the other piece of motion information in the pair of different motion information may be L1 motion information.

[0481] For example, the combined motion information to which the highest priority is added may be a combination of L0 motion information having a merge index 0 and L1 motion information having a merge index 1. When the motion information having a merge index 0 is not L0 motion information or when the motion information having a merge index 1 is not L1 motion information, the combined motion information may be neither generated nor added. Next, the combined motion information to which the next priority is added may be a combination of L0 motion information having a merge index 1 and L1 motion information having a merge index 0. The subsequent detailed combination may conform to other combinations in the field of video encoding / decoding.

[0482] Here, when the combined motion information overlaps with other motion information already present in the merge list, the combined motion information may not be added to the merge list.

[0483] Step 4) When the number of pieces of motion information in the merge list is less than N, the motion information of the zero vector may be added to the merge list.

[0484] The zero-vector motion information may be motion information in which a motion vector is a zero vector.

[0485] The number of pieces of zero vector motion information may be one or more. The reference picture indexes of one or more pieces of zero vector motion information may be different from each other. For example, the value of the reference picture index of the first zero vector motion information may be 0. The value of the reference picture index of the second zero vector motion information may be 1.

[0486] The number of pieces of zero-vector motion information may be the same as the number of reference pictures in the reference picture list.

[0487] The reference direction of the zero-vector motion information may be bidirectional. Both motion vectors may be zero vectors. The number of pieces of zero-vector motion information may be the smaller one of the number of reference pictures in the reference picture list L0 and the number of reference pictures in the reference picture list L1. Optionally, when the number of reference pictures in the reference picture list L0 and the number of reference pictures in the reference picture list L1 are different from each other, the reference direction as a unidirectional direction may be used for a reference picture index applicable only to a single reference picture list.

[0488] The encoding apparatus 100 and / or the decoding apparatus 200 may then add zero-vector motion information to the merge list while changing the reference picture index.

[0489] When the zero-vector motion information overlaps with other motion information already present in the merge list, the zero-vector motion information may not be added to the merge list.

[0490] The order of the above steps 1) to 4) is only exemplary and may be changed. In addition, some of the above steps may be omitted according to predefined conditions.

[0491] Method for deriving a candidate list of motion vector prediction in AMVP mode

[0492] The maximum number of motion vector predictor candidates in the motion vector predictor candidate list may be predefined. The predefined maximum number may be indicated by N. For example, the predefined maximum number may be 2.

[0493] A plurality of pieces of motion information (ie, prediction motion vector candidates) may be added to the prediction motion vector candidate list in the order of the following steps 1) to 3).

[0494] Step 1) An available spatial candidate among the spatial candidates may be added to the prediction motion vector candidate list. The spatial candidate may include a first spatial candidate and a second spatial candidate.

[0495] The first spatial candidate may be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate may be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.

[0496] The plurality of pieces of motion information of the available spatial candidates may be added to the prediction motion vector candidate list in the order of the first spatial candidate and the second spatial candidate. In this case, when the motion information of the available spatial candidate overlaps with other motion information already present in the prediction motion vector candidate list, the motion information of the available spatial candidate may not be added to the prediction motion vector candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, the motion information of the second spatial candidate may not be added to the prediction motion vector candidate list.

[0497] The maximum number of pieces of motion information added may be N.

[0498] Step 2) When the number of pieces of motion information in the prediction motion vector candidate list is less than N and the temporal candidate is available, the motion information of the temporal candidate may be added to the prediction motion vector candidate list. In this case, when the motion information of the available temporal candidate overlaps with other motion information already in the prediction motion vector candidate list, the motion information of the available temporal candidate may not be added to the prediction motion vector candidate list.

[0499] Step 3) When the number of pieces of motion information in the motion vector predictor candidate list is less than N, zero vector motion information may be added to the motion vector predictor candidate list.

[0500] The zero-vector motion information may include one or more pieces of zero-vector motion information. Reference picture indexes of the one or more pieces of zero-vector motion information may be different from each other.

[0501] The encoding apparatus 100 and / or the decoding apparatus 200 may sequentially add a plurality of pieces of zero-vector motion information to the prediction motion vector candidate list while changing the reference picture index.

[0502] When the zero-vector motion information overlaps with other motion information already present in the motion vector predictor candidate list, the zero-vector motion information may not be added to the motion vector predictor candidate list.

[0503] The above description of the zero-vector motion information made in conjunction with the merge list can also be applied to the zero-vector motion information, and its repeated description will be omitted.

[0504] The order of steps 1) to 3) described above is merely exemplary and may be changed. In addition, some of the steps may be omitted according to predefined conditions.

[0505] Fig.11 Transformation and quantization processes according to examples are shown.

[0506] like Fig.11 As shown in , the quantized level may be generated by performing a transform and / or quantization process on the residual signal.

[0507] The residual signal may be generated as a difference between the original block and the predicted block. Here, the predicted block may be a block generated via intra prediction or inter prediction.

[0508] The transform may include at least one of a primary transform and a secondary transform. A transform coefficient may be generated by performing a primary transform on the residual signal, and a secondary transform coefficient may be generated by performing a secondary transform on the transform coefficient.

[0509] The first transformation may be performed using at least one of a plurality of predefined transformation methods, for example, the plurality of predefined transformation methods may include discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loewenstein transform (KLT), etc.

[0510] A secondary transform may be performed on transform coefficients generated by performing the first transform.

[0511] The transform method applied to the first transform and / or the second transform may be determined based on at least one of the encoding parameters for the target block and / or the neighboring blocks. Alternatively, transform information indicating the transform method may be transmitted by the encoding device to the decoding device 200 using a signal.

[0512] The quantization level may be generated by performing quantization on a result generated by performing the first transform and / or the second transform or by performing quantization on a residual signal.

[0513] The quantized levels may be scanned based on at least one of an upper right diagonal scan, a vertical scan, and a horizontal scan according to at least one of an intra prediction mode, a block size, and a block form.

[0514] For example, the coefficients of the block may be scanned using an upper right diagonal scan to change the coefficients into a 1D vector form. Alternatively, depending on the size of the transform block and / or the intra prediction mode, a vertical scan that scans the 2D block format coefficients in a column direction or a horizontal scan that scans the 2D block format coefficients in a row direction may be used instead of the upper right diagonal scan.

[0515] The scanned quantized levels may be entropy encoded, and the bitstream may include the entropy encoded quantized levels.

[0516] The decoding device 200 may generate quantized levels by entropy decoding a bitstream. The quantized levels may be arranged in the form of 2D blocks via inverse scanning. Here, as the inverse scanning method, at least one of upper right diagonal scanning, vertical scanning, and horizontal scanning may be performed.

[0517] Inverse quantization may be performed on the quantized level. A secondary inverse transform may be performed on a result generated by performing the inverse quantization, depending on whether the secondary inverse transform is performed. In addition, a first inverse transform may be performed on a result generated by performing the secondary inverse transform, depending on whether the first inverse transform is to be performed. A reconstructed residual signal may be generated by performing the first inverse transform on a result generated by performing the secondary inverse transform.

[0518] Fig.12 is a configuration diagram of an encoding device according to an embodiment.

[0519] The encoding device 1200 may correspond to the encoding device 100 described above.

[0520] The encoding apparatus 1200 may include a processing unit 1210, a memory 1230, a user interface (UI) input device 1250, a UI output device 1260, and a storage 1240 communicating with each other through a bus 1290. The electronic device 1200 may further include a communication unit 1220 connected to a network 1299.

[0521] The processing unit 1210 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in the memory 1230 or the storage 1240. The processing unit 1210 may be at least one hardware processor.

[0522] The processing unit 1210 may generate and process a signal, data, or information input to, output from, or used in the encoding device 1200, and may perform checks, comparisons, determinations, etc. related to the signal, data, or information. In other words, in an embodiment, generation and processing of data or information and checks, comparisons, and determinations related to data or information may be performed by the processing unit 1210.

[0523] The processing unit 1210 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180 and a reference picture buffer 190.

[0524] At least some of the inter prediction unit 110, the intra prediction unit 120, the switch 115, the subtractor 125, the transform unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transform unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190 may be program modules and may communicate with an external device or system. The program module may be included in the encoding apparatus 1200 in the form of an operating system, an application program module, or other program modules.

[0525] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the encoding device 1200.

[0526] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to the embodiments or for implementing abstract data types according to the embodiments.

[0527] The program modules may be implemented using instructions or codes executed by at least one processor of the encoding device 1200 .

[0528] The processing unit 1210 can execute instructions or codes in the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transform unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transform unit 170, the adder 175, the filter unit 180 and the reference picture buffer 190.

[0529] The storage unit may represent the memory 1230 and / or the storage 1240. Each of the memory 1230 and the storage 1240 may be any one of various types of volatile or non-volatile storage media. For example, the memory 1230 may include at least one of a read-only memory (ROM) 1231 and a random access memory (RAM) 1232.

[0530] The storage unit may store data or information used for the operation of the encoding apparatus 1200. In an embodiment, data or information of the encoding apparatus 1200 may be stored in the storage unit.

[0531] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, and the like.

[0532] The encoding device 1200 may be implemented in a computer system including a computer-readable storage medium.

[0533] The storage medium may store at least one module required for the operation of the encoding apparatus 1200. The memory 1230 may store at least one module, and may be configured such that the at least one module is executed by the processing unit 1210.

[0534] Functions related to communication of data or information of the encoding device 1200 may be performed through the communication unit 1220 .

[0535] For example, the communication unit 1220 may transmit the bitstream to the decoding apparatus 1300 which will be described later.

[0536] Fig.13 is a configuration diagram of a decoding device according to an embodiment.

[0537] The decoding device 1300 may correspond to the decoding device 200 described above.

[0538] The decoding apparatus 1300 may include a processing unit 1310, a memory 1330, a user interface (UI) input device 1350, a UI output device 1360, and a storage 1340 communicating with each other through a bus 1390. The decoding apparatus 1300 may further include a communication unit 1320 connected to a network 1399.

[0539] The processing unit 1310 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in the memory 1330 or the storage 1340. The processing unit 1310 may be at least one hardware processor.

[0540] The processing unit 1310 may generate and process a signal, data, or information input to, output from, or used in the decoding device 1300, and may perform checks, comparisons, determinations, etc. related to the signal, data, or information. In other words, in an embodiment, generation and processing of data or information and checks, comparisons, and determinations related to data or information may be performed by the processing unit 1310.

[0541] The processing unit 1310 may include an entropy decoding unit 210 , an inverse quantization unit 220 , an inverse transform unit 230 , an intra prediction unit 240 , an inter prediction unit 250 , an adder 255 , a filter unit 260 , and a reference picture buffer 270 .

[0542] At least some of the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra prediction unit 240, the inter prediction unit 250, the adder 255, the filter unit 260, and the reference picture buffer 270 of the decoding apparatus 1300 may be program modules and may communicate with an external device or system. The program modules may be included in the decoding apparatus 1300 in the form of an operating system, an application program module, or other program modules.

[0543] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device that can communicate with the decoding device 1300.

[0544] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to the embodiments or for implementing abstract data types according to the embodiments.

[0545] The program modules may be implemented using instructions or codes executed by at least one processor of the decoding device 1300 .

[0546] The processing unit 1310 may execute instructions or codes in the entropy decoding unit 210 , the inverse quantization unit 220 , the inverse transform unit 230 , the intra prediction unit 240 , the inter prediction unit 250 , the adder 255 , the filter unit 260 , and the reference picture buffer 270 .

[0547] The storage unit may represent the memory 1330 and / or the storage 1340. Each of the memory 1330 and the storage 1340 may be any one of various types of volatile or non-volatile storage media. For example, the memory 1330 may include at least one of the ROM 1331 and the RAM 1332.

[0548] The storage unit may store data or information used for the operation of the decoding device 1300. In an embodiment, data or information of the decoding device 1300 may be stored in the storage unit.

[0549] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, and the like.

[0550] The decoding device 1300 may be implemented in a computer system including a computer-readable storage medium.

[0551] The storage medium may store at least one module required for the operation of the decoding device 1300. The memory 1330 may store at least one module, and may be configured such that the at least one module is executed by the processing unit 1310.

[0552] Functions related to communication of data or information of the decoding device 1300 may be performed through the communication unit 1320 .

[0553] For example, the communication unit 1320 may receive a bitstream from the encoding apparatus 1200 .

[0554] In the following embodiments, when encoding and decoding are performed using inter prediction, a method for deriving inter prediction information of a target block using inter prediction information of a neighboring block will be described.

[0555] When encoding and decoding are performed using inter prediction, inter prediction information of a target block in a target picture may be searched for in previously encoded and / or previously decoded pictures in order to remove temporal redundancy from the video.

[0556] Since the inter-frame prediction information of the target block is derived using the inter-frame prediction information of the neighboring blocks, the amount of information required for the inter-frame prediction can be reduced. Here, the amount of information can be the number of bits.

[0557] As a method for deriving inter-frame prediction information of a target block using inter-frame prediction information of a neighboring block, an AMVP mode and a merge mode can be used. In the AMVP mode and the merge mode, an AMVP candidate list and a merge candidate list can be configured respectively using neighboring blocks that are adjacent in time and neighboring blocks that are adjacent in space. Each candidate in such a list can be inter-frame prediction information or a part of inter-frame prediction information.

[0558] In the configuration of the list, inter prediction information of available neighboring blocks may be used as candidates to fill the list.

[0559] When inter prediction information of a neighboring block does not exist, or when inter prediction information of a neighboring block cannot be used, inter prediction information of the neighboring block cannot be used as a candidate.

[0560] The maximum number of candidates in each list may be predefined. When multiple pieces of inter-frame prediction information of available neighboring blocks cannot fill the predefined maximum number of candidates in the list, zero vector motion information may be added to the list.

[0561] Since the correlation between the candidates in the list (ie, the inter-frame prediction information) and the inter-frame prediction information of the target block is high, the encoding performance can be improved.

[0562] On the contrary, for candidates having low correlation with the inter prediction information of the target block (eg, zero vector motion information), the number of bits to be signaled may be increased when deriving the inter prediction information. Since the number of bits to be signaled increases, encoding performance may decrease.

[0563] In an embodiment, for a target block having no inter prediction information, the encoding device 1200 and the decoding device 1300 may use the inter prediction information of a neighboring block to generate inter prediction information for the target block, and may add the generated inter prediction information as a candidate to a corresponding list.

[0564] In an embodiment, inter-frame prediction information having high correlation may be added as a candidate to the list instead of inter-frame prediction information having low correlation. Since inter-frame prediction information having high correlation is added as a candidate to the list, encoding efficiency may be improved.

[0565] In an embodiment, each of the encoding device 1200 and the decoding device 1300 may use multiple pieces of inter-frame prediction information of multiple neighboring blocks to configure a list for a target block. The multiple neighboring blocks may include temporal neighboring blocks and spatial neighboring blocks. Each of the encoding device 1200 and the decoding device 1300 may use multiple pieces of inter-frame prediction information of multiple neighboring blocks to add inter-frame prediction information having a higher correlation with the inter-frame prediction information of the target block to the list. By using and adding the inter-frame prediction information, the number of bits used to indicate an index of the inter-frame prediction information, etc. can be reduced, and encoding performance can be improved.

[0566] In an embodiment, when inter-frame prediction information of a target block does not exist, each of the encoding device 1200 and the decoding device 1300 may generate inter-frame prediction information of the target block by combining a plurality of inter-frame prediction information of neighboring blocks of the target block. Each of the encoding device 1200 and the decoding device 1300 may add the generated inter-frame prediction information as a candidate to a list. Instead of inter-frame prediction information having a low correlation with the inter-frame prediction information of the target block (for example, zero vector motion information), inter-frame prediction information having a high correlation with the inter-frame prediction information of the target block is added to the list as a candidate, thereby enabling more efficient encoding to be performed in the derivation of the inter-frame prediction information.

[0567] In the process for recursively partitioning a CU, the CU may be divided into four square blocks of the same size or two blocks of the same size. When the CU is divided into two partition blocks, the CU may be divided horizontally or vertically. Alternatively, the CU may be divided into three partition blocks, and may be divided horizontally or vertically. For example, when the CU is divided vertically, the ratio of the width of the partition blocks generated from the partitioning may be 1:2:1. Similarly, when the CU is divided horizontally, the ratio of the height of the partition blocks may be 1:2:1.

[0568] The partitioning of a block in inter-frame prediction may mean that the encoding efficiency obtained when performing inter-frame prediction using motion information of each partition block generated from the partitioning is higher than the encoding efficiency obtained when performing inter-frame prediction using one piece of motion information of an unpartitioned block. In other words, when a block is partitioned, it is very likely that two or four partition blocks will have different motion information.

[0569] When the target block is a block partitioned from an upper layer block, when deriving motion information of a spatially adjacent block in a merge mode for the target block, motion information of another partition block in the upper layer block may be used. In other words, the motion information of the target block may be derived in the same manner as the motion information of other partition blocks. In this case, (even if the partition block is generated by partitioning from the upper layer block,) the two partition blocks have the same motion information, so encoding performance may be reduced.

[0570] Each of the encoding device 1200 and the decoding device 1300 may allocate a smaller number of bits to a candidate with a higher priority among the candidates in the list. The allocated bits may be a value indicating the index of the corresponding candidate. In other words, the index indicating the candidate with a higher priority may be signaled using a smaller number of bits than the number of bits indicating the index of the candidate with a lower priority.

[0571] Each of the encoding device 1200 and the decoding device 1300 may be configured to assign a higher priority to a candidate expected or estimated to have a higher encoding performance when configuring the list.

[0572] Here, assigning a higher priority may mean 1) assigning a smaller number of bits, 2) assigning a smaller index, or 3) the candidate being included in the list with a higher priority over other candidates in the list.

[0573] Furthermore, each of the encoding device 1200 and the decoding device 1300 may be configured to assign a lower priority to a candidate expected or estimated to have lower encoding performance when configuring the list.

[0574] Here, assigning a lower priority may mean 1) assigning a larger number of bits, 2) assigning a larger index, 3) the candidate is included in the list with a lower priority to follow other candidates in the list, or 4) the candidate is not included in the list. By such a list configuration, encoding performance can be improved.

[0575] When the list is configured using the motion information of the spatial neighboring blocks for the partitioned CU, if the spatial neighboring blocks are blocks partitioned from the upper layer CU including the partitioned CU, each of the encoding device 1200 and the decoding device 1300 may include the motion information of the spatial neighboring blocks in the list with a lower priority. Alternatively, if the spatial neighboring blocks are blocks partitioned from the upper layer CU including the partitioned CU, the encoding device 1200 may not include the motion information of the spatial neighboring blocks in the list.

[0576] In other words, each of the encoding device 1200 and the decoding device 1300 may not include motion information that is highly likely to have low encoding performance in the list, and may assign a lower priority to the motion information that is highly likely to have low encoding performance. Through such exclusion and assignment, each of the encoding device 1200 and the decoding device 1300 prevents motion information that is expected or estimated to have low encoding performance from being selected, or reduces the possibility of such motion information being selected, thereby improving encoding performance.

[0577] Fig.14 is a flowchart of an inter-frame prediction method according to an embodiment.

[0578] The inter prediction method may be performed by the encoding apparatus 1200 and / or the decoding apparatus 1300 .

[0579] For example, the encoding device 1200 may perform the inter-frame prediction method according to the embodiment to compare the efficiencies of multiple prediction methods for the target block, and may perform the inter-frame prediction method according to the embodiment to generate a reconstructed block for the target block.

[0580] The target block may be any one of the various blocks described above. For example, the target block may be a coding unit (CU), a prediction unit (PU), or a transform unit (TU).

[0581] For example, the decoding apparatus 1300 may perform the inter prediction method according to the embodiment in order to generate a reconstructed block for the target block.

[0582] Below, the processing unit may be the processing unit 1210 of the encoding device 1200 and / or the processing unit 1310 of the decoding device 1300 .

[0583] At step 1410 , the processing unit may derive inter prediction information for a target block.

[0584] The inter prediction information may include 1) motion vector, 2) reference picture list, 3) reference picture index, 4) merge flag, 5) merge index, 6) advanced motion vector prediction (AMVP) index, 7) illumination compensation (IC) flag, and 8) overlapped block motion compensation (OBMC) flag.

[0585] The IC flag may be a flag indicating whether IC is to be applied.

[0586] The OBMC flag may be a flag indicating whether OBMC is to be applied.

[0587] The processing unit may use at least one method to derive the inter-frame prediction information.

[0588] The at least one method may include 1) a merge mode, 2) an AMVP mode, 3) a method for deriving inter prediction information based on a subblock, and 4) a method for deriving inter prediction information in the decoding apparatus 1300 .

[0589] The processing unit may derive the inter-frame prediction information using the at least one piece of information.

[0590] The at least one piece of information may include 1) inter prediction information of spatially neighboring blocks, 2) inter prediction information of temporally neighboring blocks, 3) combined inter prediction information, 4) unified candidate list, 5) adaptive candidate list depending on block shape, and 6) adaptive candidate list depending on block partitioning state.

[0591] At step 1420, the processing unit may perform inter prediction on the target block using the derived inter prediction information.

[0592] Inter prediction may include motion compensation and / or motion correction.

[0593] The processing unit may perform inter-prediction using at least one of compensation and / or correction.

[0594] At least one of the compensation and / or correction may include 1) motion compensation, 2) IC, 3) OBMC, 4) bidirectional optical flow (BIO), 5) affine space motion compensation, and 6) motion vector correction in the decoding device 1300 .

[0595] Use merge mode to derive inter prediction information

[0596] The processing unit may use the merge mode to derive inter-frame prediction information. In an embodiment, the merge mode may be replaced with an AMVP mode or a specific inter-frame prediction mode using a list, etc. In other words, the use of the merge mode described in the embodiment to derive inter-frame prediction information may also be applied to deriving inter-frame prediction information using the AMVP mode or a specific inter-frame prediction mode.

[0597] The processing unit may configure a merge candidate list. The number of merge candidates in the merge candidate list may be N. N may be a positive integer. For example, the merge candidate may be inter-frame prediction information and may include a motion vector and a reference picture list.

[0598] The processing unit may configure the merge candidate list using one or more of the inter prediction information of the spatial neighboring blocks, the inter prediction information of the temporal neighboring blocks, and the combined inter prediction information. Here, the processing unit may add multiple pieces of inter prediction information to the merge candidate list in a specific order.

[0599] When configuring a merge candidate list, the processing unit may add multiple pieces of inter-frame prediction information of neighboring blocks as merge candidates to the merge candidate list. Here, the processing unit may add multiple pieces of inter-frame prediction information of neighboring blocks to the merge candidate list in a specific order of the neighboring blocks. 1) When the inter-frame prediction information of the neighboring block does not exist or 2) when the inter-frame prediction information of the neighboring block is the same as the inter-frame prediction information present in the merge candidate list (that is, when the inter-frame prediction information of the neighboring block is already included in the merge candidate list), the processing unit may not add the inter-frame prediction information of the neighboring block to the merge candidate list. In other words, when multiple pieces of inter-frame prediction information of two neighboring blocks are the same as each other, the inter-frame prediction information of the neighboring block with a lower priority may not be added to the merge candidate list.

[0600] When inter prediction information of one of the neighboring blocks is not added to the merge candidate list, the processing unit may add the combined inter prediction information to the merge candidate list instead of the inter prediction information that is not added. For example, when inter prediction information of a specific neighboring block does not exist, or when inter prediction information of a specific neighboring block is the same as inter prediction information in the merge candidate list, the processing unit may derive combined inter prediction information for the specific neighboring block, and may add the derived combined inter prediction information to the merge candidate list.

[0601] The processing unit may generate combined inter prediction information by combining two or more pieces of inter prediction information for neighboring blocks of the target block with each other.

[0602] The processing unit may configure an inter-frame prediction information palette. The inter-frame prediction information palette may be a list having N pieces of inter-frame prediction information. N may be a positive integer. Here, the processing unit may 1) add the inter-frame prediction mode of the target block to the inter-frame prediction information palette, and 2) manage the inter-frame prediction modes in the inter-frame prediction information palette according to a specific order and method. For example, when the inter-frame prediction information palette is filled with information, the processing unit may manage the inter-frame prediction information palette in a first-in-first-out (FIFO) manner.

[0603] When the inter-frame prediction information of the target block is the same as the inter-frame prediction information existing in the inter-frame prediction information palette (that is, when the inter-frame prediction information of the target block is already included in the inter-frame prediction information palette), the processing unit may not add the inter-frame prediction information of the target block to the inter-frame prediction information palette.

[0604] When the inter prediction information of the target block is identical to the inter prediction information present in the inter prediction information palette, the processing unit may move the inter prediction information of the inter prediction information palette that is identical to the inter prediction information of the target block to the position of the first inter prediction information in the inter prediction information palette. In other words, the processing unit may assign a specific priority (such as the highest priority) to the inter prediction information of the inter prediction information palette that is identical to the inter prediction information of the target block, and may adjust the positions of the plurality of inter prediction information present in the inter prediction information palette based on the assigned priorities.

[0605] The processing unit may initialize the inter prediction information palette for all blocks in the target picture as a unit of each picture. In other words, the blocks in the target picture may share a single inter prediction information palette with each other.

[0606] The processing unit may use the inter prediction information present in the inter prediction information palette as the inter prediction information of the temporally neighboring block.

[0607] Fig.15 Spatial neighboring blocks of a target block according to an example are shown.

[0608] exist Fig.15 In , A to K may indicate respective spatially neighboring blocks.

[0609] The inter-frame prediction information of the spatial neighboring blocks can be Fig.15 The inter-frame prediction information of the block existing at any one of the positions corresponding to A to K.

[0610] Hereinafter, the term “inter prediction information of block X” may be understood as “inter prediction information corresponding to position X”.

[0611] For example, the size of the neighboring block may be M×N. M and N may each be at least one of 2, 4, 8, 16, 32, 64, and 128.

[0612] The left neighboring block may be a neighboring block adjacent to the left side of the target block, and may be one or more of block A, block B, block C, block D, and block E.

[0613] The upper neighboring block may be a neighboring block adjacent to the upper side of the target block, and may be one or more of block G, block H, block I, block J, and block K.

[0614] The upper left neighboring block may be a neighboring block adjacent to the upper left corner of the target block and may be block F.

[0615] Such a spatially neighboring block may be a block adjacent to a boundary of the target block or a block not adjacent to the boundary of the target block.

[0616] Fig.16 Temporal neighboring blocks of a target block according to an example are shown.

[0617] exist Fig.16 , L to W may represent respective temporally adjacent blocks.

[0618] The temporally adjacent block may be a block in a previous picture. The previous picture may be a previously reconstructed col picture. The previous picture may be a picture that has been encoded or decoded before the target picture is encoded or decoded.

[0619] The previous picture may be a picture having a picture order count (POC) greater than a picture order count (POC) of the target picture.

[0620] The position of the temporally adjacent block in the previous picture may be the same as the position of the target block in the target picture. Alternatively, the position of the temporally adjacent block in the previous picture may correspond to the position of the target block in the target picture. Alternatively, the position of the temporally adjacent block in the previous picture may correspond to at least one of the position of the lower right portion of the target block, the central portion of the target block, and a specific position of the target block.

[0621] Optionally, the temporal neighboring block may be a block adjacent to the col block. For example, the temporal neighboring block may be a block adjacent to the lower right vertex of the col block.

[0622] Alternatively, the temporally adjacent block may be a temporally preceding block in the target picture. The temporally preceding block may be a block that has been encoded or decoded before the target block is encoded or decoded.

[0623] The temporal neighboring block may be a specific neighboring block referenced in a process of configuring a merge candidate list. Here, the specific neighboring block may be a neighboring block corresponding to the inter prediction information included in the merge candidate list.

[0624] The inter prediction information of the temporally neighboring blocks may be the inter prediction information of the blocks arranged at a specific position in the previous picture. Here, the specific position may be the position of the target block in the target picture.

[0625] The inter prediction information of the temporally neighboring blocks may be inter prediction information of a block arranged at a specific position in the target picture. Here, the specific position may be a position of a spatially neighboring block of the target block in the target picture.

[0626] Use combined inter prediction information to configure the merge candidate list

[0627] Fig.17 The generation of combined inter prediction information of the upper right neighboring blocks according to an example is shown.

[0628] Fig.18 The generation of combined inter prediction information of upper neighboring blocks according to an example is shown.

[0629] The processing unit may configure a merge candidate list using the combined inter prediction information. The combined inter prediction information may replace the inter prediction information or motion information of the neighboring block, and may then be added to the merge candidate list as a new merge candidate.

[0630] The processing unit may generate combined inter-frame prediction information by combining a plurality of inter-frame prediction information related to the target block. For example, the inter-frame prediction information related to the target block may be inter-frame prediction information for a neighboring block of the target block. The inter-frame prediction information related to the target block may be neighboring inter-frame prediction information. The neighboring inter-frame prediction information may be inter-frame prediction information of a neighboring block.

[0631] The combined inter prediction information may include motion information. The motion information of the combined inter prediction information may include at least one of a reference picture list, a reference picture index, an inter prediction indicator, a motion vector, a motion vector candidate, a motion vector candidate list, and a picture order count (POC).

[0632] In an embodiment, the term “inter prediction information” may be replaced with “motion information” and “motion vector”, and “combined inter prediction information” may be replaced with “combined motion information” and “combined motion vector”.

[0633] When generating partial information of the combined inter prediction information, the processing unit may select neighboring inter prediction information to be used from among a plurality of pieces of neighboring inter prediction information.

[0634] For example, the processing unit may use the partial information of the selected neighboring inter prediction information as the partial information of the combined inter prediction information. In other words, the processing unit may assign the value of the partial information of the selected neighboring inter prediction information to the partial information of the combined inter prediction information.

[0635] For example, the partial information may be an IC flag or an OBMC flag.

[0636] For example, the processing unit may generate partial information of the combined inter prediction information using adjacent inter prediction information selected in relation to a combination of motion vectors from among a plurality of pieces of adjacent inter prediction information.

[0637] For example, the processing unit may generate partial information of the combined inter prediction information using combined adjacent inter prediction information (scaled reference) for a motion vector among a plurality of pieces of adjacent inter prediction information.

[0638] For example, refer to Fig.15 When multiple inter-frame prediction information of blocks B and J are used to generate combined inter-frame prediction information for block F, the IC flag and / or OBMC flag of block B or block J can be used as the IC flag and / or OBMC flag of block F.

[0639] The processing unit may generate a motion vector of the combined inter-frame prediction information by combining motion vectors of a plurality of pieces of neighboring motion information. The neighboring motion information may be motion information of a neighboring block. In addition, the neighboring motion vector may be a motion vector of a neighboring block.

[0640] For example, the neighboring motion information may be a block (such as Fig.15 The spatial neighboring blocks A to K shown in Fig.16 Motion information of temporal neighboring blocks L to W) shown in .

[0641] The neighboring motion information may be motion information of each block that is not adjacent to the target block. The blocks that are not adjacent to the target block may be blocks that are adjacent to the neighboring blocks of the target block.

[0642] The neighboring motion information may be the motion information of a block having the above-mentioned specific relationship with the target block, and the block having the specific relationship may also be a block not adjacent to the target block. For example, the block having the specific relationship may be a block adjacent to the neighboring block of the target block. The neighboring block may be inserted between the block having the specific relationship and the target block.

[0643] The combined inter-frame prediction information may be a result obtained by selecting one piece from a plurality of pieces of motion information of a plurality of neighboring blocks. Here, one piece of the plurality of pieces of motion information may be selected as the combined inter-frame prediction information according to a specified condition. For example, the combined inter-frame prediction information may be a result of calculation, selection, combination, and transformation using a plurality of pieces of motion information of a plurality of neighboring blocks.

[0644] For example, the combined inter prediction information may be motion information of a neighboring block for which the difference between the POC of the target picture and the POC of the reference picture for the neighboring block is minimal.

[0645] For example, the combined inter prediction information may be specific motion information present in the merge candidate list.

[0646] The combined inter prediction information may be a result obtained by selecting and combining one or more pieces of neighboring motion information from a plurality of pieces of neighboring motion information. Here, the combined inter prediction information may be selected according to a specified condition.

[0647] For example, the motion vector of the combined inter-frame prediction information may be a motion vector of a specific neighboring block among a plurality of neighboring blocks. Here, the specific neighboring block may be a neighboring block having the smallest difference between the POC of the target picture and the POC of the reference picture for the neighboring block among the plurality of neighboring blocks. The motion vector of the combined inter-frame prediction information may be a result of a formula using a plurality of neighboring motion vectors. The neighboring motion vector may be a motion vector of a neighboring block. The neighboring motion vector of a neighboring block may include a plurality of motion vectors.

[0648] When the combined inter prediction information is generated, the processing unit may generate unidirectional combined inter prediction information or bidirectional combined inter prediction information. Here, the unidirectional combined inter prediction information may be forward (L0) inter prediction information or backward (L1) inter prediction information, and the bidirectional combined inter prediction information may be forward inter prediction information and backward inter prediction information.

[0649] The unidirectional combined inter prediction information may be a combination of 1) multiple pieces of bidirectional inter prediction information of neighboring blocks, 2) multiple pieces of unidirectional prediction information of neighboring blocks, and 3) multiple pieces of L0 inter prediction information or L1 inter prediction information among the multiple pieces of combined inter prediction information.

[0650] For example, the processing unit may generate L0 (L1) directional inter prediction information by combining two or more temporally adjacent blocks, and may add the generated unidirectional prediction information to the merge candidate list.

[0651] For example, the combined inter prediction information may be a result of combination of pieces of L0 (L1) directional inter prediction information of two previous blocks and L0 (L1) directional inter prediction information of one temporally neighboring block, and then may be added to the merge candidate list.

[0652] For example, the combined inter prediction information may be a result of combining a plurality of pieces of L0 (L1) direction inter prediction information of two specific neighboring blocks referred to in the merge candidate list configuration process, and may be added to the merge candidate list.

[0653] The bidirectional combined inter-frame prediction information may be a combination of the forward inter-frame prediction information and the backward inter-frame prediction information described above.

[0654] The motion vector of the combined inter-frame prediction information may be a combination of neighboring motion vectors. For example, the neighboring motion information may be a combination of multiple pieces of motion information of multiple neighboring blocks A to W. For example, the motion vector of the combined inter-frame prediction information may be an average value, a maximum value, a minimum value, or a median value of multiple neighboring motion vectors, and may be a combination of one or more of the average value, the maximum value, the minimum value, and the median value.

[0655] The average value can be obtained by dividing the sum of the combined motion vectors by the number of combined motion vectors. For example, the average value of motion vector (4,6) and motion vector (6,10) can be (5,8).

[0656] For example, the motion vector of the combined inter-frame prediction information may be a weighted average of a plurality of adjacent motion vectors, or may be a combination using changes among a plurality of adjacent motion vectors.

[0657] For example, Fig.17 As shown in , for the target block 1710, there may be an upper left neighboring block 1720, an upper neighboring block 1730, an upper right neighboring block 1740, a left neighboring block 1750, and a lower left neighboring block 1760. By combining the motion information 1721 of the upper left neighboring block 1720 and the motion information 1731 of the upper neighboring block 1730, combined inter-frame prediction information corresponding to the motion information 1741 of the upper right neighboring block 1740 may be generated. Such combined inter-frame prediction information may be used as the motion information 1711 of the target block 1710.

[0658] For example, Fig.18 As shown in , by combining the motion information 1721 of the upper left neighboring block 1720 and the motion information 1741 of the upper right neighboring block 1740, combined inter-frame prediction information corresponding to the motion information 1731 of the upper neighboring block 1730 can be generated. Such combined inter-frame prediction information can be used as the motion information 1711 of the target block 1710.

[0659] As described above, the motion vector combining the inter-frame prediction information may be the result of a weighted combination of neighboring motion vectors. A higher weight may be assigned to a motion vector of a neighboring block having a higher correlation with the target block.

[0660] The motion vector of the combined inter-frame prediction information may be a result of a weighted combination of neighboring motion vectors based on a block size (ie, a weighted combination based on a block size). The weighted combination based on a block size may be represented by the following equation 2:

[0661] [Equation 2]

[0662]

[0663] It may be a left neighboring motion vector of the target block. The left neighboring motion vector of the target block may be a motion vector of a neighboring block adjacent to the left side of the target block.

[0664] “ " may represent the width of the target block and may be the weight of the left neighboring motion vector for the target block.

[0665] It may be an upper neighboring motion vector of the target block. The upper neighboring motion vector of the target block may be a motion vector of a neighboring block adjacent to the upper side of the target block.

[0666] “ " may represent the height of the target block and may be the weight of the upper neighboring motion vector for the target block.

[0667] The motion vector combining the inter prediction information may be a result of weighted combination of neighboring motion vectors based on the POC (ie, POC-based weighted combination).

[0668] For example, as the POC of the reference picture for the neighboring motion information is closer to the POC of the target picture, the weight for the neighboring motion vector may be greater.

[0669] The combination using the change may be to generate inter prediction information of a block before the combined block or a block after the combined block using the change between two or more motion vectors.

[0670] For example, refer to Fig.15 , the processing unit may derive the motion vector of block K by using a combination of changes between the motion vector of block I and the motion vector of block J. The motion vector of block K may be obtained by adding the difference between the motion vector of block J and the motion vector of block I to the motion vector of block J.

[0671] Alternatively, the motion vector combining the inter-frame prediction information may be a result of an extrapolation-based combination of neighboring motion vectors.

[0672] For example, the extrapolation-based combination of two neighboring motion vectors may be represented by the following Equation 3:

[0673] [Equation 3]

[0674]

[0675] It may be the first neighboring motion vector. It may be the second neighboring motion vector.

[0676] Scaling may be applied to the inter prediction information used to generate the combined inter prediction information.

[0677] For example, when the neighboring motion information indicates bi-directional prediction, the processing unit may scale the motion vector of L0 based on the motion vector of L1. The processing unit may generate combined inter prediction information by combining the scaled motion vector of L0 with the motion vector of L1.

[0678] For example, when the neighboring motion information indicates bidirectional prediction, the processing unit may scale the motion vector of L1 based on the motion vector of L0. The processing unit may generate combined inter prediction information by combining the scaled motion vector of L1 with the motion vector of L0, and may add the generated combined inter prediction information to the merge candidate list.

[0679] When reference pictures for pieces of neighboring motion information used to generate combined inter prediction information are different from each other, the processing unit may change the inter prediction information by applying scaling to the inter prediction information.

[0680] For example, refer to Fig.15 When the neighboring blocks used for combination are block B and block J, and the reference pictures for block B and block J are different from each other, the processing unit may scale the motion vector of block J according to the temporal distance between block B and the reference picture for block B.

[0681] For example, refer to Fig.15 When the neighboring blocks used for combination are block B and block J, and the reference pictures for block B and block J are different from each other, the processing unit may scale the motion vector of block B according to the temporal distance between block J and the reference picture of block J.

[0682] When scaling, the processing unit may select neighboring motion information as a reference for scaling.

[0683] The processing unit may select neighboring motion information as a reference for scaling based on the POC. The processing unit may select the following motion information as reference motion information from a plurality of neighboring motion information: for which the POC of the reference picture for the motion information is closer to the POC of the target picture.

[0684] When combined inter prediction information is generated, the processing unit may determine whether to perform combining based on a specified condition.

[0685] When the combined inter-frame prediction information is generated, the processing unit may determine whether to perform the combination based on the similarity between the multiple inter-frame prediction information or multiple motion information for combination. For example, when the similarity is less than a predefined threshold, the processing unit may not perform the combination. Alternatively, when the similarity is greater than a predefined threshold, the processing unit may not perform the combination.

[0686] Here, the similarity may indicate a value or result of a formula using a plurality of pieces of motion information.

[0687] The processing unit may generate combined inter prediction information by combining multiple pieces of neighboring motion information based on the direction of the reference picture for the neighboring motion information. For example, the processing unit may generate combined inter prediction information by combining neighboring motion vectors in the same direction.

[0688] As described above, the processing unit may generate combined inter prediction information by combining a plurality of pieces of inter prediction information of a plurality of blocks. Here, each of the plurality of blocks may be a block that satisfies a specified condition.

[0689] exist Fig.15 In each block shown in , the inter-frame prediction information of the specific block may be replaced by the combined inter-frame prediction information. The processing unit may derive the first inter-frame prediction information for the position on the left side of the specific block, and derive the second inter-frame prediction information for the position on the right side of the specific block, and may generate the combined inter-frame prediction information by combining the first inter-frame prediction information with the second inter-frame prediction information. The generated combined inter-frame prediction information may replace the inter-frame prediction information of the specific block. Here, the inter-frame prediction information for the position on the left side of the specific block may be the inter-frame prediction information of the block located on the left side of the specific block. Here, the inter-frame prediction information for the position on the right side of the specific block may be the inter-frame prediction information of the block located on the right side of the specific block. In addition, such derivation, combination and generation may also be applied to a part of the inter-frame prediction information, such as motion information and motion vector.

[0690] For example, when first inter prediction information for a left position is derived, if inter prediction information of only one block located on the left side of a specific block is available, the processing unit may perform combination using the available inter prediction information.

[0691] For example, when second inter prediction information for a right position is derived, if inter prediction information of only one block located on the right side of a specific block is available, the processing unit may perform combination using the available inter prediction information.

[0692] For example, when pieces of inter-frame prediction information of a plurality of blocks located on the left side of a specific block are available, the processing unit may derive first inter-frame prediction information by combining the available pieces of inter-frame prediction information.

[0693] For example, when a plurality of pieces of inter-frame prediction information of a plurality of blocks located on the right side of a specific block are available, the processing unit may derive second inter-frame prediction information by combining the available plurality of pieces of inter-frame prediction information.

[0694] For example, when the first inter-frame prediction information for the left position is derived, if multiple inter-frame prediction information of multiple blocks located on the left side of a specific block are available, the processing unit may select specific inter-frame prediction information from the multiple inter-frame prediction information as the first inter-frame prediction information, and the selected inter-frame prediction information may be used for the generation of combined inter-frame prediction information.

[0695] For example, if multiple pieces of inter-frame prediction information of multiple blocks located on the left side of a specific block are available, the processing unit may select a block having the shortest temporal distance from the target picture from the multiple blocks. The processing unit may select the inter-frame prediction information of the selected block as the first inter-frame prediction information. The temporal distance between pictures may be a difference between the sequential positions of the display pictures.

[0696] For example, when the second inter-frame prediction information for the right position is derived, if multiple inter-frame prediction information of multiple blocks located on the right side of a specific block are available, the processing unit may select specific inter-frame prediction information from the multiple inter-frame prediction information as the second inter-frame prediction information, and the selected inter-frame prediction information may be used for the generation of combined inter-frame prediction information.

[0697] For example, if multiple pieces of inter-frame prediction information of multiple blocks located on the right side of a specific block are available, the processing unit may select a block having the shortest temporal distance from the target picture from the multiple blocks. The processing unit may select the inter-frame prediction information of the selected block as the second inter-frame prediction information.

[0698] For example, in Fig.15 In the blocks shown in , when multiple pieces of inter-frame prediction information of multiple blocks located on the left side of the specific block are available, the processing unit can generate combined inter-frame prediction information by combining the multiple pieces of inter-frame prediction information. The generated combined inter-frame prediction information can replace the inter-frame prediction information of the specific block. In addition, when multiple pieces of inter-frame prediction information of multiple blocks located on the right side of the specific block are available, the processing unit can generate combined inter-frame prediction information by combining the multiple pieces of inter-frame prediction information.

[0699] For example, in Fig.16 In a specific block among block N, block O, block P and block Q of the same picture, when multiple inter-frame prediction information of multiple co-located blocks in different multiple previous pictures is available, the processing unit can generate combined inter-frame prediction information by combining the multiple available inter-frame prediction information. The generated combined inter-frame prediction information can replace the inter-frame prediction information of the specific block. Such a co-located block can be a col block for the specific block. In other words, the position of the co-located block in the different previous pictures can be the same as the position of the specific block in the previous picture.

[0700] For example, the combined inter-frame prediction information for block N may be a combination of multiple pieces of inter-frame prediction information for co-located blocks L and M. For example, the combined inter-frame prediction information for block O may be a combination of multiple pieces of inter-frame prediction information for co-located blocks U and R. For example, the combined inter-frame prediction information for block P may be a combination of multiple pieces of inter-frame prediction information for co-located blocks V and S. For example, the combined inter-frame prediction information for block Q may be a combination of multiple pieces of inter-frame prediction information for co-located blocks W and T.

[0701] For example, the processing unit may generate combined inter prediction information by combining one or more pieces of inter prediction information of one or more spatially neighboring blocks with one or more pieces of inter prediction information of one or more temporally neighboring blocks.

[0702] Fig.19 The generation of combined inter prediction information of neighboring blocks according to an example is shown.

[0703] exist Fig.19 In the example, the target CU may indicate a target block.

[0704] exist Fig.19 , block AL, block A, block AR, block L, and block LB may be an upper left neighboring block, an upper neighboring block, an upper right neighboring block, a left neighboring block, and a lower left neighboring block of the target block, respectively.

[0705] The upper neighboring block may refer to the rightmost (or leftmost) block among multiple neighboring blocks above the target block. The left neighboring block may refer to the lowermost (or uppermost) block among multiple neighboring blocks on the left side of the target block.

[0706] Block A and block AR may be understood as a plurality of blocks adjacent to or located above the target block. Block L and block LB may be understood as a plurality of blocks adjacent to or located on the left side of the target block.

[0707] The processing unit may generate first adjacent inter-frame prediction information by combining multiple inter-frame prediction information of block A and block AR, wherein the multiple inter-frame prediction information may be understood as upper inter-frame prediction information or upper motion vector. The processing unit may generate second adjacent inter-frame prediction information by combining multiple inter-frame prediction information of block L and block LB. These multiple inter-frame prediction information may be understood as left inter-frame prediction information or left motion vector. The processing unit may generate combined inter-frame prediction information by combining the first adjacent inter-frame prediction information with the second adjacent inter-frame prediction information.

[0708] When the inter prediction information of block AL is not available, the processing unit may replace the inter prediction information of block AL with the generated combined inter prediction information. In addition, the processing unit may add the generated combined inter prediction information for block AL as a new merge candidate to the merge candidate list.

[0709] When the above-mentioned first neighboring inter-frame prediction information, second neighboring inter-frame prediction information, and combined inter-frame prediction information are acquired, the aforementioned combining method and the combining method to be described below may be used.

[0710] When only one of the multiple pieces of inter-frame prediction information of block A and block AR is available, the processing unit may use the available inter-frame prediction information among the multiple pieces of inter-frame prediction information as the first adjacent inter-frame prediction information. In addition, when only one of the multiple pieces of inter-frame prediction information of block L and block LB is available, the processing unit may use the available inter-frame prediction information among the multiple pieces of inter-frame prediction information as the second adjacent inter-frame prediction information.

[0711] The processing unit may select any one of the multiple inter-frame prediction information of block A and block AR, and may use the selected inter-frame prediction information as the first adjacent inter-frame prediction information. The processing unit may select any one of the multiple inter-frame prediction information of block L and block LB, and may use the selected inter-frame prediction information as the second adjacent inter-frame prediction information.

[0712] The processing unit may select inter-frame prediction information whose POC is closer to the POC of the target picture among the multiple pieces of inter-frame prediction information of block A and block AR as the first adjacent inter-frame prediction information. Here, the POC of the inter-frame prediction information may be the POC of the motion information for the inter-frame prediction information.

[0713] The processing unit may select inter-frame prediction information whose POC is closer to the POC of the target picture among the multiple pieces of inter-frame prediction information of block L and block LB as the second neighboring inter-frame prediction information.

[0714] Fig. 20 The generation of inter prediction information for block AL according to an example is shown.

[0715] MV0 and MV1 may refer to respective motion vectors used to combine inter-frame prediction information.

[0716] The processing unit may generate combined inter prediction information by combining multiple pieces of inter prediction information of block L and block LB. The combined inter prediction information may replace the inter prediction information of block AL and may be added to a merge candidate list for the target block as a merge candidate.

[0717] Fig.21 The generation of inter prediction information for block AR according to an example is shown.

[0718] The processing unit may generate combined inter prediction information by combining multiple pieces of inter prediction information of block A and block L. The combined inter prediction information may replace the inter prediction information of block AR and may be added to the merge candidate list for the target block as a merge candidate. For example, the aforementioned combination may be a combination based on extrapolation.

[0719] Fig. 22 The generation of inter prediction information of a target CU according to an example is shown.

[0720] The processing unit may generate combined inter prediction information by combining multiple inter prediction information of block L and non-neighboring blocks. The combined inter prediction information may be added as a merge candidate to a merge candidate list for the target block. For example, the aforementioned combination may be a combination based on extrapolation.

[0721] Use the inter prediction information in the merge candidate list to generate the combined inter prediction information

[0722] The processing unit may generate combined inter prediction information by combining M pieces of inter prediction information in the merge candidate list, and may use the generated combined inter prediction information for inter prediction, or may add the generated combined inter prediction information to the merge candidate list.

[0723] Here, directions of motion vectors of the combined pieces of inter prediction information may be identical to each other.

[0724] Here, M may be an integer of 2 or more, and may be less than or equal to the number of pieces of inter prediction information in the merge candidate list.

[0725] For example, when there are three inter-frame prediction information in the merge candidate list and the value of M is 2, the combinations that can be used may be (first inter-frame prediction information, second inter-frame prediction information), (first inter-frame prediction information, third inter-frame prediction information) and (second inter-frame prediction information, third inter-frame prediction information), and the processing unit can generate combined inter-frame prediction information by combining multiple inter-frame prediction information according to these combinations.

[0726] For example, when there are four inter-frame prediction information in the merge candidate list and the value of M is 4, the combination that can be used may be (first inter-frame prediction information, second inter-frame prediction information, third inter-frame prediction information and fourth inter-frame prediction information), and the processing unit can generate combined inter-frame prediction information by combining multiple inter-frame prediction information according to the combination.

[0727] Configuration of merge candidate list

[0728] As described above, the processing unit may add multiple inter prediction information of neighboring blocks as merge candidates to the merge candidate list when configuring the merge candidate list. Here, the processing unit may add multiple inter prediction information of neighboring blocks to the merge candidate list in a specific order of the neighboring blocks.

[0729] Return to reference Fig.15 , the processing unit may use multiple pieces of inter-frame prediction information of specific spatial neighboring blocks as merge candidates when configuring the merge candidate list. The specific spatial neighboring blocks may be block A, block B, block F, block J, and block K.

[0730] The processing unit may add 1) combined inter prediction information, 2) sub-block based motion information derivation mode (e.g., optional temporal motion vector prediction (ATMVP) mode, spatial-temporal motion vector prediction (STMVP) mode, etc.), and 3) affine spatial motion information derivation mode to the merge candidate list in a specific order.

[0731] For example, the processing unit may configure the merge candidate list in the order (B, J, K, A, ATMVP, F, combined inter-frame prediction information).

[0732] For example, the processing unit may configure the merge candidate list in the order of (B, J, K, A, first combined inter prediction information, ATMVP, F, second combined inter prediction information). Here, the first combined inter prediction information and the second combined inter prediction information may be different from each other in terms of neighboring blocks to be referenced to generate the first combined inter prediction information and the second combined inter prediction information, the number of neighboring blocks, and the directionality of the combined inter prediction information.

[0733] Below, “(α, β, γ, δ, ε)” may indicate the order of blocks, and may mean that blocks of symbols appearing earlier in the brackets are processed earlier than blocks of symbols appearing later. The statement that a merge candidate list is configured in the order of “(α, β, γ, δ, ε)” may mean that a task of configuring the merge candidate list is performed in the order of block α, block β, block γ, block δ, and block ε in the configuration of the merge candidate list, and may also mean that blocks are processed in the order of the listed blocks.

[0734] Here, the processing unit may perform the following tasks 1) to 5) on each block in the order of the blocks.

[0735] 1) The processing unit may determine whether to add the inter prediction information of the corresponding block to the merge candidate list.

[0736] 2) If it is determined to add the inter prediction information of the block to the merge candidate list, the processing unit may add the inter prediction information of the block to the merge candidate list.

[0737] 3) (if it is determined not to add the inter prediction information of the block to the merge candidate list), the processing unit may determine whether to derive combined inter prediction information for the corresponding block.

[0738] 4) If it is determined to derive the combined inter-frame prediction information, the processing unit may derive the combined inter-frame prediction information.

[0739] 5) The processing unit may determine whether to add the combined inter prediction information to the merge candidate list.

[0740] 6) If it is determined to add the combined inter prediction information to the merge candidate list, the processing unit may add the combined inter prediction information to the merge candidate list.

[0741] When tasks 1) to 5) are performed on one block, task 1) may be performed on the subsequent block.

[0742] For example, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0743] For example, the processing unit may configure the merge candidate list in the order of (J, B, A, K, F).

[0744] Fig.23 A case where a CU having the same width and height is divided vertically is shown.

[0745] Fig.24 A case where a CU having the same width and height is split horizontally is shown.

[0746] Fig.25 It shows the case where a CU whose width is greater than its height is split vertically.

[0747] Fig.26 It shows the case where a CU whose height is greater than its width is split horizontally.

[0748] The comparison between width and height and the direction of the partitioning can be used to determine the order of neighboring blocks.

[0749] The processing unit may determine a scheme for configuring the merge candidate list based on the shape of the target block.

[0750] The scheme for configuring the merge candidate list may include an order of neighboring blocks required for configuring the merge candidate list. The order of neighboring blocks may be an order of availability testing of multiple inter-frame prediction information for neighboring blocks and addition of the multiple inter-frame prediction information.

[0751] When the height of the target block is greater than the width of the target block, the processing unit may configure the merge candidate list in the order of (J, B, K, A, F), and when the height of the target block is less than or equal to the width of the target block, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0752] When the height of the target block is greater than the width of the target block, the processing unit may configure the merge candidate list in the order of (J, B, A, K, F), and when the height of the target block is less than or equal to the width of the target block, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0753] For example, when the height of the target block is greater than the width of the target block, the processing unit may configure the merge candidate list in the order of (J, K, B, A, F), when the height of the target block is less than the width of the target block, the processing unit may configure the merge candidate list in the order of (B, A, J, K, F), and when the height and width of the target block are equal to each other, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0754] For example, when the height of the target block is greater than the width of the target block, the processing unit may configure the merge candidate list using multiple pieces of inter-frame prediction information of spatially adjacent blocks located above the target block. When the height of the target block is greater than the width of the target block, the processing unit may configure the merge candidate list in the order of (F, G, H, I, J, K).

[0755] For example, when the height of the target block is smaller than the width of the target block, the processing unit may configure the merge candidate list using multiple pieces of inter-frame prediction information of spatially neighboring blocks located on the left side of the target block. When the height of the target block is smaller than the width of the target block, the processing unit may configure the merge candidate list in the order of (A, B, C, D, E, F).

[0756] The processing unit may determine a scheme for configuring the merge candidate list based on a state of partitioning of the target block.

[0757] The state of the partition may indicate the direction of the partition. The partition state of the target block may be the type or direction of the partition used to generate the target block. Alternatively, the partition state of the target block may be the type or direction of the partition applied to the upper layer block of the target block.

[0758] For example, when the target block is obtained as a result of vertical splitting, the processing unit may configure the merge candidate list in the order of (J, B, K, A, F), and when the target block is not obtained as a result of vertical splitting, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0759] For example, when the target block is obtained as a result of vertical splitting, the processing unit may configure the merge candidate list in the order of (J, B, A, K, F), and when the target block is not obtained as a result of vertical splitting, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0760] For example, depending on which of vertical partitioning, horizontal partitioning, and quad partitioning has been used to obtain the target block, the processing unit may select one of different orders of neighboring blocks, and may configure the merge candidate list in the selected order.

[0761] For example, when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (J, K, B, A, F). When the target block is obtained as a result of horizontal partitioning, the processing unit may configure the merge candidate list in the order of (B, A, J, K, F). When the target block is obtained as a result of quad partitioning, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0762] For example, when the target block is obtained as a result of vertical splitting, the processing unit may configure the merge candidate list using multiple pieces of inter-frame prediction information of spatially adjacent blocks located above the target block. When the target block is obtained as a result of vertical splitting, the processing unit may configure the merge candidate list in the order of (F, G, H, I, J, K).

[0763] For example, when the target block is obtained as a result of horizontal splitting, the processing unit may configure the merge candidate list using multiple pieces of inter-frame prediction information of spatially adjacent blocks located on the left side of the target block. When the target block is obtained as a result of horizontal splitting, the processing unit may configure the merge candidate list in the order of (A, B, C, D, E, F).

[0764] The processing unit may determine a scheme for configuring the merge candidate list based on both the shape of the target block and the division state of the target block.

[0765] For example, when the height of the target block is greater than the width of the target block or when the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical division, the processing unit may configure the merge candidate list in the order of (J, B, K, A, F). In other cases, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0766] For example, when the height of the target block is greater than the width of the target block or the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical division, the processing unit may configure the merge candidate list in the order of (J, B, A, K, F). In other cases, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0767] For example, when the height of the target block is greater than the width of the target block or when the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical division, the processing unit may configure the merge candidate list in the order of (J, K, B, A, F). When the height of the target block is less than the width of the target block or when the target block is obtained as a result of horizontal division, the processing unit may configure the merge candidate list in the order of (B, A, J, K, F). When the height and width of the target block are equal to each other and when the target block is obtained as a result of four division, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0768] For example, when the height of the target block is greater than the width of the target block, or the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical division, the processing unit may configure the merge candidate list using multiple pieces of inter-frame prediction information of spatially adjacent blocks located above the target block. When the height of the target block is greater than the width of the target block, or the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical division, the processing unit may configure the merge candidate list in the order of (F, G, H, I, J, K).

[0769] For example, when the height of the target block is smaller than the width of the target block or when the target block is obtained as a result of horizontal splitting, the processing unit may configure the merge candidate list using multiple pieces of inter-frame prediction information of spatially neighboring blocks located on the left side of the target block. When the height of the target block is smaller than the width of the target block or when the target block is obtained as a result of horizontal splitting, the processing unit may configure the merge candidate list in the order of (A, B, C, D, E, F).

[0770] The processing unit may determine a scheme for configuring the merge candidate list based on the location of the target block.

[0771] The position of the target block may be the relative position of the target block in the upper block. By dividing the upper block, a plurality of partition blocks may be generated, and the target block may be one of the plurality of partition blocks. The position of the target block may be the position of the target block in the upper block, or the position of the target block between the plurality of partition blocks. The division may be two divisions or four divisions.

[0772] For example, in Fig.23 In the embodiment, the processing unit may apply a unified configuration method to a merge candidate list for a first partition CU and apply an adaptive configuration method to a merge candidate list for a second partition CU.

[0773] The processing unit may determine a scheme for configuring the merge candidate list based on whether the combined inter prediction information exists. In the configuration of the merge candidate list, when the combined inter prediction information exists, the processing unit may adjust the priority of the combined inter prediction information. Here, the priority may refer to the position of the combined inter prediction information in the merge candidate list, the index of the combined inter prediction information, or the order in which the combined inter prediction information is added to the merge candidate list.

[0774] For example, when the merge candidate list is configured in the order of (B, J, K, A, F), if block B has combined inter prediction information, the processing unit may configure the merge candidate list in the order of (J, K, A, F, B). In other words, the processing unit may assign the lowest priority to the combined inter prediction information. Optionally, in the configuration of the merge candidate list, the processing unit may add the combined inter prediction information to a position after the multiple pieces of inter prediction information in the merge candidate list. In other words, the processing unit may add the combined inter prediction information to the merge candidate list with a lower priority to be after the multiple pieces of inter prediction information of the neighboring blocks.

[0775] For example, when the merge candidate list is configured in the order of (B, J, K, N, F) and block N has a temporal neighboring block, if block F has combined inter prediction information, the processing unit may configure the merge candidate list in the order of (B, J, K, F, N). In other words, the processing unit may assign a priority lower than the priority of multiple pieces of inter prediction information of spatially neighboring blocks and higher than the priority of inter prediction information of temporally neighboring blocks to the combined inter prediction information.

[0776] Optionally, in the configuration of the merge candidate list, the processing unit may add the combined inter prediction information to a position in the merge candidate list after the inter prediction information of the spatially neighboring block and before the inter prediction information of the temporally neighboring block. Optionally, the processing unit may assign a higher priority to the combined inter prediction information than the inter prediction information of the temporally neighboring block. Optionally, in the configuration of the merge candidate list, the processing unit may add the combined inter prediction information to a position in the merge candidate list before the inter prediction information of the temporally neighboring block.

[0777] For example, when there are multiple pieces of combined inter-frame prediction information, the order of the multiple pieces of combined inter-frame prediction information may remain unchanged according to the above-mentioned order determination scheme.

[0778] The processing unit may determine a method for configuring the merge candidate list based on a depth of the target block. The depth of the target block may be at least one of a QT depth based on a quadtree (QT) partition and a BT depth based on a binary tree (BT) partition.

[0779] For example, the processing unit may determine a scheme for configuring the merge candidate list based on whether the depth of the target block falls within a specific range.

[0780] For example, when the BT depth of the target block is less than or equal to n, the processing unit may configure the merge candidate list using at least one of the above schemes of configuring the merge candidate list based on the shape of the target block and the above schemes of configuring the merge candidate list based on the partition state of the target block. For example, n may be 1.

[0781] In an embodiment, when the QT depth of the target block is equal to or greater than n and the BT depth of the target block is less than or equal to m, the processing unit may configure the merge candidate list using at least one of the above-mentioned scheme of configuring the merge candidate list based on the shape of the target block and the above-mentioned scheme of configuring the merge candidate list based on the partition state of the target block. For example, n may be 3 and m may be 1.

[0782] The processing unit may determine a scheme for configuring the merge candidate list based on the position of the target block and the depth of the target block.

[0783] For example, when the BT depth of the target block is less than or equal to n and the target block is a block located at the bottom of the partition block generated by horizontal splitting, the processing unit may configure the merge candidate list using at least one of the above-mentioned scheme for configuring the merge candidate list based on the shape of the target block and the above-mentioned scheme for configuring the merge candidate list based on the split state of the target block. Here, n may be 1.

[0784] Configuration of AMVP candidate list

[0785] The processing unit may use the AMVP mode to derive inter prediction information for the target block.

[0786] The processing unit may configure an AMVP candidate list. The number of AMVP candidates in the AMVP candidate list may be N. N may be a positive integer. For example, the AMVP candidate list may include two AMVP candidates.

[0787] For example, such an AMVP candidate may be inter prediction information or motion information. Alternatively, the AMVP candidate may include a motion vector or a reference picture list.

[0788] The processing unit may configure the AMVP candidate list using one or more of the inter prediction information of the spatially neighboring blocks, the inter prediction information of the temporally neighboring blocks, and the combined inter prediction information.

[0789] The processing unit may derive one AMVP candidate from the inter-frame prediction information of the neighboring block to the left of the target block, and may derive one AMVP candidate from the inter-frame prediction information of the neighboring block above the target block. When the AMVP candidate list is not filled with candidates, the processing unit may derive additional AMVP candidates from the inter-frame prediction information of the temporally neighboring blocks.

[0790] The processing unit may derive an AMVP candidate using multiple pieces of inter-frame prediction information of neighboring blocks in a specific order of the neighboring blocks, and may add the derived AMVP candidate to an AMVP candidate list. Here, the AMVP candidate list may be configured differently according to the order of the neighboring blocks, and the prediction efficiency and encoding efficiency of encoding and decoding using the AMVP candidate list may vary according to the order of the neighboring blocks.

[0791] The processing unit may derive the AMVP candidate according to a specific order of the left neighboring blocks. Here, deriving the AMVP candidate according to a specific order of the left neighboring blocks may mean deriving the AMVP candidate using multiple pieces of inter-frame prediction information of the neighboring blocks selected in a specific order.

[0792] In the derivation of the AMVP candidate in a specific order of neighboring blocks, when the inter-frame prediction information of the neighboring block at the current sequential position is available, the processing unit may use the inter-frame prediction information of the neighboring block at the current sequential position to derive the AMVP candidate. When the inter-frame prediction information of the neighboring block at the current sequential position is not available, the processing unit may use the inter-frame prediction information of the neighboring block at the subsequent sequential position to derive the AMVP candidate. In other words, the processing unit may use the inter-frame prediction information of the previous neighboring block that appears first among the neighboring blocks and has available inter-frame prediction information to derive the AMVP candidate.

[0793] The fact that inter prediction information of a neighboring block is unavailable may mean that at least one of the following conditions 1) to 3) is satisfied.

[0794] 1) When there is no inter-frame prediction information of neighboring blocks

[0795] 2) The case where the neighboring block and the target block are included in different slices, tiles or pictures

[0796] 3) The case where the AMVP candidate derived using inter-frame prediction information is the same as another AMVP candidate already included in the AMVP list, that is, the case where the AMVP candidate derived using inter-frame prediction information is a duplicate AMVP candidate

[0797] For example, the processing unit may derive the AMVP candidate in the order of (A, B) when deriving the AMVP candidate using the left neighboring block. In other words, when the inter-frame prediction information of block A is available, the AMVP candidate derived by the inter-frame prediction information of block A may be included in the AMVP list, and when the inter-frame prediction information of block A is not available and the inter-frame prediction information of block B is available, the AMVP candidate derived by the inter-frame prediction information of block B may be included in the AMVP list.

[0798] For example, the processing unit may derive the AMVP candidate in the order of (B, A) when deriving the AMVP candidate using the left neighboring block.

[0799] For example, the processing unit may derive AMVP candidates in the order of (A, B, C, D, E) when deriving AMVP candidates using left neighboring blocks.

[0800] The processing unit may derive AMVP candidates in a specific order of the upper neighboring blocks.

[0801] For example, the processing unit may derive the AMVP candidates in the order of (K, J, F) when deriving the AMVP candidates using the upper neighboring blocks.

[0802] For example, the processing unit may derive the AMVP candidates in the order of (K, F, J) when deriving the AMVP candidates using the upper neighboring blocks.

[0803] When inter prediction information of neighboring blocks does not exist or is unavailable, in order to derive the AMVP candidate, the processing unit may use the combined inter prediction information instead of the inter prediction information of the neighboring blocks.

[0804] When the AMVP candidate list is configured using inter prediction information of spatially neighboring blocks, the processing unit may determine a scheme for configuring the AMVP candidate list based on the shape of the target block.

[0805] For example, when the height of the target block is greater than the width of the target block, the processing unit may configure the AMVP list using at least one of the following schemes 1) to 3).

[0806] 1) The processing unit may derive an AMVP candidate from the inter-frame prediction information of the upper neighboring block of the target block, and then may derive an AMVP candidate from the inter-frame prediction information of the left neighboring block of the target block.

[0807] 2) The processing unit may configure the AMVP candidate list using only multiple pieces of inter-frame prediction information of the upper neighboring blocks of the target block.

[0808] 3) The processing unit may configure the AMVP candidate list in the order of (J, B, K, A, F) or in the order of (J, B, A, K, F).

[0809] For example, when the height of the target block is smaller than the width of the target block, the processing unit may configure the AMVP list using at least one of the following schemes 4) to 6).

[0810] 4) The processing unit may derive an AMVP candidate from the inter-frame prediction information of the left neighboring block of the target block, and then may derive an AMVP candidate from the inter-frame prediction information of the upper neighboring block of the target block.

[0811] 5) The processing unit may configure the AMVP candidate list using only pieces of inter-frame prediction information of left neighboring blocks of the target block.

[0812] 6) The processing unit may configure the AMVP candidate list in the order of (B, J, K, A, F) or in the order of (B, A, J, K, F).

[0813] For example, when the height of the target block is equal to the width of the target block, the processing unit may configure the AMVP candidate list using at least one of the following schemes 7) and 8).

[0814] 7) The processing unit may configure the AMVP candidate list using at least one of the above solutions 1) to 6).

[0815] 8) The processing unit may configure the AMVP candidate list using multiple pieces of inter-frame prediction information of the left neighboring blocks of the target block and multiple pieces of inter-frame prediction information of the upper neighboring blocks of the target block.

[0816] The processing unit may determine a scheme for configuring the AMVP candidate list based on a partition state of the target block.

[0817] For example, when the target block is obtained as a result of vertical partitioning, the processing unit may derive an AMVP candidate from inter-frame prediction information of an upper neighboring block of the target block, and then may derive an AMVP candidate from inter-frame prediction information of a left neighboring block of the target block.

[0818] For example, when the target block is obtained as a result of horizontal splitting, the processing unit may derive an AMVP candidate from inter-frame prediction information of a left neighboring block of the target block, and then may derive an AMVP candidate from inter-frame prediction information of an upper neighboring block of the target block.

[0819] For example, the processing unit may configure the AMVP candidate list using the above-described scheme for determining the merge candidate list based on the partition state of the target block.

[0820] The processing unit may determine a scheme for configuring the AMVP candidate list based on both the shape of the target block and the division state of the target block.

[0821] For example, when the height of the target block is greater than the width of the target block or the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical division, the processing unit may derive an AMVP candidate from inter-frame prediction information of an upper neighboring block of the target block, and may then derive an AMVP candidate from inter-frame prediction information of a left neighboring block or an upper neighboring block of the target block.

[0822] For example, when the height of the target block is smaller than the width of the target block or the height and width of the target block are equal to each other, and when the target block is obtained as a result of horizontal splitting, the processing unit may derive an AMVP candidate from inter-frame prediction information of a left neighboring block of the target block, and may then derive an AMVP candidate from inter-frame prediction information of a left neighboring block or an upper neighboring block of the target block.

[0823] The processing unit may determine a scheme for configuring the AMVP candidate list based on the location of the target block.

[0824] The position of the target block may be the relative position of the target block in the upper block. By dividing the upper block, a plurality of partition blocks may be generated, and the target block may be one of the plurality of partition blocks. The position of the target block may be the position of the target block in the upper block, or the position of the target block between a plurality of partition blocks. The division may be two divisions or four divisions.

[0825] For example, in Fig.23 In the embodiment, the processing unit may apply a unified configuration method to an AMVP candidate list for a first partition CU, and apply an adaptive configuration method to an AMVP candidate list for a second partition CU.

[0826] The processing unit may determine a method for configuring the AMVP candidate list based on a depth of the target block. The depth of the target block may be at least one of a QT depth based on a quadtree (QT) partition and a BT depth based on a binary tree (BT) partition.

[0827] For example, the processing unit may determine a scheme for configuring the AMVP candidate list based on whether the depth of the target block falls within a specific range.

[0828] For example, when the BT depth of the target block is less than or equal to n, the processing unit may configure the AMVP candidate list using at least one of the above-mentioned scheme of configuring the AMVP candidate list based on the shape of the target block and the above-mentioned scheme of configuring the AMVP candidate list based on the partition state of the target block. For example, n may be 1.

[0829] For example, when the QT depth of the target block is equal to or greater than n and the BT depth of the target block is less than or equal to m, the processing unit may configure the AMVP candidate list using at least one of the above-mentioned scheme of configuring the AMVP candidate list based on the shape of the target block and the above-mentioned scheme of configuring the AMVP candidate list based on the division state of the target block. For example, n may be 3 and m may be 1.

[0830] The processing unit may determine a scheme for configuring the AMVP candidate list based on the location of the target block and the depth of the target block.

[0831] For example, when the BT depth of the target block is less than or equal to n and the target block is a block located at the bottom of the partition block generated by horizontal partitioning, the processing unit may configure the AMVP candidate list using at least one of the above-mentioned scheme of configuring the AMVP candidate list based on the shape of the target block and the above-mentioned scheme of configuring the AMVP candidate list based on the partition state of the target block. Here, n may be 1.

[0832] The details described in the above configuration of the merge candidate list and the AMVP candidate list can be applied to each other. For example, the features described in the derivation and addition of one of the merge candidate and the AMVP candidate can also be applied to the derivation and addition of the other candidate. Repeated descriptions will be omitted here.

[0833] Use sub-blocks to derive inter-frame prediction information for the target block

[0834] When deriving inter prediction information of the target block, the processing unit may use the inter prediction information of the subblock of the target block. In other words, the processing unit may use the inter prediction information corresponding to the unit of the subblock when deriving the inter prediction information of the target block.

[0835] The processing unit may divide the target block into a plurality of sub-blocks, and may derive inter-frame prediction information for each of the plurality of sub-blocks. For example, the processing unit may divide the target block into N sub-blocks, and may derive N pieces of inter-frame prediction information for the N sub-blocks. N may be a positive integer.

[0836] Fig. 27 Sub-blocks of temporally neighboring blocks and sub-blocks of a target block according to an example are shown.

[0837] When inter-frame prediction information of a temporally neighboring block is available, the processing unit may partition the temporally neighboring block into multiple sub-blocks, and may use multiple pieces of inter-frame prediction information of multiple sub-blocks of the temporally neighboring block to derive multiple pieces of inter-frame prediction information of multiple sub-blocks of the target block.

[0838] The processing unit may use multiple pieces of inter-frame prediction information of sub-blocks of temporally adjacent blocks to derive multiple pieces of inter-frame prediction information of sub-blocks of the target block. Here, the position of the sub-block of the target block within the target block and the position of the sub-block of the temporally adjacent block within the temporally adjacent block may be the same as each other.

[0839] For example, the processing unit may Fig.16 The temporal neighboring blocks N shown in FIG. 1 are partitioned into 4×4 temporal sub-blocks, and the inter-frame prediction information of each temporal sub-block can be used to derive multiple pieces of inter-frame prediction information of the 4×4 sub-blocks of the target block.

[0840] For example, the processing unit may Fig.16 The temporal neighboring block N shown in is partitioned into 2N×N temporal sub-blocks, and the inter-frame prediction information of each temporal sub-block can be used to derive multiple inter-frame prediction information of the 2N×N sub-blocks of the target block.

[0841] Fig.28 Spatial neighboring blocks of a target block and sub-blocks of the target block according to an example are shown.

[0842] exist Fig.28 , sub-blocks are indicated by upper case letters “A” to “P”, and spatially neighboring blocks are indicated by lower case letters “a” to “h”.

[0843] The processing unit may divide the target block into a plurality of subblocks, and may derive a plurality of pieces of inter prediction information of the subblocks of the target block using a plurality of pieces of inter prediction information of spatially neighboring blocks for the subblocks of the target block.

[0844] Here, the spatially adjacent blocks for the subblocks of the target block may include 1) additional subblocks adjacent to the corresponding subblocks of the target block and 2) blocks adjacent to the corresponding subblocks of the target block and also spatially adjacent blocks of the target block. In addition, the spatially adjacent blocks may be blocks that are encoded and / or decoded before the corresponding subblocks are encoded and / or decoded.

[0845] The processing unit may use multiple pieces of inter-frame prediction information of the spatial neighboring blocks of the sub-block to derive the inter-frame prediction information of the sub-block. The processing unit may use multiple pieces of inter-frame prediction information of multiple spatial neighboring blocks of the sub-block to derive the inter-frame prediction information of the sub-block. For example, the inter-frame prediction information of the sub-block may be an average value of multiple pieces of inter-frame prediction information of multiple spatial neighboring blocks of the sub-block.

[0846] For example, the processing unit may derive inter prediction information of subblock A using: 1) inter prediction information of spatial neighboring block d, 2) inter prediction information of spatial neighboring block e, or 3) an average of multiple inter prediction information of spatial neighboring block d and spatial neighboring block e.

[0847] For example, the processing unit may derive inter-frame prediction information of subblock K using an average value of multiple pieces of inter-frame prediction information of one or more blocks of block J, block F, block G, and block H, which are spatial neighboring blocks of subblock K.

[0848] When deriving inter prediction information of a subblock of a target block, the processing unit may simultaneously use a temporal neighboring block for the subblock of the target block and a spatial neighboring block for the subblock of the target block.

[0849] Use bilateral matching to derive inter-frame prediction information

[0850] Fig.29 The use of bilateral matching to derive inter-frame prediction information according to an example is shown.

[0851] The processing unit may use bilateral matching to derive inter-frame prediction information.

[0852] When performing bilateral matching, the processing unit may configure an initial motion vector candidate list for the target block, and may use at least one of one or more initial motion vector candidates included in the configured initial motion vector candidate list as the initial motion vector.

[0853] For example, the processing unit may use the AMVP mode to configure an initial motion vector candidate list for the target block. The AMVP candidate in the AMVP candidate list in the AMVP mode may be one or more initial motion vector candidates in the initial motion vector candidate list. The processing unit may add the AMVP candidate included in the AMVP candidate list to the initial motion vector candidate list.

[0854] For example, the processing unit may use the merge mode to configure an initial motion vector candidate list for the target block. The merge candidate in the merge candidate list in the merge mode may be one or more initial motion vector candidates in the initial motion vector candidate list. The processing unit may add the merge candidate in the merge mode to the initial motion vector candidate list.

[0855] For example, the processing unit may configure a frame rate up conversion (FRUC) unidirectional motion vector for the target block as the initial motion vector candidate list. The processing unit may add the FRUC unidirectional motion vector for the target block to the initial motion vector candidate list.

[0856] For example, the processing unit may configure the motion vectors of the neighboring blocks of the target block as the initial motion vector candidate list. The processing unit may add the motion vectors of the neighboring blocks of the target block to the initial motion vector candidate list.

[0857] For example, the processing unit may configure the combination of the above motion vectors as an initial motion vector candidate list. The number of combinations of motion vectors may be N or more. N may be a positive integer. The processing unit may add the combination of the above motion vectors to the initial motion vector candidate list.

[0858] For example, the processing unit may use a motion vector for at least one of a direction of the reference picture list L0 and a direction of the reference picture list L1 when configuring the initial motion vector candidate list. The processing unit may add a motion vector for at least one of a direction of the reference picture list L0 and a direction of the reference picture list L1 to the initial motion vector candidate list.

[0859] The processing unit may derive an initial motion vector for the target block when performing bilateral matching. The processing unit may derive the initial motion vector using an initial motion vector candidate list.

[0860] When performing bilateral matching, the processing unit may use the initial motion vector candidate list to derive a bidirectional motion vector that best matches the initial motion vector indicating block and the relative block with each other.

[0861] The initial motion vector indicating block may be a block indicated by the initial motion vector. The relative block may be a block existing on the same track as the initial motion vector indicating block in a direction opposite to the direction of the initial motion vector indicating block. In other words, the direction of the initial motion vector indicating block and the direction of the relative block may be opposite to each other, and the track of the initial motion vector indicating block and the track of the relative block may be the same as each other.

[0862] For example, Fig.29 As shown in , the processing unit may perform bilateral matching on a target block 2911 in a target picture 2910. When a motion vector present in the initial motion vector candidate list is MV0 in a reference picture reference 0 2920, the processing unit may derive a motion vector MV1, wherein 1) the motion vector MV1 exists in a reference picture reference 1 2930 present in a direction opposite to that of MV0, 2) the motion vector MV1 exists on the same track as MV0, and 3) the motion vector MV1 indicates a block 2931 that best matches the block 2921 indicated by MV0.

[0863] In other words, when the processing unit derives motion vector MV0 using the initial motion vector for the target block and determines motion vector MV1 according to MV0, 1) the direction of MV1 may be opposite to that of MV0, and the motion trajectory of MV1 may be the same as that of MV0.

[0864] The processing unit may refine the initial motion vector.

[0865] For example, the processing unit may search for blocks adjacent to the block indicated by the derived MV0, and may also search for blocks adjacent to the block indicated by the derived MV1. The processing unit may improve the initial motion vector so that the best matching block is indicated among the blocks adjacent to the block indicated by MV0 and the blocks adjacent to the block indicated by MV1.

[0866] When bilateral matching is performed, the processing unit may derive inter prediction information based on the sub-blocks. The inter prediction information may include motion information and / or motion vectors.

[0867] When deriving motion information based on sub-blocks, the processing unit may derive an initial motion vector for a sub-block using the above-described scheme for deriving an initial motion vector for a block.

[0868] When performing bilateral matching, the processing unit may define a degree of matching between the blocks. That is, the processing unit may use one of the specifically defined schemes when determining the degree of matching between the blocks.

[0869] For example, the processing unit may determine that two blocks are best matched to each other when the sum of absolute differences (SAD) between the two blocks is minimized. That is, the processing unit may determine that as the SAD between the two blocks is smaller, the two blocks are better matched to each other.

[0870] For example, the processing unit may determine that two blocks best match each other when the sum of absolute transform differences (SATD) between the two blocks is minimized. That is, the processing unit may determine that as the SATD between the two blocks is smaller, the two blocks match each other better.

[0871] Use template matching to derive inter-frame prediction information

[0872] Fig.30 The use of template matching mode to derive inter-frame prediction information according to an example is shown.

[0873] exist Fig.30 , a target block 3011 in a target screen 3010 is shown.

[0874] The processing unit may use template matching to derive inter-frame prediction information.

[0875] When performing template matching, the processing unit may use neighboring blocks of the target block as a template. The size and position of the template may be set based on a predefined scheme.

[0876] For example, the processing unit may use the neighboring block 3013 adjacent to the upper side of the target block 3011 as a template.

[0877] For example, the processing unit may use the neighboring block 3012 adjacent to the left side of the target block 3011 as a template.

[0878] For example, the processing unit may use a neighboring block 3013 adjacent to the upper side of the target block 3011 or a neighboring block 3012 adjacent to the left side of the target block 3011 as a template.

[0879] In an embodiment, when performing template matching, the processing unit may use the template to search for the block in the reference picture. Fig.30 , "reference 0 (3020)" may be a reference picture.

[0880] When performing template matching, the processing unit may use the template to derive inter-frame prediction information.

[0881] The processing unit may search for a block corresponding to the template in the reference picture. The shape and size of the block corresponding to the template may be the same as the shape and size of the template. The block corresponding to the template may be a block in the reference picture that best matches the template.

[0882] The processing unit may derive a motion vector indicating a block corresponding to the template. In other words, the processing unit may derive a motion vector indicating a block in the reference picture that best matches the template.

[0883] For example, in Fig.30 , a block 3022 corresponding to the template found for a neighboring block 3013 adjacent to the top of the target block and a block 3021 corresponding to the template found for a neighboring block 3012 adjacent to the left of the target block are depicted.

[0884] The processing unit may derive the motion vector of the target block using the motion vector of the block corresponding to the template. The processing unit may set the motion vector of the block corresponding to the template as the motion vector of the target block.

[0885] Here, the direction of the reference picture in which the block is found and the direction of the reference picture indicated by the motion vector of the target block may be opposite to each other.

[0886] The processing unit may improve the motion vector. For example, the processing unit may search for blocks adjacent to the block indicated by the current template and the derived motion vector, and may improve the motion vector so that the motion vector indicates the block that best matches the template among the found neighboring blocks. In other words, the improved motion vector may be a motion vector for the block that best matches the template among the blocks adjacent to the block indicated by the derived motion vector.

[0887] When performing template matching, the processing unit may derive inter-frame prediction information based on the sub-blocks. The inter-frame prediction information may include motion information and motion vectors.

[0888] When deriving motion information based on sub-blocks, the processing unit may derive motion vectors for sub-blocks using the above-described scheme for deriving motion vectors for blocks.

[0889] When performing template matching, the processing unit may define a degree of matching between blocks. That is, the processing unit may use one of the specifically defined schemes when determining the degree of matching between blocks.

[0890] For example, the processing unit may determine that when the SAD between two blocks is the smallest, the two blocks are best matched to each other. That is, the processing unit may determine that as the SAD between two blocks is smaller, the two blocks are better matched to each other.

[0891] For example, the processing unit may determine that the two blocks are best matched to each other when the SATD between the two blocks is minimized. That is, the processing unit may determine that as the SATD between the two blocks is smaller, the two blocks are better matched to each other.

[0892] Motion compensation and motion correction for inter-frame prediction

[0893] In inter prediction for the target block, the processing unit may perform inter prediction using at least one of motion compensation, IC, OBMC, BIO, affine spatial motion prediction and motion compensation, and motion vector correction in the decoding apparatus 1300 .

[0894] The processing unit may perform motion compensation to perform inter-prediction for the target block.

[0895] The processing unit may use the derived inter prediction information to generate a prediction block for the target block. Motion compensation may be unidirectional motion compensation or bidirectional motion compensation.

[0896] In an example, the processing unit may perform motion compensation using blocks in one picture present in one reference picture list L0.

[0897] In an example, the processing unit may perform motion compensation by combining a plurality of blocks in one picture present in one reference picture list L0.

[0898] In an example, the processing unit may perform motion compensation by combining a plurality of blocks in a plurality of pictures present in one reference picture list L0.

[0899] In an example, the processing unit may perform motion compensation by combining one block in one picture present in the reference picture list L0 with one block in one picture present in the reference picture list L1.

[0900] In an example, the processing unit may perform motion compensation by combining a plurality of blocks in one picture present in the reference picture list L0 with a plurality of blocks in one picture present in the reference picture list L1.

[0901] In an example, the processing unit may perform motion compensation by combining a plurality of blocks in a plurality of pictures present in the reference picture list L0 with a plurality of blocks in a plurality of pictures present in the reference picture list L1.

[0902] The processing unit may perform illumination compensation (IC) to perform inter prediction on the target block.

[0903] When performing motion compensation, the processing unit may compensate for a brightness change and / or an illumination change between a reference picture including a reference block for motion compensation and a target picture including a target block.

[0904] For example, the processing unit may approximate the variation between neighboring samples of the target block and the reference block as N or more linear models, and may perform illumination compensation by applying the linear models to the motion compensation block. N may be a positive integer.

[0905] Fig.31 An application of OBMC according to an example is shown.

[0906] The processing unit may perform OBMC to perform inter prediction for the target block.

[0907] The processing unit may generate a prediction block by combining the first block with the second block. The first block may be a block generated by compensation using inter-frame prediction information of the target block. The second block may be one or more blocks generated by compensation using one or more inter-frame prediction information of one or more neighboring blocks adjacent to the target block. Here, the one or more neighboring blocks adjacent to the target block may include a left neighboring block adjacent to the left side of the target block, a right neighboring block adjacent to the right side of the target block, an upper neighboring block adjacent to the top of the target block, and a lower neighboring block adjacent to the bottom of the target block.

[0908] The processing unit may perform OBMC on each sub-block of the target block.

[0909] The processing unit may generate a prediction block by combining the first block with the second block. The first block may be a block generated by compensation using inter-frame prediction information of a sub-block of the target block. The second block may indicate one or more blocks generated by compensation using one or more inter-frame prediction information of one or more neighboring sub-blocks adjacent to the sub-block of the target block. The inter-frame prediction information may include a motion vector. Here, one or more neighboring sub-blocks adjacent to the corresponding sub-block of the target block may include a left neighboring sub-block adjacent to the left side of the sub-block of the target block, a right neighboring sub-block adjacent to the right side of the sub-block of the target block, an upper neighboring sub-block adjacent to the upper side of the sub-block of the target block, and a lower neighboring sub-block adjacent to the lower side of the sub-block of the target block.

[0910] The processing unit may perform OBMC only on a specific sub-block among the sub-blocks of the target block.

[0911] For example, the processing unit may perform OBMC only on sub-blocks adjacent to an inner boundary of the target block.

[0912] In an example, the processing unit may perform OBMC on all sub-blocks of the target block.

[0913] For example, the processing unit may perform OBMC only on sub-blocks adjacent to the inner left boundary of the target block.

[0914] In an example, the processing unit may perform OBMC only on sub-blocks adjacent to an inner right boundary of the target block.

[0915] In an example, the processing unit may perform OBMC only on sub-blocks adjacent to an inner upper boundary of the target block.

[0916] In an example, the processing unit may perform OBMC only on sub-blocks adjacent to an inner lower boundary of the target block.

[0917] exist Fig.31 , an example is shown in which a CU as a target block is divided into PU1 and PU2, and OBMC is applied to the following sub-blocks among the sub-blocks of the CU: 1) sub-blocks adjacent to the upper boundary, 2) sub-blocks adjacent to the left boundary, and 3) sub-blocks adjacent to the boundary between PUs.

[0918] Fig.32 A sub-PU in ATMVP mode according to an example is shown.

[0919] The target block may be divided into a plurality of sub-blocks. The target block may be a target CU, and the sub-block may be a sub-PU. The processing unit may perform OBMC on all sub-blocks of the target CU.

[0920] The processing unit may generate a prediction block by combining the first block with the second block. The first block may be a block generated by compensation using inter-frame prediction information of each sub-block of the target block. The second block may be a block generated by compensation using multiple inter-frame prediction information of neighboring sub-blocks adjacent to the sub-block of the target block. The inter-frame prediction information may be a motion vector. Here, the neighboring sub-blocks adjacent to the corresponding sub-block of the target block may include a left neighboring sub-block adjacent to the left side of the sub-block of the target block, a right neighboring sub-block adjacent to the right side of the sub-block of the target block, an upper neighboring sub-block adjacent to the upper side of the sub-block of the target block, and a lower neighboring sub-block adjacent to the lower side of the sub-block of the target block.

[0921] The processing unit may perform affine spatial motion prediction and compensation to perform inter-frame prediction for the target block.

[0922] In an example, the processing unit may generate motion vectors for respective pixels in a target block by applying an affine transformation equation to a first motion vector at an upper left position of the target block and a second motion vector at an upper right position of the target block, and may perform motion compensation using the generated motion vectors.

[0923] In an example, the processing unit may generate motion vectors for each sub-block in the target block by applying an affine transformation equation to a first motion vector at an upper left position of the target block and a second motion vector at an upper right position of the target block, and may perform motion compensation using the generated motion vectors.

[0924] In order to provide the first motion vector and the second motion vector, at least one of the following schemes 1) to 3) may be used.

[0925] 1) The first motion vector and the second motion vector may be transmitted from the encoding apparatus 1200 to the decoding apparatus 1300 (through a bitstream).

[0926] 2) For each of the first motion vector and the second motion vector, a difference between its corresponding motion vector and a neighboring motion vector may be transmitted from the encoding apparatus 1200 to the decoding apparatus 1300 .

[0927] 3) The first motion vector and the second motion vector may be derived using affine motion vectors of neighboring blocks of the target block without being transmitted.

[0928] The processing unit may execute the BIO to perform inter prediction for the target block.

[0929] The processing unit may use optical flow using unidirectional blocks to derive a motion vector for the target block.

[0930] The processing unit may derive a motion vector of the target block using a bidirectional optical flow using a block existing in a picture before the target picture and a block existing in a picture after the target picture in a display order. Here, the block existing in the picture before the target picture and the block existing in the picture after the target picture may be a block having opposite motion and similar to the target block.

[0931] The processing unit may perform motion vector correction on the decoding apparatus 1300 to perform inter prediction on the target block. The processing unit may perform correction on the motion vector of the target block using the motion vector transmitted to the decoding apparatus 1300.

[0932] Fig.33 is a flowchart illustrating a target block prediction method and a bitstream generation method according to an embodiment.

[0933] The target block prediction method and the bitstream generation method according to this embodiment may be performed by the encoding device 1200. This embodiment may be a part of the target block encoding method or the video encoding method.

[0934] In step 3310, processing unit 1210 may derive inter-frame prediction information. Step 3310 may be similar to the above reference Fig.14 The described step 1410 corresponds.

[0935] At step 3320, processing unit 1210 may perform inter-frame prediction using the derived inter-frame prediction information. Step 3320 may be similar to the above referenced Fig.14 The described step 1420 corresponds.

[0936] In step 3330, the processing unit 1210 may generate a bitstream.

[0937] The bitstream may include information about the encoded target block. For example, the information about the encoded target block may include transformed and quantized coefficients of the target block.

[0938] The bitstream may include information for deriving inter prediction information and information for inter prediction.

[0939] In an example, the bitstream may include an indicator indicating a method for deriving inter prediction information.

[0940] For example, the bitstream may include an index indicating one of the candidates present in the list of the encoding device 1200 or the decoding device 1300 .

[0941] The processing unit 1210 may perform entropy encoding on information for deriving inter prediction information and information for inter prediction, and may generate a bitstream including a plurality of pieces of entropy-encoded information.

[0942] The processing unit 1210 may store the generated bitstream in the memory 1240. Alternatively, the communication unit 1220 may transmit the bitstream to the decoding apparatus 1300.

[0943] Fig.34 is a flowchart illustrating a target block prediction method using a bitstream according to an embodiment.

[0944] The target block prediction method using a bitstream according to this embodiment may be performed by the decoding apparatus 1300. This embodiment may be a part of a target block decoding method or a video decoding method.

[0945] In step 3410 , the communication unit 1320 may acquire a bitstream. The communication unit 1320 may receive a bitstream from the encoding apparatus 1200 .

[0946] The bitstream may include information about the encoded target block. For example, the information about the encoded target block may include transformed and quantized coefficients of the target block.

[0947] The bitstream may include information for deriving inter prediction information and information for inter prediction.

[0948] For example, the bitstream may include an indicator indicating a method for deriving inter-frame prediction information.

[0949] For example, the bitstream may include an index indicating one of the candidates present in the list of each of the encoding device 1200 and the decoding device 1300 .

[0950] The processing unit 1310 may store the acquired bit stream in the memory 1340 .

[0951] The processing unit 1310 may obtain information for deriving inter-frame prediction information and information for inter-frame prediction by performing entropy decoding on a plurality of pieces of entropy-encoded information in a bitstream.

[0952] In step 3420, processing unit 1310 may derive inter-frame prediction information. Step 3420 may be similar to the above reference Fig.14 The described step 1410 corresponds.

[0953] At step 3430, processing unit 1310 may perform inter-frame prediction using the derived inter-frame prediction information. Step 3430 may be similar to the above referenced Fig.14 The described step 1420 corresponds.

[0954] In the embodiments described above, although the method has been described based on a flowchart as a series of steps or units, the present disclosure is not limited to the order of the steps, and some steps may be performed in an order different from the order of the steps described or performed simultaneously with other steps. In addition, those skilled in the art will understand that the steps shown in the flowchart are not exclusive and may also include other steps, or one or more steps in the flowchart may be deleted without departing from the scope of the present disclosure.

[0955] The embodiments described above according to the present disclosure may be implemented as programs that can be run by various computer devices and may be recorded on a computer-readable storage medium. The computer-readable storage medium may include program instructions, data files, and data structures, either individually or in combination. The program instructions recorded on the storage medium may be specially designed or configured for the present disclosure, or may be known or available to a person of ordinary skill in the field of computer software. Examples of computer-readable storage media may include all types of hardware devices that are specially configured to record and run program instructions, such as magnetic media (such as hard disks, floppy disks, and tapes), optical media (such as compact disks (CD)-ROMs and digital versatile disks (DVDs)), magneto-optical media (such as floppy disks, ROMs, RAMs, and flash memories). Examples of program instructions include machine codes (such as codes created by a compiler) and high-level language codes that can be executed by a computer using an interpreter. The hardware device may be configured to operate as one or more software modules to perform the operations of the present disclosure, or vice versa.

[0956] As described above, although the present disclosure has been described based on specific details (such as detailed components and a limited number of embodiments and drawings), the specific details are only provided for easy understanding of the present disclosure, the present disclosure is not limited to these embodiments, and those skilled in the art will practice various changes and modifications based on the above description.

[0957] Therefore, it should be understood that the spirit of the present embodiment is not limited to the above-described embodiments and that the appended claims and their equivalents and modifications thereto fall within the scope of the present disclosure.

Claims

1. A video decoding method, comprising: Deriving inter-frame prediction information for a target block; and performing prediction for the target block using the inter-frame prediction information, Wherein, a first list including a plurality of first candidates is configured for the target block, performing said prediction using said first list, candidates in the second list including a plurality of second candidates are used to add candidates to the first list, The inter prediction information for the target block is added to the second list.

2. The video decoding method according to claim 1, wherein: The plurality of first candidates are generated based on a plurality of neighboring blocks adjacent to the target block, and The plurality of neighboring blocks include a leftmost block among blocks adjacent to an upper side of the target block and an uppermost block among blocks adjacent to a left side of the target block.

3. The video decoding method according to claim 1, wherein: The plurality of first candidates are generated by adding prediction information of a plurality of neighboring blocks adjacent to the target block to the first list according to a specific order, According to the specific order, the prediction information of the upper neighboring block is added to the first list first, and the prediction information of the left neighboring block is added to the first list second. The upper neighboring block is the rightmost block among the blocks adjacent to the upper side of the target block, and The left neighboring block is a lowermost block among blocks adjacent to the left side of the target block.

4. The video decoding method according to claim 1, wherein: One candidate in the first list is generated based on prediction information of a plurality of neighboring blocks adjacent to the target block.

5. The video decoding method according to claim 4, wherein: The one candidate is an average value of two pieces of prediction information of two neighboring blocks.

6. The video decoding method according to claim 4, wherein: The one candidate is generated based on prediction information of three neighboring blocks.

7. The video decoding method according to claim 1, wherein: The inter-frame prediction information includes a motion vector, and The motion vector is generated by improving the initial motion vector.

8. The video decoding method according to claim 7, wherein: The improvement is performed using the Sum of Absolute Differences (SAD) between two reference blocks.

9. A video encoding method, comprising: Deriving inter-frame prediction information for a target block; and generating an indicator indicating a method for deriving the inter-frame prediction information, Wherein, a first list including a plurality of first candidates is configured for the target block, performing prediction using the inter prediction information corresponding to one of the plurality of first candidates, candidates in the second list including a plurality of second candidates are used to add candidates to the first list, The inter prediction information for the target block is added to the second list.

10. The video encoding method according to claim 9, wherein: The plurality of first candidates are generated based on a plurality of neighboring blocks adjacent to the target block, and The plurality of neighboring blocks include a leftmost block among blocks adjacent to an upper side of the target block and an uppermost block among blocks adjacent to a left side of the target block.

11. The video encoding method according to claim 9, wherein: The plurality of first candidates are generated by adding prediction information of a plurality of neighboring blocks adjacent to the target block to the first list according to a specific order, According to the specific order, the prediction information of the upper neighboring block is added to the first list first, and the prediction information of the left neighboring block is added to the first list second. The upper neighboring block is the rightmost block among the blocks adjacent to the upper side of the target block, and The left neighboring block is a lowermost block among blocks adjacent to the left side of the target block.

12. The video encoding method according to claim 9, wherein: One candidate in the first list is generated based on prediction information of a plurality of neighboring blocks adjacent to the target block.

13. The video encoding method according to claim 12, wherein: The one candidate is an average value of two pieces of prediction information of two neighboring blocks.

14. The video encoding method according to claim 12, wherein: The one candidate is generated based on prediction information of three neighboring blocks.

15. The video encoding method according to claim 9, wherein: performing prediction using the inter prediction information decoded for the target block using the indicator, The inter-frame prediction information includes a motion vector, and The motion vector is generated by improving the initial motion vector.

16. The video encoding method according to claim 15, wherein: The improvement is performed using the Sum of Absolute Differences (SAD) between two reference blocks.

17. A computer-readable medium storing a bit stream, wherein: The bit stream is generated by an encoding device through the video encoding method of claim 9.

18. A method for transmitting a bit stream, wherein: The bit stream is generated by an image encoding device, wherein the method comprises: Send bitstream, Wherein, the bitstream includes coding information and a merge index for a target block; wherein the encoding information is used to perform decoding on the target block, wherein inter-frame prediction information for the target block is derived, wherein the prediction for the target block is performed using the merge index and the inter-frame prediction information, Wherein, a first list including a plurality of first candidates is configured for the target block, wherein the merge index indicates one of the plurality of first candidates to be used for the prediction, wherein the prediction is performed using the first list, wherein the candidates in the second list including a plurality of second candidates are used to add the candidates to the first list, The inter-frame prediction information for the target block is added to the second list.

Citation Information

Patent Citations

  • Method for processing video signal and device for same

    WO2016175550A1